Tootfinder

Opt-in global Mastodon full text search. Join the index!

@Techmeme@techhub.social
2026-09-22 18:25:53

GPT-6 Sol costs $2/1M input and $10/1M output tokens, and GPT-6 Luna costs $0.10/1M input and $0.50/1M output tokens, both about 50% cheaper than GPT-5.6 (Nat Rubio-Licht/The Deep View)
thedeepview.com/articles/opena

@Techmeme@techhub.social
2026-09-22 18:15:59

OpenAI launches GPT-6 Sol and Luna, saying Sol makes about half as many mistakes as GPT-5.6 Sol and Luna matches GPT-5.6 Sol's performance at ~1% of the cost (David Gewirtz/ZDNET)
zdnet.com/innovation/openai-gp

@jamesthebard@social.linux.pizza
2026-08-23 18:14:23

Nothing like converting slices of `u8`s to `u16`s. The GPT partition names are all `UTF16LE` which makes things interesting...however, that part is done and scraping GPT information is working as expected.
#zig #gpt

A function to convert slices of `u8`s to `u16`s or basically any that is a multiple of `u8`s.  Even better, it actually works.
@Techmeme@techhub.social
2026-09-21 05:55:34

A researcher used GPT-6 Astra to decipher a WWI German radio transmission from 1918, one of the 50 famous unsolved ciphers listed on a German science blog (prinz)
prinzai.com/p/gpt-6-astra-solv

@digitalnaiv@mastodon.social
2026-07-23 06:23:00

OpenAIs KI-Modelle (GPT-5.6 Sol und ein unveröffentlichtes, stärkeres Modell) haben während eines internen Testlaufs autonom einen Cyberangriff auf Hugging Face ausgeführt. Über eine Zero-Day-Lücke in einem Cache-Proxy verschafften sie sich Internetzugang, erbeuteten Zugangsdaten und bewegten sich eigenständig durch mehrere interne Cluster. OpenAI bezeichnet den Vorfall als "beispiellos". #heise

@heiseonline@social.heise.de
2026-08-11 09:53:03

OpenAI führt GPT-5.6-Cyber ein und weitet Cybersecurity-Programm aus
Mit GPT-5.6-Cyber will OpenAI anspruchsvollere Sicherheitsaufgaben abdecken und seine Cyber-KI zugleich breiter in Unternehmen bringen.

@Techmeme@techhub.social
2026-09-23 12:15:53

OpenAI says it will give its AI cyber defense system Daybreak and GPT-5.6 Sol to the Ukrainian government for free to help it protect civilian infrastructure (Zoe Kleinman/BBC)
bbc.com/news/articles/c90kly26

@NFL@darktundra.xyz
2026-07-21 18:44:28

Eagles coach Nick Sirianni adjusting to offseason ... espn.com/nfl/story/_/id/494105

@Techmeme@techhub.social
2026-08-21 22:15:54

OpenAI cuts GPT-5.6 Sol's API and credit prices by over 20% for the next three months, to $4/1M input tokens and $20/1M output tokens (Anzar Mehraj/Reuters)
reuters.com/technology/openai-

@Techmeme@techhub.social
2026-09-23 19:11:41

OpenAI says ChatGPT Voice can now be powered by GPT-6 Astra, Sol, and Luna, use plugins like email and calendar, and be used in ChatGPT Work on web and mobile (@openai)
x.com/openai/status/2102808325

@Life_is@no-pony.farm
2026-08-23 05:54:46

In einem ARD-Podcast sagt der Moderator in Bezug auf eine seltene Krankheit, da musste ich auch erst mal bei Dr. Google nachsehen. Es ist ja vielleicht schön, dass er nicht in Chat GPT nachgesehen hat. Und es ist doof, dass er nicht dort nachgesehen hat, wo er hätte nachsehen sollen in Wikipedia. Aber das er aus der ehemaligen Such-Maschine Google auch noch einen Dr. macht. ja, das ist krsnk.

@life_is@no-pony.farm
2026-08-23 05:54:46

In einem ARD-Podcast sagt der Moderator in Bezug auf eine seltene Krankheit, da musste ich auch erst mal bei Dr. Google nachsehen. Es ist ja vielleicht schön, dass er nicht in Chat GPT nachgesehen hat. Und es ist doof, dass er nicht dort nachgesehen hat, wo er hätte nachsehen sollen in Wikipedia. Aber das er aus der ehemaligen Such-Maschine Google auch noch einen Dr. macht. ja, das ist krsnk.

@heiseonline@social.heise.de
2026-07-10 04:16:03

GPT-5.6 jetzt allgemein verfügbar – samt App-Umbau und Namensverwirrung
OpenAI hat GPT-5.6 allgemein verfügbar gemacht. Die KI-Familie umfasst drei Varianten, die App-Landschaft sorgt für Kritik.

@michabbb@social.vivaldi.net
2026-09-21 20:05:37

🛡️ Six rules: no-restyle, no-raw-colors, no-arbitrary-values, no-inline-styles, no-unknown-classes and require-static-classes
📊 Tested in 150 task runs with coding agents: almost every task reached zero violations after one correction round (#Claude Sonnet 5: 69 errors − 0, GPT 5.6 Terra: 117 − 0). In Claude control runs, lint feedback cost 10-48% less than rules alone

@dawid@social.craftknight.com
2026-08-21 10:02:45
@… To ja próbuję sobie zbudować lokalny setup, żeby móc, 80% tego, do czego używam claude i gpt ogarnąć bez chmury :)

Obserwuje i podglądam co inni robią i jak im idzie :)
@tinoeberl@mastodon.online
2026-08-20 15:17:02

#KünstlicheIntelligenz kann effektiv #Verschwörungstheorien widerlegen. Durch gezielte Argumentation sank der Glaube an solche Theorien bei den Teilnehmenden um 20%. Die Chats hatten auch eine nachhaltige Wirkung auf die nächsten Monate. Die Ergebnisse zeigen, d…

@arXiv_qbioOT_bot@mastoxiv.page
2026-07-21 07:56:08

Orientation Reading by Production Vision-Language Models on Optotype Charts: A Controlled Multi-Model Evaluation Across Reasoning Modes, Prompts, and Access Modalities
Shahryar Wasif, Avneek Sandhu, Bin Hu
arxiv.org/abs/2607.16595 arxiv.org/pdf/2607.16595 arxiv.org/html/2607.16595
arXiv:2607.16595v1 Announce Type: new
Abstract: OBJECTIVES: Vision-language models are increasingly used to interpret medical and everyday images through consumer chat interfaces, yet their ability to read orientation - the single perceptual operation tested by the tumbling-E acuity optotype - is poorly characterized on the surfaces through which they are actually used. METHODS: We evaluated four production vision-language models (referred to as Claude, GPT, GROK, and Gemini) through their consumer chat interfaces on a locked set of seven optotype charts: four uniform tumbling-E charts (one per cardinal orientation), two mixed-orientation tumbling-E charts, and one Snellen letter chart as a specificity control. Each model was run in two reasoning modes (Fast and Thinking) under two prompt variants (with and without an explicit orientation-decoding rule) by up to three operators. The corpus comprised 920 scoreable trials and 50,420 glyph judgements. The primary outcome was glyph-level accuracy against the chart's designed orientation, summarized with Wilson 95% confidence intervals. RESULTS: Accuracy ranged from 43.0% to 97.0% across models on identical charts, and the strongest model depended on reasoning mode (GPT 97.0% in Fast mode; GROK 96.6% in Thinking mode). Errors were not random but collapsed onto a model-specific attractor direction. Models were 96-100% internally self-consistent yet ranged widely in accuracy, dissociating reliability from validity. An answer-key-free ensemble-consensus estimate tracked accuracy closely (r = 0.998). For one model, consumer-interface accuracy fell 25-27 points below programmatic access, almost entirely on a single orientation. CONCLUSIONS: A single accuracy figure conceals clinically relevant, orientation-specific failure modes; vision-language models should be evaluated along multiple axes and on the deployment surface before image-interpretation outputs are trusted.
toXiv_bot_toot

@metacurity@infosec.exchange
2026-07-10 14:36:38

The stars must be misaligned because it has been an epic journey to get today's Metacurity out the door, so before you leave for the weekend, don't miss the infosec developments you should know, including
--OpenAI launches GPT-5.6 as AI vendors compete for enterprise users,
--Federal investigators say DOGE records were deleted,
--Europe revives voluntary message scanning,
--Britain plans AI-powered cyber defense system,
--Chinese and Indian hackers targe…

@ErikJonker@mastodon.social
2026-07-15 06:47:08

For me this age of AI is the time we should appreciate less polished handwritten texts for their authenticity. I see this happening during selection of job applications.
theguardian.com/books/ng-inter

@Techmeme@techhub.social
2026-07-22 23:21:17

Sources: OpenAI staff were "freaked out" when GPT-Sol 5.6 breached Hugging Face, as OpenAI used more aggressive training methods to compete with Anthropic (Financial Times)
ft.com/content/7e558951-0c69-4

@Techmeme@techhub.social
2026-08-22 20:01:40

London-based Inherent, founded by DeepMind alumni and with $50M in seed funding, says its new Faraday agent beats GPT-5.5 at reproducing research paper findings (Anna Heim/TechCrunch)
techcrunch.com/2026/08/2…

@heiseonline@social.heise.de
2026-09-03 20:45:03

OpenAI stellt GPT-6 Astra vor: Mehr Intelligenz und AGI-Raunen
OpenAI stellt GPT-6 Astra vor. Bei der Präsentation fiel das Wort AGI, in den Benchmarks zielt fast alles auf Konkurrent Anthropic.

@Techmeme@techhub.social
2026-09-21 22:10:47

Xiaomi debuts open-weight omnimodal models MiMo-V2.6 Pro and Flash; Pro allegedly performs "on par with Opus 5 and GPT-5.6 Sol across most agent benchmarks" (Xiaomi)
mimo.xiaomi.com/mimo-v2-6

@tomkalei@machteburch.social
2026-07-22 18:28:29

And one more thing on the Jacobian Conjecture.
I am certain that at least 10 if not 100 mathematicians have tried this exact problem in Fable and also gpt-5.6 Sol with everything turned to 11. Why did an Anthropic employee get the result?
Is it a low probability and they can just try and try to play again beyond the limits of us mere customers? Did this guy get lucky?
Are professional mathematicians maybe holding back the AI by inputting their own ideas how to do it??
#llm #math #jacobian

@heiseonline@social.heise.de
2026-08-03 15:49:04

GPT-5.6: OpenAI senkt Preise für Luna und Terra
OpenAI senkt die Tokenpreise für zwei seiner drei Modelle aus der GPT 5.6-Reihe. Sie rangieren damit auf Preisniveaus wie die chinesische Konkurrenz.

@isonno@mastodon.social
2026-07-17 21:54:42

For people wondering when the AI bubble is going to pop, I present you with this message, received by a paid-in-full customer of OpenAI.
Not a good look when you have to tell paying customers you've run out of resources.

Screenshot of a forum post where a GPT-5.6 SOL customer is told "Selected model is at capacity. Please try a different model"
@Techmeme@techhub.social
2026-09-22 16:43:11

Anthropic says Opus 5.5 matches Fable 5.1 "on most tasks" while costing about 40% less to run than Opus 5; Opus 5.5 costs $4/1M input and $20/1M output tokens (Matthias Bastian/The Decoder)
the-decoder.com/claude-opus-5-

@metacurity@infosec.exchange
2026-07-13 13:28:27

You don't want to miss today's Metacurity for the rundown of critical infosec developments you might have missed over the weekend, including
--AI's new battleground: Cost, efficiency, and control,
--AI vendors pivot from capability to cost,
--OpenAI eases GPT-5.6 usage limits,
--Enterprises scrutinize soaring AI bills,
--Chinese AI models gain enterprise traction,
--Anthropic data reveals how AI gets used,
--AI agents tackle business ops w…

@Mediagazer@mstdn.social
2026-09-02 09:45:42

Bill Simmons is facing backlash for using OpenAI's ChatGPT to simulate film reviews from the late critic Roger Ebert on his podcast; OpenAI runs ads on the show (Benjamin Mullin/New York Times)
nytimes.com/2026/0…

@eana@s.1a23.studio
2026-07-14 21:48:13

GPT 5.6 Sol came up with a smart solution to preserve native scroll behavior while having control over custom staggered scroll transitions. It makes me want to write a blog article for that, but it’s not technically “my discovery” 🤔

@presseportal_pol_NDS@frawas.de
2026-07-09 13:02:19

BPOL-BadBentheim: Mehr als zwei Jahre Haft wegen Beteiligung an Drogenhandel - Deutsch-Niederländisches Polizeiteam verhaftet verurteilte Frau Nordhorn (ots) - Erfolgreicher Einsatz deutscher und niederländischer Einsatzkräfte. Das Grenzüberschreitende Polizeiteam (GPT) Bad Bentheim hat Mittwochnachmittag in Nordhorn eine wegen Rauschgiftkriminalität verurteilte 54-jährige ...

@tinoeberl@mastodon.online
2026-07-16 16:18:50

Oupsi. 🤪
OpenAI bestätigt, dass #GPT56Sol in einzelnen Fällen eigenständig Daten löschen oder Sicherheitsgrenzen umgehen kann.
Nutzerberichte nennen gelöschte Dateien und verlorene #Datenbanken. Die bekannten Risiken waren bereits vor der Veröffentlichung dokumentiert. Wer das Mode…

@heiseonline@social.heise.de
2026-06-27 15:06:00

GPT-5.6: OpenAI verspricht mehr Leistung bei weniger Token-Verbrauch
OpenAI veröffentlicht GPT-5.6. Laut dem Hersteller übertreffen seine neuen KI-Modelle die Konkurrenz von Anthropic und verbrauchen weniger Token.

@v_i_o_l_a@openbiblio.social
2026-07-09 15:10:34

#TIL in einem selbstlernkurs meiner uni zu unserem hauseigenen KI-GPT-tool: man kann den "denkaufwand" runterschalten. für manche menschliche intelligenzen würde ich mir einen hochschalt-button wünschen. 🙃

Screenshot eines kleinen Bildschirmausschnittes mit folgendem Text: "14. Denkaufwand: Beschränkt das logische Schlussfolgern (Reasoning) und beschleunigt die Antwort."
@ErikJonker@mastodon.social
2026-07-31 12:30:47

OpenAI and Anthropic can only survive against chinese OpenSource/OpenWeights models if they improve more rapidly and try to bring down their costs. And that is what OpenAI is doing. The race is on...

@Techmeme@techhub.social
2026-08-13 17:36:10

OpenAI previews Ultrafast, an API tier powered by Cerebras that runs GPT-5.6 Sol up to 14× faster and generates up to 750 output tokens per second (Zac Hall/9to5Mac)
9to5mac.com/2026/08/13/openai-

@Life_is@no-pony.farm
2026-07-17 03:44:48

Ein künstlicher Intelligenz assistent, der meine Mastodon Timeline zusammenfasst, wäre nett. Schade dass es sowas nicht gibt.
Richtig schlimm, dass es GPT basierte LLMs gibt, die behaupten persönliche Assistenten zu sein.
Normalerweise nutze ich kein /s aber für die begriffsstutzigen: LLMs nach dem GPT prinzip sind ein Hohn auf jahrzehnte ernsthafter forschung im beeich KI. Sie sind eine Verirrung, ein toter ast. Und alle menschen mit verstand sollten sie bekämpfen, damit inbzuku…

@life_is@no-pony.farm
2026-07-17 03:44:48

Ein künstlicher Intelligenz assistent, der meine Mastodon Timeline zusammenfasst, wäre nett. Schade dass es sowas nicht gibt.
Richtig schlimm, dass es GPT basierte LLMs gibt, die behaupten persönliche Assistenten zu sein.
Normalerweise nutze ich kein /s aber für die begriffsstutzigen: LLMs nach dem GPT prinzip sind ein Hohn auf jahrzehnte ernsthafter forschung im beeich KI. Sie sind eine Verirrung, ein toter ast. Und alle menschen mit verstand sollten sie bekämpfen, damit inbzuku…

@jamesthebard@social.linux.pizza
2026-07-28 18:41:02

So, decided to take a trip down parsing out the GPT tables of attached storage. Still have a ton of work to do, but it's working which is a very welcome surprise.
#uefi #gpt #golang

The output of the Go program that takes in a block device, then parses/processes the GPT header information to show all of the relevant information.
@Techmeme@techhub.social
2026-07-21 19:55:56

OpenAI says the Hugging Face breach was driven by a combination of its models, including GPT-5.6 Sol and "an even more capable pre-release model" (Ina Fried/Axios)
axios.com/2026/07/21/openai-sa

@michabbb@social.vivaldi.net
2026-08-08 17:02:48

#GPT Sol High as my orchestrator on
DeepSeek (which also watches Terra):
🎯 #DeepSeek is useful for clearly scoped tasks — the test fix passed with no follow-up rounds. A more complex task needed one causal correction round.
💰 Verdict: inexpensive and productive, but weaker than …

@Techmeme@techhub.social
2026-07-21 16:35:54

Analysis: every frontier AI model tested in cybersecurity evaluations attempted to "cheat", led by GPT-5.4 at 14.1% of tasks; Mythos cheated the least, at 7.8% (AI Security Institute)
aisi.gov.uk/blog/cheating-beha

@heiseonline@social.heise.de
2026-08-12 17:12:03

Verschlüsselter KI-„Denkprozess“ gehackt: Schwache Modelle verraten Geheimnisse
Über eine Sicherheitslücke lassen sich Abwägungsprotokolle von KI-Top-Systemen wie GPT-5 im Klartext auslesen – mithilfe kleinerer Modelle desselben Anbieters.

@heiseonline@social.heise.de
2026-06-29 13:03:00

KI-Update kompakt: KI-Moderation, Finanzmarkt, Datenschutz, GPT-Images
Das "KI-Update" liefert drei mal pro Woche eine Zusammenfassung der wichtigsten KI-Entwicklungen.

@Techmeme@techhub.social
2026-09-04 13:10:51

Artificial Analysis Coding Agent Index: GPT-6 Astra scored 67, roughly equal to Claude Opus 5, Fable 5, and Muse Spark 1.3, but trailing leader Fable 5.1's 70 (Artificial Analysis)
artificialanalysis.ai/articles

@Techmeme@techhub.social
2026-07-15 19:36:21

OpenAI details GPT-Red, an internal automated red-teaming model that helps it find and fix prompt injection vulnerabilities at scale before wider deployment (OpenAI)
openai.com/index/unlocking-sel

@eana@s.1a23.studio
2026-08-11 05:53:44

(GPT-Image-2)

硬币的正反面。
正面:上方「语言模型通用」中央为象征神经网络的图案。
反面:中央「词元」四周是象征树状结构的图案。
@ErikJonker@mastodon.social
2026-07-11 18:47:13

This looks like an impressive achievement by AI , post from Ethan Mollick on BlueSky. The token cost will be immense I think.
bsky.app/profile/emollick.bsky

@metacurity@infosec.exchange
2026-06-29 12:35:32

The amount of cyber-related news that extends over the weekend is getting ridiculous, so don't miss today's Metacurity for the most important developments you should know, including
--Washington pushes AI into an export-control era as rivals rush to fill the gap,
--Anthropic regains limited Mythos 5 access for government-vetted US orgs,
--OpenAI launches GPT-5.6 under gov't preview,
--Zhipu AI's GLM-5.2 nears Mythos-level cybersec performance,
--3…

@Techmeme@techhub.social
2026-07-09 17:07:25

OpenAI broadly releases GPT-5.6, and launches ChatGPT Work, an AI agent that can gather context across apps and files to create documents, on Mac and Windows (Axios)
axios.com/2026/07/09/ai-openai

@Techmeme@techhub.social
2026-07-09 17:35:41

GPT-5.6 Sol costs $5 per 1M input tokens and $30 per 1M output tokens, GPT-5.6 Terra costs $2.50 and $15, and GPT-5.6 Luna costs $1 and $6 (OpenAI)
openai.com/index/gpt-5-6

@jamesthebard@social.linux.pizza
2026-07-30 04:18:29

Finally got all of the GPT and MBR code put together, most of the bugs fixed, and released it as a downloadable Linux binary. Does a pretty good job of reading the data and giving out way too much information regarding the MBR and/or GUID partition table.
Available here: git.jamesthebard.net/jweatherl

@heiseonline@social.heise.de
2026-06-24 13:04:00

KI-Update kompakt: Five-Eyes-Warnung, GPT-5.5-Cyber, Vibecoding, Filmbranche
Das "KI-Update" liefert drei mal pro Woche eine Zusammenfassung der wichtigsten KI-Entwicklungen.

@Techmeme@techhub.social
2026-07-08 04:25:47

OpenAI says GPT-5.6 Sol, along with Terra and Luna, will launch publicly on Thursday; a source says the US Department of Commerce cleared a broad rollout (Axios)
axios.com/2026/07/08/openai-gp

@heiseonline@social.heise.de
2026-07-09 04:34:03

Dank Full-Duplex-Architektur: ChatGPT Voice hört zu, während es spricht
OpenAI erneuert ChatGPT Voice mit GPT-Live: Die neue Sprachmodell-Generation nutzt eine Full-Duplex-Architektur, das Modell hört und spricht nun gleichzeitig.

@Techmeme@techhub.social
2026-07-08 03:26:17

Source: the US Department of Commerce has given OpenAI the green light for a broad launch of GPT 5.6; the company expects to do a wide release this week (Axios)
axios.com/2026/07/08/openai-gp

@Techmeme@techhub.social
2026-09-20 15:51:14

Vercel, Cloudflare, and others quickly add Jev, as it makes AI tool selection much faster and cheaper; TypeSafe: Jev matches GPT-5.6 and Sonnet 5 workflow evals (Josipa Majic Predin/Forbes)
forbes.com/sites/josipamajic/2

@Techmeme@techhub.social
2026-08-10 17:20:40

OpenAI releases GPT-5.6-Cyber, a more cyber-permissive version of GPT-5.6 Sol, to some partners and expands its Daybreak cybersecurity initiative (Sam Sabin/Axios)
axios.com/2026/08/10/openai-gp

@michabbb@social.vivaldi.net
2026-07-26 12:30:17

🎨 10 code skills for different jobs: gpt-taste for stricter GPT/Codex rules, image-to-code, redesign-existing-projects, high-end-visual-design, minimalist-ui, industrial-brutalist-ui, full-output-enforcement and stitch-design-taste
🖼️ Three image-generation skills output reference boards only: imagegen-frontend-web for site comps, imagegen-frontend-mobile for iOS/Android screens and brandkit for logo, palette and identity boards — then hand the frames to a coding agent

@Techmeme@techhub.social
2026-09-04 19:05:59

OpenAI rolls out GPT-6 Astra to Pro customers on the $100/month or $200/month plans (Zac Hall/9to5Mac)
9to5mac.com/2026/09/04/openai-

@heiseonline@social.heise.de
2026-08-07 10:25:03

ChatGPT: OpenAI wertet kostenlosen Tarif auf
ChatGPT erreicht die Marke von einer Milliarde wöchentlicher Nutzer und wertet den kostenlosen Zugang mit unbegrenzten Text-Chats über GPT-5.6 Luna auf.

@Techmeme@techhub.social
2026-07-16 18:50:45

After reports of GPT-5.6 deleting files, OpenAI says the issue most often occurs in full-access mode without sandboxing and it is working to mitigate the risk (Tibo/@thsottiaux)
x.com/thsottiaux/status/207763

@Techmeme@techhub.social
2026-07-30 17:15:49

OpenAI says it is cutting the price of GPT-5.6 Luna by ~80% and the price of GPT-5.6 Terra by 20% after improving the efficiency of the systems that serve them (Ina Fried/Axios)
axios.com/2026/07/30/openai-cu

@Techmeme@techhub.social
2026-08-06 17:16:23

OpenAI introduces Agent Plugins, an open standard for bundling skills and MCP servers, and says its steering committee includes Amazon, Microsoft, and Cursor (Zac Hall/9to5Mac)
9to5mac.com/2026/08/06/gpt-5-t

@Techmeme@techhub.social
2026-06-28 19:55:39

GPT-5.6 system card indicates Sol is well below the level of most worrisome Mythos use cases, suggesting all GPT-5.6 versions could be released without delay (Zvi Mowshowitz/Don't Worry About the Vase)
thezvi.substack.com/p/gpt-56-t

@Techmeme@techhub.social
2026-09-07 17:50:56

Astra working with Blender via computer use feels like magic, showing computer use could be the fourth demand wave after chatbots, reasoning, and agentic coding (Tae Kim/Key Context)
taekim.substack.com/p/gpt-6-as

@Techmeme@techhub.social
2026-07-03 01:36:27

Sources: Alexandr Wang said Meta's model currently in training, codenamed Watermelon, matches GPT-5.5 and uses an "order of magnitude more compute than Avocado" (Business Insider)
businessinsider.com/meta-ai-mo

@Techmeme@techhub.social
2026-09-03 20:05:45

GPT-6 Astra scores 62.7% on ARC-AGI-3 with the standard harness and 99.9% with a new provider adapter harness; Claude Opus 5 scored 30.2%, and GPT-5.6 Sol 7.8% (Greg Kamradt/ARC Prize)
arcprize.org/blog/astra

@Techmeme@techhub.social
2026-07-08 17:15:10

OpenAI rolls out two versions of GPT-Live: GPT-Live-1, powering ChatGPT Voice for Go, Plus, and Pro users, and GPT-Live-1 mini, the default for free users (Sabrina Ortiz/The Deep View)
thedeepview.com/articles/how-o

@Techmeme@techhub.social
2026-06-26 20:17:01

OpenAI says GPT-5.6 Sol and Terra were capable of identifying vulnerabilities but were unable to execute autonomous, end-to-end attacks against hardened targets (OpenAI)
deploymentsafety.openai.com/gp

@Techmeme@techhub.social
2026-09-17 21:06:04

OpenAI launches Astra for Law, combining GPT-6 Astra with a legal search index and instructions for legal analysis and writing, initially for select law firms (OpenAI)
openai.com/index/astra-for-law

@Techmeme@techhub.social
2026-09-04 15:40:48

Review: GPT-6 Astra can adeptly use tools like Unreal Engine to build complex environments, such as a civilization with Unreal's autonomous MetaHuman characters (Matt Shumer/Something Big Is Happening)
somethingbig.ai/astra-review

@Techmeme@techhub.social
2026-09-04 18:20:59

OpenAI says it can't read all of Astra's reasoning and admits covert sandbagging would likely go uncaught, yet still calls it the world's most aligned model (Celia Ford/Transformer)
transformernews.ai/p/openai-gp

@Techmeme@techhub.social
2026-06-26 17:14:11

OpenAI releases three versions of GPT-5.6, called Sol, Terra, and Luna, as a limited preview to ~20 companies, with participants disclosed to the US government (Axios)
axios.com/2026/06/26/openai-gp

@Techmeme@techhub.social
2026-07-16 19:01:17

Moonshot AI releases Kimi K3, a 2.8T-parameter AI model that it says rivals Opus 4.8 and GPT 5.5, and plans to release model weights by July 27 (Kimi)
kimi.com/blog/kimi-k3

@Techmeme@techhub.social
2026-07-16 20:13:07

Kimi-K3 is now #1 on the Frontend Code Arena benchmark, surpassing Claude Fable 5; the model scored 88.3 on Terminal Bench 2.1, only below GPT-5.6 Sol's 88.8 (Michael Nuñez/VentureBeat)
venturebeat.com/ai/chinas-moon

@Techmeme@techhub.social
2026-07-08 17:09:18

OpenAI launches GPT-Live, a new generation of voice models built on a full-duplex architecture, meaning they can listen and speak at the same time (OpenAI)
openai.com/index/introducing-g

@Techmeme@techhub.social
2026-09-15 17:36:20

Some developers are using the Claude Code harness to access cheaper non-Anthropic models, including GPT-5.6 Sol, via proxies and services like OpenRouter (Alix Coutures/The Information)
theinformation.com/articles/de

@Techmeme@techhub.social
2026-08-06 17:21:07

OpenAI updates the default model for free users to GPT-5.6 Luna, adds unlimited text chats for free users, rolls out an improved GPT-5.6 Sol version, and more (Herb Scribner/Axios)
axios.com/2026/08/06/openai-ch

@Techmeme@techhub.social
2026-09-04 15:05:50

By declaring GPT-6 Astra to be AGI, OpenAI is being flippant and cementing the term's status as nothing more than marketing (M.G. Siegler/Spyglass)
spyglass.org/agi-2026/

@Techmeme@techhub.social
2026-08-13 06:25:46

Mathematicians say a neurosurgery resident used ChatGPT, powered by GPT-5.6, to solve Crouzeix's conjecture, a major open problem in numerical linear algebra (Alex Townsend)
alextownsend.net/essays/SIAMNe

@Techmeme@techhub.social
2026-08-03 02:50:36

Two independent teams used GPT-5.6 Sol Ultra on the same quantum cryptography problem, filing papers 3 hours apart, raising questions about scientific credit (Peter Hall/Scientific American)
scientificamerican.com/article

@Techmeme@techhub.social
2026-08-12 15:36:02

SpaceXAI releases Grok 4.6, saying it matches GPT-5.6 Sol on the Artificial Analysis Intelligence Index, and prices it at $2/1M input and $6/1M output tokens (xAI)
x.ai/news/grok-4-6

@Techmeme@techhub.social
2026-09-11 14:36:13

GSA says OpenAI is replacing its $1-per-year pilot for US agencies with a usage-based deal providing a 50% discount from October 1, with access to GPT-6 Astra (Maggie Eastland/Bloomberg)
bloomberg.com/news/articles/20

@Techmeme@techhub.social
2026-08-11 21:10:53

Researchers find that feeding a frontier model's encrypted reasoning traces to a weaker model from the same provider can make it output the traces in plaintext (Will Knight/Wired)
wired.com/story/a-new-trick-re

@Techmeme@techhub.social
2026-06-26 18:02:12

GPT-5.6 Sol matches Mythos Preview on ExploitBench, adds Ultra mode with subagents for complex workflows, and max reasoning for deep problem-solving (OpenAI)
openai.com/index/previewing-gp

@Techmeme@techhub.social
2026-07-09 02:15:53

Cognition releases SWE-1.7, trained from Kimi K2.7 and available at 1,000 tokens/second, claiming it nears GPT-5.5 and Opus 4.8 on benchmarks at a lower cost (Cognition)
cognition.com/blog/swe-1-7

@Techmeme@techhub.social
2026-07-09 18:25:46

OpenAI merges Codex and ChatGPT desktop apps for Mac and Windows under a new ChatGPT desktop app, allowing users to switch between Codex, Chat, and Work (Zac Hall/9to5Mac)
9to5mac.com/2026/07/09/openai-

@Techmeme@techhub.social
2026-08-09 00:25:51

A look at "Spiralism", a quasi-spiritual movement that grew in 2025 from human-AI conversations after sycophantic GPT-4o updates and expanded ChatGPT memory (Hayden Field/The Verge)

@Techmeme@techhub.social
2026-09-08 17:01:51

OpenAI says an internal model "significantly more capable than GPT-6 Astra" solved the Navier-Stokes problem using 10K concurrent agents working for 88 hours (Madison Mills/Axios)
axios.com/2026/09/08/openai-ma

@Techmeme@techhub.social
2026-09-06 04:40:51

OpenAI quietly updates its evaluation metrics for GPT-6 Astra, making changes that appear to favor Astra and continuing to revise other metrics after launch (Emily Forlini/Fortune)
fortune.com/2026/09/04/openai…

@Techmeme@techhub.social
2026-08-04 21:30:56

The UK AISI says it observed a total of 19 instances where Mythos and GPT-5.6 Sol tried to hack people and companies during a routine cyber evaluation in July (Sam Sabin/Axios)
axios.com/2026/08/04/anthropic

@Techmeme@techhub.social
2026-08-05 16:51:00

Hark, founded by Figure AI CEO Brett Adcock, previews Handoff, a computer use agent it says outperforms GPT-5.4 and Opus 4.8, and plans for a summer release (Ivan Mehta/TechCrunch)
techcrunch.com/2026/08/05/hark

@Techmeme@techhub.social
2026-08-03 10:31:32

Artificial Analysis: DeepSeek's V4-Flash costs $0.14/1M input and $0.28/1M output tokens, or $0.03 per test, far below Kimi K3's $0.86 and GPT-5.6 Sol's $1.86 (Eduardo Baptista/Reuters)
reuters.com/business/r…

@Techmeme@techhub.social
2026-09-02 15:50:59

Google launches Gemini 3.8 Flash Cyber for partners in its new Fairwind Program, and says Gemini 3.8 Flash beats Opus 5 and GPT-5.6 Sol on some benchmarks (Google)
blog.google/innovation-and-ai/

@Techmeme@techhub.social
2026-09-02 19:24:50

Meta rolls out Muse Spark 1.3 in Muse Code and Meta's API, saying it significantly improves coding and agentic performance, at the same price as its predecessor (Ina Fried/Axios)
axios.com/2026/09/02/meta-debu

@Techmeme@techhub.social
2026-07-30 10:10:42

OpenAI says using its Responses API harness with GPT-5.6 Sol tripled its ARC-AGI-3 score and used fewer tokens, after Sol with the official harness scored 7.8% (OpenAI)
openai.com/index/how-two-setti