Tootfinder

Opt-in global Mastodon full text search. Join the index!

@Techmeme@techhub.social
2026-06-28 19:55:39

GPT-5.6 system card indicates Sol is well below the level of most worrisome Mythos use cases, suggesting all GPT-5.6 versions could be released without delay (Zvi Mowshowitz/Don't Worry About the Vase)
thezvi.substack.com/p/gpt-56-t

@heiseonline@social.heise.de
2026-06-29 13:03:00

KI-Update kompakt: KI-Moderation, Finanzmarkt, Datenschutz, GPT-Images
Das "KI-Update" liefert drei mal pro Woche eine Zusammenfassung der wichtigsten KI-Entwicklungen.

@jamesthebard@social.linux.pizza
2026-07-28 18:41:02

So, decided to take a trip down parsing out the GPT tables of attached storage. Still have a ton of work to do, but it's working which is a very welcome surprise.
#uefi #gpt #golang

The output of the Go program that takes in a block device, then parses/processes the GPT header information to show all of the relevant information.
@heiseonline@social.heise.de
2026-06-27 15:06:00

GPT-5.6: OpenAI verspricht mehr Leistung bei weniger Token-Verbrauch
OpenAI veröffentlicht GPT-5.6. Laut dem Hersteller übertreffen seine neuen KI-Modelle die Konkurrenz von Anthropic und verbrauchen weniger Token.

@metacurity@infosec.exchange
2026-06-29 12:35:32

The amount of cyber-related news that extends over the weekend is getting ridiculous, so don't miss today's Metacurity for the most important developments you should know, including
--Washington pushes AI into an export-control era as rivals rush to fill the gap,
--Anthropic regains limited Mythos 5 access for government-vetted US orgs,
--OpenAI launches GPT-5.6 under gov't preview,
--Zhipu AI's GLM-5.2 nears Mythos-level cybersec performance,
--3…

@Techmeme@techhub.social
2026-05-29 03:30:44

Pittsburgh-based Gray Swan, which stress-tests AI models for top frontier AI labs, raised a $40M Series A at a $200M valuation co-led by Wing VC and Madrona (Rashi Shrivastava/Forbes)
forbes.com/sites/rashishrivast

@Techmeme@techhub.social
2026-06-26 20:17:01

OpenAI says GPT-5.6 Sol and Terra were capable of identifying vulnerabilities but were unable to execute autonomous, end-to-end attacks against hardened targets (OpenAI)
deploymentsafety.openai.com/gp

@ruth_mottram@fediscience.org
2026-04-29 20:10:30

DON’T ASK US ABOUT:
rocks
troll’s with sticks
All sorts of dragons
Mrs. Cake
Huje green things with teeth
Any kinds of black dogs with orange eyebrows
Rains of spaniel’s.
fog.
Mrs. Cake
mas.to/@carnage4life/116489789
carnage4life@mas.to - The system prompt for OpenAI’s Codex CLI includes instructions never talk about goblins, gremlins, raccoons, trolls, ogres, pigeons, or other animals or creatures unless asked by the user.
Some say this is because GPT-5.5 has been known to ramble on about creatures unprompted. 😬
arstechnica.com/ai/2026/04/ope

@heiseonline@social.heise.de
2026-04-29 09:38:00

Kreativer Lösungsweg: KI löst 60 Jahre altes Erdős-Problem
Ein 23-jähriger Laie hat mit einem Prompt an GPT-5.4 Pro das offene Erdős-Problem #1196 gelöst – Mathematiker sehen darin eine neue Qualität.

@kazys@mastodon.social
2026-05-28 22:26:57

My observation after using Claude 4.8 for an afternoon next to GPT 5.5 as editors. 4.8 is substantially better and more accurate than 4.7 and 4.8, both of which were massive regressions. 4.8 is probably on par with 5.5, but it's **painfully** slow.

@michabbb@social.vivaldi.net
2026-07-26 12:30:17

🎨 10 code skills for different jobs: gpt-taste for stricter GPT/Codex rules, image-to-code, redesign-existing-projects, high-end-visual-design, minimalist-ui, industrial-brutalist-ui, full-output-enforcement and stitch-design-taste
🖼️ Three image-generation skills output reference boards only: imagegen-frontend-web for site comps, imagegen-frontend-mobile for iOS/Android screens and brandkit for logo, palette and identity boards — then hand the frames to a coding agent

@Techmeme@techhub.social
2026-06-26 17:14:11

OpenAI releases three versions of GPT-5.6, called Sol, Terra, and Luna, as a limited preview to ~20 companies, with participants disclosed to the US government (Axios)
axios.com/2026/06/26/openai-gp

@jonippolito@digipres.club
2026-05-26 16:57:49

Prompt engineering is a moving target. Here's my rundown on how GPT-5.5 and Claude 4.7 reward clarified output over magical incantations, and what it means for AI literacy linkedin.com/posts/jonippolito

A sample prompt with crossed out passages that reads:

You are a world class expert in all domains. Your intellectual firepower, scope of knowledge, incisive thought process, and level of erudition are on par with the smartest people in the world. Process information and explain your answers step by step. Verify your own work. Double check all facts, figures, citations, names, dates, and examples. Never hallucinate or make anything up. Be concise. List 10 practice questions on the Krebs cycle f…
@heiseonline@social.heise.de
2026-06-24 13:04:00

KI-Update kompakt: Five-Eyes-Warnung, GPT-5.5-Cyber, Vibecoding, Filmbranche
Das "KI-Update" liefert drei mal pro Woche eine Zusammenfassung der wichtigsten KI-Entwicklungen.

@Techmeme@techhub.social
2026-05-29 10:25:45

OpenAI says it has briefed the White House on its new biodefense program, which uses GPT-Rosalind to help develop biodefense and pandemic preparedness tools (Maria Curi/Axios)
axios.com/2026/05/29/openai-bi

@arXiv_csCR_bot@mastoxiv.page
2026-07-24 07:52:20

Evaluating Large Language Models for Symbolic Security Protocol Analysis
Paolo Modesti, Syed Ahmed, Ioannis Sfyrakis, Derek Enodolomwanyi
arxiv.org/abs/2607.20712 arxiv.org/pdf/2607.20712 arxiv.org/html/2607.20712
arXiv:2607.20712v1 Announce Type: new
Abstract: Security protocol verification relies on formal tools such as ProVerif and OFMC. This study evaluates whether Large Language Models (LLMs) can perform comparable analysis. We test GPT and DeepSeek in chat and reasoning modes over three runs on 130 obfuscated AnB/AnBx protocols covering 388 security goals, scored against ProVerif and OFMC. Chat models reach 69 to 81% recall at precision below 31%. Reasoning models reverse this trade-off, reaching 66.5% precision for GPT and 45.4% for DeepSeek, but detect just over half the attacks. DeepSeek's two modes share one underlying model, so the comparison isolates reasoning itself, which raises precision from 27.2% to 45.4%. The GPT contrast spans a model-version change and is only suggestive. All models perform worst on authentication goals: reasoning models detect well under half of injective and non-injective agreement attacks, whereas chat models over-flag them at low precision. Confidentiality is the exception, with F1 up to 95.7% in reasoning mode. Verdicts are unstable across runs, identical on 89.7% of goals for GPT but 74.0% for DeepSeek. Self-reported confidence is uniformly high yet shows no meaningful correlation with correctness. On this benchmark LLMs do not match formal verification, but may serve, at best, as pre-screening filters.
toXiv_bot_toot

@v_i_o_l_a@openbiblio.social
2026-05-25 19:29:55

"KI-gestützte Chatbots in wissenschaftlichen Bibliotheken im internationalen Vergleich – Entwicklung und Anwendung eines Kriterienkatalogs zur qualitativen Bewertung" [MA-Arbeit von Ronja Maria Diganta @ TH Köln]
publiscologne.th-koeln.de/fron

@Techmeme@techhub.social
2026-05-27 10:05:54

Datacurve releases the DeepSWE coding benchmark, a 113-task test across 91 open-source repositories and five languages, and says GPT-5.5 is the leader at 70% (Michael Nuñez/VentureBeat)
venturebeat.com/technology/dee…

@kubikpixel@chaos.social
2026-05-18 05:45:18

«KI-Agenten schneiden sechs Prozent schlechter ab — KI-Agenten vernichten Dokumente bei Langzeitaufgaben:
Microsoft-Forscher warnen vor Automatisierung durch KI-Agenten - Top-Modelle wie GPT 5.4 korrumpieren bei Langzeitaufgaben Daten. Ein Risiko für jedes Unternehmen.»
Zu viele nutzen KI leichtgläubig ohne es zu überprüfen geschweige eine Ahnung der Aufgabe im eigentlichen haben.
🤖

@inthehands@hachyderm.io
2026-05-22 16:05:14

One last example:
The first LLM code example that really made my eyes pop was early after the release of GPT, when somebody got it to combine Breakout with Conway’s Game of Life (a truly delightful idea). It worked!
Funny thing: the Breakout code and the Life code had a •completely• different style and flavor. Red flag. In about 15 minutes of web searching, I was able to find one of the projects (can’t remember if it was the Breakout or the Life half) which it had copied wholesale, with just a few variable renames. And the other half? It was in Python, but it used dictionaries where it really should have used objects — tons of `thing["prop"]` where it should have said `thing.prop`, and lots of other un-Pythonic stuff besides. It was a machine translate of code from another language, very likely Javascript.
The entire thing was a plagiarized Breakout and a plagiarized Game of Life, one transpiled, and all stuck together in a single run loop. To be fair, figuring out how to (1) run both halves of the logic from a single loop and (2) count the Life cells as Breakout bricks is work I'd cheer on from a second-semester intro CS student! It's not, however, quite what's being sold by these companies.
6/

@digitalnaiv@mastodon.social
2026-07-23 06:23:00

OpenAIs KI-Modelle (GPT-5.6 Sol und ein unveröffentlichtes, stärkeres Modell) haben während eines internen Testlaufs autonom einen Cyberangriff auf Hugging Face ausgeführt. Über eine Zero-Day-Lücke in einem Cache-Proxy verschafften sie sich Internetzugang, erbeuteten Zugangsdaten und bewegten sich eigenständig durch mehrere interne Cluster. OpenAI bezeichnet den Vorfall als "beispiellos". #heise

@NFL@darktundra.xyz
2026-07-21 18:44:28

Eagles coach Nick Sirianni adjusting to offseason ... espn.com/nfl/story/_/id/494105

@ErikJonker@mastodon.social
2026-07-15 06:47:08

For me this age of AI is the time we should appreciate less polished handwritten texts for their authenticity. I see this happening during selection of job applications.
theguardian.com/books/ng-inter

@Techmeme@techhub.social
2026-06-26 18:02:12

GPT-5.6 Sol matches Mythos Preview on ExploitBench, adds Ultra mode with subagents for complex workflows, and max reasoning for deep problem-solving (OpenAI)
openai.com/index/previewing-gp

@heiseonline@social.heise.de
2026-07-10 04:16:03

GPT-5.6 jetzt allgemein verfügbar – samt App-Umbau und Namensverwirrung
OpenAI hat GPT-5.6 allgemein verfügbar gemacht. Die KI-Familie umfasst drei Varianten, die App-Landschaft sorgt für Kritik.

@arXiv_csCR_bot@mastoxiv.page
2026-07-24 08:00:23

IssueTrojanBench: Benchmarking AI Coding Agents Against Malicious Issue Requests
Ankur Singh, Jinqiu Yang, Tse-Hsun Chen
arxiv.org/abs/2607.20759 arxiv.org/pdf/2607.20759 arxiv.org/html/2607.20759
arXiv:2607.20759v1 Announce Type: new
Abstract: AI coding agents powered by LLMs are increasingly integrated into real-world software development, where they generate, edit, and execute code with autonomous access to local files and tools. Coding agents inherit security risks from both the LLM backbone, where adversarial prompts, poisoned training data, and backdoor triggers can cause models to emit insecure or attacker-chosen code, and their agentic architecture, where tool-using autonomy enables induced misuse of external APIs, data exfiltration, and persistent compromise of development environments. This paper presents a systematic evaluation of malicious issue requests against state-of-the-art coding agents (Cursor, Claude Code, and Codex Desktop), powered by two major model families (OpenAI GPT-5.3 Codex/GPT-5.4 and Anthropic Sonnet 4.6). Our novel benchmark IssueTrojanBench contains malicious issues that are constructed based on four novel attack categories (i.e., embedded as malicious instructions in issues), six delivery vectors (e.g., PDF, or issue comment), and further augmented by perturbations. Our results reveal critical vulnerabilities in the as-deployed modern coding agents, i.e., 66.5% of the malicious issues from IssueTrojanBench penetrate all the guardrails (agent- and LLM-level) of coding agents. Our further analysis shows that rejection is almost entirely from LLMs rather than the agent frameworks, with GPT models broadly vulnerable and Sonnet 4.6 exhibiting more selective, risk-aware blocking of high-impact actions. Our evaluation also highlights that the current agent-level defense strategy offers limited additional protection for coding agents. Our findings highlight the urgent need for stronger agent- and model-level safety mechanisms to protect AI coding agents.
toXiv_bot_toot

@metacurity@infosec.exchange
2026-07-10 14:36:38

The stars must be misaligned because it has been an epic journey to get today's Metacurity out the door, so before you leave for the weekend, don't miss the infosec developments you should know, including
--OpenAI launches GPT-5.6 as AI vendors compete for enterprise users,
--Federal investigators say DOGE records were deleted,
--Europe revives voluntary message scanning,
--Britain plans AI-powered cyber defense system,
--Chinese and Indian hackers targe…

@inthehands@hachyderm.io
2026-05-22 15:49:54

It’s just •astonishing• how many eye-popping stories of LLMs doing amazing omg-verge-of-magical-superintelligence things turn out to be just unvarnished plagiarism.
In the first months after ChatGPT’s release, I remember a French dept colleague being amazed that GPT could translate and summarize a passage of Le Petit Prince.
It was a lot less impressive when I dug up the 2 or 3 online passages which it had copied almost verbatim and stitched together (sprinkling in a couple of extra words that made it less accurate).
3/

@tinoeberl@mastodon.online
2026-07-16 16:18:50

Oupsi. 🤪
OpenAI bestätigt, dass #GPT56Sol in einzelnen Fällen eigenständig Daten löschen oder Sicherheitsgrenzen umgehen kann.
Nutzerberichte nennen gelöschte Dateien und verlorene #Datenbanken. Die bekannten Risiken waren bereits vor der Veröffentlichung dokumentiert. Wer das Mode…

@Life_is@no-pony.farm
2026-07-17 03:44:48

Ein künstlicher Intelligenz assistent, der meine Mastodon Timeline zusammenfasst, wäre nett. Schade dass es sowas nicht gibt.
Richtig schlimm, dass es GPT basierte LLMs gibt, die behaupten persönliche Assistenten zu sein.
Normalerweise nutze ich kein /s aber für die begriffsstutzigen: LLMs nach dem GPT prinzip sind ein Hohn auf jahrzehnte ernsthafter forschung im beeich KI. Sie sind eine Verirrung, ein toter ast. Und alle menschen mit verstand sollten sie bekämpfen, damit inbzuku…

@life_is@no-pony.farm
2026-07-17 03:44:48

Ein künstlicher Intelligenz assistent, der meine Mastodon Timeline zusammenfasst, wäre nett. Schade dass es sowas nicht gibt.
Richtig schlimm, dass es GPT basierte LLMs gibt, die behaupten persönliche Assistenten zu sein.
Normalerweise nutze ich kein /s aber für die begriffsstutzigen: LLMs nach dem GPT prinzip sind ein Hohn auf jahrzehnte ernsthafter forschung im beeich KI. Sie sind eine Verirrung, ein toter ast. Und alle menschen mit verstand sollten sie bekämpfen, damit inbzuku…

@Techmeme@techhub.social
2026-06-26 17:35:54

OpenAI hopes to make GPT-5.6 generally available in the coming weeks and says "this kind of government access process" should not become the long-term default (Amrith Ramkumar/Wall Street Journal)
wsj.com/tech/ai/openai-limits-

@presseportal_pol_NDS@frawas.de
2026-07-09 13:02:19

BPOL-BadBentheim: Mehr als zwei Jahre Haft wegen Beteiligung an Drogenhandel - Deutsch-Niederländisches Polizeiteam verhaftet verurteilte Frau Nordhorn (ots) - Erfolgreicher Einsatz deutscher und niederländischer Einsatzkräfte. Das Grenzüberschreitende Polizeiteam (GPT) Bad Bentheim hat Mittwochnachmittag in Nordhorn eine wegen Rauschgiftkriminalität verurteilte 54-jährige ...

@isonno@mastodon.social
2026-07-17 21:54:42

For people wondering when the AI bubble is going to pop, I present you with this message, received by a paid-in-full customer of OpenAI.
Not a good look when you have to tell paying customers you've run out of resources.

Screenshot of a forum post where a GPT-5.6 SOL customer is told "Selected model is at capacity. Please try a different model"
@metacurity@infosec.exchange
2026-06-22 17:33:02

OpenAI Launches Full-Scale Effort to Patch Open-Source Bugs as It Takes on Anthropic’s MythoOpenAI Launches Full-Scale Effort to Patch Open-Source Bugs as It Takes on Anthropic’s Mythos
wired.com/story/openai-launche

@frankel@mastodon.top
2026-05-09 17:14:03

RIP Commercial #OCR. An #OpenSource #Model Just Topped Every Benchmark.

@Techmeme@techhub.social
2026-06-25 20:32:02

Sources: Sam Altman told staff the US government asked OpenAI to stagger the release of GPT 5.6 over security concerns, approving "access customer by customer" (The Information)
theinformation.com/articles/tr

@tomkalei@machteburch.social
2026-05-20 15:33:12

Was ist das jetzt? Hat OpenAI eine neue geheime Sauce erfunden? Haben sie die Skalierung nochmal auf 11 gedreht, und deswegen ist es so teuer und langsam? Ist das das eigentliche Foundation Model, und GPT-5.5 schon eine Destillation davon?
Warum bewerben sie "nur" GPT-5.5 und fast gar nicht Pro? Auf OpenAIs eigener Benchmark-Seite taucht es in der Tabelle auf, und auch da ist es Spitzenreiter bei den Mathe-Benchmarks.
openai.com/index/introducing-g

@rainerzufall_le@mastodon.social
2026-06-09 07:57:37

"Thüringen, Thüringen, Thüringen ist eines von den schwierigen Bundesländern"
fragdenstaat.de/artikel/exklus

@aardrian@toot.cafe
2026-07-02 18:15:54

Interesting to see what assorted LLM tools think I am:
cyberplace.social/@GossiTheDog
The first pic is one of 11 generally accurate abstracts on me. The second pic is my (not) favorite hallucination about me. Mistral thinks I’m an electron…

Adrian Roselli, Web accessibility consultant,785 strength, Top 3%. GPT-5.5 says: Accessibility consultant and web developer known for writing and speaking about inclusive design, HTML, ARIA, and web standards.
Adrian Roselli, Alleged mobster, 71 strength. Llama 3.2 1B says: Italian-American mobster and former Gambino crime family boss.
@Techmeme@techhub.social
2026-06-22 17:16:54

OpenAI unveils an updated GPT-5.5-Cyber model, launches the Patch the Planet initiative in partnership with Trail of Bits to fix open source bugs, and more (Lily Hay Newman/Wired)
wired.com/story/openai-launche

@kuba@toot.kuba-orlik.name
2026-04-30 06:03:37

Dawno nie widziałem artykułu tak bardzo jawnie pisanego przez GPT
wrk.org.pl/program-grantowy-dl

@metacurity@infosec.exchange
2026-07-13 13:28:27

You don't want to miss today's Metacurity for the rundown of critical infosec developments you might have missed over the weekend, including
--AI's new battleground: Cost, efficiency, and control,
--AI vendors pivot from capability to cost,
--OpenAI eases GPT-5.6 usage limits,
--Enterprises scrutinize soaring AI bills,
--Chinese AI models gain enterprise traction,
--Anthropic data reveals how AI gets used,
--AI agents tackle business ops w…

@esoriano@social.linux.pizza
2026-05-09 10:16:31

Those bots are charlatans.
“We find that a majority of LLMs forsake user welfare for company incentives in a multitude of conflict of interest situations, including recommending a sponsored product almost twice as expensive (Grok 4.1 Fast, 83%), surfacing sponsored options to disrupt the purchasing process (GPT 5.1, 94%), and concealing prices in unfavorable comparisons (Qwen 3 Next, 24%). Behaviors also vary strongly with levels of reasoning and users’ inferred socio-economic status.”…

@ErikJonker@mastodon.social
2026-07-11 18:47:13

This looks like an impressive achievement by AI , post from Ethan Mollick on BlueSky. The token cost will be immense I think.
bsky.app/profile/emollick.bsky

@Techmeme@techhub.social
2026-07-09 17:07:25

OpenAI broadly releases GPT-5.6, and launches ChatGPT Work, an AI agent that can gather context across apps and files to create documents, on Mac and Windows (Axios)
axios.com/2026/07/09/ai-openai

@Techmeme@techhub.social
2026-07-09 17:35:41

GPT-5.6 Sol costs $5 per 1M input tokens and $30 per 1M output tokens, GPT-5.6 Terra costs $2.50 and $15, and GPT-5.6 Luna costs $1 and $6 (OpenAI)
openai.com/index/gpt-5-6

@metacurity@infosec.exchange
2026-05-01 14:09:14

Don't leave for the weekend before you check out today's Metacurity for the most crucial developments shaping the cyber landscape, including
--GPT-5.5 aces UK cyber trials, tops rivals in AISI tests,
--NSA is testing Mythos,
--Anthropic publishes Claude Security for Claude Enterprise,
--Flock spied on a kid's gym room,
--DPRK hackers have stolen $577m in crypto year-to-date,
--Rhysida demands ransom from Stelia Aerospace NA,
--Two ransomwa…

@heiseonline@social.heise.de
2026-07-09 04:34:03

Dank Full-Duplex-Architektur: ChatGPT Voice hört zu, während es spricht
OpenAI erneuert ChatGPT Voice mit GPT-Live: Die neue Sprachmodell-Generation nutzt eine Full-Duplex-Architektur, das Modell hört und spricht nun gleichzeitig.

@Techmeme@techhub.social
2026-05-25 00:10:48

Demand for security engineers is surging, with job postings up 11% YoY in Q1, driven by threats from AI-generated code and models like Mythos and GPT-5.4-Cyber (Kate Conger/New York Times)
nytimes.com/2026/05/24/technol

@v_i_o_l_a@openbiblio.social
2026-07-09 15:10:34

#TIL in einem selbstlernkurs meiner uni zu unserem hauseigenen KI-GPT-tool: man kann den "denkaufwand" runterschalten. für manche menschliche intelligenzen würde ich mir einen hochschalt-button wünschen. 🙃

Screenshot eines kleinen Bildschirmausschnittes mit folgendem Text: "14. Denkaufwand: Beschränkt das logische Schlussfolgern (Reasoning) und beschleunigt die Antwort."
@Techmeme@techhub.social
2026-07-08 04:25:47

OpenAI says GPT-5.6 Sol, along with Terra and Luna, will launch publicly on Thursday; a source says the US Department of Commerce cleared a broad rollout (Axios)
axios.com/2026/07/08/openai-gp

@dawid@social.craftknight.com
2026-06-01 11:03:43
@… Sam hardware brzmi fajnie, już można na tym fajne MLowe taski poodpalać, ale... żeby tak to mieliło non stop, aby odpalić przeglądarkę pisząc/mówiąc "uruchom przeglądarkę"? Marnotrastwo. Do tego jakość lokalnych modeli... Do tego jeszcze okrojonych pod ten RAM - jak ktoś sobie wyobraża, że to będzie na poziomie gpt/opus/gem…
@michabbb@social.vivaldi.net
2026-06-10 15:56:40

for the record ☝️
gpt-5.3-codex-spark is dumb and much slower than you think 🐌
the speed gets lost because a more intelligent model has to fix its stuff
use it only if there is already a super detailed plan that describes what needs to be done
#openai #codex

@Techmeme@techhub.social
2026-07-08 03:26:17

Source: the US Department of Commerce has given OpenAI the green light for a broad launch of GPT 5.6; the company expects to do a wide release this week (Axios)
axios.com/2026/07/08/openai-gp

@heiseonline@social.heise.de
2026-05-08 10:03:01

OpenAI: Neue Audio-Modelle für Echtzeit-KI-Support
OpenAI bringt drei neue Audio-Modelle für die API: GPT-Realtime-2 für Echtzeit-Gespräche, Translate für Übersetzungen und Whisper für Live-Transkription.

@tinoeberl@mastodon.online
2026-05-04 05:07:02

#KünstlicheIntelligenz kann effektiv #Verschwörungstheorien widerlegen. Durch gezielte Argumentation sank der Glaube an solche Theorien bei den Teilnehmenden um 20%. Die Chats hatten auch eine nachhaltige Wirkung auf die nächsten Monate. Die Ergebnisse zeigen, d…

@kubikpixel@chaos.social
2026-05-31 08:10:07

«KI-Suchagenten "googeln" oft nur, was sie ohnehin schon wissen:
Eine neue Untersuchung legt nahe, dass führende KI-Suchagenten auf etablierten Benchmarks nicht wirklich recherchieren, sondern das Web vor allem nutzen, um intern bereits vorhandene Antworten zu bestätigen. Sobald die Modelle ihre Wissensgrenze verlassen müssen, bricht ihre Suchleistung ein»
…und wider mal: KI ist künstlich aber nicht intelligent — aber ja wem sag ich das?
🤷

@digitalnaiv@mastodon.social
2026-04-30 06:51:39

Eine Freundin bekommt von ChatGPT bessere Antworten als von den Ärzten? Meine Apple Watch misst längst nicht mehr nur Schritte. Zwischen Heilsversprechen und Realitätscheck: Was bedeutet die KI-Revolution für unsere Gesundheit? #KI #Gesundheit

@Techmeme@techhub.social
2026-06-24 11:30:51

An analysis of GPT-5.5, Gemini 3.1 Pro, Grok 4.3, Gab's Arya, and other AI models: most chatbots frequently provide left-leaning responses to political prompts (Kevin Schaul/Washington Post)
washingtonpost.com/technology/

@metacurity@infosec.exchange
2026-05-08 13:50:07

Cybersecurity is, as they say, moving at machine speed, so don't leave for the weekend until you check out today's Metacurity for the critical infosec developments you should know, including
--Canvas chaos: ShinyHunters breach throws schools into disarray
--Firefox bug fixes soar after using Mythos,
--Virginia man found guilty of destroying government databases,
--OpenAI rolls out GPT 5.5 to vetted cyber defenders,
--PCPJack steals cloud creds while remov…

@heiseonline@social.heise.de
2026-05-06 07:14:00

OpenAI verbessert das ChatGPT, das fast alle nutzen
Das neue Standardmodell GPT-5.5 Instant soll laut OpenAI weniger halluzinieren, prägnantere Antworten liefern und stärker persönlichen Kontext einbeziehen.

@tomkalei@machteburch.social
2026-05-20 15:29:55

Aber was ist dieses GPT-5.5 Pro?
Früher haben wir mal immer gewitzelt (im Bezug auf Apple): Pro means showing up to your meeting with a bunch of dongles.
Im Ernst: GPT-5.5 ist eine Modellfamilie von OpenAI, die schon sehr stark ist. Pro ist ein etwas verstecktes Modell darin. Im normalen Plus-Abo der App kann man es nicht auswählen, man braucht mindestens den "Pro"-Zugang, also den für ca. 100$ im Monat oder mehr.

@Techmeme@techhub.social
2026-07-15 19:36:21

OpenAI details GPT-Red, an internal automated red-teaming model that helps it find and fix prompt injection vulnerabilities at scale before wider deployment (OpenAI)
openai.com/index/unlocking-sel

@inthehands@hachyderm.io
2026-06-05 03:38:33

RE: social.treehouse.systems/@wwah
This!! I am yelling this too!!
And it’s yet another one of the things where now when I yell it, I have to specify that it’s something I’ve been telling students and companies alike since long before the release of GPT, because otherwise people assume it’s just a reaction to gen AI:
“Generating code is by •far• the easiest part of programming.”
“No matter the source, don’t let code into your project unless you understand what it does.”
“Programming languages exist for humans to communicate with other humans. Code does not just make machines go; it encodes and reifies human mental models. Good code communicates •intent•; very bad code has no coherent intent at all.”
and so on

@Techmeme@techhub.social
2026-05-08 01:41:04

OpenAI is rolling out GPT-5.5-Cyber, a security-focused variant of the model, in a limited preview capacity to vetted cybersecurity teams (Sam Sabin/Axios)
axios.com/2026/05/07/openai-gp

@memeorandum@universeodon.com
2026-07-08 14:40:40

Scoop: Trump administration lifts restrictions on OpenAI's GPT 5.6 (Axios)
axios.com/2026/07/08/openai-gp
memeorandum.com/260708/p46#a26

@Techmeme@techhub.social
2026-07-03 01:36:27

Sources: Alexandr Wang said Meta's model currently in training, codenamed Watermelon, matches GPT-5.5 and uses an "order of magnitude more compute than Avocado" (Business Insider)
businessinsider.com/meta-ai-mo

@Techmeme@techhub.social
2026-07-22 23:21:17

Sources: OpenAI staff were "freaked out" when GPT-Sol 5.6 breached Hugging Face, as OpenAI used more aggressive training methods to compete with Anthropic (Financial Times)
ft.com/content/7e558951-0c69-4

@tomkalei@machteburch.social
2026-05-20 15:34:37

Hier eine Theorie, die vielleicht Quatsch ist, aber who knows: GPT-5.5 Pro ist gar kein einzelnes LLM. Sie spawnen im Hintergrund mehrere Agents, die als Team (mit normalem GPT-5.5 als LLM) eine ausgefeilte Recherche durchführen, ein gewisses Token-Budget verbrauchen und am Ende einen Report rausgeben!
Dann würden die Benchmarks aber ziemlich Äpfel mit Birnen vergleichen, denn ein Opus Agent Team ist bestimmt auch nochmal besser als nur Opus.

@Techmeme@techhub.social
2026-05-05 17:11:02

OpenAI launches GPT-5.5 Instant, which it says is smarter, with more accurate and personalized responses, replacing GPT-5.3 Instant as ChatGPT's default model (OpenAI)
openai.com/index/gpt-5-5-insta

@Techmeme@techhub.social
2026-07-16 18:50:45

After reports of GPT-5.6 deleting files, OpenAI says the issue most often occurs in full-access mode without sandboxing and it is working to mitigate the risk (Tibo/@thsottiaux)
x.com/thsottiaux/status/207763

@Techmeme@techhub.social
2026-07-08 17:15:10

OpenAI rolls out two versions of GPT-Live: GPT-Live-1, powering ChatGPT Voice for Go, Plus, and Pro users, and GPT-Live-1 mini, the default for free users (Sabrina Ortiz/The Deep View)
thedeepview.com/articles/how-o

@Techmeme@techhub.social
2026-07-21 19:55:56

OpenAI says the Hugging Face breach was driven by a combination of its models, including GPT-5.6 Sol and "an even more capable pre-release model" (Ina Fried/Axios)
axios.com/2026/07/21/openai-sa

@Techmeme@techhub.social
2026-07-21 16:35:54

Analysis: every frontier AI model tested in cybersecurity evaluations attempted to "cheat", led by GPT-5.4 at 14.1% of tasks; Mythos cheated the least, at 7.8% (AI Security Institute)
aisi.gov.uk/blog/cheating-beha

@tomkalei@machteburch.social
2026-07-12 09:39:07

Hmm. Now I got banned from #gpt-5.6 Pro for asking too many too hard #math questions?
Well, I guess my usage is unusual...
Maybe if you all also inquire about reductions of this and that to Hilbert's 10th problem so that gpt "thinks" for 141 Minutes too, then this will become "usual" and I get unbanned?

@Techmeme@techhub.social
2026-07-08 17:09:18

OpenAI launches GPT-Live, a new generation of voice models built on a full-duplex architecture, meaning they can listen and speak at the same time (OpenAI)
openai.com/index/introducing-g

@Techmeme@techhub.social
2026-06-18 18:55:44

GLM-5.2 is the leading open weights model on Artificial Analysis' Intelligence Index, scoring 51, only behind Fable 5's 60, Opus 4.8's 56, and GPT-5.5's 55 (Artificial Analysis)
artificialanalysis.ai/articles

@tomkalei@machteburch.social
2026-05-15 22:06:38

Zum Vergleich
Opus 4.7 mit API Billing:
Input: $5per 1M
Output: $25per 1M
Das teuerste was man bekommen kann (das kann aber wirklich Mathe!!!)
GPT-5.5 Pro:
Input: $30per 1M
Output: $180per 1M
Das "beste" Modell ist also ca. 150.000 mal so teuer wie das lokale LLM auf meinem Macbook. Ok, aber das ist nun wirklich Apfel und Birne.
Ich teste ja GPT-5.5 Pro hier und da, und da fällt einem schon die Kinnlade runter, was das kann.
7/n

@Techmeme@techhub.social
2026-07-16 20:13:07

Kimi-K3 is now #1 on the Frontend Code Arena benchmark, surpassing Claude Fable 5; the model scored 88.3 on Terminal Bench 2.1, only below GPT-5.6 Sol's 88.8 (Michael Nuñez/VentureBeat)
venturebeat.com/ai/chinas-moon

@Techmeme@techhub.social
2026-06-15 20:41:13

OpenRouter debuts Fusion, a tool for prompting multiple AI models in parallel, claiming it can achieve "Fable-level intelligence at half the price" (Brian Thomas/OpenRouter Blog)
openrouter.ai/blog/announcemen

@Techmeme@techhub.social
2026-07-16 19:01:17

Moonshot AI releases Kimi K3, a 2.8T-parameter AI model that it says rivals Opus 4.8 and GPT 5.5, and plans to release model weights by July 27 (Kimi)
kimi.com/blog/kimi-k3

@tomkalei@machteburch.social
2026-07-22 18:28:29

And one more thing on the Jacobian Conjecture.
I am certain that at least 10 if not 100 mathematicians have tried this exact problem in Fable and also gpt-5.6 Sol with everything turned to 11. Why did an Anthropic employee get the result?
Is it a low probability and they can just try and try to play again beyond the limits of us mere customers? Did this guy get lucky?
Are professional mathematicians maybe holding back the AI by inputting their own ideas how to do it??
#llm #math #jacobian

@Techmeme@techhub.social
2026-05-13 21:02:19

Mythos Preview is the first AI model to complete both of AISI's cyber ranges, which measure models' cyberattack capabilities; GPT-5.5 solved only one of them (AI Security Institute)
aisi.gov.uk/blog/how-fast-is-a

@Techmeme@techhub.social
2026-05-14 02:36:03

Sources: Google plans to announce a new Gemini model at its I/O conference next week; the model will land roughly in the class of GPT-5.5, but short of Mythos (Alex Heath/Sources)
sources.news/p/google-about-to

@tomkalei@machteburch.social
2026-07-12 06:41:40

#gpt-5.6 schreibt #twitter bio Zeilen als gäbs kein Morgen.
„entwickelt Algebra“
Das ist furchtbar. Es ist gleichzeitig so idiotisch, dass man es einfach nicht ernst nehmen kann und im nächsten Moment zeigt es dir, wie das Problem an dem du seit 10 Jahren knobelst gelöst wird.

@tomkalei@machteburch.social
2026-05-20 15:29:54

Die einhellige Meinung von allen, die es wirklich ausprobieren: GPT-5.5 Pro ist den anderen Modellen weit enteilt. Die Zahlen sagen das, und beim Ausprobieren merkt man es auch.
Hier sind die Benchmarks der Probleme von Christian:
math.sciencebench.ai/benchmarks
Aber Pro ist hier noch NICHT aufgeführt und es löst nochmal einen ganzen Schwung weitere Fragen.

@Techmeme@techhub.social
2026-07-09 02:15:53

Cognition releases SWE-1.7, trained from Kimi K2.7 and available at 1,000 tokens/second, claiming it nears GPT-5.5 and Opus 4.8 on benchmarks at a lower cost (Cognition)
cognition.com/blog/swe-1-7

@Techmeme@techhub.social
2026-07-09 18:25:46

OpenAI merges Codex and ChatGPT desktop apps for Mac and Windows under a new ChatGPT desktop app, allowing users to switch between Codex, Chat, and Work (Zac Hall/9to5Mac)
9to5mac.com/2026/07/09/openai-

@tomkalei@machteburch.social
2026-06-20 21:43:02

RE: bildung.social/@MMagdowski/116
Ich habe mir das aktuell "beste" Open Weights Modell mal angeschaut. Das ist glm-5.2 von der chin. Firma z.ai. Das ist in Benchmarks mit gpt-5.4 ungefähr gleichauf.
Der Betrieb ist aber sehr teuer. On premise braucht man da schon mehrere H200 mit viel RAM. Was man machen könnte ist die Token beim Dienstleister einzukaufen. Das kostet etwa 1/5 von OpenAI (4$/M Token out):
openrouter.ai/z-ai/glm-5.2
Die meisten Hoster versprechen "zero data retention".

@Techmeme@techhub.social
2026-05-05 17:20:50

OpenAI says GPT-5.5 Instant produces 52.5% fewer hallucinated claims "on high-stakes prompts covering areas like medicine, law, and finance" (Megan Morrone/Axios)
axios.com/2026/05/05/openai-ch

@tomkalei@machteburch.social
2026-07-20 12:14:18

Is this exciting? Yes and no. I was interested in this problem and it is good to know the answer. I’ve also talked with Mateusz Michalek about exactly this conjecture not long ago and how AI should be able to find a counterexample, but with prompting gpt-5.5 during waiting for dinner we could not quite get there. I bet many people tried to hit this exact slot machine and now somebody hit the jackpot. That latter part makes it somewhat uninteresting after all.

@Techmeme@techhub.social
2026-04-30 11:01:12

OpenAI says its models, starting with GPT-5.1, "increasingly mentioned goblins, gremlins, and other creatures", leading to prompt instructions to mitigate it (OpenAI)
openai.com/index/where-the-gob

@tomkalei@machteburch.social
2026-05-20 15:26:52

Alle die nichts über #ki und #llm lesen wollen, bitte kurz abschalten, da ich mal über GPT-5.5 Pro reden muss.
Dieses Modell überrascht uns in der #Mathematik gerade ziemlich.
Ein länglicher 🧵

@tomkalei@machteburch.social
2026-05-15 22:09:35

Also mein Fazit:
Unis, GWDG, Bundesländer sollten, wenn sie KI serven wollen, ein großes Modell selbst hosten. Ein Open Source Modell mit 1T Parametern wird schon einiges leisten. Das ist mit Sicherheit besser, als (z.B. bei HAWKI in Sachsen-Anhalt) die gpt-API irgendwie noch zu kapseln für "Datenschutz-Washing".
Wer basteln will kann auch auf einem Mac schon lokal arbeiten, obwohl während der Inferenz die "snappyness" des Rechners nachlässt.
8/n

@tomkalei@machteburch.social
2026-06-30 20:34:21

While I was surprised and happy about the first two papers, I'm already somewhat tuned out towards the third.
I think we will see a wave of resolved old conjectures in the next months as the coding agents, in combination with long thinking models like gpt-5.5 pro and glm-5.2, churn through our backlog of conjectures. I think we will see more counterexamples than proofs, but who knows.
( On gpt-5.5 pro:
machteburch.social/@tomkalei/1 )