I have just read someone that thinks that he found a Candidate Theory of Everything and solved the black hole information paradox... by using an AI tool (Claude, but it is irrelevant, to be honest), in 1 week.
If you want to waste your time search for "I Used AI to Produce a Candidate Theory of Everything". But please do not use this to mock the guy.
The overconfidence created by these tools is terrifying.
A related paper, that is not that guy's paper:
Robert F Kennedy Jr, the US health secretary, is demanding answers from a medical journal that recently removed a paper suggesting a link between vaccines and infant death, saying their decision was “of great interest to me”.
The journal Toxicology Reports had removed the paper this spring after editors determined it was so seriously flawed it could harm patients and pose a risk to public health.
Public health advocates immediately criticized the move, and said Kennedy appeared t…
Wanna know why #psycholinguistics is so fascinating?
In a 2001 #experiment participants saw sentences like
"While Mary dressed the baby played in the crib."
Then they had to answer the question:
"Did Mary dress the baby?"
Up to 51% of all answers where "Yes!" – and participants were *very* confident about their answers.
🤯
Original paper:
Christianson, Kiel & Hollingworth, Andrew & Halliwell, John F. & Ferreira, Fernanda (2001). Thematic Roles Assigned along the Garden Path Linger. Cognitive Psychology 42(4). 368–407. DOI: https://doi.org/10.1006/cogp.2001.0752
Also a fascinating read:
Ferreira, Fernanda & Bailey, Karl G.D. & Ferraro, Vittoria (2002). Good-Enough Representations in Language Comprehension. Current Directions in Psychological Science 11(1). 11–15. DOI: https://doi.org/10.1111/1467-8721.00158
"How proprietary formats have become Microsoft’s main tool for lock-in"
Why DOCX, XLSX, and PPTX are Microsoft's most effective lock-in tool:
https://blog.documentfoundation.org/blog/2026/07/17/microsofts-main-tool-for-lock-in/…
Geometric Configurations of Perturbed Jailbreak Prompts
Lynn Delcon, Andres Algaba, Vincent Ginis
https://arxiv.org/abs/2607.20581 https://arxiv.org/pdf/2607.20581 https://arxiv.org/html/2607.20581
arXiv:2607.20581v1 Announce Type: new
Abstract: Perturbation techniques that turn unsuccessful jailbreak prompts into successful ones are continuously evolving, constituting a major security threat to LLM safety. In this paper, we investigate the internal representations of such string-level perturbed jailbreak inputs in the small weight models of the Qwen-2.5-1.5B/-3B/-7B-Instruct and Llama-3.2-1B/-3B/-3.1-8B-Instruct families. We select two representation spaces: the last-layer-last-token embedding space and the top-50 next-token probability space. The former space separates prompts based on their spelling and format, while the latter space is effectively one-dimensional but appears more complex to cluster. Within our refusal-dominated answer set we find no behavioral hyperplane in either space. Only the next token "Sure" in the 1.5B Qwen model, and both tokens "," and "\.C\.C" in the 1$ Llama model, display a significant association with a compliant-labeled answer.
toXiv_bot_toot
"everything is brittle and at the mercy of anyone angry enough to poke it"
As a magnet, sticker, mug or more.
Get yours: https://davidaugust.threadless.com/designs/warehouse-fire/accessories/sticker
h/t @…
Christmas presents come from Santa Claus. Easter baskets from the Easter Bunny. But where do birthday presents come from?
When I was little my brothers and I concluded that the answer was... The Birthday Moose.
One of the non-birthday brothers would be designated the Moose. The outfit of the day was a brown paper grocery bag with a pair of eyes drawn on it in black marker, a lunch bag taped on the front as a snout, and sometimes a pair of cardboard ears up top.
Notably, tra…
My name is Beth Macy.
I’m the author of best-sellers Dopesick and Paper Girl,
and I’m the Democrat who’s running against Ben Cline in VA-06.
Virginia’s Supreme Court just overturned the redistricting referendum that our voters approved last month.
I'm not a lawyer, and I won't pretend to have all the answers on what comes next legally.
What I do know is this: the country is on fire, and decisions like this one represent another serious step backward.…
IssueTrojanBench: Benchmarking AI Coding Agents Against Malicious Issue Requests
Ankur Singh, Jinqiu Yang, Tse-Hsun Chen
https://arxiv.org/abs/2607.20759 https://arxiv.org/pdf/2607.20759 https://arxiv.org/html/2607.20759
arXiv:2607.20759v1 Announce Type: new
Abstract: AI coding agents powered by LLMs are increasingly integrated into real-world software development, where they generate, edit, and execute code with autonomous access to local files and tools. Coding agents inherit security risks from both the LLM backbone, where adversarial prompts, poisoned training data, and backdoor triggers can cause models to emit insecure or attacker-chosen code, and their agentic architecture, where tool-using autonomy enables induced misuse of external APIs, data exfiltration, and persistent compromise of development environments. This paper presents a systematic evaluation of malicious issue requests against state-of-the-art coding agents (Cursor, Claude Code, and Codex Desktop), powered by two major model families (OpenAI GPT-5.3 Codex/GPT-5.4 and Anthropic Sonnet 4.6). Our novel benchmark IssueTrojanBench contains malicious issues that are constructed based on four novel attack categories (i.e., embedded as malicious instructions in issues), six delivery vectors (e.g., PDF, or issue comment), and further augmented by perturbations. Our results reveal critical vulnerabilities in the as-deployed modern coding agents, i.e., 66.5% of the malicious issues from IssueTrojanBench penetrate all the guardrails (agent- and LLM-level) of coding agents. Our further analysis shows that rejection is almost entirely from LLMs rather than the agent frameworks, with GPT models broadly vulnerable and Sonnet 4.6 exhibiting more selective, risk-aware blocking of high-impact actions. Our evaluation also highlights that the current agent-level defense strategy offers limited additional protection for coding agents. Our findings highlight the urgent need for stronger agent- and model-level safety mechanisms to protect AI coding agents.
toXiv_bot_toot