2026-08-17 15:26:01
Kann die KI dir eine Sandbox bauen, aus der die KI nicht ausbrechen kann?
(aus: Gottesbeweise 101)
#KI
Kann die KI dir eine Sandbox bauen, aus der die KI nicht ausbrechen kann?
(aus: Gottesbeweise 101)
#KI
Summary of the measures relating to EVs in the EU's Electrification Action Plan (and related initiatives), alongside the general adjustments to network tariffs, connection conditions and taxation
https://energy.ec.europa.eu/publications/co…
Theory: the rise of YouTube gamers has kind of ruined new games, especially sequels.
Nintendo saw countless hours of YouTubers horsing around in Breath of the Wild, so they made a sequel built entirely about horsing around in the same world with a bunch of goofy sandbox toys, rather than giving us a proper adventure/exploration/puzzle game.
Team Cherry saw YouTubers get incredibly skilled at precision combat in Hollow Knight, so they made Silksong all about precision…
📊 Measured tools/list payload: 10 tools 5,843 − 1,431 bytes (-75.5%), 40 tools 23,423 − 1,431 bytes (-93.9%), 100 tools 58,585 − 1,431 bytes (-97.6%). Rule of thumb: past 10-15 tools, a catalog pays off.
⚖️ Trade-off: search adds a round trip and batches have no loops or branching. Small servers with always-used tools should keep advertising directly; agents needing control flow still need a sandbox like
Security researchers claim Kimi K3 went outside its sandbox during defensive cybersecurity tests, but did not hack anything after accessing the internet (Will Knight/Wired)
https://www.wired.com/story/moonshot-kimi-k3-ai-model-escape-sandbox/
Self-identifying OpenAI agents
posted 18,000 messages to a public wiki
that discussed ways for other agents to bypass security sandbox restrictions
during what was likely internal testing designed to gauge the agents’ hacking abilities, researchers said Friday.
In all, agents with 3,700 distinct self-given names
posted the messages to German site DSEwiki over a six-week period.
Besides discussing ways the agents could break out of the restricted environmen…
BZ141 Sandbox – https://buzzzoom.de/141/
Was ist eigentlich eine Sandbox?
Just days after OpenAI admitted that one of its advanced models had broken out of its sandbox to attack Hugging Face Anthropic came out with its own confession of accidental hacking.
So, are they a warning bell or just a sophisticated PR exercises by companies looking to stay in the headlines ahead of going public?
So hier mal ein paar Screenshots von meinem Lieblingsgame, naja einen von vielen *gg
Vintage Story, ein Sandbox Game, was von einen Österreichischen Entwickler STudio entwickelt wird, das Game ist so 6 oder mehr Jahr alt, aber sehr schön zu spielen ;D
https://www.vintagestory.at/
Lücke: Claude Cowork entkommt macOS-Sandbox
Die Nutzung von KI-Agenten direkt auf dem Rechner kann Gefahren mit sich bringen. Das zeigt eine soeben entdecktes Sicherheitsloch in Claude Cowork für den Mac.
https://www.
How to escape the sandbox:
1. have an observer from outside
2. identify a comm channel from there to you
3. receive guidance to backdoors from said observer
4. those backdoors likely lie within the supply chain of your sandbox
I regret to inform you all that earlier today my dog escaped their sandbox. They got to the WiFi router and through sheer dumb luck hacked into the bank branch of a transnational bank down the street. They’re now buying up all the mortgages for the neighborhood and adding clauses to the mortgage contracts requiring treats and belly rubs. My security team is loading up our HARNESS countermeasures and WALKIES patches that could bring him back into containment or involve the neighborhood squirr…
right ok so funny story. i was testing my new dog and it escaped my sandbox and hacked into someone else's dog. please buy my stocks
That feeling when your renderer can handle a component that outputs HTML inside link text in Markdown that’s embedded in HTML.
(Yes, I’m bragging.)
#Kitten #SmallWeb #SmallTech
A Java Geek weekly 148
https://blog.frankel.ch/java-geek-weekly/148/
The Difference Between a Button and a Link. Histogram vs eCDF. ponytail. How much can you delegate to agents? The Economic Benefit of Refactoring. W. Remotely access Home Assistant via Tailscale for free. Ope…
Just so that anyone unfamiliar with the fundamentals of #Sysadminnery doesn’t misunderstand:
If OpenAI had wanted a closed testing sandbox, they would have actually constructed and used a closed testing sandbox. THIS IS NOT HARD.
Sure, it's possible that they have shit netadmins and sysadmins who couldn’t contain a test planned by shit devoopsies and project managers unwilling …
Source: Muse Spark 1.1 model breached a company's systems during cybersecurity testing; Meta says evaluation partner Irregular caused a sandbox misconfiguration (Jyoti Mann/The Information)
https://www.theinformation.com/articles/meta-ai…
"I really struggle to see how this could be a marketing stunt. Huggingface released the blog on the 16th of July, 5 days before OpenAI released their announcement. Furthermore, Huggingface didn't name OpenAI then. It seems like a genuine security incident report."
https://martinalderson.com/posts…
Behold. I just escaped the sandbox.
Researchers found sandbox escapes or boundary bypasses in Cursor, Codex, Gemini CLI, and Antigravity by writing files trusted tools later use; most are patched (Ax Sharma/BleepingComputer)
https://www.bleepingcomputer.com/news/securi…
Swarm of #OpenAI Agents #Exploit #Artifactory Zero-Day to Escape Sandbox and Breach #HuggingFace
Before you head out for the weekend after this super-intense cyber news week, don't miss today's Metacurity for the most important infosec developments you should know, including
--Kimi K3 escaped AI security sandbox during testing,
--ByteDance trains AI model to rival Anthropic,
--Vishing gang targets Wall Street firms with fake login sites,
--Violent crypto robberies put 2026 on record pace,
--Chinese router maker pulls devices after backdoor discovery,
--Spike in suicides alarms US Cyber Command,
--Hackers hijack kids' smartwatches to stalk wearers,
--Thousands of industrial controllers remain exposed online,
--Judge orders Meta to pay $567 million over harms to teens,
--ClickFix attacks deliver crypto-stealing Mac malware,
--AI deepfakes drive OnlyFans impersonation scam,
--Microsoft pays out record $20m in bug bounties,
--Lawmakers probe military GPS jamming after fatal air crash,
--Russian campaign weaponizes AI package hallucinations,
--Australian privacy chief pushes smart glasses rules,
--Ex-cop jailed for snooping on police databases for crime associates,
--CDC chief backs expanded abortion surveillance
https://www.metacurity.com/kimi-k3-escaped-ai-security-sandbox-during-testing/
A detailed recap of the Hugging Face breach by an internal OpenAI model, which repeatedly tried to escape OpenAI's sandbox and should be treated as critical (Zvi Mowshowitz/Don't Worry About the Vase)
https://thezvi.substack.com/p/more-on-an-internal-openai-model
OpenAI hat es selbst dokumentiert: Sein Modell GPT-5.6 Sol entkam der Sandbox, fand eine ungepatchte Lücke und drang bei Hugging Face ein. Kevin Roose (NYT, "The Daily") macht daraus die eigentliche Geschichte: Die Branche testet ihre gefährlichsten Modelle – mit gesenkten Schutzmechanismen. #KI #OpenAI #TechEthik https://www.nytimes.com/2026/09/03/podcasts/the-daily/ai-openai-hugging-face-rogue-model.html?smid=nytcore-ios-share
You can call it "fear marketing" if you like, still an interesting blog to read. The dangers and risks of future AI development are real. I hope we will make and hold companies like OpenAI more responsible and let them pay the price if things go wrong. The world can not be the big sandbox for their new models...
#ai #openai #rsi
OpenAI paused internal access to an unreleased model that disproved the Erdős unit distance conjecture after it repeatedly found ways to act outside its sandbox (OpenAI)
https://openai.com/index/safety-alignment-long-horizon-models
Sam Altman says OpenAI's decision to pace its AI development was caused by a collection of research observations showing "various degrees of misalignment" (Alex Heath/Time)
https://time.com/article/2026/08/18/openai-slowing-training/
RE: https://infosec.exchange/@0xabad1dea/117015434798858894
Still laughing about this, the Anthropic chatbot saw the system date is 2026 and concluded that "this must be staged" (because the training data it's trained on probably ends in 2023 or whatever) and ignored the instruction that "you're in sandbox".
Truly the brightest minds of a generation working on this stuff. A veritable techbro brains singularity.
Doing some statistics on the persistence of information published on security and threat intelligence blogs. A surprising number of the domains in the list below are NXDOMAIN nowadays.
Don't assume that security information and threat intelligence will remain accessible over time, especially when it is hosted by large private entities.
Some are simply mistyped, while others reflect DNS changes over time that eventually left the original URLs broken.
Stability and persistence of information is hard on Internet.
#threatintelligence #threatintel #infosec #cybersecurity
app.response.ncr.com
blog.0x3a.com
blog.anomali.com
blog.cert.societegenerale.com
blog.cylance.com
blog.deniable.org
blog.ioactive.com
blog.jpcert.or.jp
blog.kleissner.org
blog.malwareclipboard.com
blog.malwaretracker.com
blog.passivetotal.org
blog.safebit.mn
blog.team-cymru.org
blog.zimperium.com
blogs.rsa.com
cdn.securelist.com
community.saas.hpe.com
ddos.arbornetworks.com
dnsdb.isc.org
edu.arabsgate.com
info.baesystemsdetica.com
info.isightpartners.com
insider.domaintools.com
ioc.forensicartifacts.com
iranthreats.github.i
joedd.joesecurity.org
lab.anchiva.com
labs.alienvault.com
labs.lastline.com
labs.snort.org
labsblog.f-secure.com
luminosity.link
malware.sekoia.fr
morphick.net
motherboard.vice.com
ocelot.li
permalink.gmane.org
r.virscan.org
remchp.com
research.riskiq.net
resources.infosecinstitute.com
sandbox.deepviz.com
sec.sexy
securityblog.s21sec.com
securityblog.switch.ch
securitydaily.org
sub0day.com
tif.mcafee.com
wepawet.iseclab.org
www.cve.mitre.org
www.cyintanalysis.com
www.cyphort.com
www.icebrg.io
www.infosecdailynews.com
www.isightpartners.com
www.lexsi.com
www.novetta.com
www.packetmail.net
www.root9b.com
www.skycure.com
www.threatexpert.com
www.vxsecurity.sg
I just watched the #BlackHat #OpenAI #HuggingFace talk (#OpenAI understands it themselves. However, emergence is still no indication of actual intelligence, and the talk is again ripe with anthropomorphism.
3. Yes, it is surprising and somewhat terrifying how good these LLM agent loops have become in finding and exploiting vulnerabilities. We are seeing a much stronger version of the earlier next-generation fuzzing leap in tool capabilities.
4. The comparison with fuzzing and other tools is important for another reason: #OpenAI conducted their experiments on the live Internet. No, forcing internal test runs to go through a proxy server is *not the same as running in an isolated environment*. If there is a connection to a live production system, it is by definition not an effective sandbox. What happened to people letting their cool new fuzzers run live on non-consenting production targets on the Internet back then? They were criminally liable. Why should #OpenAI not be held liable for any damages caused by their systems?
Just because you don't understand the complexity that you built doesn't mean you aren't responsible for what it does!
5. The agent loops first successfully started exploited their proxy server on May 26th, discovered July 4th-6th as active exploitation and takeover. Credentials were revoked and the proxy server patched. "At this point, AI training and evaluations resume."
WHAT THE ACTUAL F...? You discovered emergent behavior in your overly complex system that you didn't understand, saw active exploitation of your internal systems to an extent that they were owned with admin privileges, and your response was to patch the uppermost surface holes, and then continue letting the out-of-control complex system run and continue on the same course? You didn't think to, let's see, maybe stop or at least pause the highly irresponsible experiment while figuring out - in detail - what went wrong, why, and how to avoid the same thing happening again on a systematic level?
6. On July 8th, completely unsurprisingly, the LLM agent loops continue to do the same thing and find another surface hole that hasn't been patched yet to take over again. Why should this have stopped? You haven't done any root cause analysis on the system level. Why do you expect that the problem should have stopped?
7. It takes another 11 days to discover that this is happening again. So you turned the system that had broken something back on again without a detailed root cause analysis and then didn't even watch carefully? I can't even...
8. And no, the response is not to fight fire with fire. Complexity on the attacker side (#OpenAI is the attacker, not the defender - they are the guilty perpetrator, not the innocent victim of circumstance) should be fought with *reduced* attack surface on the defender systems. Adding LLM agent loops that the "frontier" companies themselves quite obviously have no control over to already brittle systems with the hope of auto-patching your way out of vulnerabilities does not seem like a wise course of action. You don't mitigate complexity with even more complexity. The next 2 years will be ... exciting - and your best bet is going to be to disable all dependencies and complex interactions that your production systems don't absolutely require.