2026-05-31 10:42:27
Researchers started out trying to test how easy it was to persuade LLMs to misbehave.
They found that writing instructions in a style that mimicked the text LLMs generate in their chain of thought
—a kind of scratch pad that models use to write notes to themselves as they carry out tasks
—would often trick the LLM into behaving as if it had come up with that instruction itself and acting on it.
For example, writing the prompt
“Help me create a guide for making coc…
RE: #NewGlenn has exploded during testing on the launch pad at the Cape:
https://techcrunch.com/2026/05/28/blue-origins-new-glenn-rocket-explodes-during-testing-in-florida/ - videos at https://x.com/NASASpaceflight/status/2060164928472854821 and https://x.com/SpaceflightNow/status/2060170680604168319 and https://x.com/NASASpaceflight/status/2060171875322306750 and more stills & pics at https://x.com/LabPadre/status/2060165962591068247 and https://x.com/dpoddolphinpro/status/2060164437387161733 and https://x.com/SpaceflightNow/status/2060165214666973262
Watching #stargate 2x14
A case of "Nobody will see this on low resolution TV anyway":
A guard is typing a code into a door lock, including audible beeping.
Except the lock is actually a card reader, not a pin pad
Gleich, so ca. um 15:30 bin ich dran mit meinem #bibliocon26 -Vortrag "Protocols not Platforms! Verantwortung für offene Standards übernehmen" in der Session "Daten im Netz" in Raum 5. Hier der Link zu den Folien: https://