PrismML releases Bonsai 2 27B, which compresses Alibaba's Qwen3.8 27B to 5.9 GB, small enough for smartphones, while retaining 98.2% of Qwen's benchmark scores (Julie Bort/TechCrunch)
https://techcrunch.com/2026/09/17/prismml-hopes-its-tin…
RE: https://eldritch.cafe/@miranda_blue/116935608086213327
This is showing one thing that I think should become general principle: Using an LLM for someone else _is rude and bad_. Presenting something untranslated that _they_ can have an LLM translate is better than doing it yourself, and not showing that input — the prompt here is not included and is a key piece of context, and the results are also bad _with no recourse_.
Information has been destroyed. Trust has been broken (or in the case of commerce, failed to be established). Even if the total amount of LLM-use were the same, this is worse than the reader using it.
PROMPTING AN LLM FOR SOMEONE ELSE IS RUDE.
#Atlassian nutzt alle Kundendaten für das Training von #KI. Wir erinnern uns an die zahlreichen Fälle, wo man aus #LLM teilweise die Trainingsdaten rekonstruieren konnte. 😜
Atlassian: Wer zahlt, bleibt verschont | D…
ChatGPT: Werbeeinblendungen kommen in weitere Länder
Werbung im eigenen Chat mit dem LLM: Das testet OpenAI gerade mit einem Teil der ChatGPT-Nutzer, auch Chatinhalte sollen einfließen.
https://www.
🤖 some #Git users may be interested into this bash script that analyses your git diff and creates meaningful commit messages using an #Ollama #LLM (local / cloud).
#MerriamWebster announces their new LLM. “There’s Artificial Intelligence, and then there’s Actual Intelligence.”
https://boingboing.net/2025/10/02/merriam-…
Am starting a series of blog posts on design and implementation of a AI Harness - based on the one from Fisk AI https://choria.io/fisk-ai
Will cover all the bits from what it is (todays post) to cover the loop, tools, memory, context, session history and more.
First post here
I think Google Health has sacrificed its math to a LLM.
Nur einmal angenommen, längere Texte schreibt man am besten mit Microsoft Word 5.5 für DOS. Für Behördenbriefe eignet sich Word für Windows 2.0 aber besser. Wenn der Brief auch noch Tabellen enthalten soll, geht das mit Word für Windows 2.0 zwar, aber viel besser eignet sich dafür Word 97. Ihr habt keine Tabellen, stattdessen aber Bilder? In einem viele Seiten langen Text? Dann greift man natürlich am besten zu Office Word 2007.
Wirrwarr bei LLM-Versionen, Symbolbild.
@… Five and a half hours of Opus 4.8 and I'm now looking at no errors whatsoever.
Honestly still unsure and uneasy about using LLM's for this, but it's hard to deny the results. This would have taken me a good week or two to do. Maybe even 3 or 4. But the real question now is: How much did it fuck up?
@… @… and given the existence of difficult to configure tools with crap documentation, an LLM is an appropriate response, and I say this neck deep in crap WorldPay documentation.
WorldPay have however …
Okay, the breakdown *by-language* of Artificial Analysis's "Omniscience" hallucination/capability benchmark is absolutely fascinating.
I feel like this is some of the first concrete evidence I've seen of my longtime lean: that the *correctness* tooling we obsess over in PLT spaces is one of the greatest possible value-adds to the LLM era. (Just look at the Rust row, holy shit.)
Choosing a good language basically shifts the quality of output by an entire cost-cat…
Idea: glue together the most horrific LLM hallucinated spaghetti code possible and release it as The_Aristocrats.exe
Why to use Grist for vibe-coding – an app being one file three things: how coherently you can build it with an LLM, how easily you can deploy, copy, and version the result, and how you secure it. https://www.getgrist.com/blog/why-is-grist-an-ideal-vibe-cod…
Replaced article(s) found for cs.LG. https://arxiv.org/list/cs.LG/new
[6/11]:
- Zero-Shot Active Feature Acquisition via LLM-Elicitation
Binyamin Perets, Natalie Mendelson, Shiran Vainberg, Yehuda Chowers, Shai Shen-Orr, Shie Mannor
Open-weight AI models, many developed by Chinese companies, are quickly expanding their role in enterprise AI systems, according to new research.
https://www.computing.co.uk…
The accuracy of LLM models varies depending on how many questions were asked a time? Structure your prompts wrong and even 'good' models fail? 😬
'The performance differences between frontier models, under realistic clinical input conditions, are smaller than the performance differences within a single model across different prompt structures'
But I would warn anyone circumventing anti-LLM protections that, as the tide turns against #LLM usage and "#AI" as a whole, that it may become politically useful for some state AGs to score political points. Those political points may come by putting you in prison, or, at the very least, dumping lawyers on you and ruining your life. An AG doesn't need to win to score points for their political party or to virtue signal. All they need to do is build a case against you. Because the CFAA is *so vague* it's easy to make something plausible enough that it won't get thrown out immediately.
So if the gentile "please don't do that" isn't enough, maybe consider this.
Also from Codeberg:
"You must not share projects that mostly consist of code written by "generative AI"-tools (including services such as *Claude*, *OpenAI Codex*).
Such projects having an unclear copyright status and furthermore have little safeguards to ensure that they do not include harmful code."
🧠 Routing strategies: LLM-as-classifier, signal-driven stage router, escalation router (weak tier first, a judge decides whether to escalate), random split for A/B tests, or a custom algorithm you write yourself
🚀 Launcher path: uv tool install "nemo-switchyard[cli]", then switchyard launch claude --model switchyard to point #ClaudeCode, Codex CLI or OpenClaw at an open-sou…
Old MacDonald's Server Farm
A-I-A-I-O
And on that farm an LLM
A-I-A-I-O
With a breakout here and a breakout there
here a break, there a break, everywhere a breakout
Old MacDonald's Server Farm
A-I-A-I-O
#AI #LLM #Security
В процесі іноді народжуються зрозумілі лише айтішникам жарти.
З *defineModel* мене познайомила саме LLM, це відносно нова фіча.
В Реакті, з якого я родом, щоб компонент перемикав state батька, треба прокидувати в нього як props
1) сам state
2) колбек на функцію, яка його змінює (бо сама функція має бути описана там же, де й state)
Vue концептуально робить те ж саме, але у нього є додаткова цукровість: можна прокинути щось як v-model (в свій кастомний компонент!), тод…
Today I braincoded some short SQL fixup while letting the LLM also generate a fix. The LLM was a tiny bit faster (I'd guess it took 60sec vs 90sec on my side), but I liked my solution more. I used a CTE to generate the fixed data and then did an update on select, while the LLM used a subselect.
Using #llm to OCR is...interesting. I've been trying it on an old hand written card; and it does very well on stuff I can only just about read; but the mistakes are way more subtle than trad OCR; e.g. one misread a 'standard 6' rack' as a 'standard 19" rack' (the other said '61 rack' which is more obviously wrong). The same one also turned 'M/coil mike'…
Claude Code today in an otherwise perfectly normal answer
"The gain would be близко to zero — [...]"
In case you are not proofreading your generated texts.
#AI #ClaudeCode #LLM
Metadatenabgleich schlägt fehl.
Mela: *verwundert* "Kann die Bibliothek denn keine konkreten Metadaten aus einer Datei extrahieren?"
LLM: "Doch, klar."
*grillenzirpen*
Mela: "Warum ziehst du dann nur die erste Seite, anstatt der Metadaten?"
LLM: "Die Metadaten stehen doch schon in der Datenbank, das wäre dann ja redundant
.."
*grillenzirpen*
*grillenzirpen*
*grillen verstummen*
Mela: "ABZU…
This is the “should I walk to car wash” question but without an LLM.
RE: s.1a23.studio/notes/aq089nipvi2xbfqc
This incite in to what my ubuntu system is doing that would have taken hours of time to research if I ever could have learned it at all had I not used an LLM. I'll continue to caveat that LLMs have many many things about them that are bad on multiple levels.
Using LLMs to fix system problems is doing something I wouldn't be able to do otherwise.
#ai
Mit dem LLM auf Distributionssuche – #linux
RE: https://norden.social/@thijs_lucas/117276776028587546
Ist KI verlässlich? Oder ist KI verlässlicher als mancher Lokaljournalismus?
Ich hab den KN-Artikel durch ein LLM laufen lassen und gebeten, zu prüfen ob und wie weit Positionen der CDU dari…
Cool idea:
"Researchers from Tracebit on Monday said they found that placing prompt injections alongside passwords, cryptographic keys, and other secrets stored on Amazon Web Services was often all that was needed to shut down attacks from AI hacking agents. The prompts direct the attacking LLM to perform an action forbidden by its guardrails, the safety barriers AI developers erect to prevent it from taking harmful actions. The LLM responds by shutting down."
"What Changes When an LLM agent Searches Your Library Catalogue?" asks Aaron Tay:
https://aarontay.substack.com/p/what-changes-when-an-llm-agent-searches
"I connected an LLM to search Primo and found it was excellent at academic database …
funniest thing in #LLM news. One man's month-long research to uncover the secrets of the Frontier Labs is another man's "hey I wrote this python deep_think script":
https://xcancel.com/_can1357/stat…
Current #LLM-driven degradation of software seems to follow an eternal pattern of humans and #technology, where the former seek the latter as a shortcut through reality and experience any resistance and complexity as loss and insult. In my publications I’ve called this “prelapsarian programming” bc it betray…
For me this age of AI is the time we should appreciate less polished handwritten texts for their authenticity. I see this happening during selection of job applications.
https://www.theguardian.com/books/ng-interactive/2026…
I don't understand how everyone can have such unconditional trust in LLM tools.
Most MCP servers are basically arbitrary code execution, sometimes even remotely.
The only thing preventing disaster is the unbroken belief that the word-predicting automaton surely won't be evil.
#LLM #AI
RE: https://mastodon.gamedev.place/@jonikorpi/117263464874714292
The push to replace UI/UX design with LLM generated outcomes is so disillusioning/shortsighted and sooner than later will land us in a clichéd, bland, semi-broken cesspit of good-eno…
Follow me for more high quality LLM training data. The more data we produce, the better informed AI models can become.
Sharing this not because it’s about AI/LLMs but because it’s solid advice and the recommendations work equally well for *people* coding web applications (funny that, almost like these things were made by people for people or something).
TL;DR: Use web and platform standards; don’t reinvent the wheel.
https://www.jimmo…
«Jadepuffer — Die erste KI-Ransomware-Attacke verbaselt nicht nur den Schlüssel:
Forscher dokumentieren den ersten LLM-gesteuerten Ransomware-Angriff. Doch die KI scheiterte an grundlegenden Fehlern. Eine Mahnung für Admins.»
Das die KI mittlerweile scheinbar "alles" hackt ist auch ein Teil ihres Marketings. Denn durch Unsicherheit lässt sich einiges besser als "Die Lösung" verkaufen.
🙄
I've only been back from holiday a day and I already feel like I'm burning out triaging the #qemu bug tracker. The flood of #llm assisted bug reports is not actually very helpful to the project even if they are real bugs. Just reporting more isn't going to help get the others fixed any faster.
Anthropic and other researchers detail how thousands of people were catfished by dating scam apps using LLM-generated replies from Claude and other models (Yael Grauer/The Verge)
@… wouldn't the deciding factor be how much a language is represented in the LLM training set vs the inherent properties of a language?
BTW it's Yossi Kreinin, not Kreinen
If your idea of a game is a text adventure like Zorg but with a LLM creating the story as you go then please reconsider.
There's already 100 of them and none of them works.
On one hand, I am heartened by seeing more and more people say things like "I've come to realize I can't trust the LLM to do the right thing", but then they shatter my heart by immediately following up with "So I make it do the tedious stuff, like writing the tests".
NO.
No. Stop it. NO.
The tests *are* the important part, they're the ones that prove your code does what it should. Be confident in the tests, and you can even afford to LLM gene…
Je viens de publier « Qu'est-ce que je dois faire signer Š mes clients pour être autorisé Š analyser leurs données confidentielles avec un LLM ? »
#LLM
Debian-Projekt vor Grundsatzentscheidung zu LLM-Einsatz
Das Debian-Projekt diskutiert über den grundsätzlichen Umgang mit LLM-Nutzung. Die Vorschläge für eine neue Generalresolution reichen von Verbot bis Erlaubnis.
⚡ One Docker command starts the full stack locally on port 3010. First startup pulls images in 2-3 minutes, then you pick inbound or outbound, name the bot, describe the use case in a few words and click Web Call.
🔑 No API keys required to start: #Dograh ships with auto-generated keys and its own LLM, TTS and STT stack. Own keys for LLM, TTS, STT or telephony can be connected later.
How I feel about #AI "Detectors"
#LLM
Bookmarked: TEI-NER Pipeline — Automatisches Entity-Linking für historische TEI-Editionen http://tei-llm.histomatiker.de/ #GND #LLM…
Insisto, ojalš unos juicios de Nüremberg contra los CEO de las IA generativas, LLM y similares.
On LLM watermarking: LLMs aren't doing anything other than statistical generation (accurately or not) _by design_.
Having them watermark stuff doesn't change anything about how they work, it's just ends up being slightly different statistics used for the stochastical generation.
In other words, they always have been and always will be watermarking their output; the only difference about the Anthropic watermarking thing is it's less computationally intensive (for them) to detect.
This is categorically different from e.g. steganography in printers.
Qwen 3.8 27B shows a 17GB open-weight general purpose model can have long context, effective tool calling, strong vision ability, and competent code generation (Simon Willison/Simon Willison's Weblog)
https://simonwillison.net/2026/Aug/16/qwen-38-27b/
Nur einmal angenommen, die Benutzung von LLMs und GenAI wäre Handwerk, einschließlich drei Jahre Lehrzeit zum Gesellen, Qualifizierung zum Handwerksmeister LLM und einer Handwerksrolle. Nach der bestandenen Gesellenprüfung ist es üblich, drei Jahre auf die Walz zu gehen um seine Erfahrungen zu schärfen. Erst danach kann man auch seinen Meister LLM-Handwerk machen. #justthinkin
I am likely to be chastised for my latest post in a (private) mailing list, but IDGAF.
"I do appreciate your transparency in making it clear that your message was from a LLM. If you didn't bother writing what you submit, I certainly will not bother reading it.”
RE: https://front-end.social/@fox/116899694179909261
I try very hard to do this on the daily. But where I don’t is when the LLM output is wrong. I’ll intentionally anthropomorphize it by saying it’s lying or it’s a liar. I’m rather hoping, sans evidence, that…
Crosslisted article(s) found for stat.AP. https://arxiv.org/list/stat.AP/new
[1/1]:
- Which Pairs to Compare for LLM Post-Training?
Jiangze Han, Vineet Goyal, Will Ma
I feel confident that not every person agrees on what "used LLM to make art" means.
Some might be obvious "write X" but other seems less obvious to me. LLM to market a book? LLM to fix an email to a publisher? LLM to make a grocery list so you have more time to do art?
I feel like all of that could count because everything in your life affects everything. Everyone could pick their own line or we could all agree on one but there is no clear answer.
Not that this is anywhere near a well hosted full frontier model. It does show that there can be creative avenues, along with the normal semiconductor and software learning curves, we won’t need those stinkin Data Centers and the chips they are buying today will probably be obsolete by the time any of these planned data centers might get built.
https://www.geeky-gadgets.com/run-llm-esp32-microcontroller/
Debian developer Antoine Le Gonidec just quit the project over its decision not to ban #LLM-assisted code contributions:
Pretending to have a "neutral stance" when faced with fascism is not neutrality, it’s active collaboration.
AI slop is fascism materialized. Stop it whenever you can!
Source:
I love that this includes 'What is Benchmarking and Why Should You Care?'
RISE-UNIBAS/humanities_data_benchmark: LLM Benchmark Suite for Humanities Data https://github.com/RISE-UNIBAS/humanities_data_benchmark
Am I an LLM? Just out of interest, I checked different variations of a piece of my very own writing I've done this weekend with GPTZero, and each time it says it's 100% AI generated, literally for every single sentence...
Such great motivation on a Monday AM! 😭
#Writing #LLM…
🖼️ Can images replace text as LLM context? A #DeepSeek paper claims one image token carries ~10 text tokens of information at close to 100% accuracy, with 59-70% cost reduction reported by PixelPipe. ThePrimeagen put it to the test. #AI
I have a little cow joke LLM agent made with my Fisk AI tool, I use it to test various.
I told it to save the jokes it already told me and always tell new ones, this is getting pretty bad after a bit lol
Je viens de publier « Septembre 2026 - je code avec des open weights pour 10 Š 30 € par mois »
#LLM
Finally, an honest explanation of AI data centers.
#ai #llm
hear me out, a LLM but trained exclusively on works by Joseph Weizenbaum, Hubert Dreyfus and Shoshana Zuboff
The thing about training an #LLM on human culture and then renting that culture back isn't that it's "Intellectual Property Theft." There's nothing wrong with sharing. There is something wrong with framing LLMs as piracy. That's a completely different concept.
When a powerful group of people takes the writing, the art, the music, various parts of a culture and then make some bland reproduction that lacks all the meaning and essence in the original, then those powerful people profit off that mess while the people who made the original stuff are erased, marginalized, and made to suffer, there is a different term for that.
The term everyone is looking for is "appropriation."
I'm not saying this is the same as other forms of cultural appropriation, but there are definitely enough shared elements that we should revisit that conversation. Perhaps some folks who thought the idea was silly when the idea first came up may feel differently now.
¿Qué sentido tiene usar un LLM para una afición? Es como hacer trampas jugando al solitario. ¿No disfrutas con el proceso de tu afición? Entonces, si lo haces, por qué tomar atajos. El disfrute estš mšs en el proceso que en el resultado.
Closing the floating action button for the LLM chat to reveal and close the floating action button for the accessibility overlay to reveal and close the floating action button for the shopping cart to reveal the floating action button for the original chat feature.
FAB archeology.
Your regular reminder:
ALL LLMs are ultimately just autocomplete. There is no thought, no creativity, no consciousness, no insight. It's just autocomplete with infinite monkeys. What turbocharged autocomplete can do is legit impressive, but it is still just turbocharged autocomplete.
Don't give it more credit than it deserves.
#AI
@… A lot harder, because the hard part of software engineering was always agreeing on what we’re trying to build, not actually building it.
And getting agreement on what we’re trying to do is a lot harder once nobody reads or engages with what everyone else is saying, but everyone is busy slinging 10 page LLM generated documents that nobody else reads anyway.
Sources: Apple trained a China-specific LLM with Alibaba's support, which would make Apple the first foreign company to offer a proprietary AI model in China (Reuters)
https://www.reuters.com/business/retail-co
heise | Lokale KI schneller machen: So zünden Sie den LLM-Turbo
Ein technischer Kniff bei lokalen LLMs kann die Generierungsrate massiv erhöhen: die Multi Token Prediction. Wir haben getestet, wie sie LLMs beschleunigt.
How many senior tech folk are honestly like 'I'm in this picture and I don't like it'?
'a lot of people who started out as developers are thrilled to get their hands dirty again with a bit of vibe coding – that using genAI makes them feel more in control of a tech stack that’s got away from them, perhaps even makes them feel younger and more dynamic'
Making progress with my OpenSCAD shelf model.
For the artwork, I vibe-coded* a converter from PNG to voxels (1x1x1 cubes) as OpenSCAD doesn't support textures.
*it took like ~2 mins to make a little Python script with a local LLM that's running on my MacBook Pro and is powered by our solar cells.
RE: https://front-end.social/@SaraSoueidan/116845563478631182
My experience is that LLM-derived writing about digital accessibility is wrong.
So not only do I abandon the page, I also mentally move the writer (LLM prompter) into a do-not-contra…
If the net result of LLMs on education is that homework becomes unworkable and is abandoned in favor of in-school work... I am OK with that.
#Homework #Education #AI
OH: "It has a certain gen-AI sais quoi."
#ai #llm
I really think for “AI” coding that local models are the future.
It’s basically free to run them on hardware you might already have, doesn’t have the environmental issues that data centers have and you don’t have to share your code or project details with anyone.
I’m a bit ambivalent about the ethics of the training data—for coding stuff that data is mostly open source code. I don’t have an issue with my MIT-licensed stuff used for this, but I know other authors do and their preferences should be honored.
I’m hoping that we get “clean” models for this in the future.
Generally I don’t think LLM-assisted programming is universally applicable to all software development but it can be a useful tool.
I also don’t think LLMs are useful for much else (coding is like a one in a million special case). ¯\_(ツ)_/¯
The code is here, if you're interested: https://github.com/madrobby/png-to-scad/
MIT launches the LLM Election Observatory, a dashboard tracking how nearly a dozen AI models tailor responses to political queries during the 2026 US midterms (Tiffany Hsu/New York Times)
https://www.nytimes…
Unconventional idea for cheap local #ai #LLM inference: repurposed PS5 chips.
The #AMD BC-250 is an ex-crypto-mining board (Zen2 CPU RDNA2 GPU, 16GB unified GDDR6) built around the same "Cyan S…
Who cleans up after the vibe-coding party?
#ai
Hugging Face says it used the open-weight GLM-5.2 hosted on its own compute for breach forensics, after US frontier model safety guardrails blocked the requests (Edward Targett/The Stack)
https://www.thestack.technology/hugging-fa