Sources: Mercor and other firms gathering data for AI labs are driving demand to buy or license internal datasets from startups shutting down or being acquired (Alix Coutures/The Information)
https://www.theinformation.com/articles/startup…
The optimization of neuroprosthetic interfaces relying on biophysical and surrogate digital twins https://www.nature.com/articles/s44385-026-00076-8 "the need to ensure computational feasibility has driven the adoption of substantial simplifications in the level of detail used t…
Læser om en super interessant sag om dataejerskab og national sikkerhed i The Continent: Namibia har droppet en kontrakt med amerikanske 6th Grain som skulle bruge satellitovervågning af marker til at lave AI-stŸttede dyrkningsplaner. Men selvom kontrakten lovede at alle data ville tilhŸre Namibia, var teksten for uklar til at man turde stole på det. GennemfŸrelse af planen kunne både betyde at landbruget i Namibia ville blive afhængig af udenlandske firmaer, og at USA ville sidde inde med v…
I also forgot to post a photo of the magnificent view of a walk I did a few weeks ago next to #Balquhidder, the "Creag-an-tuirc" viewpoint
https://www.walkhighlands.co.uk/lochlomond
GateTruth: Auditing the Rigor of RTL Design Benchmarks via Mutation Testing
Meet Bhadra
https://arxiv.org/abs/2608.12635 https://arxiv.org/pdf/2608.12635 https://arxiv.org/html/2608.12635
arXiv:2608.12635v1 Announce Type: new
Abstract: Benchmarks for evaluating large language models on register-transfer-level (RTL) hardware design have proliferated rapidly, yet none reports having applied mutation testing, an established hardware-verification technique for quantifying testbench quality, to ask whether its own testbenches are trustworthy. A testbench that never fails is not evidence of a correct design; it may simply never stimulate the logic that is actually broken. We introduce GateTruth, a mutation-testing engine and methodology for auditing RTL benchmark testbench rigor: inject a deterministic, seeded set of semantic mutants into a reference design and measure what fraction the testbench catches.
We validate the methodology against our own 68-task, dual-track reference suite -- 60 specification-to-RTL generation tasks and 8 agentic-repair tasks, scored through a pinned, deterministic synthesis-to-timing flow with correctness enforced as a strict gate -- certifying that 46 of 60 Track A testbenches kill at least 95% of injected mutants under sequential, reproducible execution; we disclose why the other 14 do not, including a Goodhart effect on testbenches revised to pass this gate. We then point the same engine, unmodified, at RTLLM v2.0, a widely adopted external benchmark: of 46 auditable designs, 72% fall below the 95% floor our own suite is held to, and three score 0% outright.
A comparable audit of NVIDIA's CVDP benchmark is structurally impossible: its public release withholds reference solutions, removing the golden RTL mutation testing requires. Auditing our own instrument also surfaced a second finding: an initially uniform 4096-token output cap silently truncated three of seven evaluated models, and re-running at 16,384 tokens moved one model from fifth place to first. We argue mutation-kill certification should become a standard reporting requirement for RTL-generation benchmarks generally.
toXiv_bot_toot
A visual imagery paradigm for #BCI strategies using imagined flickering patterns https://www.nature.com/articles/s41598-026-41324-6 (table correction at
What about privacy guarantees? Network connectivity loss? Availability after bankruptcy? Network-adaptive cloud processing for visual neuroprostheses https://arxiv.org/abs/2602.13216 Network-adaptive cloud preprocessing for visual neuroprostheses