Tootfinder

Opt-in global Mastodon full text search. Join the index!

@ErikJonker@mastodon.social
2026-07-15 06:47:08

For me this age of AI is the time we should appreciate less polished handwritten texts for their authenticity. I see this happening during selection of job applications.
theguardian.com/books/ng-inter

@arXiv_csCR_bot@mastoxiv.page
2026-07-24 07:52:20

Evaluating Large Language Models for Symbolic Security Protocol Analysis
Paolo Modesti, Syed Ahmed, Ioannis Sfyrakis, Derek Enodolomwanyi
arxiv.org/abs/2607.20712 arxiv.org/pdf/2607.20712 arxiv.org/html/2607.20712
arXiv:2607.20712v1 Announce Type: new
Abstract: Security protocol verification relies on formal tools such as ProVerif and OFMC. This study evaluates whether Large Language Models (LLMs) can perform comparable analysis. We test GPT and DeepSeek in chat and reasoning modes over three runs on 130 obfuscated AnB/AnBx protocols covering 388 security goals, scored against ProVerif and OFMC. Chat models reach 69 to 81% recall at precision below 31%. Reasoning models reverse this trade-off, reaching 66.5% precision for GPT and 45.4% for DeepSeek, but detect just over half the attacks. DeepSeek's two modes share one underlying model, so the comparison isolates reasoning itself, which raises precision from 27.2% to 45.4%. The GPT contrast spans a model-version change and is only suggestive. All models perform worst on authentication goals: reasoning models detect well under half of injective and non-injective agreement attacks, whereas chat models over-flag them at low precision. Confidentiality is the exception, with F1 up to 95.7% in reasoning mode. Verdicts are unstable across runs, identical on 89.7% of goals for GPT but 74.0% for DeepSeek. Self-reported confidence is uniformly high yet shows no meaningful correlation with correctness. On this benchmark LLMs do not match formal verification, but may serve, at best, as pre-screening filters.
toXiv_bot_toot

@NFL@darktundra.xyz
2026-07-21 18:44:28

Eagles coach Nick Sirianni adjusting to offseason ... espn.com/nfl/story/_/id/494105

@Techmeme@techhub.social
2026-07-09 18:25:46

OpenAI merges Codex and ChatGPT desktop apps for Mac and Windows under a new ChatGPT desktop app, allowing users to switch between Codex, Chat, and Work (Zac Hall/9to5Mac)
9to5mac.com/2026/07/09/openai-

@rainerzufall_le@mastodon.social
2026-06-09 07:57:37

"Thüringen, Thüringen, Thüringen ist eines von den schwierigen Bundesländern"
fragdenstaat.de/artikel/exklus