OpenAI says it plans to let third-party groups conduct technical safety evaluations of its AI models during the training, evaluation, and deployment phases (Rachel Metz/Bloomberg)
https://www.bloomberg.com/news/articles/2026…
Leaky Language Models: Stealing Architecture and Inference Optimizations via Per-Token Timing
Sadegh Majidi, Niloofar Mireshghallah, Kazem Taram
https://arxiv.org/abs/2607.20723 https://arxiv.org/pdf/2607.20723 https://arxiv.org/html/2607.20723
arXiv:2607.20723v1 Announce Type: new
Abstract: This work presents LeakyLMs, a set of attacks that leak proprietary model, architecture, and deployment information from production language models. LeakyLMs is the first to demonstrate that key model and deployment details can be inferred using only token generation timing, even when interacting through remote APIs. LeakyLMs introduces two core attacks. The first attack targets inference optimizations and deployment strategies. For example, our attack detects whether a provider uses speculative decoding, a widely deployed inference-time optimization, and further identifies the context length of the draft model used in the pipeline. Our measurements show that Google Gemini Flash 2.5 uses speculative decoding with a draft context window of approximately 128K tokens. The second attack recovers key architectural properties, including the number of transformer layers, hidden dimension size, and number of attention heads. To achieve this, LeakyLMs builds a detailed and accurate model of token-generation timing on modern NVIDIA GPUs, characterizing how latency scales with model configuration and hardware parameters. The attack then performs a search over the architecture space using this timing model. In experiments with Llama models, the near-correct architectural configuration appears in the top-10 guesses more than 90% of the time.
toXiv_bot_toot
Psychedelics align brain activity with context https://www.nature.com/articles/s41586-026-10910-z Under psilocybin, "the separation between internal models and sensory context on which predictive processing depends" dissolves.
Brain-wide reconfiguration of burst firin…
Replaced article(s) found for cs.CR. https://arxiv.org/list/cs.CR/new
[1/2]:
- Facade: High-Precision Insider Threat Detection Using Deep Contextual Anomaly Detection
Alex Kantchelian, et al.
https://arxiv.org/abs/2412.06700 https://mastoxiv.page/@arXiv_csCR_bot/113627226016881063
- Enhancing Membership Inference Attacks on Diffusion Models from a Frequency-Domain Perspective
Puwei Lian, Yujun Cai, Songze Li, Bingkun Bao
https://arxiv.org/abs/2505.20955 https://mastoxiv.page/@arXiv_csCR_bot/114584229384238830
- Cryptographic Choreographies
Sebastian M\"odersheim, Simon Lund, Alessandro Bruni, Marco Carbone, Rosario Giustolisi
https://arxiv.org/abs/2602.12967 https://mastoxiv.page/@arXiv_csCR_bot/116079470877038972
- TALUS: FIPS-204-Exact Threshold ML-DSA via Boundary Clearance
Leo Kao, Raymond Chang
https://arxiv.org/abs/2603.22109 https://mastoxiv.page/@arXiv_csCR_bot/116283533730322814
- SafeHarbor: Defining Precise Decision Boundaries via Hierarchical Memory-Augmented Guardrail for ...
Liu, Ying, Zhang, Zou, Zhang, Yang, Zhang, Peng
https://arxiv.org/abs/2605.05704 https://mastoxiv.page/@arXiv_csCR_bot/116537968700983890
- AI Security Policy Should Assess Systems, Not Only Models
Michael A. Riegler, Inga Str\"umke
https://arxiv.org/abs/2605.09504 https://mastoxiv.page/@arXiv_csCR_bot/116560756584469721
- Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models
Cheng Li, Jiexiong Liu, Yixuan Chen, Yi Li
https://arxiv.org/abs/2607.13093 https://mastoxiv.page/@arXiv_csCR_bot/116928605403267286
- From Neural Intent to Cryptographic Authorization: Securing AI-Driven Enterprise Workflows
Jiasi Weng, Jian Weng, Minrong Chen, Ming Li, Jia-Nan Liu, Zhi Li, Yue Zhang
https://arxiv.org/abs/2607.15596 https://mastoxiv.page/@arXiv_csCR_bot/116951166143038362
- Towards an Automated Test of LLM Security Knowledge
Shufan Chai, Liangliang Sun, Jessica Staddon
https://arxiv.org/abs/2607.18496 https://mastoxiv.page/@arXiv_csCR_bot/116962579240952448
- DynaMark: A Reinforcement Learning Framework for Dynamic Watermarking in Industrial Machine Tool ...
Navid Aftabi, Abhishek Hanchate, Satish Bukkapatnam, Dan Li
https://arxiv.org/abs/2508.21797 https://mastoxiv.page/@arXiv_eessSY_bot/115128274334268839
- CLOAK: Contrastive Guidance for Latent Diffusion-Based Data Obfuscation
Xin Yang, Omid Ardakanian
https://arxiv.org/abs/2512.12086 https://mastoxiv.page/@arXiv_csLG_bot/115729063749487767
toXiv_bot_toot
SpaceXAI releases Grok 4.7, which it says is better at verifying its own work and managing longer context, available for $2/1M input and $6/1M output tokens (xAI)
https://x.ai/news/grok-4-7
PhantomSeal: Proactive Deepfakes Defense with Identity/Context Protection and Forensic Tracing
Liangqin Ren, Zeyan Liu, Ye Wang, Yuxin Chen, Fengjun Li, Bo Luo
https://arxiv.org/abs/2607.20564 https://arxiv.org/pdf/2607.20564 https://arxiv.org/html/2607.20564
arXiv:2607.20564v1 Announce Type: new
Abstract: Deepfakes, especially face-swapping attacks, pose significant challenges to authenticity, security, and ethics across science, engineering, and society. While most existing detection/tracing approaches operate post hoc, proactive defenses that aim to intervene before deepfake generation remain limited in terms of real-world effectiveness. In this paper, we present PhantomSeal, the first proactive defense to simultaneously protect both the identity and the context of users' images from being used in face-swapping attacks, while supporting forensic tracing. We present a novel cloaking technique that embeds a selected identity as a stealthy identifier. This mechanism steers the deepfake generation process toward producing content that resembles the chosen cloak identity, thereby preventing successful face-swapping while enabling effective feature-based forensic analysis. The effectiveness and robustness of PhantomSeal is demonstrated in extensive experiments across different face-swapping architectures and models. For example, it reduces the attack success rate of SimSwap, an advanced deepfake model, to 0.30%, and correctly identifies 97.97% of manipulated content. Codes can be found at https://github.com/LiangqinRen/PhantomSeal
toXiv_bot_toot
Over 30 crypto companies, including Coinbase and Block, say frontier AI safety guardrails hinder legitimate security work while attackers use stronger tools (Shaurya Malwa/CoinDesk)
https://www.coindesk.com/tech/2026/08/13/bitcoin-firm…
Anthropic is partnering with Accenture to embed evaluators within Anthropic, including red teaming models and conducting alignment assessments (Anthropic)
https://www.anthropic.com/news/accenture-embedded-evaluation