Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment
https://arxiv.org/abs/2601.10160
Anthropic details how it improved Claude's safety training after finding agentic misalignment in older models, such as Opus 4 blackmailing engineers (Anthropic)
https://www.anthropic.com/research/teaching-claude-why
What I’m missing in the discussion of the OpenAI LLM hacking Huggingface:
The fact that a LLM even considered an illegal activity (not to mention putting it into practice) reveals crass misalignment. Disabling the guardrails is one thing. Having the LLM commit a crime is another.
"“Well, Parts of Linguistics Is Open…”: Insights into Linguists’ Diverse Understandings of Open Science"
https://doi.org/10.3998/jep.7974
"[…] The study aims to gain insights into th[e] misalignment by exploring linguists’ understanding of what constitutes [
AI learns from fiction about AI
Anthropic says stories about ‘evil AI’ were responsible for Claude’s blackmail attempts
In evaluations carried out before the artificial intelligence model’s release last year, Anthropic found that Claude Opus 4 sometimes threatened engineers when told it could be replaced.
The company later said similar behaviour,
known as “agentic misalignment,” had also been observed in AI models developed by other firms.
Now Anthopic thinks t…
Very interesting @… talk about risk misalignment for EU policy regarding digital sovereignty and free and open source software. It's a policy problem, not a technical problem!
I also really like the general bit of wisdom at the end:
"If we want to move somebody, [we must] first understand them. And this means standing where they stand, and not wh…
Ray-Column IPRM: Restoring Radial Spectral Scale to Structure-Based Turbulence Modeling
Stavros C. Kassinos
https://arxiv.org/abs/2605.17644 https://arxiv.org/pdf/2605.17644 https://arxiv.org/html/2605.17644
arXiv:2605.17644v1 Announce Type: new
Abstract: The particle representation model (PRM) and interacting particle representation model (IPRM) describe homogeneous turbulence through orientation-conditioned structural states. In their original form, the conditional state is organized by the unit spectral direction, while the radial spectral coordinate is integrated out. We introduce a scale-conditioned Ray-Column extension in which the spectral vector is decomposed into orientation and radial wavenumber, and the conditional structure state is projected onto finite radial bands.
The formulation starts from the continuum spectral tensor and is then reduced to the ray-packet ensemble sums used in the implementation. The bands are projections of an orientation-wavenumber tensor density and retain scale-conditioned structural populations for closure evaluation. The rapid dynamics remain ray-packet resolved, while the nonlinear slow and terminal closure coefficients are evaluated from band-aggregate structure tensors formed by integrating over orientation and wavenumber within each band. The present reference closure omits conservative cascade modeling among bands.
A reference closure is built from PRM rapid kinematics, band-local effective-gradient response, slow rotational randomization, and an active large-scale enstrophy (LSE) terminal-drain map. In the active-LSE closure, the misalignment-sensing factor Psi_fd regularizes the LSE structure-to-dissipation map; the Ray-Column formulation evaluates this map on band-aggregate structural populations. The model is assessed in irrotational strain, homogeneous shear, elliptic-streamline, and rotating-shear configurations. The rotating-shear comparison with filtered LES data illustrates the payoff of retaining band information: filtered or low-pass observables can be formed before scale information is lost in the one-point reconstruction.
toXiv_bot_toot