Geometric Configurations of Perturbed Jailbreak Prompts
Lynn Delcon, Andres Algaba, Vincent Ginis
https://arxiv.org/abs/2607.20581 https://arxiv.org/pdf/2607.20581 https://arxiv.org/html/2607.20581
arXiv:2607.20581v1 Announce Type: new
Abstract: Perturbation techniques that turn unsuccessful jailbreak prompts into successful ones are continuously evolving, constituting a major security threat to LLM safety. In this paper, we investigate the internal representations of such string-level perturbed jailbreak inputs in the small weight models of the Qwen-2.5-1.5B/-3B/-7B-Instruct and Llama-3.2-1B/-3B/-3.1-8B-Instruct families. We select two representation spaces: the last-layer-last-token embedding space and the top-50 next-token probability space. The former space separates prompts based on their spelling and format, while the latter space is effectively one-dimensional but appears more complex to cluster. Within our refusal-dominated answer set we find no behavioral hyperplane in either space. Only the next token "Sure" in the 1.5B Qwen model, and both tokens "," and "\.C\.C" in the 1$ Llama model, display a significant association with a compliant-labeled answer.
toXiv_bot_toot
IssueTrojanBench: Benchmarking AI Coding Agents Against Malicious Issue Requests
Ankur Singh, Jinqiu Yang, Tse-Hsun Chen
https://arxiv.org/abs/2607.20759 https://arxiv.org/pdf/2607.20759 https://arxiv.org/html/2607.20759
arXiv:2607.20759v1 Announce Type: new
Abstract: AI coding agents powered by LLMs are increasingly integrated into real-world software development, where they generate, edit, and execute code with autonomous access to local files and tools. Coding agents inherit security risks from both the LLM backbone, where adversarial prompts, poisoned training data, and backdoor triggers can cause models to emit insecure or attacker-chosen code, and their agentic architecture, where tool-using autonomy enables induced misuse of external APIs, data exfiltration, and persistent compromise of development environments. This paper presents a systematic evaluation of malicious issue requests against state-of-the-art coding agents (Cursor, Claude Code, and Codex Desktop), powered by two major model families (OpenAI GPT-5.3 Codex/GPT-5.4 and Anthropic Sonnet 4.6). Our novel benchmark IssueTrojanBench contains malicious issues that are constructed based on four novel attack categories (i.e., embedded as malicious instructions in issues), six delivery vectors (e.g., PDF, or issue comment), and further augmented by perturbations. Our results reveal critical vulnerabilities in the as-deployed modern coding agents, i.e., 66.5% of the malicious issues from IssueTrojanBench penetrate all the guardrails (agent- and LLM-level) of coding agents. Our further analysis shows that rejection is almost entirely from LLMs rather than the agent frameworks, with GPT models broadly vulnerable and Sonnet 4.6 exhibiting more selective, risk-aware blocking of high-impact actions. Our evaluation also highlights that the current agent-level defense strategy offers limited additional protection for coding agents. Our findings highlight the urgent need for stronger agent- and model-level safety mechanisms to protect AI coding agents.
toXiv_bot_toot
Ricci flow with metric torsion on surfaces of positive Euler characteristic
Shubham Dwivedi
https://arxiv.org/abs/2609.15880 https://arxiv.org/pdf/2609.15880 https://arxiv.org/html/2609.15880
arXiv:2609.15880v1 Announce Type: new
Abstract: We study an adapted Ricci flow of connections with metric torsion on surfaces with positive Euler characteristic. We first prove that there do not exist any nontrivial solitons of the flow on the $2$-sphere thus confirming a conjecture of Branding--Kr\"oncke (J. Geom. Anal. 27.3 (2017), arXiv:1606.09121). We give an explicit family of torsion data for which the corresponding global solutions fail to converge on $\mathbb{S}^2$. Nevertheless, we provide several sufficient conditions for the convergence of the flow to a stationary point. We first prove that the normalized adapted Ricci flow always converges on $\mathbb{RP}^2$, which completely answers a question in the paper of Branding and Kr\"oncke. Using this, we deduce that the flow converges on $\mathbb{S}^2$ whenever the initial metric and the torsion one-form are antipodally symmetric. We also prove a {\L}ojasiewicz--Simon gradient inequality for the flow and use it to prove convergence to a stationary point provided the solution is close to an arbitrary stationary point.
toXiv_bot_toot
DefoEye: Python-Based Software for Facilitating Time-Series InSAR Analysis of Sentinel-1 Remote-Sensing Data
Alireza Taheri Dehkordi, Hossein Hashemi, Amir Naghibi
https://arxiv.org/abs/2608.04915 https://arxiv.org/pdf/2608.04915 https://arxiv.org/html/2608.04915
arXiv:2608.04915v1 Announce Type: new
Abstract: Many existing time-series Interferometric Synthetic Aperture Radar (TS-InSAR) software tools have limitations, including restricted geographic applicability, commercial licensing, and incomplete end-to-end processing support. Although GMTSAR avoids some of these constraints, it still requires substantial manual intervention and C-shell commands, lacks a user-friendly graphical interface, and omits important steps such as interferogram network pruning and anchoring of unwrapped interferograms. This paper introduces DefoEye (v1), an open-source Python-based software package that wraps GMTSAR and provides a unified, user-friendly TS-InSAR workflow for Sentinel-1 data. DefoEye supports parallel job execution, interferogram network pruning, and multiple anchoring options. Its performance was evaluated from 2020 to 2024 in four regions with different geological settings, deformation mechanisms, and atmospheric and climatic conditions. In Bologna, Italy; Gotland, Sweden; and Houston, USA, DefoEye results were compared with observations from 10 GNSS stations and showed strong agreement, with RMSE values of 4.3-11.9 mm and Pearson correlation coefficients of 0.63-0.95. In Karaj, Iran, where GNSS observations were unavailable, DefoEye was compared with other widely used processing tools and achieved similarly close agreement, with an RMSE of 4.8 mm/yr and a Pearson correlation coefficient of 0.98. These results demonstrate that DefoEye provides reliable TS-InSAR products for geological, hydrological, and environmental applications.
toXiv_bot_toot
"How proprietary formats have become Microsoft’s main tool for lock-in"
Why DOCX, XLSX, and PPTX are Microsoft's most effective lock-in tool:
https://blog.documentfoundation.org/blog/2026/07/17/microsofts-main-tool-for-lock-in/…
AI-Generated Interactive Fiction for Educational Use: A Pilot Study of Perceived Comprehensibility, Coherence, and Engagement
Finn Rogosch, Andreas Schrader
https://arxiv.org/abs/2608.10818 https://arxiv.org/pdf/2608.10818 https://arxiv.org/html/2608.10818
arXiv:2608.10818v1 Announce Type: new
Abstract: Generative artificial intelligence (AI) can produce educational content at scale, including interactive and narrative learning experiences, but technical generation alone is not sufficient: scenarios that are confusing, narratively inconsistent, or unengaging are unlikely to be useful in practice. This paper presents a pilot user-centred evaluation of AI-generated interactive fiction (IF) for educational use in higher education. Using a previously described domain-agnostic pipeline and a shared STEM content base, we generated a controlled pool of scenarios and asked participants (N = 22, STEM higher-education) to play one generated episode and rate it on narrative clarity, story-content coherence, engagement, and length acceptance. A free-text prompt captured open feedback. Narrative clarity and length acceptance were rated positively, engagement sat near the neutral mid-point of the scale, and story-content coherence was the weakest dimension by a clear margin. Qualitative feedback points to quiz integration as the bottleneck. Artificial in-fiction motivation for quiz prompts and abrupt setting changes were reported. Feedback also pointed to missing story-level consequences for wrong answers. From these observations, we derive concrete design implications that can inform larger follow-up studies, including later work on learning effectiveness.
toXiv_bot_toot
❝Most journals and conferences require authors to ensure correctness and integrity of their papers and take responsibility for them. This is particularly relevant given potential AI generated or heavily AI-assisted submissions. When authors could not answer questions about the technical parts, and sometimes even basic questions about the paper, it is difficult to see how they could have verified the paper’s contents.❞
jfc
https://medium.com/@TmlrOrg/asking-authors-about-their-own-papers-3d2e04e5dee0
RE: https://hachyderm.io/@zwarich/116896925301582827
I am delighted to see this, and look forward to studying the paper a bit.
Back in my “Paul is tinkering around with programming on his own” days, radix sort was one of the first algorithms to fascinate me. When I started taking academic CS classes, I was baffled by fixation on •comparison• in sorting. What about radix sort?!?
Well, the obvious answer is “you can’t use radix for everything; comparison is more generic.” But…what if you made radix more generic too? This paper seems to take that little thought and run to grandiose places with it.