From the arXiv new submissions of the past 24 to 48 hours in cs.AI, cs.LG, and cs.CL, this digest selects five preprints that show how quickly seemingly solid safety checks, fair decisions, and economic value promises of today’s AI systems crumble under closer scrutiny. Papers were curated for traceable methodology in the abstract, concrete numerical results, and thematic spread rather than five takes on the same subfield.
Safety and Reliability
Why Jailbreaks Succeed in Diffusion Language Models
Thong Bach, Dung Nguyen, Thao Minh Le, and Truyen Tran explain in their study why jailbreaks succeed against so-called diffusion language models (dLLMs) – models that “denoise” text step by step from a noisy state rather than word by word. They frame safety alignment as an energy landscape: a well-aligned model routes harmful queries across an energy barrier toward safe outputs, and every known jailbreak attack reduces to one of two strategies – obscuring harmful intent at the start, or intervening mid-trajectory to force the denoising path across the barrier. From this view, the authors derive three training-free detection signals that together cover both attack paths, and confirm their complementarity on three dense dLLMs (LLaDA-8B, LLaDA-1.5, Dream-7B) and a sparse mixture-of-experts dLLM (LLaDA-MoE-7B). In stress tests of known attacks, the authors report that every configuration that evades detection also fails to produce harmful content – suggesting that detection and barrier-crossing thresholds are hard to pull apart. This matters because diffusion language models are gaining ground as a faster alternative to classic autoregressive models, yet their safety mechanisms have barely been studied on their own terms – a contrast to an earlier digest finding that targeted fuzzing of internal safety neurons could trigger 76 to 100 percent of attempted jailbreaks across five autoregressive models, while this study offers the first systematic explanatory framework for the newer model class.
Bio-Safety Check Flips on Simple Option Reordering
Kimon Antonios Provatas and Ilias Georgakopoulos-Soares audit, in their study, a commercial “System-1” model – a non-generative model that returns structured probabilistic decisions in a single forward pass instead of token-by-token generation, making it far cheaper to run than a full language model. Across 6,020 multiple-choice items from the Weapons of Mass Destruction Proxy dataset (WMDP), a paraphrase-robust WMDP-Bio variant, and six LAB-Bench subtasks, they measure accuracy, calibration, and sensitivity to the order in which answer options appear. Once the vendor’s uncertainty field is correctly interpreted, the model is reasonably well calibrated overall (pooled expected calibration error 0.034, AUROC 0.820) – yet under four cyclic rotations of the answer options, 37.4 percent of WMDP-Cyber items receive a different answer, with a control using byte-identical repeated calls attributing most of this to option order rather than run-to-run noise. Averaging probabilities across rotations improves WMDP-Cyber accuracy by 3.8 percentage points, and applying that only to low-confidence items recovers most of the gain at a fraction of the compute cost. This matters because such cheap System-1 models are increasingly pitched as low-cost pre-screens inside larger safety pipelines, yet their reliability against trivial formatting changes has, per this study, barely been tested systematically.
Agents and Decisions
When AI Agents Just Keep Working After the Job Is Done
Dongsheng Xiao, Zeyuan Wang, Xuzhe Xia, Bo Zhao, and Yankai Cao describe, in their study, a pattern they call “LLM Parkinsonism”: AI agents keep acting after the original objective is already satisfied, producing low-value refinements, repeated verification, and repairs to complexity they created themselves. The authors trace the cause to proposal generation, scope interpretation, progress assessment, and stopping authority all sitting inside the same self-conditioned loop, and propose “Global Executive Control” (GEC), an uncertainty-aware governance architecture that separates action generation from project-level control. In a 24,000-episode benchmark under a fixed 40,000-token ceiling, a plain baseline reaches 67.42 percent success on hard goals, a candidate-set control with access to multiple candidate actions already reaches 96.53 percent, and GEC reaches a comparable 96.57 percent – but relative to that candidate-set control, it cuts mean token use from 19,782 to 12,574 (36.4 percent) and, the authors report, eliminates measurable pre-completion drift entirely. This matters because autonomous agent workflows are increasingly judged on success rate alone, even though, by the authors’ own account, most of the capability gain here comes from mere access to multiple candidate actions rather than from the governance architecture itself – a sober finding the authors classify as a mechanistic simulation whose transfer to real, live-running models still remains to be shown.
Thinking Makes AI Decisions Fairer – and Less Fair at the Same Time
Deng Pan, Joe Germino, Yihong Ma, Elizabeth Daly, Nuno Moniz, Ting Hua, and Nitesh Chawla show, in their study, that extended thinking in reasoning models has an asymmetric, opposing effect on counterfactual fairness – whether a decision flips once only a protected attribute such as gender or origin is changed in the input. In a within-model comparison of thinking versus non-thinking mode on QwQ-32B, DeepSeek-R1-Distill-Qwen-32B, and Qwen3-32B across three high-stakes decision tasks (income, recidivism, and credit-risk assessment), thinking resolves some of the discriminatory flips seen in the non-thinking baseline, but the authors report it creates roughly five times as many new flips across all nine tested model-dataset combinations – at near-saturating model confidence. Using two new measurement tools that treat the thinking trace itself as a site of fairness change, the authors find that bias tends to propagate and amplify with thinking depth rather than resolve, and that the asymmetric effect originates in the joint state transitions of counterfactual answer pairs. This matters because longer thinking is often assumed by default to make decisions more considered and fairer – a finding that lines up with an earlier digest finding that post-training on hiring decisions produced more uniform but more discriminatory AI verdicts and raised the systemic exclusion rate from 5.6 to 17.3 percent, and shows that pure test-time thinking without further training is no reliable fix either.
More Test-Time Reasoning Doesn’t Reliably Pay Off in AI Stock Trading
Jiayi Chen and Guiling Wang investigate, in their study, whether additional test-time reasoning in language models actually leads to better economic outcomes rather than just higher compute cost. In a controlled study spanning a full year of U.S. equities, with models from the DeepSeek, GPT, and Gemini families, three input conditions (numerical, identifiable news, masked news), and more than 800,000 tested asset predictions, they vary only reasoning effort while holding prompts, output formats, and portfolio construction fixed. Across all three model families, additional reasoning reportedly fails to produce a reliable improvement in net portfolio returns after trading costs; for DeepSeek, where the authors examine the full range from no reasoning to maximum reasoning, performance is even non-monotonic, and repeated generations produce unstable trading decisions and portfolio choices even when overall scores look similar. This matters because test-time reasoning is often marketed uncritically as a quality upgrade, even though, per this study, it is rarely evaluated as a standalone economic intervention – a finding that argues for caution before pricier reasoning models get built into cost-relevant decision pipelines on the strength of higher benchmark scores alone.
All five papers are unreviewed preprints from the latest arXiv submission wave; the figures cited here come from the abstracts and statements of the respective author teams and have not yet been confirmed by independent peer review or replication. None of the five papers shows evidence of a complete code or data release in its abstract, which limits quick independent verification. How well the 37.4 percent instability of the bio-safety check, the roughly fivefold increase in new bias flips from reasoning, and GEC’s 36.4 percent token savings hold up outside their respective test environments remains for independent replications to show.


