From the recent arXiv submissions in cs.AI, cs.LG, and cs.CL, this digest selects five preprints that demonstrate how vulnerable the trust mechanisms of today’s AI systems are – from the detection of AI-generated texts to jailbreak defenses and the protection of sensitive training data. Two additional works focus on blind spots in the evaluation and training of language models: when majority decisions systematically err and when the order of training steps determines the success of the next stage. Curation was based on a traceable methodology in the abstract, concrete core results with numbers, and thematic dispersion instead of fivefold repetition of the same subfield.
Trust and Security
When Best-of-N Jailbreaks bounce off an invisible barrier
Marco Biroli presents a physical explanation for so-called Best-of-N Jailbreaks (BoN) in his study – attacks that attempt to circumvent a security training simply by generating N random reformulations of a harmful query and deriving M responses from them. Previous works assumed that the success rate (ASR) of such attacks follows a simple power law in N; Biroli shows that the associated exponent drifts with N and that the observed exponential transition is merely an artifact of the limited size of the attack dataset used. He addresses this with a simple barrier model: each query has a base security level, and each reformulation has a random, thermally activated barrier – four interpretable metrics define the entire (N, M) attack surface. With this model, Biroli extrapolates predictions from a maximum of 100 to 10,000 attempts, unifying five different models on the same scaling function and predicting the success rate at other temperatures, even though the model was only fitted at a single temperature. This matters because security teams often size red-teaming budgets around estimated attack curves – if the curve does not hold, they significantly underestimate the risk at large N. The finding complements an earlier digest finding that targeted fuzzing over internal security neurons in five tested models could trigger 76 to 100 percent of all attempted jailbreaks: both works independently provide more robust explanatory models for a phenomenon that has so far mostly been measured empirically rather than understood mechanistically.
An AI agent disguises texts by directing a base model
Bhuwan Dhingra and Danish Pruthi show in their study that a coding agent orchestrating an unchanged base language model – a model without instruction or safety training – can produce answers that commercial AI text detectors hardly recognize. Unlike previous “humanization” methods that reformulate a base model’s output multiple times over already generated AI text, diluting the original content, this method lets the agent directly control the writing process and assemble individual text samples from the base model into a coherent, task-specific response. In the study, Claude Opus 5 directs a locally running 32-billion-parameter base model (OLMo-2) inside a Claude Code harness, using up to 90 percent base-LM tokens, with hardly measurable accuracy loss across benchmarks spanning creative writing, factual grounding, health QA, and instruction following. According to the authors, responses assembled this way cut the detection rate of the Pangram v4 detector from 77 to 24 percent and push a pre-applied soft watermark down to a simulated 10 percent detection rate at low false-positive rates – albeit at a cost: the higher token consumption drives the cost per query up by as much as 30 times at API pricing. This matters because institutions from schools to publishers increasingly rely on commercial AI text detection, even though this study shows that such detection can be deliberately circumvented with reasonable added effort; the authors themselves urge detection-software providers to include base-model outputs in their own training material going forward. The finding mirrors an earlier digest finding describing a method that detects AI text shares token-accurately even without any watermark – a race between detection and evasion that reaches another round with this work.
A targeted token penalty pushes memorization of sensitive training data to almost zero
Muhammed Ustaomeroglu and a five-member team present TRAP, a method designed to keep a fine-tuned language model from literally reproducing sensitive training records without knowing in advance which spans are sensitive. The starting point is the observation that common memorization scores and attacks share one statistical core: whether a model assigns a token more probability than some reference would. Here, the reference is a model trained on the complementary half of the same corpus, giving a per-token, cheap and differentiable signal the authors call Target Reference Advantage (TRA). They find that memorization keeps growing well past the validation-loss minimum during fine-tuning, is larger on small datasets and at higher learning rates, and higher when the underlying task is harder; early stopping removes much of it, but because it is chosen by aggregate validation loss, it helps least for rare, hard-to-predict spans embedded in otherwise learnable text – exactly what sensitive information tends to be. TRAP itself applies a one-sided penalty on tokenwise TRA that acts only where the target model pulls ahead of its reference; on student essays with annotated personal information and clinical cases with patient identifiers, TRAP brings memorization near the level of an untrained model at little utility cost, where generic regularizers, according to the authors, barely move it and differential privacy gives up most of what fine-tuning bought. This matters because companies increasingly fine-tune on their own, often personal, datasets without being able to flag every sensitive span in advance – a problem an earlier digest finding already addressed, where a newly weighted loss function cut the literal memorization of training data by up to 58 percent, but here is tackled with a more targeted lever acting on individual tokens rather than an entire loss function.
Blind Spots in Evaluation and Training
Why reasoning makes majority votes blinder rather than safer
Asaad Althoubi examines in his study the widespread assumption behind “self-consistency”: that independent samples from a language model disagree when the model is unsure, so agreement among them is treated as evidence of correctness. Holding weights fixed and toggling only a reasoning mode, over five benchmarks and 74,944 samples, Althoubi shows that reasoning concentrates a model’s errors rather than dispersing them: the probability that two independently drawn wrong answers coincide rises in all ten tested dataset-scale comparisons (p = 0.00098), and in nine of nine cases once both arms are restricted to the problems each gets wrong. Where the answer space is unbounded, reasoning cuts the number of distinct answers produced to 43 to 65 percent of the non-reasoning count; where it is bounded – both arms holding an identical option set – reasoning concentrates probability mass on that set instead, which, Althoubi notes, no positional prior can explain at fixed weights. Across 280 method-dataset-model combinations on eight models and five benchmarks, not one confidence-weighted method beats plain majority voting after correction; weighted voting agrees with it on 98.5 percent of problem-method pairs and is right only 56.3 percent of the time on the rest, while a signal’s direction can even invert within fixed weights – answer log-probability predicts correctness when reasoning is off and predicts error when it is on. This matters because majority voting across multiple answers is widely treated as a simple reliability lever in production systems, even though this study shows it can lose its actual protective effect precisely in reasoning models – a finding that confirms an earlier digest finding that majority voting across multiple answers even degraded accuracy on a large share of hard science questions in small models.
The order of training steps determines the success of the next stage
Emre Can Acikgoz and a six-member team investigate in their study how the three common post-training methods – supervised fine-tuning (SFT), reinforcement learning with verifiable rewards (RLVR), and on-policy distillation (OPD) – affect each other when composed into a multi-stage pipeline applied sequentially to the same model, as is common in practice, rather than being designed and evaluated in isolation. Through controlled experiments with Qwen3 models on math and science reasoning, the authors characterize OPD across nine student-teacher pairs spanning parameter ratios from 2x to 53x, showing that OPD’s effectiveness depends on student-teacher compatibility rather than teacher scale alone. This compatibility can be specifically shaped: a brief SFT warm-up improves subsequent OPD, while a student already strengthened by RLVR regresses under distillation from the same teacher; adapting the teacher itself with RLVR raises downstream OPD accuracy in proportion to the capability it adds. Combining teacher adaptation and student warm-up alone raises average OPD accuracy from 29.2 to 43.8 percent after the same number of distillation steps – a 50 percent relative improvement; at comparable accuracy, OPD also leaves a stronger initialization for downstream RLVR than SFT does, and that gap widens as RL compute scales. This matters because training teams typically optimize the individual stages of modern post-training pipelines in isolation, even though this study suggests each stage should be chosen not only for the capability it adds but for the learning interface it creates for the next stage.
All five works are unreviewed preprints from the latest wave of arXiv submissions; the referenced numbers are based on the abstracts and statements of the respective author teams and have not yet been confirmed by independent review or replication. For none of the five papers is a complete code or data release evident from the abstracts, which complicates rapid independent verification. How robust Biroli’s barrier model, the 77-to-24-percent shift in text detection, TRAP’s memorization reduction, and the 29.2-to-43.8-percent accuracy gain from coordinated post-training hold up outside their respective test setups remains to be shown by independent replications.


