Research

Four AI Papers: Claude Masters ARC-AGI, Robot Practices High

6 min read

TL;DR Too Long; Didn’t read

Four new arXiv preprints from the past days show: The most striking finding is that a visual harness raises the efficiency value of Claude Opus 5.0 on the puzzle benchmark ARC-AGI-3 from 40.68 to a perfect value of 100.00, with 57.4 percent fewer actions than human participants. A second study allows a robot agent to increase its success rate in object manipulation from 28.6 to 95.0 percent through repeated practice in simulation and real-world trials, subsequently passing all 30 physical test runs. A third preprint shows with a new cybersecurity benchmark that no open language model delivers more than 42 percent accurate command-line instructions for security tools. A fourth paper demonstrates that reinforcement learning with verifiable rewards allows uncontrolled language drift in thought traces, while classical fine-tuning prevents it.

A magnifying glass hovers over a stack of academic papers, from which a robotic arm places a glowing puzzle piece into a flawless grid of squares. Image generated with GPT Image 2

Key takeaways

  • A visual agent harness raises Claude Opus 5.0 on the puzzle benchmark ARC-AGI-3 from 40.68 to the perfect value of 100.00.
  • Through simulated practice, a robot agent's success rate increases from 28.6 to 95.0 percent, confirmed in 30 practical tests.
  • No open language model translates queries into command-line instructions for security tools with more than 42 percent accuracy.
  • Reinforcement learning with verifiable rewards allows unlimited language drift in thought traces, while pure fine-tuning demonstrably prevents it.

From the arXiv new submissions of the past 24 to 48 hours in cs.AI, cs.LG, and cs.CL, this digest selects four preprints that demonstrate how far agents can go through practice and the right toolset – and where control and safety have not yet kept pace with this speed. Two papers focus on agents that improve themselves through repetition and better perception, while two others address blind spots in examining cybersecurity tool usage and training with verifiable rewards. Curation was based on traceable methodology in the abstract, concrete core results with numbers, and thematic diversity instead of fourfold repetition of the same subfield.

Agents in Practice

A visual harness elevates Claude to a perfect puzzle outcome

Qiushi Han and a four-member team of researchers, including those from MIT, present a visual harness with VISTA – a framework of tools and rules that grants an existing model additional capabilities without retraining it. VISTA provides a multimodal model with a lossless visual memory: instead of discarding previous screen views, it retains them in their original form, allowing the model to selectively refer back to past observations and actively reorder its visual input while thinking. On the benchmark ARC-AGI-3 – a collection of interactive visual puzzles specifically designed to be non-memorizable – this approach reportedly raises the efficiency score of Claude Opus 5.0 from 40.68 to a perfect score of 100.00, solving all 25 public games with 57.4 percent fewer actions than human first players. This matters because it shows that much of what agents currently lack in visual continuity apparently does not stem from the underlying model but from the framework surrounding it – a finding that connects to an earlier digest finding that Claude Opus 4.8 achieved gold medal level at the Linguistics Olympiad without any specialized training: In both cases, the outcome is determined less by the model itself than by the task environment and its interaction with it.

A robot agent improves from 28.6 to 95.0 percent success rate

Yen-Jen Wang and a nine-member team, including those from UC Berkeley, present Reconstruct, Practice, Go Real (RPG), a method by which a robot execution system improves itself without changing the weights of the underlying model. RPG identifies existing manipulation skills in a given dataset, builds corresponding practice tasks in simulation, and diagnoses based on execution feedback, privileged simulator state, and example videos why a trial fails. From these diagnoses, the system develops new, reusable symbolic skills, refines existing ones, and adjusts the system prompt before testing each change in its own test runs. For 22 manipulation tasks, the success rate reportedly increases from 28.6 percent after the first practice round to 95.0 percent after 15 rounds – significantly ahead of the comparison methods ASPIRE (75.5 percent) and a GPT-6-Astra-Pro-powered agent (60.0 percent); after a one-time calibration, the frozen system subsequently passes all 30 physical test runs on real hardware. This matters because it shows a way to improve robotic skills through practice rather than through expensive retraining – an approach that differs from an earlier digest finding where a single agent controlled various robot platforms without any specialized training: There, generality counted without adaptation, while here, targeted learning progress toward a specific task does.

Safety and Training

No open model reliably translates security commands into the command line

Pengfei Li and a three-member team present KaliBench, a benchmark that tests how reliably language models translate natural language instructions into executable commands for cybersecurity tools on Kali Linux – a Linux distribution specialized for penetration testing. The dataset includes 8,504 request-command pairs across 1,642 tools, 23 capability dimensions, and five phases of a security test, constructed through a pipeline with deterministic unification of command forms and a multi-stage evaluation consisting of automated assessment, isolated test execution, and human oversight. In 24 configurations of general and security-specialized open-weight models, according to the authors, none achieves more than 42 percent exact command accuracy in the unrestricted test mode, underscoring the difficulty of precise command-line commands without tool hints; targeted fine-tuning and reinforcement learning with verifiable rewards derived from KaliBench subsequently elevate an 8-billion-parameter model to a level comparable to a 685-billion-parameter mixture-of-experts model. This matters because security teams are increasingly examining whether routine penetration-testing tasks can be delegated to language models, even though simple flag confusions can render a command unusable – this study provides the first fine-grained, reproducible benchmark instead of mere end-to-end success rates.

Training pressure allows uncontrolled language drift in thought traces

Michael Sullivan and Alexander Koller investigate in their study why reasoning models increasingly drift into unusual, barely comprehensible language while thinking – a phenomenon that complicates the observability of their thought traces by humans. The authors first theoretically show that reinforcement learning with verifiable rewards (RLVR) allows unlimited language drift, while supervised fine-tuning (SFT) does not, because it binds the model more closely to predefined, human-readable text examples. Empirically, the drift occurs specifically when RLVR is trained on novel tasks for which the base model does not yet master the target behavior – the model then develops its own linguistic shortcuts, according to the authors, instead of relying on already learned, comprehensible thought patterns. As a third step, they formally prove that language drift cannot be contained without simultaneously limiting the expected reward – observability of the thought trace and training performance thus stand, at the frontier of model development under RLVR, in an irreconcilable conflict of objectives. This matters because providers of reasoning models increasingly cite readable thought traces as a safety feature, even though this finding suggests that the very training that brings the strongest capability leaps systematically undermines this readability; the finding complements an earlier digest finding that merely the language in which a model thinks could reduce its willingness to simulate a nuclear strike from 93 to 17 percent: Both works show how strongly the language of the thought trace shapes a model’s behavior, without users necessarily noticing it.

All four works are unreviewed preprints from the latest arXiv submission wave; the referenced numbers are based on the abstracts and statements of the respective author teams and have not yet been confirmed by independent review or replication. For VISTA, RPG, and KaliBench, the authors promise project pages or parts of code and data; for the study on language drift, no publication of material is apparent from the abstract. How robust the perfect ARC-AGI-3 score, the 95 percent success rate of the robot agent, the 42 percent upper limit at KaliBench, and the proven conflict of objectives between language drift and training performance hold up outside their respective test environments remains to be shown by independent replications.

Frequently asked questions

Are the four presented papers peer-reviewed?

No, all four works are unreviewed preprints from the latest arXiv submission wave. The referenced figures come from the abstracts and statements of the respective author teams and have not yet been confirmed by independent review or replication.

Do the authors provide code or data?

For VISTA, RPG, and KaliBench, the authors announce project pages as well as partial code and data intended to facilitate independent verification. For the study on language drift in RLVR, no material publication is evident from the abstract.

What does Claude's ARC-AGI result have to do with the underlying model?

According to the authors of VISTA, the model itself remains unchanged – all progress comes from the visual harness, which losslessly stores previous observations and makes them selectively retrievable. This suggests that a large part of today's agent weaknesses lies more in the lack of a framework around a model than in its core capabilities.

Does the language drift finding mean that reasoning models become fundamentally less reliable?

The study shows, according to its authors, a structural conflict of objectives: language drift cannot be contained without simultaneously limiting the achieved reward and thus the model's capability. This does not automatically mean that every reasoning model becomes less reliable, but rather that readable thought traces and maximum training performance under RLVR must be weighed against each other.

Sources (4)
  1. VISTA: A Visual Harness for Reasoning in an Interactive World
  2. Reconstruct, Practice, Go Real: Guided Self-Improvement for Embodied Agents
  3. KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards
  4. On Language Drift during RLVR Post-Training

Your AI update for the work week

Once a week, the most important AI news – plus one practical tip to try right away. No spam, unsubscribe anytime.

← Back to the blog