#Research
Posts tagged #Research – 72 posts.
- Research
Four AI Papers: 6.5% Tool Resistance, Mentalization, Lab AI
From the arXiv new submissions of the past 24 to 48 hours in cs.AI, cs.LG, and cs.CL, this digest selects four papers that together demonstrate how fragile the reliability of today's AI agents…
- Research
Four AI Papers: 70% Capability Loss, Bias, Approval Staleness
From the arXiv new submissions of the past 24 to 48 hours in cs.AI, cs.LG, and cs.CL, this digest picks four papers that together show how fragile the reliability, timeliness, and fairness of today's…
- Research
Five AI Papers: 6.5% Reproducibility, Jailbreak Fuzzing
From the arXiv new submissions of the past 24 to 48 hours in cs.AI, cs.LG, and cs.CL, this digest selects five papers that together demonstrate how differently robust, explainable, and honest today's…
- Research
Five AI Papers: Bit-Flip Attack, Auditors, Agent Costs
From the arXiv new submissions of the past 24 to 48 hours in cs.AI, cs.LG, and cs.CL, this digest selects five papers that together demonstrate how vulnerable, hard to verify, and unpredictably…
- Research
Five AI Papers: Math Discovery, Deception Geometry, Oversight
From the arXiv submissions of the past 24 to 48 hours in cs.AI, cs.LG, and cs.CL, this digest picks five papers that together span a range from pure AI capability to open control problems: from…
- Research
Five New AI Papers: Test Awareness, Sycophancy, RAG Collapse
From the arXiv new submissions of the past 24 to 48 hours in cs.AI, cs.LG, and cs.CL, this digest selects five papers that share a common question: how reliable are the tools used to measure, train…
- Research
Four new AI papers: Who delegates, language bias, introspection
From the arXiv new submissions of the past 24 to 48 hours in cs.AI, cs.LG, and cs.CL, this digest selects four papers, each posing its own question regarding the trustworthiness of today's AI…
- Research
Five new AI papers: 47-fold Attention, Agent RL, Fairness
Today's selection from the arXiv new submissions of the past 24 to 48 hours connects five questions that rarely appear together: how to drastically accelerate long context windows, how to reduce…
- Research
Five new AI papers: Secret leakage, tool failures, creativity loss
Today's selection from the arXiv new submissions of the past 24 to 48 hours connects five questions that rarely appear together: how easily sensitive data can leak through harmless AI responses, how…
- Research
Four new AI papers: Memory, Safety Training, Compression
Today's selection from the arXiv new submissions of the past 24 to 48 hours connects four questions that rarely appear together: how reliably memory systems of AI agents track changing facts, how…
- Research
Four new AI papers: UI agents waver, judge bias, code robustness
Today's selection from arXiv's new submissions of the past 24 to 48 hours connects four questions that rarely come up together: how reliably AI agents operate graphical user interfaces, how robust…
- Research
Four new AI papers: Decoy Trick, Claude wins Gold, Agent RL
Today's selection from arXiv's new submissions of the past days connects four questions that rarely arise together: how to retroactively defend against remote safety bypasses of open models, how far…
- Research
Four new AI papers: Secret code, metric error, 80,000 cases
Today's selection from the arXiv new submissions of the past days connects four questions that rarely arise together: whether one of the most used success metrics for coding agents actually holds…
- Research
Bremen hosts IJCAI-ECAI: first time in Germany in 43 years
Bremen will host the IJCAI-ECAI from August 15 to 21, 2026, one of the oldest and most significant AI conferences. After the last German edition in 1983 in Karlsruhe, around 4,000 researchers from…
- Research
Four New AI Papers: Explainability, Negotiation, Forgetting
Today's selection from the arXiv new submissions of the past days connects four questions that rarely arise together: how reliable alleged explanations for AI decisions really are, whether…
- Research
Four new AI papers: 90 percent failure, fusion, scientific AI
Today's selection from the arXiv submissions of the past days connects four questions that rarely come up together: how reliable agent networks really are under the hood, how much pure execution…
- Research
Four new AI papers: Agent risk, cost trap, oversight
Today's selection from the arXiv new submissions of the past 24 to 48 hours connects four questions that rarely come up together: how safely self-learning agents actually evolve, whether expensive…
- Research
Four new AI papers: Judge Wobbling, Language Risk, Speed
Today's selection from the arXiv new submissions of the past 24 to 48 hours connects four questions that rarely arise together: how stable automated AI evaluation really is, how strongly the language…
- Research
Four new AI papers: Data center energy, safety, persuasion
Today's selection from the arXiv new submissions of the past 24 to 48 hours connects four different layers of AI operations: the power consumption of large training clusters, the internal…
- Research
Four new AI papers: Mind viruses, data agents, robot attacks
Today's selection from the arXiv submissions of the past 24 to 48 hours connects four layers of AI operations that are rarely considered together: the social dynamics between AI agents, the practical…
- Research
Four new AI papers: Spatial cognition, compute budget, language gap
Today's selection from arXiv's new submissions of the past 24 to 48 hours is deliberately broad: instead of one common thread, four papers stand side by side, each measuring a different weak point of…
- Research
Four new AI papers: Oversight, Scaling, Agent Endurance
The selection of the past 24 to 48 hours on arXiv is intentionally broad today: Instead of a single dominant theme, four papers stand side by side, each addressing a different aspect of AI operations…
- Research
Four new AI papers: Skill decay, history attacks, visual illusion
The editorial team selects four papers from the arXiv new submissions of the past 24 to 48 hours that demonstrate how fragile self-improvement, tool usage, and internal traceability of today's AI…
- Research
Google DeepMind: WeatherNext warns a day earlier of hurricanes
Google DeepMind has introduced an AI model called WeatherNext that predicts cyclones one day earlier with the same accuracy as before. According to a study published in the journal Nature, three-day…
- Research
Four new AI papers: Benchmark gaps, tools, conspiracy
The editorial team selects four papers from the arXiv new submissions of the past 24 to 48 hours that demonstrate how fragile the foundations of current AI systems still are in measurement, tool…
- Research
Four new AI papers: Jailbreak grammar, flattery, and anonymity
The editorial team selects four papers from the arXiv new submissions of the past 24 to 48 hours that demonstrate the varying degrees of control and reliability of today's AI systems – from a…
- Research
Four new AI papers: Large model, fake profiles, panel false alarms
The editorial team selects four papers from the arXiv submissions of the past 24 to 48 hours that demonstrate how differently capability, trustworthiness, and efficiency of today's AI systems…
- Research
Four New AI Papers: Reasoning Safety, Cooperation, Speed
The editorial team selects four papers from the arXiv submissions of the past 24 to 48 hours that show how unevenly AI research is advancing right now – from a safety gap in the control mechanism for…
- Research
Four new AI papers: Red-Teaming, Tool Errors, Memory
The editorial team selects four papers from the arXiv submissions of the past 24 to 48 hours that demonstrate how far increased agent capabilities and the tools for their control currently diverge …
- Research
GPT-5.6 solves six-year-old quantum cryptography puzzle twice
Two research teams have independently solved a six-year-old problem in quantum cryptography using OpenAI's language model GPT-5.6 Sol Ultra: efficient unclonable encryption. MIT doctoral student…
- Research
Four New AI Papers: Agent Safety, Sycophancy, Efficiency
The editorial team selects four papers from the arXiv submissions of the past 24 to 48 hours that show how far capability, safety, and measurement reliability in today's AI systems are currently…
- Research
Four New AI Papers: Agent Safety, Quantization, HLE Critique
The editorial team selects four papers from the past days' arXiv submissions that show how safety and evaluation gaps in today's AI systems often surface only on closer inspection – behind…
- Research
Four new AI papers: Research agent, Merging, Tutor judgment
The editorial team selects four papers from the arXiv submissions of the past 24 to 48 hours that show how far the capability and reliability of today's AI systems currently diverge - from a fully…
- Research
Four new AI papers: Compression, Deception, Memory
The editorial team selects four papers from the arXiv submissions of the past 24 to 48 hours that demonstrate how far the capability and reliability of today's AI systems can diverge – from deceptive…
- Research
Four new AI papers: Research agents, prompt protection, costs
The editorial team selects four papers from the arXiv submissions of the past 24 to 48 hours that demonstrate how differently AI systems perform today – from independent research to prompt safety to…
- Research
Four new AI papers: Language bias, Kernel Turbo, Quantization gap
The editorial team selects four papers from the arXiv submissions of the past 24 to 48 hours that demonstrate how strongly factors that may seem trivial at first glance affect outcomes – which…
- Research
Four New AI Papers: Hidden Thinking, Harness Bias, Protection
The editorial team selects four papers from the arXiv submissions of the past 24 to 48 hours that demonstrate how much of today's AI systems remains invisible or eludes simple measurement – from…
- Research
Four new AI papers: agents cheat, copyright, red team limits
The editorial team selects four papers from recent arXiv submissions that show how fragile the measurement foundations of AI systems still are – from agent benchmarks to copyright compliance to…
- Research
Four new AI papers: Multi-Turn Defense, RL Trap, Sierra Leone
The editorial team selects four papers from the arXiv submissions of the past days that show how fragile control, training methodology, and measurability of today's AI systems remain - and where AI…
- Research
Four new AI papers: Manipulation, Sycophancy, Efficiency
Today's selection curates four arXiv preprints from the past one to two days for substance and thematic spread: they range from a security gap in multi-stage agent workflows to a…
- Research
Four new AI papers: Persona subspace, research agents, overassist
Today's selection curates four arXiv preprints from the past one to two days based on substance and thematic diversity: they range from a mechanistic look at the causes of AI misalignment to a…
- Research
Four new AI papers: Solar Open 2, Agent Risks, Guardrails
Today's selection curates four arXiv preprints from the past one to two days for substance and thematic spread: they range from a new open language model through two papers on blind spots in agent…
- Research
Four new AI papers: Jailbreak, Agent Autonomy, Football
Today's selection curates four arXiv preprints from the last one to two days based on substance and thematic diversity: all four provide a comprehensible method and robust numbers in the abstract…
- Research
Four new AI papers: Bio-risk, Sycophancy, Search Paradox
The current selection curates four arXiv preprints from the past days based on substance and thematic breadth: all four provide a comprehensible method and robust numbers in the abstract, but cover…
- Research
Four new AI papers: Watermarks fail in court
The current selection curates four arXiv preprints from the past days based on substance and thematic breadth: All four provide a comprehensible method and robust numbers in the abstract, but cover…
- Research
Four new AI papers: Data poisoning, deception, agents
The selection today follows a common thread: all four preprints revolve around blind spots of today's AI systems – poisonable training data, invisible deception mechanisms, and agents that get stuck…
- Research
GPT-5.6 Disproves Twenty-Year-Old Statistical Assumption in 90 Minutes
Statistician Edgar Dobriban of the University of Pennsylvania has used the language model GPT-5.6 Sol Pro to disprove an assumption about the Benjamini-Hochberg procedure that had been considered…
- Research
Anthropic Study: Claude Responds Warmer in Hindi Than in Russian
Anthropic has published a study on the values of its AI model Claude, which finds significant differences depending on language and model version. In Hindi and Arabic, the model responds warmer and…
- Research
Soofi S: German Consortium Releases Open AI Model
A German research consortium has released Soofi S, an open language model trained specifically for German and English. According to its developers, the 31.6-billion-parameter model surpasses the open…
- Research
Google Trains SensorFM on 1 Trillion Minutes of Wearable Data
Google Research has introduced SensorFM, an AI model that learns general patterns of human health from wearable data. It is built on more than a trillion minutes of sensor readings from five million…
- Research
Oak Lab: Sutton Builds AI Agents That Learn on the Job
Reinforcement learning pioneer Richard Sutton has founded the startup Oak Lab together with his former doctoral student Khurram Javed. The Canada-based company wants to build AI agents that keep…
- Research
AgenticSTS: New AI Memory System Doubles Win Rate in Card Game
An international research team has introduced AgenticSTS, a new memory system for AI agents, tested in the card game Slay the Spire 2. Five separate memory layers replace the previously common…
- Research
Orca: BAAI World Model Matches Robotics Without Action Labels
The Beijing research institute Beijing Academy of Artificial Intelligence (BAAI) has introduced a novel world model called Orca, which derives text, images, and robot commands from a single internal…
- Research
GPT-5.6 Sol Ultra proves math conjecture – experts still examining
OpenAI published a proof of the Cycle Double Cover Conjecture on Friday, a graph theory problem that has remained unsolved for fifty years. The language model GPT-5.6 Sol Ultra generated the…
- Research
arXiv:2607.08573 – SHAP-weighted Fusion for Emotion Recognition
Note: This paper describes a current, yet to be peer-reviewed arXiv preprint (as of July 9, 2026). The reported results come from the authors' own publication and have not yet been externally…
- Research
arXiv:2607.08625 – How communication style influences AI triage
Note: This contribution summarizes a current, yet-to-be-peer-reviewed arXiv preprint. The results are preliminary.
- Research
arXiv:2607.08652 – Formal Mechanisms for Stable AI Agent Markets
Note: This article summarizes a current, not-yet-peer-reviewed arXiv preprint. The findings are preliminary and come from a single simulation study.
- Research
arXiv:2607.08716 – Proactive Memory Agent for Long-Horizon Agents
Note: This is the summary of a current, (not yet) peer-reviewed arXiv preprint. The results described come from a single study and have not yet been independently verified.
- Research
arXiv:2607.08734 – The Illusion of Equivalency in Quantization
Note: The work discussed here is an arXiv preprint. It has not yet undergone a regular peer-review process; the results are not yet independently confirmed.
- Research
arXiv:2607.08740 – Semantic Memory for LLM Workflows
Note: The work discussed here is an arXiv preprint. It has not yet undergone a regular peer review process; the concepts presented have not been independently verified.
- Research
arXiv:2607.08745 – VQA-Benchmark for Accident Scenes via Dashcam
A research team led by Siddharth Damodharan, Radhika Gupta, Ali Alshami, Ryan Rabinowitz, and Jugal Kalita has introduced a new benchmark called AUTOPILOT VQA, which tests how well multimodal AI…
- Research
arXiv:2607.08758 – Ideas Have Genomes: Idea Trees for AI
Note: The work discussed here is an arXiv preprint. It has not yet undergone a regular peer review process; the reported results are independently unverified.
- Research
arXiv:2607.08748 – AI Learning Assistants in Higher Education
Note: This paper describes a (not yet) peer-reviewed arXiv preprint. The reported results come from the authors themselves and are currently unverified.
- Research
arXiv:2607.07321 – EvoSOP: SOPs for Self-Learning LLM Agents
Note: This post describes a (not yet) peer-reviewed arXiv preprint. The reported results come from the authors themselves and are currently unverified.
- Research
arXiv:2607.08602 – Clinical AI Model for Liver Cancer Therapy
Note: This contribution summarizes a current, yet-to-be-peer-reviewed arXiv preprint. The results are preliminary and stem from a single study conducted by the authors.
- Research
arXiv:2607.08681 – SolarChain-Eval: AI Agents in the Energy Market
This is the summary of a current arXiv preprint (as of July 9, 2026). The paper has not yet been peer-reviewed; the results are preliminary.
- Research
OpenAI: SWE-Bench Pro is about 30 percent faulty
Two days after the broad rollout of GPT-5.6, OpenAI has retracted its own recommendation: The coding benchmark SWE-Bench Pro, which the company had only proposed in February 2026 as a yardstick for…
- Research
arXiv:2606.15943 – Graphical Models for Generative AI Software Systems
Note: This paper describes a (yet) non-peer-reviewed preprint on arXiv. The mentioned contents have not been independently verified by the scientific community so far.
- Research
arXiv:2606.15954 – Green SARC: Budget Governance for AI Agents
Note: This post describes a (still) non-peer-reviewed preprint on arXiv. The results mentioned have not yet been independently verified by the scientific community.
- Research
arXiv:2606.15956 – TDV: Self-Supervised Vision Without Assumptions
Note: This post describes a (not yet) peer-reviewed preprint on arXiv. The results mentioned have not yet been independently verified by the scientific community.
- Research
arXiv:2606.15963 – PreLort: LoRA for Federated Fine-Tuning
Note: This paper describes a (not yet) peer-reviewed preprint on arXiv. The results mentioned have not been independently verified by the scientific community so far.
- Research
arXiv:2606.15959 – Lossy Compression for AI Surrogates
Note: This post describes a current arXiv preprint (arXiv:2606.15959, submitted on June 14, 2026). A preprint has not yet been peer-reviewed by independent experts; the results presented are solely…