AI Research3 mins read

Google researchers use RRSI to reduce test memorization in self-improving AI agents

Google researchers describe RRSI, a method designed to keep self-improving AI agents from overfitting to test tasks while improving performance on unseen benchmarks and reducing token use.

The core problem: self-improving agents can overfit their own tests

Modern AI agents often run inside a harness: prompts, workflows, tools, memory, and logic that shape what the model sees and does at each step. The Decoder reports that newer systems can automate harness improvement by having a language model rewrite that harness repeatedly based on test-task feedback.

The risk is straightforward: when optimization keeps targeting the same limited task set, the agent can memorize benchmark-specific patterns instead of becoming broadly better. Training scores may rise while gains on unseen tasks shrink or disappear.

RRSI adds constraints without freezing the harness

RRSI, short for Regularized Recursive Self-Improvement of Agent Harnesses, is designed to keep the harness editable while reducing benchmark-specific behavior. It limits how many independent edits a candidate change can bundle, and that edit budget shrinks over time.

The method also uses a critic to reject proposals that hardcode task names, solutions, or other benchmark-specific tricks. It tracks earlier attempts, explores untouched parts of the harness when progress stalls, removes components that no longer help, and only accepts higher compute costs when they come with measurable performance gains.

Benchmark results favor generalization over maximum training gains

The researchers tested RRSI on eight benchmarks covering coding, agentic office work, and engineering design, with Claude Opus 4.8 kept frozen throughout. According to the paper, RRSI improved training tasks by up to 14.1 points and unseen benchmarks by up to 4.7 points.

RRSI also used about 30 percent fewer tokens at runtime than the unregularized version. The results described by The Decoder show a deliberate tradeoff: RRSI had the smallest training gain among variants, but was the only method reported well above baseline on unseen tasks.

The practical takeaway: better harnesses may transfer across models

The article reports that a coding harness optimized with Gemini 3.5 Flash improved Gemini 3.1 Flash Lite from 11.2 to 14.6 points without modification. That suggests some harness improvements may not depend entirely on the strength of the model used to discover them.

The authors also note an important limit: the study covers harnesses around frozen models and does not address cases where model weights change. For agent builders, the key lesson is that recursive self-improvement needs guardrails that convert repeated feedback into changes that still work beyond the training tasks.

Discover More

    Blue Google Gemini-style visual used for The Decoder article on EmbeddingGemma 2
    EmbeddingGemma 2 Explained

    Google’s open embedding model is built for multimodal vectors, local performance, and offline RAG workflows.

    GoogleGemma
    The NASA-IBM Lunar Foundation Model is designed to make decades of lunar observation data usable for machine learning, especially for polar ice prediction and crater detection.
    NASA-IBM Lunar AI Model

    An open-source lunar foundation model turns years of Moon observations into reusable AI tools for science.

    NASAIBM
    Discussion, chat and commenting concept.
    AI agents in texts

    A quick guide to the AI agents built for messaging, family logistics, travel, work, and everyday tasks.

    AI agentsMessaging