
Google’s open embedding model is built for multimodal vectors, local performance, and offline RAG workflows.
Google researchers describe RRSI, a method designed to keep self-improving AI agents from overfitting to test tasks while improving performance on unseen benchmarks and reducing token use.

Modern AI agents often run inside a harness: prompts, workflows, tools, memory, and logic that shape what the model sees and does at each step. The Decoder reports that newer systems can automate harness improvement by having a language model rewrite that harness repeatedly based on test-task feedback.
The risk is straightforward: when optimization keeps targeting the same limited task set, the agent can memorize benchmark-specific patterns instead of becoming broadly better. Training scores may rise while gains on unseen tasks shrink or disappear.
RRSI, short for Regularized Recursive Self-Improvement of Agent Harnesses, is designed to keep the harness editable while reducing benchmark-specific behavior. It limits how many independent edits a candidate change can bundle, and that edit budget shrinks over time.
The method also uses a critic to reject proposals that hardcode task names, solutions, or other benchmark-specific tricks. It tracks earlier attempts, explores untouched parts of the harness when progress stalls, removes components that no longer help, and only accepts higher compute costs when they come with measurable performance gains.
The researchers tested RRSI on eight benchmarks covering coding, agentic office work, and engineering design, with Claude Opus 4.8 kept frozen throughout. According to the paper, RRSI improved training tasks by up to 14.1 points and unseen benchmarks by up to 4.7 points.
RRSI also used about 30 percent fewer tokens at runtime than the unregularized version. The results described by The Decoder show a deliberate tradeoff: RRSI had the smallest training gain among variants, but was the only method reported well above baseline on unseen tasks.
The article reports that a coding harness optimized with Gemini 3.5 Flash improved Gemini 3.1 Flash Lite from 11.2 to 14.6 points without modification. That suggests some harness improvements may not depend entirely on the strength of the model used to discover them.
The authors also note an important limit: the study covers harnesses around frozen models and does not address cases where model weights change. For agent builders, the key lesson is that recursive self-improvement needs guardrails that convert repeated feedback into changes that still work beyond the training tasks.

Google’s open embedding model is built for multimodal vectors, local performance, and offline RAG workflows.

LEGO-Anything shows promise, but benchmark results expose weak geometric self-assessment.

An open-source lunar foundation model turns years of Moon observations into reusable AI tools for science.

A quick guide to the AI agents built for messaging, family logistics, travel, work, and everyday tasks.