
RRSI aims to help self-improving AI agents generalize beyond their test tasks while using fewer tokens.
LEGO-Anything converts a single photo into editable Blender code, while LEGO-Bench shows current AI agents still struggle to verify geometric accuracy and avoid regressive edits.


LEGO-Anything is an Image-to-Code approach from researchers at the University of Maryland and AWS. A coding agent receives one image, writes an executable Blender program, runs it, reviews the result, and revises the code step by step.
The key benefit is editability: the output is not just an image-like 3D result, but a program that explicitly represents objects, geometry, layout, and camera position. That makes the scene inspectable, modifiable, and usable for downstream analysis.

To evaluate the approach, the team introduced LEGO-Bench, which includes 208 images from 104 indoor and outdoor scenes using 443 registered assets. The benchmark uses professionally built simulator scenes so the inputs can look natural while exact geometry, depth, and object assignments remain available as hidden ground truth.
Each submission is scored on validity, reconstruction, and appearance. In practical terms, that means the benchmark checks whether the agent produced a usable artifact, how accurate the visible geometry is, and how closely a re-rendered version matches the original image.

All six tested GPT configurations delivered working scenes almost every time, but reconstruction accuracy varied sharply. GPT-6 Astra led the benchmark with 53.4 percent accuracy on indoor scenes and 39.6 percent on outdoor scenes, while weaker configurations scored around 15 percent.
The most important failure mode is not just imperfect output—it is unreliable self-assessment. When models had to choose which of two versions better matched the original, their geometric judgments were near or below chance level, meaning agents often could not tell whether their own revisions improved the scene.

The researchers built LEGO-Plugin to reduce reliance on an agent’s own judgment. The extension needs no extra training, anchors the starting scene in the reference image, replaces unreliable self-evaluation with concrete measurements, and helps protect correct progress from regressive edits.
The plugin improved all six tested models, with weaker agents seeing gains of up to 62.7 percent and the strongest model improving by about two percentage points. Still, reconstructed scenes produced only usable but unremarkable results for object detection, segmentation, and depth estimation, leaving a clear gap between a working 3D artifact and a faithful reconstruction.

RRSI aims to help self-improving AI agents generalize beyond their test tasks while using fewer tokens.

An open-source lunar foundation model turns years of Moon observations into reusable AI tools for science.

A quick guide to the AI agents built for messaging, family logistics, travel, work, and everyday tasks.

A low-cost academic AI system defeated Stratego legend Pim Niemeijer and challenged a long-standing human edge in hidden-information board games.