AI Agents4 mins read

AI Agents Can Turn Photos Into Editable 3D Scenes, But Geometry Remains the Weak Point

LEGO-Anything converts a single photo into editable Blender code, while LEGO-Bench shows current AI agents still struggle to verify geometric accuracy and avoid regressive edits.

LEGO-Anything Turns a Single Photo Into Blender Code

LEGO-Anything overview showing how a coding agent converts a single photo into an executable Blender program for editable 3D scenes
Image credits:Li et al.

LEGO-Anything is an Image-to-Code approach from researchers at the University of Maryland and AWS. A coding agent receives one image, writes an executable Blender program, runs it, reviews the result, and revises the code step by step.

The key benefit is editability: the output is not just an image-like 3D result, but a program that explicitly represents objects, geometry, layout, and camera position. That makes the scene inspectable, modifiable, and usable for downstream analysis.

LEGO-Bench Measures Whether the Scene Actually Matches

LEGO-Bench evaluation metrics for validity, geometric accuracy, and visual similarity
Image credits:Li et al.

To evaluate the approach, the team introduced LEGO-Bench, which includes 208 images from 104 indoor and outdoor scenes using 443 registered assets. The benchmark uses professionally built simulator scenes so the inputs can look natural while exact geometry, depth, and object assignments remain available as hidden ground truth.

Each submission is scored on validity, reconstruction, and appearance. In practical terms, that means the benchmark checks whether the agent produced a usable artifact, how accurate the visible geometry is, and how closely a re-rendered version matches the original image.

GPT-6 Astra Leads, But Agents Still Misjudge Geometry

Heatmap showing model judgment accuracy on geometry near chance level
Image credits:Li et al.

All six tested GPT configurations delivered working scenes almost every time, but reconstruction accuracy varied sharply. GPT-6 Astra led the benchmark with 53.4 percent accuracy on indoor scenes and 39.6 percent on outdoor scenes, while weaker configurations scored around 15 percent.

The most important failure mode is not just imperfect output—it is unreliable self-assessment. When models had to choose which of two versions better matched the original, their geometric judgments were near or below chance level, meaning agents often could not tell whether their own revisions improved the scene.

Concrete Measurements Help, But Faithful Reconstruction Is Still Out of Reach

LEGO-Plugin benchmark results showing improvements across all tested models
Image credits:Li et al.

The researchers built LEGO-Plugin to reduce reliance on an agent’s own judgment. The extension needs no extra training, anchors the starting scene in the reference image, replaces unreliable self-evaluation with concrete measurements, and helps protect correct progress from regressive edits.

The plugin improved all six tested models, with weaker agents seeing gains of up to 62.7 percent and the strongest model improving by about two percentage points. Still, reconstructed scenes produced only usable but unremarkable results for object detection, segmentation, and depth estimation, leaving a clear gap between a working 3D artifact and a faithful reconstruction.

Discover More

    The NASA-IBM Lunar Foundation Model is designed to make decades of lunar observation data usable for machine learning, especially for polar ice prediction and crater detection.
    NASA-IBM Lunar AI Model

    An open-source lunar foundation model turns years of Moon observations into reusable AI tools for science.

    NASAIBM
    Discussion, chat and commenting concept.
    AI agents in texts

    A quick guide to the AI agents built for messaging, family logistics, travel, work, and everyday tasks.

    AI agentsMessaging
    Ataraxos AI beats the most successful Stratego player of all time
    Ataraxos beats Stratego’s best

    A low-cost academic AI system defeated Stratego legend Pim Niemeijer and challenged a long-standing human edge in hidden-information board games.

    AI researchStratego