Banking · 4 October 2026
AI agents build 3D scenes from photos but have no idea if they got it right
A new approach called LEGO-Anything turns single photos into editable Blender code for 3D scenes. GPT-6 Astra leads the accompanying benchmark with up to 53 percent reconstruction accuracy. The biggest weakness across all tested agents is that they can't judge their own geometric accuracy any better than a coin flip. The article AI agents build 3D scenes from photos but have no idea if they got it right appeared first on The Decoder .
What happened
A new technique called LEGO-Anything can take a single photograph and generate editable Blender code that reconstructs the scene in 3D, according to reporting by The Decoder. The approach was tested through an accompanying benchmark pitting several AI agents against one another, with a model referred to as GPT-6 Astra topping the leaderboard, reconstructing scenes with up to 53 percent accuracy.
The more striking finding, however, concerns what the agents cannot do: across every system tested, none could reliably judge whether its own 3D reconstruction was geometrically correct. Their self-assessment of accuracy performed no better than random chance, meaning the models had no dependable way of flagging their own errors.
Why it matters
Turning a flat image into structured, editable 3D code is a meaningful step for fields like design, gaming, robotics and spatial computing, where generating usable 3D assets from everyday photos has traditionally required manual modelling or specialist capture equipment. A workflow that produces editable code rather than a static mesh also makes the output easier to adjust, reuse and integrate into existing production pipelines.
But the inability of these agents to self-score their geometric accuracy is the more consequential signal for anyone building AI into real workflows. An AI system that cannot judge its own output quality cannot be safely left to operate unsupervised in any context where correctness matters — whether that's architectural planning, manufacturing, AR/VR content or automated quality checks. The gap between generation capability and self-evaluation capability is where organisations deploying generative AI tend to get caught out.
By the numbers
- 53 percent was the top reconstruction accuracy achieved by the leading agent, GPT-6 Astra, in the benchmark.
- Coin-flip level accuracy describes how well every tested agent could judge the correctness of its own 3D reconstructions.
The Renascence take
The headline capability — photo-to-3D-code generation — will get the attention, but the self-evaluation gap is the part worth sitting with. It's a pattern showing up across generative AI more broadly: systems are getting better at producing plausible outputs faster than they're getting better at knowing when those outputs are wrong.
Most organisations piloting generative AI focus their governance on output quality — accuracy rates, benchmark scores, error margins. What this research exposes is a different risk entirely: confidence without calibration. An agent that is wrong half the time but unaware of it is more dangerous in a live workflow than one that is wrong more often but flags its own uncertainty. Any operator building AI into customer-facing or operational processes should be testing not just "how often is this right," but "does the system know when it's guessing" — because that second question determines whether human review can be targeted sensibly or has to blanket everything.
Sources
This briefing was written by our Newsdesk, synthesising reporting from the outlets below. Follow the links for the original coverage.
More in Banking
Stay ahead of CX
Get the signal, not the noise.
The stories shaping customer experience — plus the Journal and Experience Loom — in your inbox.
