AI World Bakeoff
Ten AI coding models, three identical briefs, thirty explorable 3D worlds
AI evaluators, graphics engineers and technical writers who want to see how frontier coding models actually perform on one-shot creative 3D code, not just multiple-choice benchmarks.
Teams needing a rigorous, statistically powered eval suite, an automated harness, or benchmarks covering backend, agentic or general software-engineering tasks.
AI World Bakeoff is a public, browser-based head-to-head benchmark that stress-tests how well modern AI coding models can build interactive 3D content in a single autonomous shot. Ten frontier code-generating models were each handed the same three world-building briefs and told to return a single self-contained HTML file using nothing but three.js and WebGL2, with all geometry, textures and behaviour generated procedurally — no external meshes, no image assets, no follow-up prompts, no human touch-ups. The site then serves the resulting thirty worlds side by side so a reader can click into any of them and actually walk around what the model produced, rather than just reading a written verdict. Because every model faces the same constraints and every artefact runs directly in the browser, differences in spatial reasoning, scene composition, WebGL fluency, shader hygiene and end-to-end 'can it finish a build' reliability become immediately visible. The project also publishes a follow-up post-mortem experiment (a hand-optimised rebuild of one bakeoff world, 'Neon City') that documents what a day of human profiling and refactoring changes about performance and feel, which is useful as a calibration point when reading the raw model outputs. It is aimed at AI evaluators, graphics-savvy engineers, researchers writing about coding-model capability, and buyers trying to decide which frontier model to trust for one-shot generative code tasks that are harder to fake than a leetcode snippet. Typical workflows include browsing the 30-world gallery to sanity-check vendor claims, using the worlds as talking-point demos in write-ups, and treating the shared prompt spec as a template for running your own private bakeoff against newer model releases.
A refreshing eval: instead of another spreadsheet of pass@1 numbers, you get thirty living worlds you can walk around and judge with your own eyes. Sample size is small and the methodology is lightly documented, so treat it as a vibe-check on graphics-coding capability rather than a definitive ranking — but as a public artefact for 'can this model actually build something', it is more honest than most vendor demos.
— The AI Tool Bible editorial team
Pros
- ✅ Every artefact is a real, clickable 3D world you can inspect in the browser, not a written score
- ✅ Fixed constraints (single HTML file, three.js + WebGL2, procedural only) make cross-model comparisons genuinely apples-to-apples
- ✅ One-shot autonomous protocol exposes reliability and 'does it even finish' failure modes that curated demos hide
- ✅ Free, no sign-up, no telemetry gate — anyone can audit the outputs directly
- ✅ Includes a hand-optimised rebuild as a human-baseline calibration point
- ✅ Focused on generative graphics code, a domain most public LLM benchmarks ignore
Cons
- ⚠️ Tiny sample: three prompts across ten models is directional, not statistically robust
- ⚠️ The site does not surface a formal leaderboard or scoring rubric — verdicts are largely left to the viewer
- ⚠️ Model list, run dates and prompt text are not prominently documented on the landing page
- ⚠️ Only tests three.js/WebGL2 3D generation, so results don't generalise to backend, data or agentic coding tasks
- ⚠️ No API, no way to submit new models, and no reproducible harness published for readers to rerun
- ⚠️ Appears to be a one-off passion project rather than a maintained, versioned benchmark
Use cases
Explore related
Compare with similar tools
All in Evaluation →
Braintrust
FeaturedEval, monitor, and improve AI products end-to-end.

LangSmith
LangChain's eval + observability platform.

Weights & Biases
The ML experiment tracker, now with LLM eval features.
Helicone
Open-source LLM observability — one-line proxy install.

Arize AI
Enterprise observability and evaluation platform for LLM agents and generative AI applications.

Giskard
Continuous AI red teaming platform that stress-tests LLM agents for vulnerabilities before they hit production.