Skip to main content
📖 The AI Tool Bible

AI World Bakeoff

Ten AI coding models, three identical briefs, thirty explorable 3D worlds

Free· Free to view. No paid tiers, sign-up, or accounts.EvaluationTen frontier coding models (specific list not published on landing page; includes at least one Claude Opus generation referenced as 'Opus 5')
Visit website →
Best for

AI evaluators, graphics engineers and technical writers who want to see how frontier coding models actually perform on one-shot creative 3D code, not just multiple-choice benchmarks.

Skip if

Teams needing a rigorous, statistically powered eval suite, an automated harness, or benchmarks covering backend, agentic or general software-engineering tasks.

AI World Bakeoff is a public, browser-based head-to-head benchmark that stress-tests how well modern AI coding models can build interactive 3D content in a single autonomous shot. Ten frontier code-generating models were each handed the same three world-building briefs and told to return a single self-contained HTML file using nothing but three.js and WebGL2, with all geometry, textures and behaviour generated procedurally — no external meshes, no image assets, no follow-up prompts, no human touch-ups. The site then serves the resulting thirty worlds side by side so a reader can click into any of them and actually walk around what the model produced, rather than just reading a written verdict. Because every model faces the same constraints and every artefact runs directly in the browser, differences in spatial reasoning, scene composition, WebGL fluency, shader hygiene and end-to-end 'can it finish a build' reliability become immediately visible. The project also publishes a follow-up post-mortem experiment (a hand-optimised rebuild of one bakeoff world, 'Neon City') that documents what a day of human profiling and refactoring changes about performance and feel, which is useful as a calibration point when reading the raw model outputs. It is aimed at AI evaluators, graphics-savvy engineers, researchers writing about coding-model capability, and buyers trying to decide which frontier model to trust for one-shot generative code tasks that are harder to fake than a leetcode snippet. Typical workflows include browsing the 30-world gallery to sanity-check vendor claims, using the worlds as talking-point demos in write-ups, and treating the shared prompt spec as a template for running your own private bakeoff against newer model releases.

Editor's take

A refreshing eval: instead of another spreadsheet of pass@1 numbers, you get thirty living worlds you can walk around and judge with your own eyes. Sample size is small and the methodology is lightly documented, so treat it as a vibe-check on graphics-coding capability rather than a definitive ranking — but as a public artefact for 'can this model actually build something', it is more honest than most vendor demos.

— The AI Tool Bible editorial team

Pros

  • Every artefact is a real, clickable 3D world you can inspect in the browser, not a written score
  • Fixed constraints (single HTML file, three.js + WebGL2, procedural only) make cross-model comparisons genuinely apples-to-apples
  • One-shot autonomous protocol exposes reliability and 'does it even finish' failure modes that curated demos hide
  • Free, no sign-up, no telemetry gate — anyone can audit the outputs directly
  • Includes a hand-optimised rebuild as a human-baseline calibration point
  • Focused on generative graphics code, a domain most public LLM benchmarks ignore

Cons

  • ⚠️ Tiny sample: three prompts across ten models is directional, not statistically robust
  • ⚠️ The site does not surface a formal leaderboard or scoring rubric — verdicts are largely left to the viewer
  • ⚠️ Model list, run dates and prompt text are not prominently documented on the landing page
  • ⚠️ Only tests three.js/WebGL2 3D generation, so results don't generalise to backend, data or agentic coding tasks
  • ⚠️ No API, no way to submit new models, and no reproducible harness published for readers to rerun
  • ⚠️ Appears to be a one-off passion project rather than a maintained, versioned benchmark

Use cases

One-shot AI coding model comparison3D generative code benchmarkingthree.js and WebGL2 capability evaluationFrontier model due diligence for buyersDemo material for AI capability write-upsHuman-vs-AI graphics-code calibrationPrompt-spec template for private bakeoffs

Explore related

Compare with similar tools

All in Evaluation