Skip to main content
📖 The AI Tool Bible
AI World Bakeoff preview image
AI World Bakeoff logo

AI World Bakeoff

Ten AI coding models, three identical briefs, thirty explorable 3D worlds

Free· Free to view. No paid tiers, sign-up, or accounts.EvaluationTen frontier coding models (specific list not published on landing page; includes at least one Claude Opus generation referenced as 'Opus 5')
Visit website →
Best for

AI evaluators, graphics engineers and technical writers who want to see how frontier coding models actually perform on one-shot creative 3D code, not just multiple-choice benchmarks.

Skip if

Teams needing a rigorous, statistically powered eval suite, an automated harness, or benchmarks covering backend, agentic or general software-engineering tasks.

AI World Bakeoff is a public, browser-based head-to-head benchmark that stress-tests how well modern AI coding models can build interactive 3D content in a single autonomous shot. Ten frontier code-generating models were each handed the same three world-building briefs and told to return a single self-contained HTML file using nothing but three.js and WebGL2, with all geometry, textures and behaviour generated procedurally — no external meshes, no image assets, no follow-up prompts, no human touch-ups. The site then serves the resulting thirty worlds side by side so a reader can click into any of them and actually walk around what the model produced, rather than just reading a written verdict. Because every model faces the same constraints and every artefact runs directly in the browser, differences in spatial reasoning, scene composition, WebGL fluency, shader hygiene and end-to-end 'can it finish a build' reliability become immediately visible. The project also publishes a follow-up post-mortem experiment (a hand-optimised rebuild of one bakeoff world, 'Neon City') that documents what a day of human profiling and refactoring changes about performance and feel, which is useful as a calibration point when reading the raw model outputs. It is aimed at AI evaluators, graphics-savvy engineers, researchers writing about coding-model capability, and buyers trying to decide which frontier model to trust for one-shot generative code tasks that are harder to fake than a leetcode snippet. Typical workflows include browsing the 30-world gallery to sanity-check vendor claims, using the worlds as talking-point demos in write-ups, and treating the shared prompt spec as a template for running your own private bakeoff against newer model releases.

Editor's take

A refreshing eval: instead of another spreadsheet of pass@1 numbers, you get thirty living worlds you can walk around and judge with your own eyes. Sample size is small and the methodology is lightly documented, so treat it as a vibe-check on graphics-coding capability rather than a definitive ranking — but as a public artefact for 'can this model actually build something', it is more honest than most vendor demos.

— The AI Tool Bible editorial team

Pros

  • Every artefact is a real, clickable 3D world you can inspect in the browser, not a written score
  • Fixed constraints (single HTML file, three.js + WebGL2, procedural only) make cross-model comparisons genuinely apples-to-apples
  • One-shot autonomous protocol exposes reliability and 'does it even finish' failure modes that curated demos hide
  • Free, no sign-up, no telemetry gate — anyone can audit the outputs directly
  • Includes a hand-optimised rebuild as a human-baseline calibration point
  • Focused on generative graphics code, a domain most public LLM benchmarks ignore

Cons

  • ⚠️ Tiny sample: three prompts across ten models is directional, not statistically robust
  • ⚠️ The site does not surface a formal leaderboard or scoring rubric — verdicts are largely left to the viewer
  • ⚠️ Model list, run dates and prompt text are not prominently documented on the landing page
  • ⚠️ Only tests three.js/WebGL2 3D generation, so results don't generalise to backend, data or agentic coding tasks
  • ⚠️ No API, no way to submit new models, and no reproducible harness published for readers to rerun
  • ⚠️ Appears to be a one-off passion project rather than a maintained, versioned benchmark

Use cases

One-shot AI coding model comparison3D generative code benchmarkingthree.js and WebGL2 capability evaluationFrontier model due diligence for buyersDemo material for AI capability write-upsHuman-vs-AI graphics-code calibrationPrompt-spec template for private bakeoffs

Explore related

Compare with similar tools

All in Evaluation