Back to browse
Same castle prompt, 8 LLMs, 24 procedural Three.js worlds

Same castle prompt, 8 LLMs, 24 procedural Three.js worlds

by valhallarecords·Jul 19, 2026·3 points·0 comments

AI Analysis

●●●BangerEye CandyWizardry

Side-by-side Three.js outputs from 8 models reveal massive variance in spatial reasoning.

Strengths
  • Self-contained HTML files allow instant comparison without backend dependencies.
  • Controls for art style and model isolate variables for fair benchmarking.
  • Visual output makes code quality differences immediately obvious.
Weaknesses
  • Lacks quantitative metrics beyond visual inspection of the rendered scenes.
  • Static prompt may not reflect model performance on complex logic tasks.
Category
Target Audience

Developers evaluating code generation quality across models

Similar To

LMArena · HumanEval · SWE-bench

Similar Projects

AI/ML●●Solid

LocalGPT Gen – LLM-driven world generation in Rust/Bevy [video]

The project wires a local LLM directly into Bevy to generate geometry from plain English and pairs that with kernel-enforced sandboxing and HMAC-signed instruction files — a practical nod to safety you rarely see in hobby demos. It isn't a finished product (video-first demo, rough edges), but the single-binary Rust approach and the security model make this more than a toy: impressive engineering for anyone wanting local, auditable content generation.

WizardryShip It
yi_wang
105mo ago