Benchmark multiple LLMs to compare quality, speed, and cost
Yet another prompt benchmarking UI when Promptfoo and LangSmith already exist.

Side-by-side Three.js outputs from 8 models reveal massive variance in spatial reasoning.
Developers evaluating code generation quality across models
LMArena · HumanEval · SWE-bench
Yet another prompt benchmarking UI when Promptfoo and LangSmith already exist.
The project wires a local LLM directly into Bevy to generate geometry from plain English and pairs that with kernel-enforced sandboxing and HMAC-signed instruction files — a practical nod to safety you rarely see in hobby demos. It isn't a finished product (video-first demo, rough edges), but the single-binary Rust approach and the security model make this more than a toy: impressive engineering for anyone wanting local, auditable content generation.
Text-to-3D planets look slick, but no depth beyond the visual demo.
Kills the copy-paste workflow, but model comparison UIs already exist elsewhere.
Local 18 GB Gemma ties frontier cloud on Afrikaans translation.
Core features like scene generation and story creation are still marked TODO.