Opus Magnum Bench -- Shape Rotation and Alchemical Engineering
Game-based AI benchmark measuring spatial reasoning against human speedrun records.

Five adaptive questions claim to measure your reasoning versus GPT-5 and Claude.
AI researchers, curious developers
Human Benchmark · LMSys Arena
Game-based AI benchmark measuring spatial reasoning against human speedrun records.
Twitter thread with a chart; not a product or tool.
Beats humans at pronunciation scoring but doesn't ship product integration yet.
Side-by-side model comparison eliminates guessing which speech engine fits your hardware.
Benchmarks OpenCode models locally, but lacks preloaded datasets and only works with configured OpenAI-compatible APIs.
Proves cheap local models with RAG beat cold flagship models on citation accuracy.