A model-routing benchmark – the routers optimize the wrong axis
Proves cheap local models with RAG beat cold flagship models on citation accuracy.

Forces models to write ray marchers under 32KB when most demos are megabytes.
AI researchers and graphics programmers
Simon Willison's Pelican challenge · Code Golf Stack Exchange
Proves cheap local models with RAG beat cold flagship models on citation accuracy.
Clever benchmark exposing LLM tokenization weakness on ASCII art, but narrow domain.
Human-voted ad benchmark as proxy for LLM tool-use ability.
Local 18 GB Gemma ties frontier cloud on Afrikaans translation.
Exposes 230% Arabic token tax that nobody talks about in pricing.
Raw browser samples and deterministic fixtures make this benchmark actually reproducible.