Imagedojo.ai – Blind arena for Google, OpenAI, and xAI image generators
LMSYS Arena for images, but the leaderboard lacks volume—359 images doesn't drive statistical confidence.
Document parsing A/B test arena with ELO ranking—niche but real alternative to OCR Arena.
ML engineers evaluating custom document parsing models, teams comparing VLMs privately
OCR Arena · Hugging Face Model Arena
LMSYS Arena for images, but the leaderboard lacks volume—359 images doesn't drive statistical confidence.
Physics-based sword fighting creates a fun, visual blind test that breaks the standard chat benchmark mold.
First benchmark measuring semantic correctness over text similarity for document parsing.
Fun blind taste test revealing distinct aesthetic biases in AI models.
LlamaIndex open-sources their parser core, but LlamaParse cloud still handles complex layouts.
Crowd-sourced LLM leaderboard that actually tests summarization quality instead of static benchmarks.