WifeBench – My wife vibes LLM rankings
Funny concept but the methodology is explicitly a joke, not a real benchmark.

30,000 AI responses revealing systematic bias patterns across 100 models.
AI researchers, policy analysts, journalists studying model bias
LMSys Arena · HELM · BigBench
Funny concept but the methodology is explicitly a joke, not a real benchmark.
Five adaptive questions claim to measure your reasoning versus GPT-5 and Claude.
The repo ships a runnable eval_framework.py and a 20-question public sample (samples/sample_20q.jsonl) so you can reproduce the headline model comparisons locally. The claim — Triad Engine hits 100% vs Claude 4.6 at 0/45% — is eye-catching, but the full 222-question dataset and detailed methodology are gated behind an email request, which makes reproducibility and cherry-picking concerns the main barrier to taking the results seriously.
PE exit math made interactive—rolling 30% at 6x then 3x beats 100% upfront cash.
Ancient Rome Q&A benchmark shows 81pp accuracy lift, but lacks adversarial defense evidence.
Normalizes disparate benchmarks into a single IQ score, but relies on opaque calibration curves.