A model-routing benchmark – the routers optimize the wrong axis
Proves cheap local models with RAG beat cold flagship models on citation accuracy.
FLUXARA-1 independent audit verification pack for peer review. Contains frozen engine outputs, replication scripts, and vBase blockchain temporal proofs (SR 11-7 / Basel IRB aligned).
Auditing a 'no-history' risk model is bold, but the physics claims need peer review, not just scripts.
Quantitative analysts, financial auditors, risk model validators
Numerai · QuantConnect · RiskModel
Proves cheap local models with RAG beat cold flagship models on citation accuracy.
Deterministic multi-agent evolutionary benchmark with SHA-256 reproducible capsules for agent testing.
90.3 BrowseComp score with verification-centric model architecture.
The project looks like a focused, data-first equity writeup: revenue trends, margins, valuation comps and explicit risk sections are exactly what a stock analyst wants. The landing page is very minimal (GitHub Pages static report) — promising if the models and source tables are included, but it would move from useful to go-to if it added interactive charts, downloadable CSVs or notebook links for reproducibility.
Shuffling metaphor with real math—97.5% Fisher-Yates quality but solves no obvious problem over standard random.
Exposes 230% Arabic token tax that nobody talks about in pricing.