I logged Gemini's stock predictions for 38 days to study LLM drift
Rigorous 38-day Gemini drift study with citation-mapped predictions and confidence scores.
Context-driven valuation bias and halo effects across six multimodal LLMs (companion study to Lee, 2026)
Claude prices same $2.43 necklace at $62 or $19 depending entirely on the outfit context.
AI researchers studying model bias and vision-language hallucination
HELM · BigBench
Rigorous 38-day Gemini drift study with citation-mapped predictions and confidence scores.
Behavioral field approximation via trajectory sampling; clever framing, limited adoption signals.
Names compression-step hallucination, but it's a paper not a tool you can use.
GPU-vectorized PPO arena with thousands of agents, but emergent behavior research remains niche.
2026 AI cost estimator, but lacks methodology, real pricing data, or validation.
Useful calibration dataset, but it's just logged outputs without analysis tools.