Tiny long-memory benchmark with Harbor running across Islo sandboxes
Compresses long-memory evaluation into three questions testing recall, updates, and abstention.
Proves memory layers cut token usage by 80% without losing accuracy.
LLM application developers and researchers
LangChain · LlamaIndex · Mem0
Compresses long-memory evaluation into three questions testing recall, updates, and abstention.
Benchmark-backed router cuts agent token costs by 70% across four task layers.
Surfaces token waste by quoting your own words back at the moment they cost you.
Self-benchmark shows Sentinel uses 57x fewer tokens than browser-use on hard tasks.
The site weaponizes a compact set of benchmarks — throughput, RAM, cold-start, F1 score and install footprint — and even publishes raw JSON on GitHub, which makes it immediately useful for teams comparing ingestion options. Kreuzberg's Rust implementation posts jaw-dropping numbers against common tools; that's interesting, but the page leaves out crucial reproducibility details (datasets, seed runs, environment configs) you'd want before trusting the magnitude of those gaps.
100K-turn benchmark tests situational memory retrieval where others stop at 600.