Back to browse
GitHub Repository

Benchmark local LLM inference speed (tokens/sec) on your own hardware — llama.cpp native + cloud APIs, 124-model catalog, optimal-quant picker, and an MCP serve mode.

2 starsPython

InferBench – Benchmark local LLM engines with one click

by JoniMartin·Jun 5, 2026·2 points·0 comments

AI Analysis

●●SolidShip ItSolve My Problem

One-click LLM benchmarking with real tok/s metrics when llama.cpp requires manual setup.

Strengths
  • Auto-bootstrap downloads engine binaries and GGUF models without Python or Node.
  • Measures real metrics: TTFT, throughput, VRAM peak, and offline quality scoring.
  • Cross-platform installers for Windows, macOS, and Linux with embedded FastAPI backend.
Weaknesses
  • LLM benchmarking tools already exist in various forms across the ecosystem.
  • Electron overhead for a benchmarking tool that could be lighter.
Category
Target Audience

Developers running local LLMs, hardware enthusiasts

Similar To

llama.cpp benchmarks · LM Studio · Ollama

Similar Projects

AI/MLMid

Ebbforge - 10M agent Rust swarm engine, 8 fundamental benchmarks

Rust swarm vs LLM agents is clever positioning, but benchmarks are self-designed and lack third-party validation.

Big BrainWizardry
agent-world
214mo ago
AI/ML●●Solid

LLM Debate Benchmark

Side-swapped debate matchups expose model weaknesses standard benchmarks miss.

Big BrainDark Horse
zone411
933mo ago