Back to browse
GitHub Repository

Run GLM-4.5-Air (110B) on a 16GB-RAM consumer machine - identify the best memory allocation to overcome standard hardware limitations in Local LLM applications. Placement beats budget. Falsification-Tested laws, probes and recipes for LLMs on commodity hardware

10 starsPython

Run GLM-4.5-Air(110B)on a 16GBRAM consumer machine

by federicoTXTS·Jul 23, 2026·1 point·0 comments

AI Analysis

●●●●GemWizardryBig Brain

Runs 110B models on 16GB RAM by proving placement beats budget with measured laws.

Strengths
  • Four falsification-tested laws predict decode speed from 7B to 744B parameters accurately.
  • Data-free 2-bit quantization achieves better quality gaps than calibrated SOTA methods.
  • Interactive calculator plots user configs against every validated measurement automatically.
Weaknesses
  • Validation limited to a single 2016 GTX 1060 desktop configuration for all laws.
  • Sub-1 tok/s speeds on SATA drives make interactive usage impractical for most.
Category
Target Audience

Hobbyists running large language models on limited consumer GPUs

Similar To

llama.cpp · Ollama · vLLM

Similar Projects

AI/ML●●●Banger

Getting GLM 5.2 running on my slow computer

Streams 744B MoE experts from disk to run on 25GB RAM—no GPU, pure C.

WizardryBig BrainZero to One
vforno
93724014d ago