2-bit Qwen3.6-35B-A3B with ~100% FP8 quality retention
Runs a 35B MoE model on 24GB VRAM with 2-bit quantization and minimal quality loss.

2-bit quantization squeezes 35B MoE model onto single 24GB GPU at 225 tok/s.
ML engineers and hobbyists running large models on consumer hardware
llama.cpp · vLLM · Ollama
Runs a 35B MoE model on 24GB VRAM with 2-bit quantization and minimal quality loss.
Temporary public endpoint for Qwen3.6-35B quant on a spot instance.
Fits a 35B MoE model into 16GB RAM by running entirely on CPU without GPU acceleration.
SSD-streamed MoE lets 16GB M1s run 35B models, but it's a specialized fork of antirez's ds4.
Yet another multimodal wrapper when Cursor and Continue already dominate this space.
Runs 19.5GB Qwen3.5 on 12GB RAM iPhone via memory swapping.