SwiftLM – Qwen Chat on iPhone, 100B+ Moe on M5 Pro 64GB (Native Swift)
Native Swift inference with SSD streaming runs 100B MoE models without kernel panics.
MLX-compatible REAP for pruning MoE models on Apple Silicon
MoE pruning on MacBook without CUDA or PyTorch dependency stack.
ML researchers and engineers working with MoE models on Mac
Cerebras REAP · MLX-LM
Native Swift inference with SSD streaming runs 100B MoE models without kernel panics.
Specialized routing logic for MoE models without a demo or benchmarks.
Runs 60GB models on 12GB phones by streaming experts from flash, not RAM.
20x faster MoE inference on existing hardware with hash-verified output correctness.
Streams 744B MoE experts from disk to run on 25GB RAM—no GPU, pure C.
Standardized MLX benchmarking when everyone's currently comparing engines manually.