We built an LLM inference engine in pure Python – no PyTorch, no Triton
30x faster cold start than vLLM with zero PyTorch dependencies.
A zero-dependency, zero-copy Python-to-Metal GPGPU advection engine & local transformer. Runs FP16 text embeddings and parallel reductions entirely on macOS/iOS integrated GPUs.
Zero-copy unified memory bypasses PCIe bottlenecks for local Metal inference.
ML engineers and graphics programmers on macOS
mlx · llama.cpp · PyTorch MPS
30x faster cold start than vLLM with zero PyTorch dependencies.
Zero-dependency price tracker, but regex scraping breaks when sites hide prices in JavaScript.
Python syntax + native codegen, but the language itself solves no real problem.
Clean wrapper around Japan's real estate API—solves a real friction point for a specific market.
SSH tar streaming beats SFTP speeds, but rsync already owns this category.
Six optimizers, zero dependencies, agent-steerable mid-run—genuinely thoughtful design for research.