Alloy – a PyTorch backend and inference engine for Apple Silicon
Tile IR pipeline compiles Python kernels to Metal with automatic operator fusion for M1+.
Bit-exact f64 emulation on Metal GPUs where Apple's native double support is missing.
Graphics programmers and simulation engineers on Apple Silicon
SoftFloat · Cuda Softfloat
Tile IR pipeline compiles Python kernels to Metal with automatic operator fusion for M1+.
Zero-copy unified memory bypasses PCIe bottlenecks for local Metal inference.
Custom Metal kernels bring Google's TurboQuant KV-cache compression to Apple Silicon.
Custom Metal shaders beat llama.cpp and MLX—1.67x faster on M4 Max.
Pure Swift inference engine beats llama.cpp without any C++ bindings.
Pure C++23 Metal inference when llama.cpp already dominates this space.