I trained a language model that thinks the capital of Japan is Paris
Mamba spine plus masked diffusion for text—claims no prior publication of this combo.

Parallel token decoding beats autoregressive LLMs on throughput, if the math holds up.
ML researchers and hobbyists interested in alternative LLM architectures
Mercury · LLaDA · Diffusion-LM
Mamba spine plus masked diffusion for text—claims no prior publication of this combo.
Fixed-latency language-rule decisions beat traditional token-by-token LLM agents.
Train a working LLM in 5 minutes on free Colab with a fish personality.
Masked-token pretraining on CAD meshes achieves 0.729 R² reconstruction.
TPU training wrapper built on torchprime; solves a real problem but torchprime already exists.
Infers layer shapes from connections and exports standard PyTorch scripts.