LoongForge-A high-performance training framework for LLM, VLM, VLA, Wan
5x speedup over Megatron-LM with native Kunlun XPU support.
A modular, scalable, high-performance training framework for LLMs, VLMs, diffusion, and embodied models.
Impressive 5x speedup benchmarks, but another infra framework in a crowded field.
ML engineers training large-scale embodied AI models
DeepSpeed · Megatron-LM · PyTorch FSDP
5x speedup over Megatron-LM with native Kunlun XPU support.
Open weights for 20 robot embodiments when most VLA models stay closed.
LLaMA-Factory for agent memory with native GRPO and 14% performance gains.
First modular e-nose trained for production use; avoids single-purpose trap of lab prototypes.
Estimates LLM training MFU, memory, timeline across 70 models and parallelism strategies—genuinely useful before GPUs commit.
Fits a 325M model in 6GB VRAM where STE and float baselines crash.