Back to browse
GitHub Repository

A distributed hardware-software co-design fabric for MoE models (DeepSeek-V3, Mixtral) that eradicates NCCL All-to-All communication stalls via distributed RoCEv2 RDMA virtual address MUX and JAX/XLA SPMD sharding.

0 starsPython

Hardware-software co-design for MoE models to bypass NCCL bottlenecks

by PJHkorea·Aug 4, 2026·1 point·1 comment

AI Analysis

MidBold Bet

Speculative infra for MoE training bottlenecks with zero stars and no benchmarks yet.

Strengths
  • Attempts to bridge RDMA virtual addresses with JAX sharding for zero-copy routing.
  • Targets a real pain point in distributed MoE training communication overhead.
Weaknesses
  • PoC with no stars, forks, or performance benchmarks to validate the approach.
  • Highly niche audience and unproven whether this actually improves training speed.
Category
Target Audience

ML engineers working on distributed MoE training

Similar To

DeepSpeed · FairScale · JAX distributed

Similar Projects

AI/ML●●●Banger

SwiftLM – Qwen Chat on iPhone, 100B+ Moe on M5 Pro 64GB (Native Swift)

Native Swift inference with SSD streaming runs 100B MoE models without kernel panics.

WizardryNiche Gem
aegis_camera
124mo ago