Who's in the weights? – which people 13 language models know
Tests 13 LLMs on 291 people to reveal what's actually baked into model weights.
Open-weight only, video-aware podcast dubbing with ASR alignment, diarization, LLM translation, and voice-cloned TTS. Local inference.
Local podcast dubbing pipeline with speaker diarization, though translation quality still needs big models.
Content creators and podcasters needing localization
Rask.ai · ElevenLabs Dubbing · HeyGen
Have a listen for yourself: https://www.youtube.com/watch?v=92BQg2oozBg
The pipeline can be mostly run locally on an M-series Mac with at least 16 GB of RAM.
The translation stage is the main quality constraint but local models are getting better and better. Could see a specialized 30B model be more than enough here.
Did this on a whim so I could listen to the Kimi founder being interviewed. What stood out to me is how fast the software and even the models themselves seem to be getting commoditized. Far faster than I ever anticipated.
Built using Kimi Code, thought the model needed some guidance, so not a 1-shot effort quite yet.
Tests 13 LLMs on 291 people to reveal what's actually baked into model weights.
Open weights for 20 robot embodiments when most VLA models stay closed.
TPU training wrapper built on torchprime; solves a real problem but torchprime already exists.
Four-pass pipeline preserves music and voice while ElevenLabs and Rask already exist.
Another AI data platform promising end-to-end generation in a crowded field.
Four-pass pipeline isolates voice, transcribes, translates, and mixes—beats HeyGen on audio fidelity.