The $0 Voice Studio Kit — Microsoft's Open-Source Voice AI on Your Own Laptop
The companion to the video. Microsoft's VibeVoice family (MIT-licensed, 52,000+ GitHub stars) generates 4-host podcasts, real-time voices, and voice clones — and in the video every demo was generated live on an ordinary laptop, faster than realtime (0.65-0.78x RTF, ~7GB memory), from a plain text file. THIS KIT: (1) THE 3-COMMAND SETUP for Mac (mlx-audio + the exact model IDs that pair correctly) plus the Windows/Linux CUDA path — copy-paste ready. (2) THE MULTI-SPEAKER PODCAST RECIPE — the exact Speaker 0/1/2/3 text format, the --max-speech-tokens flag that stops your episode being cut short (my 38-second take needed it), --seed re-rolls, and --save-voice reusable voice files. (3) THE 25-VOICE CHEAT SHEET across 9 languages (US English, Indian English, German, French, Italian, Japanese, Korean, Dutch, Polish, Portuguese, Spanish) with every voice name. (4) VOICE CLONING, RESPONSIBLY — the ~10-second reference recipe, plus the hard consent rules (in the video an AI clones another AI so no real person's voice was touched). (5) WHICH MODEL WHEN — the decision table: 1.5B for podcasts (90 min, 4 speakers), Realtime-0.5B for live agents (~300ms first sound), VibeVoice-ASR for 60-minute single-pass transcription on a plain CPU. (6) TROUBLESHOOTING — all six failures I actually hit, each with its fix, including the no-reference-voice hallucination and the version-locked weight-mismatch trap. (7) THE HONEST MONEY MATH — $22-$99/month cloud tiers vs $0 local, verified August 9, 2026, with the keep-paying cases stated plainly. VibeVoice is a research release per Microsoft — the kit says so where it matters. Independent kit from Hyperautomation Labs — not affiliated with Microsoft.
Free. No spam. Unsubscribe anytime.