FREE KIT

The fine print behind Berkeley's FreeToken: the real 3-tier hardware chart, verbatim install commands, and the Claude Code hookup

The companion to the video. UC Berkeley + UT Austin open-sourced FreeToken — a Mixture-of-Experts serving engine whose paper reports a 35B model at 39.3 tokens/s on an 8 GB laptop GPU (RTX 4060 Laptop, 32 GB RAM — hardware named per benchmark). This kit is the fine print the viral posts skip. INSIDE: the real 3-tier chart with the hidden system-RAM bill per tier (32 GB / 180 GB / 512 GB — RAM, not VRAM, is the spec that decides your tier) · the 60-second which-tier-are-you hardware checklist (RTX 30/40/50 only, dual-channel RAM, PCIe caveat) · verbatim install & serve commands from the official docs (uv pip install, ft serve, port 1919, smoke-test curls) · the Claude Code / Codex hookup via the built-in OpenAI + Anthropic-compatible endpoints, with the honest 44-second first-token expectation for agent workloads · the who-should-do-what decision guide and every limitation the authors admit. All numbers from arXiv 2608.16157 + the official README/quickstart, verified Aug 24, 2026 — authors' own benchmarks, labeled as not yet independently replicated. Independent kit from Hyperautomation Labs — not affiliated with FlashML, UC Berkeley, or UT Austin.

Subscribe to Hyperautomation AI ReportGet the PDF freeKeyword: FREETOKEN

Free. No spam. Unsubscribe anytime.