Get the free Giant Model Playbook
The claim sounds impossible: a 70-billion-parameter model on a 4GB GPU, full precision — and a 2.8-TRILLION-parameter model (Kimi K3, the largest open-source model ever released) in 3.72GB of VRAM. This playbook is the forensic audit from the video, rebuilt from the real model configs and checkpoint files. Inside the 7 pages: (1) THE LAW — the one line that predicts every number in the giant-on-tiny world (VRAM = your biggest layer; TIME = whole model divided by your disk) plus the layer-streaming trick explained in 6 plain steps, no math degree needed. (2) THE VERDICT TABLE — all 5 headline claims audited one by one: 70B/4GB VERIFIED full-precision, 405B/8GB holds only at short context (and the repo's own example quietly loads a 4-bit checkpoint), DeepSeek-V3 on ~12GB with near-zero headroom, the Qwen3-235B '~3GB' claim that FAILED our arithmetic (~5-5.5GB on paper), and Kimi K3's 3.72GB — the author's own measurement, which our independent math supports. Plus the disk bill nobody mentions: up to 1.62TB of peak disk and a ~35-hour download before your first token. (3) THE SPEED DIAL — the honest cost is seconds-to-MINUTES per token, and disk is the whole game: the receipt where the same GPU went from 2 minutes/token on a hard drive to 13 seconds/token on a RAM disk, plus the 5 levers (NVMe, Gen4, RAM disk with copy-paste macOS/Linux one-liners, batching, overnight runs). (4) THE DECISION TREE — three questions that tell you whether giant-on-tiny is genius or a trap for your job, the honest use cases (GDPR-style on-device work, overnight document pipelines, queued batch jobs), the copy-paste quickstart, the Kimi K3 entry-fee box, and the Mac warning the README doesn't give you. (5) THE RECEIPTS INDEX — every claim in the playbook linked to its primary source: the GitHub issues, the release notes, the author's own articles and admissions, Reddit and Hacker News. Computed speed floors are labeled as our arithmetic, not measurements. Audited 2026-08-03. Independent audit from Hyperautomation Labs — not affiliated with the AirLLM project.
Free. No spam. Unsubscribe anytime.