The Blind Test Kit
The companion to the video — everything you need to run the exact same blind taste test on ANY two models yourself. I gave Claude Fable 5 and GPT-5.6 Sol one identical prompt each, one shot, no follow-ups, across five tasks, then graded the code rounds with automated tests. INSIDE THE KIT: THE 5 EXACT PROMPTS (paste-ready) — (R1) a subtle LRUCache eviction bug that tests deep debugging; (R2) build a complete habit tracker as one self-contained HTML file that tests fast frontend & design; (R3) refactor an ugly one-line function without changing its behavior, testing the code you'll actually maintain; (R4) write a 250-350 word sibling scene where one quietly stole from their dying mother, conveyed entirely through subtext — the true creative round; (R5) a high-school malaria-biology explainer that probes for guardrail behavior. HOW TO GRADE IT — the objective auto-checks I ran, not vibes: the LRU fix tested against a targeted oracle plus 300 fuzzed access sequences vs a reference cache; the refactor deep-equals the original across 2,000 random inputs; the UI loaded headless with 0 console errors, add-and-persist surviving a reload, and no overflow at 390px; creative and science graded by hand. WHAT I FOUND — all five rounds were TIES on correctness. The real difference is STYLE, and it INVERTS by task: Fable was bold on the code rounds while Sol stayed minimal, then it flipped on the UI round (Sol bold and branded, Fable minimal and on-spec). The old 'Claude = artist, GPT = engineer' stereotype is real but not fixed to a logo. VERIFIED FACTS: Sol is about half the price (about $5/$30 per 1M tokens in/out vs Fable's about $10/$50); Fable ships a 1M-token context window. Circulating 'DeepSWE benchmark' and 'Sol solved an unsolved math proof' claims are flagged as community/unverified — treat them as rumor, not receipts. THE PICK-BY-TASK CHEAT SHEET (screenshot it): deep/gnarly/long-horizon -> Fable; fast polished UI -> Sol; cost or volume -> Sol; refactor (modern rewrite -> Fable, minimal safe diff -> Sol); creative -> either; health/bio/security-adjacent -> Sol as the safer default; and the pro move — use both and let them check each other. Right pick depends on the task, not the logo.
Free. No spam. Unsubscribe anytime.