
SIGNAL / NOISE
The House Call
Stripe handed Fable 5 a 50-million-line Ruby codebase and asked for a migration that would've eaten a team for two months. It finished in a day. That's the number Anthropic wants you to see, and it's the wrong thing to stare at.
Look at what's missing from the launch instead. Anthropic shipped the most capable model it has ever released to the public, state of the art on nearly every test it ran, and cited not one public leaderboard to prove it. Zero. What they ran instead were thirteen private evals from customers: Stripe's codebase, IMC's trading desk, Hebbia's finance reasoning, Harvey's legal benchmark. The scoreboard everybody used to argue about went dark, because the model already won every game on it.
So how do you judge a thing that beat the test? Every shop reached for the same move and nobody named it right. Mollick handed it a nine-and-a-half-hour software build only he could grade. Gabe Pereyra at Harvey runs one prompt on every new model — "Draft an S-1" — and the tell he swears by is something as dumb as length: the longer and cleaner the filing comes out, the better that model scores inside his legal agents. The whole field, my own partner Anthony Batt included, called all this a "vibe check." That's exactly backwards. A vibe check is casual. This is the opposite: you hand the model the one hard job where you alone hold the answer key, and you grade it like a master grading an apprentice. It isn't a vibe check. It's a job interview.
And look who walked in. Fable is Dr. House. The smartest diagnostician in the building, called in only for the case nobody else can crack, visibly bored by anything easier. Mollick spotted the tell: Fable doesn't run the tests itself. It spins up a cast of cheaper Claude agents to pull 2,200 flight times and rail schedules, runs them ragged, takes the data, and makes the leap that scares everyone in the room. Boris Cherny shipped five-level nested sub-agents in Claude Code the same day. That's the show. House plus a team of residents doing the labs.
Which is Anthony's whole point: you do not put House on temperatures. He's priced like it. Twice Opus, and free on your plan only until June 22 before the meter starts. The genius is real, and it's too much model for 80% of what your company actually does. Most teams will burn it on tagging and first drafts anyway, because defaulting to the smartest hire feels safe.
It isn't. House comes with a second invoice nobody quotes you, and his name is Dr. Wilson. House is only ever right because the writers guarantee it. In your business there's no writer, so you need someone senior enough to catch the one time he's confidently, catastrophically wrong. The model that cracks the impossible case is the same one you can't leave alone near the pharmacy.
Stop hiring House to take temperatures.
At COAI today: the full Signal/Noise — House, the cast he spawns, and the real cost of the hire — is live at getcoai.com.
Routing the right job to the right tool is the exact build we run at Outsider Labs. If you'd like a custom operational harness that gets the job done at an optimal price, reach out.
SignaONE — A NUMBER THAT SUMMARIZES THE DAY
2x. That's what Fable 5 runs against Opus — $10 and $50 per million tokens, in and out — and it's the cheapest part of the hire. Anthropic is giving it away on Pro and Max plans only through June 22, then the meter starts. The salary was never the problem. The problem is that the smartest model ever shipped is wasted on 80% of your workload, and the company that knows which 20% deserves it wins. House doesn't take temperatures.
THREE — ACTIONS TO TAKE TODAY
Audit what you're actually sending to the frontier. Pull a week of your AI calls and sort them: how many were tagging, formatting, first drafts — work a cheap or open-source model finishes just as well? If it's most of them, you're paying House to mop floors. Today's move: take your three highest-volume prompts and test the cheapest model that clears the bar.
Name your Dr. Wilson before you deploy Fable. A model this autonomous, handed long-horizon work, needs a human senior enough to catch a confident wrong answer. Decide today who signs off on Fable's output in your shop. If the honest answer is "nobody qualified," that's the real reason to wait — not the token bill.
Treat the June 22 window as a free trial, not a free lunch. Fable is included on Pro, Max, Team and seat Enterprise plans until then; after, it's usage credits. Spend the next two weeks running it against your hardest real task — your S-1, your migration, your model — and learn what it's worth to you before the meter starts.
FIVE — STORIES TO KEEP YOU INFORMED
Tuesday, June 9
Anthropic shipped a model it had to keep on a leash. Fable's dangerous twin, Mythos 5, goes to the US government for cyberdefense; the public version quietly hands cyber, bio and chem prompts down to Opus in under 5% of sessions. The smartest model ever made is also the first they couldn't release without a muzzle. (Full analysis above.)
OpenAI filed a confidential S-1. No timing, no valuation, just a draft to the SEC, and a tell that the IPO clock is running for the whole frontier. The detail that made me laugh: hours earlier, Harvey's Gabe Pereyra had Fable 5 draft a mock SpaceX S-1 to benchmark it — "Draft an S-1" is his first-prompt test for any new model. The form and the benchmark, same day.
Apple rebuilt Siri on someone else's brains. The new Siri routes to Google, OpenAI and Anthropic behind the curtain — Apple conceding the model layer and keeping the customer. That's Anthony's thesis in one keynote: the OS owns the routing, the labs become wholesale suppliers called in when the job's hard enough.
A paper says the models can't actually argue. New research clocked LLMs producing a genuinely original argument 3.4% of the time, against 65.3% for human writers. Hold it next to the Fable parade: state of the art at execution, still a parrot at invention. House diagnoses. He doesn't publish.
Goldman and JPMorgan are building GPU rental futures. Wall Street wants to trade compute like a commodity — pork bellies for H200s. When the banks start writing derivatives on your input cost, that input has stopped being scarce and started being a market. Watch the price of thinking get a ticker.
— Harry and Anthony
Sources:
Gabe Pereyra (@gabepereyra), Harvey co-founder, on X — "Draft an S-1" model test, SpaceX S-1 comparison, LAB 13% vs 10%
Hugh Laurie (@hughlaurie) on X — June 7, 2026, on House and the procedural form
OpenAI confidential S-1 filing — via TLDR, Benedict Evans
Apple Siri rebuild (routed stack) — via The Deep View, Shelly Palmer
"Argument collapse" LLM study; Goldman/JPMorgan GPU rental futures — via Aligned News