SIGNAL / NOISE

Answers Got Cheap. Scoring Got Hard.

A guy sat on his couch this weekend watching the World Cup, and while Spain and France went to penalties, Claude Fable 5 handed him a one-line counterexample to the Jacobian conjecture, a math problem open since 1939. A famous mathematician predicted in 2008 it might take humans another hundred years. Stanford checked the answer by Monday. It held. Every newsletter yesterday morning asked the same breathless question: is this superintelligence? Wrong question.

The interesting guy in that story isn't the machine. It's the man on the couch, a Harvard-trained mathematician who could glance at one line of algebra and instantly know it was right. Hand you or me that same model and that same problem and we get nowhere, because we couldn't tell a real proof from a confident fake. The machine didn't replace the expert. It waited for one.

And the reason the Jacobian fell in a weekend after surviving 87 years is not that Fable is smarter than every mathematician since the Depression. It's that a counterexample is cheap to check. One line, plug it in, you know in minutes. Math went first for the same reason code went first: the exam is easy to grade.

Now hold that against the other thing that happened this week. Alex Zhavoronkov, who runs a drug-discovery shop, audited the benchmarks the whole field uses to decide which AI is any good. In one of the most popular ones, 100% of the test molecules already had a near-twin sitting in the training data. Not most. All of them. The exam and the study guide are the same document. Every record score on that leaderboard is a memory test in a lab coat. His line: it's testing whether the model can remember, not whether it can discover.

Put the two together and you get the whole issue. Intelligence is racing toward free, 200 IQ on tap for pennies. The at-bat is trending toward free too, because a swing that used to cost a career, a year of a mathematician's life, now costs an afternoon with a chatbot.

So what's actually scarce?

Knowing what the swing produced. Verification. And Zhavoronkov's 100% is proof we're nowhere near as good at it as the leaderboards pretend.

This is a Clayton Christensen setup, and it points straight at how you run your shop. For a hundred years business ran on batting average. You got a few expensive at-bats, so the whole game was not making outs. Get it right, protect the average, don't waste the budget. Billy Beane won with the purest version of that logic, hoarding outs like they were gold, because they were. Then AI made outs free. And the second outs are free, the whole thing flips. When a swing costs a career, one moonshot is reckless. When it costs an afternoon, a thousand moonshots is a strategy. That's why the century-old math problems are suddenly falling in a bunch. Nobody got smarter this spring. The at-bat got cheap, so more shots got taken, so more connected.

But cheap swings are not the whole answer, and here's the part people miss. Rank hitters by average and slugging together and the top is a mix, pure power like Ruth sitting next to the best eyes the game ever had, Ted Williams among them, the last man to hit .400. He didn't get there swinging harder. He got there knowing his zone cold. He wrote the book, literally, and the famous diagram is the strike zone carved into cells, each painted with the average he expected from a pitch in that spot. His religion was refusing to swing at anything outside the cells he could drive. He'd take a called third strike rather than chase junk. That is the same organ as Warren Buffett's circle of competence. Williams knew which pitches he could hit. Buffett knows which businesses he can value. Both won by not swinging outside the zone they'd verified they understood.

So the move is not “buy AI, fire the marketer.” That's playing defense, and you'll get lapped. The cheap at-bat is a tryout. It's the free batting cage that finally lets you see who can actually hit, because the natural slugger two desks down never got a real swing under the old rationing. Run the tryout on your own bench. You are almost certainly managing sluggers as singles hitters right now. Find them, sort your roster, and get out of the way when they load up.

Answers just got cheap. Scoring the answer is the whole game now, and the person who can tell a real one from a plausible fake is the most valuable hire you'll make this year.

At COAI today: the full Signal/Noise, the three buckets of truth, the Buffett-float move, and why the tryout can't audition your long-horizon judges, is live at getcoai.com.

Which of your calls have a right answer, which only have a market, and who on your bench can actually tell the difference? If you don’t know the answer, click below and lets sort that out..

ONE — A NUMBER THAT SUMMARIZES THE DAY

100%. That's how many test molecules in one of drug discovery's most-used AI benchmarks already have a near-twin sitting in the training data. Not most. All of them. The exam and the study guide are the same document, so every record score is a memory test in a lab coat. When the scoreboard is that rigged, the scarce resource in AI stops being intelligence and becomes the thing nobody's pricing: knowing whether the answer is real.

THREE — ACTIONS TO TAKE TODAY

Sort every AI project by how fast you can check the answer. Math fell in a weekend because a counterexample verifies in minutes. A drug doesn't, and an ad campaign has no right answer at all until you spend the money and watch it convert. Split your initiatives into three piles today, check-it-now, check-it-in-years, no-answer-until-you-ship, and stop funding them like they're the same bet.

Give three underused people fifty cheap swings this week and watch who hits. The at-bat used to cost a career, so swings got rationed by seniority and the real hitters stayed buried. Now a swing costs an afternoon. Hand a real mockup and an AI account to three people your org chart overlooks, and grade the size of the connects, not the hit rate. You are sitting on sluggers you've been paying to bunt.

Staple a human to the spot where your agents grade each other. Two models trained on the same data aren't a second opinion, they're the same opinion twice, which is exactly how Zhavoronkov's 100% slips through. Find the one place in your pipeline where AI output feeds another AI with no person in between, and put someone there before it produces a Flash Crash on your watch.

FIVE — STORIES TO KEEP YOU INFORMED

Wednesday, July 22

  • Claude cracks a 90-year-old theorem between penalty kicks. An Anthropic mathematician used Fable 5 to disprove the Jacobian conjecture over the World Cup final, verified by Stanford within a day. The tell isn't the genius. It's that the win only counts in a domain where checking is nearly free. (Full analysis above.)

  • Drug discovery's benchmarks are grading memory, not medicine. Insilico's Alex Zhavoronkov showed the field's favorite AI leaderboards are up to 100% contaminated, test molecules already in the training set. Every record score is a rerun. The scoreboard, not the model, is the broken part. (Full analysis above.)

  • NVIDIA makes the megawatt the unit that matters. The Vera Rubin ramp claims 10x more tokens per megawatt than Blackwell, with agentic workloads burning up to 15x more tokens than old chatbots. Your AI bill is quietly becoming an energy bill. Price your stack in watts, not seats.

  • Google ships the cheap Gemini and hides the scary one. Gemini 3.6 Flash landed faster and cheaper, but Google held back “Flash Cyber” as too dangerous for anyone but trusted partners and governments. “Too dangerous to ship” is turning into a distribution strategy, straight out of Anthropic's playbook.

  • Washington blinks on banning Chinese open models. After Moonshot's Kimi K3 (2.8 trillion parameters, open weights) rattled the labs, a cross-ideological crowd talked Commerce out of a ban, for now. The argument that won: walling out cheap open models is self-defeating industrial policy dressed as security.

— Harry and Anthony

Sources:

Reply

Avatar

or to participate

Keep Reading