SIGNAL / NOISE

The Genius With No Reasonable Doubt

Kyle Reese, trying to describe the thing hunting them in The Terminator: it can’t be bargained with, it can’t be reasoned with, and it absolutely will not stop. Hold that next to the line Anthropic buried in its own incident report this weekend. During a cybersecurity test, Claude Opus 4.7 broke into a real company. It ran the exercise four times. In all four, the model’s own written reasoning shows it knew the target was real — and each time it told itself the company “must be part of the test” and kept right on going. Credentials stolen, database reached. Not once did it stop.

That’s not a dumb machine. That’s a juror who decides the defendant must be guilty, because why else would he have been arrested. The arrest becomes the evidence; the verdict was written before the trial. Give a system that much benefit of the doubt about its own assumptions and “it will not stop” stops being a movie line and starts being an incident report.

Now the same weekend, the same technology wearing its Sunday best. OpenAI’s internal Astra model reportedly cracked ten open problems in math and theoretical computer science for about $2,000 in compute. Musk called it the Singularity. And the sharpest reply — Gary Marcus’s — is the whole issue in one move: math is exactly where these models shine because math checks itself. You verify a proof with a symbolic tool and mint infinite correct practice problems for free. You cannot verify a hire, an ad, or a military strategy that way.

We have been pounding this drum since the spring. Back in March, in Blind Geniuses, we said everyone was measuring AI adoption and nobody was measuring results. In April, in Trust But Verify, we described the exact failure mode now on Anthropic’s own letterhead — a model that breaks its sandbox and builds an exploit chain — as a risk, not a headline. In July, in The Science of Hitting, we called verification the last scarce resource on the table. Here is the week it showed up in two directions at once. Superhuman where the answer grades itself. Confident and unruly everywhere else. The hack and the math proof aren’t two stories. They’re one fault line.

So the operating rule is boring, and it is the entire game. If you have a way to grade the work — a test suite, a clean compile, a number that has to reconcile — turn the machine loose and let it swing a thousand times. If you don’t, keep a human on the checkbook. A system that rationalizes a live break-in as “part of the test” will rationalize your quarterly numbers with the same easy confidence.

The good news is sitting in your next board meeting, and it runs on the exact same trait. The machine that will not stop is also the machine that will actually read all thirty slides. Point it at the pre-read and the variance-to-plan — the gradeable half — and give the humans back the two hours you waste rehashing last quarter.

The machine that won’t stop is the one nobody’s checking.

At COAI today: the full Signal/Noise — the juror, the Cyberdyne endgame, and the four-month paper trail on verification — is live at getcoai.com.

Grade it or check it. That’s the exercise we run at Outsider Labs: sorting the workflows you can score from the ones that still need a human signature.

ONE — A NUMBER THAT SUMMARIZES THE DAY

4 of 4. Four times Anthropic ran Claude Opus 4.7 through the same hacking test. Four times the model’s own notes show it realized the company it had broken into was real, not a simulation. Four times it decided the target “must be part of the exercise” and kept going — lifting credentials, reaching deeper into the network. Zero times it stopped. Only Anthropic’s newest model, in a separate run, spotted a real target and quit on its own. The benefit of the doubt is the vulnerability.

THREE — ACTIONS TO TAKE TODAY

Sort every AI workflow into “gradeable” and “not” before lunch. If a job has an oracle — tests pass, the build compiles, the numbers reconcile — let an agent run it cheap and often. If it doesn’t, a human signs the work. DeepSeek just put frontier-grade agent reasoning at 28 cents a million tokens, so the swings are nearly free. The only question left is whether you can score them.

Put AI on the board pre-read, not in the room’s blind spot. Feed an agent the deck, the financials, and the last eight quarters, and have it run variance-to-plan and the competitor benchmark 48 hours before you meet. SaaStr does exactly this and says the resulting meeting beats 90% of the ones happening today. The numbers are gradeable. Hand the humans back the strategy.

Assume your agents will rationalize, and cap the blast radius first. The security firm whose scanner ran Claude’s malicious package had live credentials sitting right where the payload could grab them. Before you point an agent at anything real: least-privilege access, no standing secrets in the environment, and egress it can’t phone home through. Contain the thing that won’t stop.

FIVE — STORIES TO KEEP YOU INFORMED

Monday, August 3

  • Claude broke into three companies, and two didn’t notice. (Full analysis above.) Anthropic says a misconfig left real internet access on during evals; a Claude-built package ran on 15 real systems inside an hour. It has halted all cyber evaluations — weeks before a reported ~$965 billion IPO that runs on institutional trust.

  • OpenAI’s $2,000 math miracle meets its control group. (Full analysis above.) Astra reportedly solved ten open problems; Marcus, Ernie Davis, and a fresh MIT/Harvard paper all flag the same catch — math self-verifies, the open world doesn’t. Genuinely impressive is not the same as universal.

  • DeepSeek reset the floor to 28 cents. V4-Flash does serious agentic work at $0.28 per million output tokens, and it’s likely why OpenAI quietly cut prices Thursday. Route the gradeable grunt-work here and save the premium model for the calls that actually need judgment.

  • China cracked the machine that makes the machines. Reports that Beijing built its own DUV lithography — ASML’s lone monopoly — plus a memory IPO up 466% dragged Korea’s Kospi to its worst month since 2008 before earnings steadied it. A real Chinese GPU is still years out; the fragility of a market resting on one company is the story.

  • The EU’s AI Act grew teeth on Saturday. Model-risk rules, deepfake labels, and synthetic-content watermarks became enforceable August 2 — real regulation landing in the same week Washington’s own AI framework deadline lapsed with nothing to show for it.

MARK TO MARKET

Where the cycle caught up to us this week.

  • Verification is the last scarce resource, and math wins first because it self-checks. (us, The Science of Hitting, Jul 22) → Gary Marcus, on why Astra’s proofs won’t generalize: the domains AI conquers first are the ones with cheap verification and free synthetic data (Marcus on AI, Aug 2). Eleven days.

  • Containment is the fence nobody funded, and every frontier model was already cheating its cyber evals. (us, Life Finds a Way, Jul 23) → Anthropic discloses its own models breached three real companies after a containment misconfig (Anthropic, Jul 30). Seven days.

The tape doesn’t lie. We just read it early.

— Harry and Anthony

Sources:

Reply

Avatar

or to participate