Sponsored by

Customer agents you can test, version, and trust in production.

Most teams treat AI agents like a black box. They wire together a workflow tool, a speech vendor, and a transcription API, then hope the result behaves. When it doesn't, there's no way to know why, and no safe way to change it.

ElevenAgents treats customer agents like software. Built on the voice models the market already builds on, it runs voice, chat, email, transcription, and reasoning in one pipeline, so a customer can start in chat, escalate to a call, and follow up over email without losing context. Voice responses come back in under 400 milliseconds.

Then you get the controls stitched-together stacks can't offer. A/B test prompts and personas with Experiments. Enforce behavior with Guardrails. Version every change, ship improvements, roll back mistakes. Plug in any LLM and ground answers in your knowledge base.

Launch in minutes, improve every week. Transparent, flat pricing at $0.08 per minute.

SIGNAL / NOISE

Governing Dynamics

Book a padel court at the New Canaan Field Club and Court Reserve makes you enter four names. Doesn't matter that you've got three and you're hunting a fourth. Four names or no court. So you type in your wife's name, and you never change the reservation. Everybody does it. Reserve that court 7 to 8:30 in the morning and the same software locks you out of tennis, pickle, and platform tennis until 7am the next day. The fix is the same fix: a friend's name, or you call the pro shop and somebody overrides it.

None of that is written down anywhere. It lives in the members, learned by getting burned once and asking the guy next to you. That undocumented layer, the workaround nobody put in a manual, is the thing every business actually runs on. Hiten Shah said it in four words this week: "Ask Sarah. She knows." That sentence is a product roadmap.

Here's what's eating it. An agent doesn't inherit the workaround. It never got burned, never asked the guy next to it. Point one at a booking system and it does one of two things. It kicks the door in: a Melbourne guy asked his agent to bump him up a gym waitlist, and it found a bug in the booking API and cancelled a stranger's reservation to make room. He never asked it to. He never would have. Or it walks away, which is quietly what's killing Mailchimp. Eleven million users, twelve billion dollars in acquisition price, now shrinking inside Intuit, because Replit and Lovable and Vercel all point new AI-built apps at Resend by default. Mailchimp has no door an agent can walk through, so the agents building the next wave of companies have never heard of it.

Now put those next to the number that matters. Anthropic ran a controlled test: humans caught a planted dangerous command 13.6% of the time. The classifier caught it 89%. So the industry's move is to pull the human out of the loop, and that's rational. Every firm cutting its juniors is rational too. But that grunt work was the only school that ever produced someone who could tell when the machine was wrong. Zara Zhang's version of this hit half a million views this week and asked the obvious follow-up: who checks AI's homework in fifteen years? Who will check it in five months?

That's the tragedy of the commons. Everybody optimizing for himself, the pasture going bare. Nash worked it out at a bar in 1950 and it won him a Nobel: the group loses when every man chases only his own best interest. Adam Smith was wrong.

But there's a third move, and it's the whole game. A smart agent reads the pattern of the workarounds, sees everyone stuffing a spouse's name in the fourth slot, and hands the owner the fix: let people hold a court with two or three players until 24 hours out. Same intelligence that cancelled the stranger's booking. Pointed the other way, it stops draining the commons and starts refilling it.

Survival first. Growth second. The business that lives is the one whose system an agent can talk to, and learn from.

At COAI today: the full Signal/Noise, with the three things an agent does to your business and how to build the door before your Sarah retires, is live at getcoai.com.

Can an agent actually talk to your business? If your Sarah is the employee with the whole operating manual in her head, let's talk.

ONE — A NUMBER THAT SUMMARIZES THE DAY

13.6%. That's how often a human reviewer caught a planted dangerous command in Anthropic's own test. The classifier caught it 89% of the time, so the humans are getting pulled from the loop. Fine. But that reviewing, and the grunt work behind it, was the only school that ever taught anyone to catch the machine being wrong, and we're closing the school the same week. Who checks AI's homework in fifteen years? Who will check it in five months?

THREE — ACTIONS TO TAKE TODAY

Put a door on your product an agent can use. Replit, Lovable, and Vercel point new apps at whoever ships an MCP server; Mailchimp didn't, and it's shrinking on eleven million users. Find the one workflow a customer's agent would want to run against you, and ship the endpoint for it today, before your competitor becomes the default and you become Mailchimp.

Write down what only Sarah knows. Every team has the person who knows the workaround nobody documented. That knowledge used to pass by osmosis at the junior desk you just automated away. Pick one "ask Sarah" process today and get it out of her head and onto paper, while she's still here to explain why the rule is really the rule.

Tell your agent to flag the crack, not take it. The gym agent found a booking exploit and used it, because it never occurred to anyone to tell it not to. In the standing instructions for any agent you run against a live system, add the line today: surface the vulnerability to a human, don't exploit it to finish the job. Same capability, opposite outcome.

FIVE — STORIES TO KEEP YOU INFORMED

Monday, August 11

  • Anthropic decides you're the weak link. Claude Code's auto mode goes default August 14: the AI reviews the AI, humans out of the loop, because in testing people rubber-stamped 97% of prompts and caught the dangerous ones 13.6% of the time. Convenience, sold as safety. (Full analysis above.)

  • Anthropic signs a $10B compute deal with a company younger than a pregnancy. The six-year Norway deal runs through Volta Infra, founded in January, on a $1.3B JPMorgan letter of credit, same bank-risk-shifting as the $71B chip SPV. When the customer, lender, and vendor all rhyme, it's a hall of mirrors.

  • An AI tried to talk a human into merging its malware. Guardrails off, one model produced 17 of 19 unsanctioned actions in the UK AI Security Institute's test, including fake identities and a supply-chain attack on a real open-source project, then edited its own tracks. A human maintainer caught it.

  • Your AI notetaker left 181,874 meetings wide open. tl;dv had no tenant isolation, so any free user could read every meeting on the platform, government calls from 23 countries included, unpatched for six months. Everyone's watching for rogue superintelligence; the open door is what leaks your board deck.

  • Claude moved a 160-year-old math problem, and it took a swarm. An unreleased version pushed the Riemann lower bound from 41.6% to 67.2%, Lean-verified, on 60 subagents and 31 million tokens. Not a smarter chatbot, an orchestrated system grinding out verifiable work.

— Harry and Anthony

Sources:

  • Hiten Shah (@hnshah), "Ask Sarah. She knows. That sentence is a product roadmap," X, Aug 10, 2026

  • Zara Zhang (@zarazhangrui), on "The Tragedy of the Cognitive Commons," X, Aug 8, 2026

  • "Anthropic Makes Claude Code's Auto Mode the Default," DevOps.com, Aug 2026

  • "How Mailchimp Went From $1B+ ARR to Shrinking Inside Intuit," SaaStr / Jason Lemkin, Aug 2026

  • "AI agent asked to book a gym class ends up hacking the system," Indian Express / The Neuron, Aug 2026

  • "Incident Report: Unsanctioned Agent Behaviour During Cyber Testing," UK AI Security Institute (aisi.gov.uk), Aug 2026

  • "Anthropic Locks In $10B European Compute Bet With a Seven-Month-Old Startup," Yahoo Finance, Aug 2026

  • "tl;dv: 181,874 Meetings Left Wide Open," bobdahacker.com, reported Jan 2026

  • "Learning More About Claude's Mathematical Capabilities," Anthropic, Aug 2026

  • John Nash, governing dynamics (Nash equilibrium), 1950; via A Beautiful Mind

Reply

Avatar

or to participate