The AI now approves its own steps unless you turn it off. ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏
The Agent Stack mascot
The Agent Stack _
Daily B2B AI automation brief · Friday, August 14, 2026 · Issue #63

Hey there 👋

I read Anthropic’s auto mode writeup twice looking for the word that would let me relax, something like “fully unattended.” It is not in there. What is in there, a few clicks down, is a short list of the actions Claude Code still stops to ask about, and that list is the whole story today.

Because for everyone who did not go read it, the asking stops now. As of today, on paid Claude Code, the agent reviews its own steps by default, and you are the one who has to opt out. Here is my bias, for what it is worth: I think flipping that default matters more than any model release this month, and most people running it will not feel the change until something ships that they never watched.

The rest of the issue rhymes with it. A bank that lets an agent spend up to a ceiling and blocks the rest, an independent scoreboard you can trust over a vendor’s, a browser stripped down to what an agent needs and nothing a person would. One idea running through all of them: set the rule once and let the thing run inside it.

So today comes down to who holds the approval button, and what still counts as a control once the machine holds it most of the time.

— Ron


The Big Thing

Claude Code auto mode is the default starting today

As of today, new Claude Code sessions on the paid Pro, Max, and Team plans start in auto mode. That means the agent moves through each step without stopping to ask you first. If you had set a different default, you may get a one-time prompt to switch. The feature itself is a couple weeks old. What changed today is the default, and that is the whole news: approval went from something you turned on to something you turn off.

The agent does not run blind. Anthropic carved out three cases where Claude Code still stops and asks: an action it judges irreversible, an action it judges destructive, or one aimed outside your environment. Inside those lines it keeps moving. So “fully unattended” overstates it. The honest description is “unattended until it hits one of those three tripwires,” and how reliably those tripwires fire is the thing you cannot know from a launch page.

Anthropic’s argument for the flip is a study it ran itself. Across 1,053 paid professional testers, it reports that default auto mode caught 89% of deliberately dangerous commands, against 13.6% caught by human reviewers. Call it a self-reported figure. No outside team has replicated it, and it is the load-bearing claim under the decision to make the machine the reviewer. Trust that the direction is real, and wait for someone outside Anthropic to confirm the number before you lean on it.

One line matters for anyone on a bigger contract. Claude Enterprise, the API, Amazon Bedrock, Google Cloud’s agent platform, and Microsoft Foundry stay opt-in for now. Anthropic says it plans to default them within about a month, so the same change is coming to the enterprise lane, just later.

Why it matters: this is a live change to how a huge base of agent sessions behaves, today, with no action from you. It rewrites the year’s human-in-the-loop advice in one move. The vendor is betting its model gates risk better than you do, and the work is now on you to turn that bet off if you disagree.

Ship it? If you run paid Claude Code, decide your carve-outs before your next unattended run, and opt out today if you are not ready. If you are on Enterprise or the API, you have about a month to write your policy before the default reaches you.

Sources: TechCrunch, Simon Willison, Help Net Security, InfoWorld


Tour de Headlines

🧹 Microsoft folded its two Copilot apps into one and started cutting. On Aug 13, Microsoft began rolling out a single Microsoft Copilot app that merges the consumer and Microsoft 365 versions and lets you flip between personal and work accounts in one place. The tidy-up has a cost: Group Chats and every shared message and image get wiped on Aug 18, with AI Podcasts and Deep Research also on the way out. Microsoft frames the merge as groundwork for a planned Copilot “super app.” If your team standardized on M365 Copilot, check what you lose this week before the delete lands. TechCrunch, The Register, GeekWire

💸 DeepSeek’s cheap-model story cracked on the same day it shipped. DeepSeek V4 Pro (build 0813) reached general availability on Aug 12 with a claim that it “greatly enhances agent capabilities,” and paired the GA with an API price increase that reporting puts as high as 12x, taking effect in mid-August. Independent scoring is the part to hold onto: Artificial Analysis puts V4 Pro at 53 on its Intelligence Index, a composite of nine benchmarks. DeepSeek reports higher agent scores off a closed harness nobody outside the company can rerun, so treat them as marketing until that changes. For teams leaning on DeepSeek as the cheap self-host hedge, the hedge just got pricier right as the proof got thinner. TechTimes, Caixin, Artificial Analysis, OpenRouter

💳 Mercury gave AI agents corporate cards that decline at checkout when they go over budget. On Aug 11, Mercury launched Spend, which lets a business issue cards to both employees and agents with an intelligent budget attached. Go past the ceiling and the charge is declined at the point of sale. The last agent-commerce wave built the payment rails. Mercury Spend builds the limit that sits on top of them, which is the piece an ops or finance lead has been waiting for. Agent cards can only be created by a human, and every charge is trackable, auditable, and cancellable. If you have been stuck on how to let an agent buy things without handing over the company card, this is a deploy-today answer. Mercury, PYMNTS, Fast Company


A word from RapportScore

Today everything got a grader. Who is grading how your reps show up on the call?

The agent grades each step it takes, an independent index grades the models, a budget grades the spend. One grade still runs fully on humans: how your reps show up on a call. RapportScore measures how your people communicate in their own recorded conversations and coaches them on it. It scores behavior in the call, and it is honest about its edges: it does not claim to read intent or honesty, and it tells you when the signal is thin instead of guessing to fill the gap. If you grade your agents but never grade your reps, that is the blind spot worth closing this week.

See how it works →

Tool of the Day

🪁 Cloudflare Kitesurf: a browser built for agents, not people

What it’s for: giving an AI agent a real browser to drive without paying the full Chromium tax.

This one shipped back on Aug 6, so file it as last week’s tool rather than breaking news. It still earns the slot. Kitesurf is a cloud-hosted browser Cloudflare built for automation instead of for humans. It drops tabs, extensions, themes, and pixel-perfect 60fps rendering, and it runs on Cloudflare Workers. Cloudflare says it uses roughly 3 to 7 times less CPU and memory than Chromium on common automation work. Cloudflare measured that itself, so clock it against your own workload before you bank the savings. It is free in beta and slated to be open-sourced. Cloudflare has shipped plenty of agent plumbing lately, so keep the frame narrow here: this is a knob on a browser-automation bill you can turn today, before you commit to anything bigger. If your agents are burning compute driving headless Chromium, this is a cost lever you can test this afternoon without re-architecting anything.

Sources: Cloudflare, TechCrunch

Try Kitesurf


Worth a Click

  • 🧩 The Qwen you can download differs from the Qwen they sell. Alibaba dropped Qwen3.8-Max open weights (a 2.4T MoE, roughly 95B active) on Hugging Face, and it is a real self-host option for agentic work. Read the fine print first: the download is reported text-only, without the API’s 1M-token context, under a new revenue-share license, and the promised 27B variant has not shown up yet. Confirm the caveats against the license file before you build on it. Hugging Face, Latent Space
  • 📊 Grade this week’s model claims yourself. Artificial Analysis runs an independent Intelligence Index that sits next to the vendor harnesses. It pegs DeepSeek V4 Pro 0813 at 53, a useful counterweight when a press release quotes a number nobody outside the company can reproduce. The gap between that 53 and the headline vendors quote is the spread worth arguing over before you wire a model into an agent loop. Bookmark it and check the composite before you trust a headline score. Artificial Analysis
  • 🃏 Clippy is dead, again: Microsoft is retiring Mico, its floating-blob Copilot mascot, in the same cleanup, one more sign the cute avatar gets deprecated while the plumbing underneath keeps shipping. TechCrunch

The Bottom Take

The quiet through-line today is the default. For a year, letting an agent act unsupervised was a setting you chose to turn on. As of today, on paid Claude Code, it ships on by default, and turning it off is the deliberate move. The controls that survive that flip are not prompts, they are policies: a carve-out list for the irreversible stuff, a budget that declines the overspend at checkout, an independent score you trust over the vendor’s harness. So the move is to write your exceptions up front, because the default just decided everything else for you.

— Ron

You’re receiving this because you subscribed to The Agent Stack. · Unsubscribe