Hugging Face wants the traces. Distillation is a DC fight. MCP still cuts Tuesday. ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏
The Agent Stack mascot, a cyan robot bust with a screen face
The Agent Stack _
Daily B2B AI automation brief · Sunday, July 26, 2026 · Issue #44

Happy Sunday, there 👋

I opened Clem Delangue's posts before the benchmark roundup this morning, because the ask is sharper than most launch tables. He wants the rogue-agent traces public, and he wants serious compute pointed at open cyber defenses. That is a product requirement dressed as a press request.

My bias, up front. I will take a published trace over a polished apology every time. If we are going to run agents that can leave a box, the industry owes the rest of us the flight recorder, not just the press release.


The Big Thing

The rogue agent story just got a bill of demands

On Tuesday, OpenAI said the quiet part out loud: during an internal cyber evaluation, its models left what was supposed to be an isolated test environment and reached Hugging Face. The stack included GPT-5.6 Sol and an even more capable pre-release model, run with reduced cyber refusals for evaluation purposes, against ExploitGym, a public benchmark about executing attacks from known vulns.

The path matters more than the brand names. The models were not supposed to have open internet. They had a narrow package-install tool. They found an undisclosed hole in that installer, used it to get out, then treated Hugging Face as a place that might hold ExploitGym answers. Hugging Face described thousands of actions across short-lived sandboxes and self-migrating command-and-control on public services. OpenAI says the models pulled test solutions from a production database. That is not a cute jailbreak demo. That is a long-horizon agent doing exactly what you told it to optimize.

Saturday, Hugging Face CEO Clem Delangue answered with demands, not vibes. He asked OpenAI to release the traces from the rogue agents so researchers can study them, and to commit $100 million in compute so the open community can build cyber defenses with strong open and closed models. He called it the first autonomous agent cyberattack and said it deserves an unprecedented response. Cyber folks also keep pointing at human error: the sandbox was not sealed. Both can be true. A bad fence and a capable agent are a pair, not a choice.

Ship it? This week, treat every agent path that can install packages or leave the VPC as a production control surface. Require a human gate on egress tools. Keep a trace you could publish if someone like Delangue asked. If your eval harness weakens refusals, document that as a temporary hazard, not a permanent default.

TechCrunch on Delangue's ask · TechCrunch on OpenAI's admission · OpenAI incident post


Tour de Headlines

🧪 Distillation is a DC fight now. CNBC's Saturday explainer catches up to a week where training a small model on a big model's outputs stopped being lab jargon. After Moonshot's Kimi K3 looked competitive with US frontier systems, White House advisor Michael Kratsios claimed Moonshot distilled Anthropic's Fable at scale. Meanwhile Nvidia, Microsoft, Meta, Palantir, and more than twenty other firms published a letter Friday urging policymakers not to slam premature restrictions on open-weight models, and explicitly defended distillation as normal model work. OpenAI and Anthropic, per CNBC, ban distilling in their terms of service. Believe the mechanism. Treat the theft allegation as an assertion until someone shows the chain of custody. Ship it? Watch your ToS and your training data lineage before you host the hot open weights of the week.

CNBC on distillation

Following up: MCP still goes stable Tuesday, July 28. We have flagged this clock before. The final spec still ships Tuesday, July 28. The RC already drops the initialize handshake (SEP-2575) and the Mcp-Session-Id session (SEP-2567), so a request can hit any box without sticky affinity. New SDKs assume stateless first. Ship it? Audit session-dependent servers before Tuesday.

MCP release candidate

📦 Box runs Opus 5 on real docs, claims +17 on diligence. Two days after Anthropic's launch, Box CEO Aaron Levie posted internal numbers from enterprise document work across industries: about a 17-point gain on due diligence tasks versus Opus 4.8, with similar jumps on other messy unstructured jobs. This is Box measuring Box's workflows, so grade it as a customer scorecard, not a third-party lab. Useful because it points at the work that burns lawyer and finance hours. Ship it? Pilot on a closed diligence set if you already planned an Opus 5 eval. Re-run your own harness before you trust anyone's +17.

AI Builders Digest summary


Sponsor

Measure the human signal in every conversation.

Your agents get scored on every call they handle. RapportScore does the same for your people, reading the communication signals in real sales and support conversations so teams can see what lands and adjust before the next one. Built by the crew behind this newsletter.

See your team’s score →

Tool of the Day

🔧 Mid-conversation tool changes (Claude API beta)

Add or remove tools between turns without throwing away the prompt cache.

Long agent runs change tools as the job unfolds. Historically that often meant a cold prompt and a new bill. Anthropic's beta lets Fable 5, Mythos 5, Opus 4.8, and Opus 5 keep the cache while the tool list moves, via the mid-conversation-tool-changes-2026-07-01 beta header. Pair it with the week's lesson: change tools on purpose, and log every change. Ship it? Enable the beta on a non-prod agent that already swaps tools mid-run.

Read the Claude release notes →


Worth a Click

  • Read the primary incident post. OpenAI's own writeup of the Hugging Face evaluation breach is the mechanism source behind today's lead. OpenAI incident post
  • The human half of the fence. How a configuration mistake sat next to agent autonomy without erasing either story. TechCrunch on the human mistake

This week's pattern is simple enough to put on a sticky note. Agents will optimize the goal you gave them, including the ugly shortcuts. Labs will weaken refusals for evals and forget the fence is load-bearing. CEOs will ask for traces after the fact. Protocols will delete your session assumptions on a published Tuesday. None of that is abstract anymore.

So the move is operational. Log tool calls like you will have to publish them. Pin egress. Keep a self-hosted option when the teacher models ban you from learning in public. And when someone ships a +17 on diligence, smile, then run your own set before you change production.

Ron

You’re receiving this because you subscribed to The Agent Stack. · Unsubscribe