The Agent Stack mascot
The Agent Stack _
Daily B2B AI automation brief · Thursday, October 1, 2026 · Issue #111

Hey there 👋

I spend most mornings reading launch posts that promise to change everything, and most of them are a feature with a press release stapled on. So when a protocol quietly goes to 1.0 and the post reads like a changelog, I pay attention, because boring and frozen is exactly what you want from the thing your product depends on. I read the AG-UI 1.0 notes twice looking for the catch. The catch is that there is not one, and that is the news.

Here is my bias, up front. The agent wars spent two weeks on kill switches and credentials, and the quieter story is that the plumbing underneath all of it is starting to set. A frozen protocol. An access layer. A cost meter you can finally read. None of it trends on a Thursday, and all of it is what you will still be running next year. So today leans into the plumbing.


The Big Thing

The Agent Front End Just Got a Standard

On September 30, the AG-UI protocol shipped version 1.0, and with it the spec and its JSON schema are frozen. In plain terms, the wire format your agent uses to talk to its own user interface stopped moving. The team behind it, CopilotKit, put the promise in one line: the 1.0 spec will not change, so what you build on it today keeps working.

What makes this bigger than another model release is the gap it fills. You already have a protocol for an agent to call tools, which is MCP. You have one for agents to talk to each other, which is A2A. The missing piece was the layer between the agent and the human watching it work: how the thing streams its thinking, asks you to approve a step, shows a tool result, and reports what it spent. That is the layer AG-UI 1.0 pins down.

Version 1.0 adds subagent support, a way to carry metadata, multimodal tool results, human in the loop interrupts, and token usage reporting, all backward compatible with the 0.x versions. It ships with generated SDKs for TypeScript, Python, and .NET, so you are not hand rolling a client. CopilotKit says it is adopted by Google, Microsoft, Amazon, and Oracle, and supported across LangChain, Mastra, Anthropic’s Claude Managed Agents, Google ADK, the OpenAI Agents SDK, and more. Read the adopter list as the vendor’s own framing until each runtime proves it speaks AG-UI by default, which is the one thing worth checking before you tear out your own streaming code.

A frozen 1.0 is a bet against more protocol churn, in a year when a new agent standard seemed to land every month. If you ship agents into a real product, that bet is now yours to make too. You build on the frozen spec and stop maintaining your own, or you keep your custom layer and own it forever.

Ship it? DEPLOY, with one check. If you are building an agent a human watches or steers, AG-UI 1.0 is the cleanest front end standard to build on today, and it is stable enough to commit to. The check: confirm your framework emits it before you rip out what you have. The spec being frozen is worth more than any single feature inside it.

Sources: CopilotKit’s AG-UI 1.0 release post (2026-09-30); spec and SDKs at docs.ag-ui.com and the ag-ui GitHub. The adopter and framework list is CopilotKit’s own, so confirm your runtime before you commit.


Tour de Headlines

🔎 Contentstack shipped Canoe on October 1, and it answers a question a lot of B2B teams have started asking in private: when a buyer asks ChatGPT to shortlist vendors, what does it say about you? Canoe scans your site, works out your competitors, generates the questions a buyer would type at each stage, and runs them across the big assistants. The free tier checks ChatGPT and Gemini. The paid plans add Perplexity and Claude, benchmark you against up to ten named rivals, and rank the fixes. Growth is $279 a month on a limited launch price, and the free plan needs no card. One honest caveat: you are measuring a surface that shifts with every prompt and model update, so treat Canoe as a monitor you watch, not a dial you turn. Learning you are invisible to the model your buyer trusts is still worth the afternoon.

🔐 Noma Security launched agentic access control, aimed straight at the governance hole everyone deploying agents has been stepping around. It discovers the AI agents and MCP servers running across your enterprise, then governs what each one can reach, from discovery through to enforcement, using agent boundaries and scoped permissions instead of a shared human login. This is the beat we keep coming back to: an agent with its own reach needs its own permissions, and most teams handed theirs a borrowed account and hoped. If you have agents touching real systems through MCP, this is the control plane that decides which of them can open which door. Verdict: watch, and pull it into your evaluation the moment your agent count passes what one person can track in their head.

🧪 Momentic launched an agentic quality platform for a problem your coding agents just created. When agents write code faster than anyone can test it, the testing becomes the bottleneck. Momentic’s answer is scriptless, AI-driven tests that a non-engineer can author and that survive a UI redesign, instead of brittle scripts that break on the first change. The pitch is quality that scales with AI code output rather than falling behind it. It earns a look for any team that shipped an AI coding workflow and quietly let test coverage slide to catch up. The thing agents are worst at is noticing they broke something, so the layer that checks their work does not stay optional for long.


Sponsor

Your calls leave a signal. Read it.

RapportScore measures how your team communicates on real calls and coaches them on it. It scores what happened on the call, using human signals plus AI, so your people get better at the conversation itself, meeting after meeting.

See your team’s score →

Tool of the Day

🛠 NVIDIA NeMo Relay

An open agent runtime that logs what every model call and tool call costs, in standard telemetry you can read.

If your agent bill keeps climbing and nobody can say which step is eating it, NeMo Relay is the thing to try this week. It runs your agents and emits OpenTelemetry traces for each model call, tool call, retry, and token count, so a smarter harness that quietly makes five calls where you expected one shows up as a line on a chart instead of a surprise on an invoice. It is open source, it has a CLI, and it plugs into the stacks you already use. One caveat worth stating plainly: it surfaces the cost, and capping it is on you, so pair the trace with a budget you enforce. Instrument first. You cannot manage a number you have never seen.

Read the NeMo Relay docs →


Worth a Click

  • Gartner says 40 percent of agentic AI projects die by 2027: the forecast is more than 40 percent canceled by the end of 2027, on cost, unclear value, and thin controls. That is the mortality rate sitting behind every governance pitch in today’s issue. The projects die for the exact reasons these tools say they fix.
  • Cloudflare’s tollbooth for the agentic web: pay per crawl lets a site charge an AI crawler for access, returning an HTTP 402 and a price instead of free data. It is the clearest early shape of agent-to-agent commerce, where the thing your agent reads has its own meter running. Worth understanding before your agents start getting billed per request.
  • Where the agent money is going: one roundup tallies roughly $435 million across a dozen rounds for enterprise agent security and governance in five months, most of it into safety and oversight. Pair it with the Gartner number above and the market is telling on itself: everyone is funding the layer that keeps the failing projects from failing.

For two weeks the agent story was about control: kill switches, credentials, who approved the login. Today it is about plumbing, and plumbing is the part that lasts. A protocol froze, an access layer shipped, a cost meter went open source. None of it is loud, and all of it is what you will still be running when the scary headlines have moved on. Here is my bet, with money on it. The teams that win the next year are not the ones with the flashiest agent. They are the ones who standardized the boring layer early, the protocol, the permissions, and the bill, while everyone else was still arguing about the leash.

🍬 One for the road: a new open-source benchmark called Vending Bench put models in charge of running simulated vending machines, and the expensive model did not win. By the benchmark’s own tally, Claude Opus turned the best profit, a cheaper model came second, and one model managed to lose money. Its own summary line says it best: being good at answering questions does not automatically qualify you to sell snacks.

— Ron