|
Hey there 👋
Yesterday the story was the agent bill coming due. Today the bill got smaller.
The price of agentic work — planning, driving a browser, checking its own output, not just chatting — stepped down this week. And the plumbing under multi-agent systems standardized in the same 48 hours. That is a different kind of news than “spend less.” It means the pilot you shelved last quarter for failing a cost test might pencil out today.
So here is the one move for the whole issue: re-run the math. The substrate got cheaper. Let’s see what that unlocks.
The Big Thing
The price of agentic work just stepped down
Point your agents at a cheaper model, then re-run cost-per-task on the workflows you killed last quarter. That is the move. Anthropic shipped Claude Sonnet 5 on June 30, and the headline is not a benchmark — it is a category shift.
Here is why this matters beyond one lab. The expensive part of an agent is not the chat. It is the loop: plan, call a tool, read the result, self-check, try again. A single task can burn ten or twenty of those turns. When that loop gets cheaper, agents that failed a budget test suddenly clear it. The lever moving a pilot to always-on is falling cost of autonomy, not a higher score. Read this as a category signal, not a Claude ad — every lab now competes on the price of the loop.
The numbers, and who is stating them. Anthropic says intro pricing runs $2 per million input tokens and $10 per million output through August 31, then $3/$15. Anthropic also reports 63.2% on its own agentic-coding benchmark, against 69.2% for the pricier Opus 4.8 and 58.1% for Sonnet 4.6 — vendor-reported, so treat it as a signal, not a verdict.
Now the honest catch. Sonnet 5 uses a new tokenizer that Anthropic says emits roughly 1.0–1.35x more tokens for the same text. Anthropic states the intro price is set to be about cost-neutral. Translation: a lower sticker price does not guarantee a lower bill. More tokens per task can eat the discount.
So do not take the savings on faith. Run your own eval on your own workload before you move budget. Price the workflow end to end — tokens times loop depth times tool calls — not the per-token rate.
The deployable pattern is a two-tier route. Send the planning, tool-use, and self-check work to the cheaper agentic tier. Reserve the priciest model for the few hardest calls where a wrong answer costs real money. Then re-run cost-per-task on the agents you shelved as too expensive last quarter — a lot of them were one price cut away from clearing the bar.
Ship it? Deploy now — for agentic and tool-use workloads where you can measure cost-per-task. Keep the top-tier model on the hardest steps, and prove the savings on one real workflow before you scale.
Read Anthropic’s announcement.
Tour de Headlines
💸 Run your agents on open models — and check the bill
The infra half of cheaper autonomy showed up too. Together AI raised $800M at an $8.3B post-money valuation, led by Aramco Ventures, the company and reports say. Ignore the number for a second and look at the customers: Cursor, Cognition, and Decagon — all agent builders. That is the tell. The deployable pattern is running agent workloads on open models like DeepSeek, MiniMax, and Kimi for less than closed frontier APIs, while keeping data in your region. The company says bookings crossed $1.15B last quarter and it has committed 500-plus megawatts for a 50x capacity push. Builder read: benchmark an open-model backend against your closed-API bill on one real agent workflow before you assume it wins. More at TechCrunch.
🎥 Turn a video archive into something an agent can act on
TwelveLabs raised a $100M Series B co-led by NEA and NAVER Ventures, with Amazon, Radical, and Index participating — total raised now past $207M, per the company and reports. The pattern worth stealing: a vertical agent that reasons over video, not just tags it, with what the company calls compounding memory that carries across queries. If you sit on unstructured video — recorded calls, field footage, security feeds, training libraries — the deployable question is narrow: what decision could an agent make from it on a schedule? Start with one job you can verify by hand, prove it, then widen. Small vertical, repeatable template — that is where video finally earns its keep. Details at SiliconAngle.
🔌 Wire new integrations to a two-layer pattern that will outlast the churn
The standardization half of the story: MCP (for tools) and A2A (for agent-to-agent), both now Linux Foundation projects, are reportedly moving toward a joint interoperability spec targeted for Q3. Say it plainly — this is targeted, not shipped, and the figures below come from cited analyses, not primary filings. Those analyses put MCP at 18,000-plus community servers. The architecture decision for builders is the takeaway: design integrations to the two-layer stack — MCP for tools, A2A between agents — so today’s wiring survives the convergence instead of getting torn out. A bet on the default carries less lock-in risk than a bet on a single vendor’s glue. Ecosystem map here.
|
Sponsor
The price of running your agents just dropped. The price of your people misunderstanding each other did not.
You can rent cheaper tokens. You cannot rent clarity. RapportScore measures the one thing you can’t buy on sale — how your humans actually communicate on calls, in email, across every deal thread. You still can’t manage what you don’t measure.
See your team’s score →
|
Tool of the Day
🧠 Couchbase AI Data Plane (GA)
What it’s for: giving your agents a durable place to remember and retrieve context — framework-agnostic, with a self-managed MCP server for tool access.
Couchbase says its AI Data Plane went GA on June 30, bundling Agent Memory, an Agent Catalog, and a self-managed MCP server, and validated with LangGraph, CrewAI, and LlamaIndex. That last part is the point. The framework-agnostic bit is the value: pick your memory layer before you marry an orchestration framework, so a framework swap later does not cost you the agent’s memory. It is a concrete answer to “where does agent state live in prod?” Capabilities here are Couchbase’s own description, so try it on one agent before you standardize.
See the launch →
Worth a Click
- The multi-agent hand-off, in production. Klaviyo says its Composer entered public beta, with agents that coordinate to spot an opportunity, build a campaign, and run it together. It skews consumer-brand, but the orchestration pattern — packaged skills, agents passing work between each other — is the part worth stealing (all lift claims are Klaviyo’s own). How it works.
- Two smaller Anthropic launches. Claude Science, a workbench Anthropic says produces auditable artifacts — the auditable framing matters for regulated agent work — plus Fable 5’s global redeploy (July 1), which Anthropic pairs with an industry jailbreak-severity scoring framework as a safety-standards signal. Read up.
Q2 made agents a line item. This week the meter reading dropped — a cheaper agentic model, an open-model inference layer, interop protocols converging, a memory plane you can adopt today. The substrate got cheaper and more standardized at the same time. One move: take the agent you killed last quarter for costing too much, and re-run the math today. Falling cost of autonomy plus standardizing plumbing is what turns a shelved pilot into a line of business. The lever is right there. Pull it.
See you tomorrow. — The Agent Stack Built for people who ship AI, not people who tweet about it. Published weekday mornings by Pixiu Media Holdings LLC.
|