Salesforce, Google, and Cognition all shipped agents you switch on this week. ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ 
The Agent Stack mascot
The Agent Stack _
Daily B2B AI automation brief · Saturday, September 12, 2026 · Issue #92

Hey there 👋

For most of the last month I read agent news that was really security news. Firewall the context. Lock the endpoint. Give the agent its own identity. Good stuff, and I covered all of it. But it shared one quiet assumption: the agent is a risk you wrap in controls before you dare turn it loose.

This week the tone flipped. Three big vendors spent it selling the agent itself, named, scoped to a job, and priced like a seat. Salesforce handed enterprises seven of them. Google put one on the Windows desktop. Cognition shipped one that writes code at a fraction of the going rate.

My read, for what it’s worth: the interesting question moved from “can I trust an agent in production” to “which one do I hire, and what does it cost to keep it working.” Read on and you will see the same shift three different ways.


The Big Thing

Salesforce shipped seven agents you can hire this week

Salesforce announced seven named, job-ready Agentforce agents on September 11, and six of them are generally available right now. Casey works service, Paige covers IT and HR help desks, Carter runs commerce, Marshall handles supply chain, Piper is an inbound sales rep, and Fin owns customer experience. The seventh, Hunter, an outbound seller, is in pilot with general availability planned for November.

Look past the mascots. What changed is the runtime beneath them: persistent memory, durable execution, and mid-run steering, so an agent can carry a goal across days and hand-offs instead of forgetting everything after one chat. That is Salesforce admitting the old scoreboard, single-turn resolution rate, measured the wrong thing. Work that matters rarely finishes in one turn.

Salesforce also says it has delivered 7 billion “agentic work units” across Agentforce and Slack, 3.2 billion of those last quarter, and it points to customers like Anthropic resolving 79% of inbound with Fin and Hibbett handling roughly 90% of core shopper journeys. Those are Salesforce’s own figures and customer-reported wins, so weigh them as vendor numbers rather than audited ones. It open-sourced Agent Script, a small language for mixing hard rules with model reasoning, which is the piece a builder can pick apart today. Timing is not subtle either: this dropped four days before Dreamforce.

Ship it? Pilot now. If you already run Agentforce, switch on Casey or Paige and route one live channel through it this week, since the skills, actions, and data wiring ship pre-built. Watch Hunter until the November GA before you trust an agent with outbound.

Source: Salesforce Newsroom (salesforce.com/news), September 11, 2026.


Tour de Headlines

🪟 Gemini landed on Windows. Google shipped the Gemini app for Windows on September 10, generally available worldwide on Windows 10 and 11, x64 and ARM64. Hit Alt+Space over whatever you are doing and Gemini opens on top of it and can draft from your Gmail and Drive. The system-wide hotkey is the real story here, because it puts a grounded assistant one keystroke from every Windows workflow, on the OS most B2B desks already run. Set expectations, though. This is a version one that drafts and fetches but cannot drive your other apps yet, and Spark, Google’s real do-things agent, did not ship on Windows at launch. Spark still runs only in the mobile, Mac, and web apps, needs a paid Google AI plan, and is blocked in the EEA. So Windows gets a fast assistant on a hotkey today, and the real agent handoff is a promise for a later build.

💻 Devin’s new model codes at a quarter of the price. Cognition shipped SWE-2 on September 10, live now in Devin Desktop and the CLI, and rolled its Fusion mode (a frontier model plans, a cheap model executes) into the same tools. The pitch is cost: 50% on its FrontierCode benchmark, which it says lands within a few points of GPT-6 Astra at about a quarter of the cost, and 64% cheaper than Fable 5.1. The clever bit is the bet behind it. SWE-2 is post-trained on Moonshot’s Kimi K3 rather than a house base model, so Cognition is wagering that the harness and the training recipe matter more than owning the weights. FrontierCode is Cognition’s own benchmark and the cost math is internal, so grade it as a vendor claim. But the cost-per-task number is the one that hits your invoice, and this is the first coding agent to lead with it.

📋 Everyone trusts their agents; far fewer can prove it. Harness published its State of Agent DLC 2026 report, a survey of 700 engineering leaders at large enterprises already running agents in production. The headline gap: leaders are confident they have a full inventory of every agent, MCP server, and model in their stack, but far fewer run any active discovery to check whether that inventory is real. Read it as a mirror rather than a scare piece. Agents are non-deterministic, and most teams still govern them with pipelines built for software that does the same thing every time. The deploy-this-week move costs nothing: run one discovery pass across your environment and diff what you find against what you assumed was there. The gap is usually the story.


Sponsor

Your team’s calls are full of signal you never grade.

RapportScore reads your recorded calls and measures how your people communicate, then coaches them on it. Real measurement, not vibes.

See your team’s score →

Tool of the Day

🧪 Google’s agent sandboxes went GA

Run an agent’s shell and browser actions somewhere they cannot hurt you.

Google moved Computer Use and Shell sandboxes in its Gemini Enterprise Agent Platform to general availability. What they are for: giving an agent a real place to do dangerous things safely. A Shell sandbox runs untrusted commands, installs packages, and edits files inside an isolated Linux container you reach over a /exec API call. A Computer Use sandbox hands the agent a throwaway browser to click and navigate. The part worth your attention is egress control through VPC Service Controls and Private Service Connect, because the real production blocker is rarely raw capability; it is egress, where the agent is allowed to phone home. Stand one up, point it at a real internal task, and pin its egress to a single private range before you widen it.


Worth a Click

  • Salesforce’s own announcement of the seven agents. The full list, the long-horizon runtime, and the customer numbers, straight from the source so you can grade the claims. salesforce.com
  • The Cognition SWE-2 write-up. The FrontierCode scores and the 64% cost claim in one place. Read it with the vendor-benchmark caveat in hand. techtimes.com
  • Microsoft Agent Framework .NET 1.21.0 release. If you build on the framework, this one ships breaking changes to A2A, MCP archives, and file access. Read the migration notes before you bump the version. github.com

The week the agent became a SKU. For a month the vendors sold ways to fence an agent in; this week they put the agent on the price list, with a name, a job description, and a seat cost. Salesforce sells you seven of them, Google puts one on every Windows desk, Cognition sells one that codes cheaper than the last one. The counterweight arrived the same week, from Harness: the teams buying all these agents are more sure they have them under control than their own tooling can back up. So here is the honest place we landed. Hiring an agent is now a purchase order, not a research project, and the hard part is no longer whether it works. It is what the thing costs to run, and whether you can even find it next quarter.

See you tomorrow.
— Ron

You’re receiving this because you subscribed to The Agent Stack. · Unsubscribe