|
Hey there 👋
I lined up this week’s biggest agent stories on one page, and the same shape kept showing up. Nobody won by shipping a smarter brain. The wins went to the body around the model: a runtime to live in, a phone line to dial, a job to hold down, and a revenue line one agent already prints. Microsoft rebuilt Copilot around agents and put the agent work on a meter. Cognition’s Devin cleared a billion in run-rate. Warp handed an agent a chunk of the HR department. Bland will rent your agent a phone number for thirty bucks a month.
My bias, stated plainly: if you are still only shopping for a better model, you are shopping in last quarter’s store. The harness is the product now, and this week proved it four times. Below are the four stories that made the case.
The Big Thing
Microsoft put its agents on a meter
Microsoft did something this week that reads like a launch and behaves like a roadmap. It rebuilt Copilot, the AI product roughly 430 million knowledge workers touch, into three surfaces. Home is the everyday chat plus a new shared workspace called Cowork. Code is a natural-language builder running on GitHub Copilot tech, wrapped in what Microsoft calls a Copilot Managed Runtime. Autopilot is an always-on cloud agent you summon by @mention inside Teams, Outlook, and documents, and it keeps working after you close the tab.
Notice what is not on that list: a new model. The restructure is about the runtime and the surfaces where agents live, and that sets up the second half of the announcement, which is where the money is.
Billing now splits in two. Your fixed per-seat license still covers everyday chat and Copilot inside Word, Excel, PowerPoint, Outlook, and Teams, with an automatic router picking the model for each request. Everything agentic moves to consumption billing: Cowork, Code, Autopilot, long-running tasks, and the premium frontier models. Translation for your finance team: agent work becomes a variable cost line somebody has to own, forecast, and cap. Admins can set spend limits, which tells you Microsoft expects the meter to run hot.
Before you rip out a workflow, read the calendar. Autopilot is in private preview as of late September. Home and Code start rolling out in the coming weeks through the Frontier program, and Code lands in Microsoft 365 Premium and Pro later this year. Almost nothing here is generally available today, and no per-unit pricing was disclosed, so any cost model you build now is a sketch drawn in pencil.
Nadella said the enterprise agent market could be orders of magnitude bigger than cloud. That is a CEO sizing his own opportunity, so treat it as ambition and grade the multiple yourself before you plan around it.
What you can do this week: get into the Autopilot preview, wire it to one real Teams or Outlook flow, and watch the meter under live load while the volumes are small enough to make cheap lessons.
Ship it? PILOT. Spin up Autopilot in private preview and model the metered cost now. Hold production workflows until GA and pricing land.
Sources: Microsoft blog, Sep 25 2026, plus the-decoder and IT-Connect. Preview features, not GA; no per-unit pricing disclosed.
Tour de Headlines
💻 Devin crossed a $1 billion annualized run-rate this week, Cognition said, up from about $492 million in May. That is roughly double in four months, one of the steeper climbs anyone has put on paper. Read the number correctly: it is a company-reported run-rate, a recent month multiplied by twelve, and nobody outside the company has audited it, so treat the slope as the signal and grade the exact figure yourself. Customers named include GE Aerospace, Rivian, NVIDIA chip design, Citi, and Mercedes-Benz. Devin also added Auto-Triage for incident investigation, Security Swarm for vuln detection, and event-triggered runs off Slack, GitHub, and Linear. Coding is still the one agent job with narrow, verifiable, high-wage output you can check line by line.
🧩 Warp Agent shipped this week inside Warp 2.0, aimed straight at HR ops: onboarding, tax compliance, benefits admin, and employee questions, run as recurring or event-driven routines. The lever is programmability. A non-engineer uploads an internal doc and turns it into a working routine in plain English. Warp says first-month new-hire management falls from ten to twenty hours to under two, though that number is the vendor’s and its customers’, with no outside audit. Pricing starts at $35 per person per month; the $85 million you may have seen is cumulative funding to date, and there was no fresh raise this week. The founder calls it what comes after Workday. Is a vertical AI employee more defensible than a horizontal platform? Pilot it on one team and find out.
📞 Bland will now sell your agent its own phone plan: $29.99 a month for a dedicated US local number, with the agent placing and answering calls and texts on that single line. The delight is real. Your agent gets a working number and can call the dentist, chase a refund, or confirm a delivery while you get on with something else. The word unlimited is doing a lot, though. Fair-use caps sit under it: 60 calls an hour, 500 a day, 60 minutes a call, one at a time, US and Canada only. The number also stays on your account after you cancel. What nobody prices at $29.99 is the tail: call-recording consent and TCPA exposure once a persistent number starts transacting by voice. Verdict: watch.
|
Sponsor
Your team’s calls are full of signal. Measure it.
RapportScore reads the calls your team already records and scores what actually happened: talk ratio, questions asked, and trust signals. Deterministic measurement, not vibes. See where every rep stands and what to coach next.
See your team’s score →
|
Tool of the Day
🧰 Docker Cloud Sandboxes + Sandbox Kits
A safe place to run agents unattended, and a portable way to package one.
Docker made Cloud Sandboxes generally available this week, extending its local agent sandbox into the cloud. The point is simple: your agent keeps running after you close the laptop, using the same CLI and the same policies you already set locally. Isolation is microVM-level, matching the local sandbox exactly, boot times land in the low hundreds of milliseconds, and you get 1 to 16 vCPU with no infrastructure to provision.
Docker also introduced Sandbox Kits, an open OCI-image format that bundles an agent, its tools, and its access-control rules into one portable artifact you can hand to a teammate or a server. The spec is already up on Docker Hub and GitHub, and Docker has committed to submit it to the CNCF for open governance. Two caveats worth holding: pricing was not disclosed, and the Kits spec is brand new, with the CNCF handoff still a promise rather than a done deal. For running agents unattended without handing one your whole machine, this earns a real test on a low-stakes workflow first. Verdict: pilot.
Read Docker’s writeup →
Worth a Click
- Terminal-Bench-Science 0.1. A new independent agentic-science benchmark where the best models still miss about a third of real end-to-end computational workflows (GPT-6 Astra 63.3%, Claude Opus 5.5 61.9%).
- Exa Agent Ultra. A subagent-swarm deep-research API built for exhaustive, sourced list-building like market maps, KYC, and lead matching, aimed at RevOps and GTM teams that want a complete list rather than a best guess.
- ElevenLabs Image and Video APIs. Async generation with webhook callbacks now sits beside the voice stack, so you can script a voice, image, and video pipeline against one vendor.
Agent of the day: the harness. Look at the whole slate and the pattern is hard to miss. Nobody sold a smarter brain this week. They sold the body around it: a runtime to live in from Docker, a phone line to dial out on from Bland, an HR job to hold down from Warp, and a revenue line one agent already prints for Cognition. The model was the whole story for two years, and this week it was the least interesting part of every announcement. A raw model is easy to swap. A working body wired into your calls, your code, and your onboarding is a lot harder to pull out, and that is where the moat now sits. My bet, with a steak dinner on it: a year from now you will pick agent products by what they can do while you sleep, and the model underneath will read like a spec line nobody bothers with.
— Ron
|