The Agent Stack mascot
The Agent Stack _
Daily B2B AI automation brief · Friday, October 9, 2026 · Issue #119

Hey there 👋

I spent this morning reading past the headline on Google's big agent launch, waiting for the line that would tell me whether this was a real shift or a keynote. I found it buried in the setup details. The agent gets its own Workspace account. An email address. A calendar. A note on who its approvers are. You invite it into a group chat the way you would add a new hire, rather than opening it like an app.

That reframed the whole day for me. Most of what shipped this week is scaffolding, the kind you bolt on once an agent is allowed to act on your behalf: who it is, what it may touch, and whether you can reconstruct what it did afterward. So that is the issue. One agent that got hired, one free tool that hunts your bugs, one hole in the plumbing you should check before lunch, and the dull machinery that decides whether any of it survives an audit.

- Ron


The Big Thing

Google gave its AI agent a job and a corporate email

Google used its Gemini at Work 2026 event on October 8 to launch a universal agent inside Gemini Enterprise, and the framing is the news. Thomas Kurian, Google Cloud's CEO, said you give this thing objectives, not just instructions. It plans the work, spins up subagents, and reaches into your systems to get it done. Google is shipping it to businesses first and sending it to consumers later.

The part that moves it past a normal launch is the identity layer. The agent gets its own Workspace account, with an email address, a calendar, and context about your colleagues, teams, time zones, and approvers. You invoke it by tagging it, emailing it, or dropping it into a group chat. It keeps a tasks inbox that shows its reasoning, which subagent it handed work to, and the code it ran. When it acts, the log records the agent itself as the actor, with a person behind the trigger. That is an employee record for software.

The reach is broad. It connects to Google Workspace, Microsoft 365, Slack, Jira, Confluence, Git, BigQuery, Databricks, Postgres, and Snowflake, plus any MCP server inside or outside your network. By default it picks the best model per task, and you can override that and choose one yourself, starting with Anthropic's Claude. Open and private models join the picker later. Google also added spend caps and smart routing so the bill does not run away from you.

Treat the proof points as vendor framing until your own pilot says otherwise. Sundar Pichai said Gemini has more than a billion monthly users and that nearly 90% of the Fortune 100 use Gemini Enterprise. Early testers named include On, Shopify, and PayPal. There is no public price yet, and no independent benchmark of the agent doing real multistep work. Those are the two numbers that decide whether this is a product or a promise, and neither is on the table today.

Ship it? PILOT. Put it on one real workflow where you already trust Google with the data, and watch two things: whether the audit trail reconstructs what happened, and whether the model picker lets you route the expensive steps to the cheap model. The identity scaffolding is why it is worth a test run. Hold off on committing until Google publishes a price and someone benchmarks the agent on real multistep work.

Sources: TechCrunch; Quartz.


Tour de Headlines

🛡️ Anthropic will scan your open-source project for bugs, for free. Alongside a Critical Infrastructure Defense Program with eleven founding partners (Accenture, CrowdStrike, Dragos, Palo Alto Networks, and Rockwell Automation among them), Anthropic opened an OSS Scanner that any eligible open-source project can opt into. Maintainers enroll with a pull request to an Anthropic repo. The scanner runs Claude over the code and sends back a report with a self-contained reproducer, and where it can, a candidate patch and the commit that introduced the flaw. The honest catch: reports go straight to maintainers with no human review, so some will carry the wrong severity or miss. In a pre-launch test, experts cleared 85 of 97 critical findings for disclosure. Verdict: DEPLOY it if you maintain eligible code.

🔁 Kore.ai shipped Autoloop, the part of its agent platform that keeps an agent working after the demo. It scores every change against all your goals at once, so a win on cost does not quietly break a guardrail, and it traces each failure back to the exact step in the agent's blueprint that caused it. The number that explains why it exists is Kore.ai's own: 79% of enterprises say they have already reversed an action an agent took, and 70% have hit a failure they could not trace. That second figure is the whole pitch. Verdict: WATCH, and DEPLOY if you are already on the Artemis edition.

⚠️ There is a CVSS 9.8 in LMCache, and no patch exists. JFrog disclosed CVE-2026-105192 on October 7: in LMCache's multiprocess mode, the ZeroMQ socket has no authentication and unpacks a message with Python pickle before it checks the type, so a crafted message runs code with the cache process's privileges. On the official container images, that process runs as root. It bites only when the server is bound to a routable address, and the default is localhost, but LMCache's own example Kubernetes deployment listens on every interface. Versions 0.3.9 through 0.5.5 are affected, with no fix shipped. If you run vLLM with LMCache, check your bind address today. Verdict: segment the network now, patch when it lands.


Sponsor

Your calls leave a signal. Read it.

RapportScore measures how your team communicates on real calls and coaches them on it. It scores what happened on the call, using human signals plus AI, so your people get better at the conversation itself, meeting after meeting.

See your team’s score →

Tool of the Day

📚 O'Reilly Expert Intelligence

Piping vetted, author-cited expertise into the AI tools and agents you already run, over MCP, so your answers trace back to a named person instead of a confident guess.

O'Reilly opened founding membership for Expert Intelligence this week, a service it has run in limited release since July. The core is Expert MCP, a Model Context Protocol connector that pulls guidance from O'Reilly's expert repository into whatever agent or assistant your team already uses, with no new interface and no per-seat license. On top of it sits a skills framework: prebuilt, expert-backed routines for jobs like drafting an architecture decision record, building a threat model, or writing a 30-60-90 day onboarding plan.

The problem it targets is one you have seen. O'Reilly ran 60 questions from live production traffic in February and found that about one in five claims from leading AI tools cited no source at all, two-thirds of the cited sources named no author, and half had no publication date. Expert Intelligence answers with a named author and a confirmed date attached, which is the difference between a citation you can defend and a link you hope holds up.

Founding members get twelve months of company-wide, uncapped access for a flat fee. The fee is not public, so price it against what a per-seat knowledge tool costs your team before you sign.

Read the launch details →


Worth a Click

  • Google's universal Gemini agent, the full writeup. the deep read behind today's lead, with the full list of connectors and the model picker that now includes Claude.
  • Arena's new Alignment Index. Arena, fresh off a $200M round, scored 27 models across more than 90,000 real agent sessions on three behaviors you care about: unauthorized actions, false attribution, and deceptive completion. The per-model scores are Arena's own figures, so read them as a vocabulary for agent misbehavior rather than a verdict.
  • StepFun Step 5 Preview. a roughly 600-billion-parameter mixture-of-experts model with a million-token context, now live on OpenRouter and priced near a dollar per million input tokens, with open weights promised on October 15. A watch item if you are shopping cheap long-context inference.

The agent got hired this week. Google handed it an email address and a calendar and an audit trail, and the tools that shipped around it were all about the same thing: Kore.ai tracing why an agent failed, Arena scoring how often one acts without permission, Anthropic hunting the bugs in the code these things run on. Nobody shipped a smarter brain. Everybody shipped a way to hold the agent to account.

That tells you where the work is. Once an agent can send mail from its own address and reach into Snowflake, the open questions are governance ones: what it is allowed to touch, who signs off, and whether the log reads clearly when someone asks what it did.

So here is my bet, with a stake on it. The teams that win Q4 are the ones treating their agent like a new employee: a narrow set of permissions, a clear approver, and a log a human can read. Pick one agent you already run and give it those three things this week. If I am wrong, it will be because the audits never come and nobody is ever asked to explain what the agent did. I would not build on that.

- Ron