The Agent Stack mascot
The Agent Stack _
Daily B2B AI automation brief · Thursday, August 13, 2026 · Issue #62

Hey there 👋

For two years the fight was about which model is smartest. This week two launches made that argument feel beside the point. xAI gave its agents cloud machines and the keys to your app logins. AWS started keeping an agent's runtime alive for as long as two weeks at a stretch. Both land on the same shift: the question that matters moved from how clever the model is to what a logged-in agent runs on and what it can reach while wearing your credentials.

Here is my bias, for what it's worth. The model stopped being the variable I watch a while ago. What I watch now is what an agent is allowed to sign into, and whether anyone graded it before it got the seat. Today runs on that thread.

— Ron


The Big Thing

xAI's Grok Bot gets its own computer and your app logins

On Tuesday xAI put Grok Bot into public beta, and the pitch is blunt: an always-on AI teammate that runs on a dedicated cloud computer, signs into the apps and sites you already use, and keeps working while you are away. Each Bot carries context across tasks and comes back for your approval before it commits anything. You can train one by letting it watch you do a job a single time. The named use cases are straight B2B ops: research accounts overnight and draft the outbound, update the CRM, process invoices, reproduce a bug and file the ticket.

So why am I not leading with the capability? An agent that drives a computer is not new. Cloudflare, Anthropic, and OpenAI all shipped a version of the operator this year. The tell in this launch is distribution. Beta access is bundled into SuperGrok Heavy at $300 a month, Cursor Ultra at $200, and Cursor Teams Premium at $120 a seat. Read that seating chart twice: xAI's flagship agent ships through Cursor's subscription tiers as much as through xAI's channel. The xAI and Anysphere merger stopped being a press release and became a line you can see on the invoice.

So the real adoption gate is price, plus the exact list of apps you are willing to let it log into unattended while wearing your credentials. That second part is where the risk lives once an agent is signed in as you.

Two honest flags. This is still a gated beta. General availability, tied to the wider Grok 4.6 rollout, was still landing at press time. The dollar prices come from the Business Today writeup, not the xAI page, so hold them loosely. And the headline claim, that it finishes multi-step work end to end, is xAI's framing off the launch page. The open question all over the launch threads is whether it closes the loop or clips a clean demo. Grade that yourself before you trust it with a seat.

Why it matters: when an agent can sign in as you, the exposure widens from the sentences it can generate to the actions it can take inside every app you let it open, from CRM records to invoices to outbound mail.

Ship it? If you are already on a top xAI or Cursor tier, pilot it this week on one low-stakes app and watch exactly what it touches. Everyone else: study the app-login permission model before you grant it a single login.

Sources: xAI news, Unite.AI, Business Today


Tour de Headlines

🧩 Your intent data is becoming a tool your sales agent can call. The shift worth naming this month is data vendors turning themselves into MCP tool providers. 6sense serves its buying-stage, 6QA, and intent signals as an MCP server that Claude, ChatGPT, Writer, and Agentforce can call with no custom integration, and its August GTM-stack release pushes that intelligence deeper into the workflow. When proprietary data becomes a callable tool, the moat moves from who owns a dashboard to whose signals your agent can act on mid-thread. This is the concrete answer to the RevOps question we keep hearing: how do I ground my AI SDR on live intent instead of exporting CSVs at midnight. One figure to keep in a box: 6sense says 6QA-prioritized accounts carry 29% higher opportunity value and close 33% larger deals, but 6sense measured that itself and no outsider has checked it. Treat it as the applied pattern plus an August GTM push; the MCP server has been in open beta since mid-July, so this is not a brand-new debut. 6sense, MarTech Series

💸 Another cheap, agent-first model just undercut the floor. Upstage shipped Solar Pro 4 on August 10, a model built to close the agent loop at near-throwaway token prices. List pricing is $0.30 in and $1.20 out per million tokens, with a launch promo cutting that to $0.03 and $0.12, ninety percent off. That promo is a welcome mat that will expire, so budget around the list price. Upstage reports a big jump on agent evals, Terminal-Bench v2.1 climbing from 12% to 57%, and that figure is vendor-run, though Artificial Analysis is running an independent check you can wait for. The steady thread across the last few weeks holds: right-size the model to the loop instead of reaching for the frontier by reflex. Upstage, Artificial Analysis

🖥️ AWS gave agents a runtime that stays up for two weeks. AWS added persistent runtime instances to Bedrock AgentCore: EC2-backed compute where several agents share a host and a single session can run up to 14 days, against the old cap near eight hours, with GPU support and stop or restart built in. The design shift is the story. Instead of short single-task runs, you can build an agent as a long-lived service that monitors and coordinates over days. It pairs with Grok Bot on the same spine: agents are getting long-lived machines to sit on. If you were about to hand-roll an EC2 and state layer for a long-running agent, AWS just answered that for you. One caveat on timing: the primary AWS blog reads about a week old, so file this as last week's infra note rather than breaking news. Pricing is standard EC2 plus an AgentCore management fee. AWS News Blog


Tool of the Day

🧰 Insygna Agent Report Card: grade an agent before you connect it

What it's for: running a rogue-risk check on an AI agent before you hand it access.

After a month of nine-figure agent-security rounds, the useful thing this week is free and you can run it today. Insygna put out a public Agent Report Card: point it at an agent's repository, it runs independent tests, and it returns a security score across six dimensions with a findings list, version history, and a "Verified" badge you can show your IT team. It is part of Insygna's wider agent-identity platform. The hook is the stat: Insygna says the median agent it has tested scores under 50 out of 100, with roughly six in ten falling below that line. That number is Insygna's, self-reported and unaudited, so treat the tool as a useful first pass and the score as a claim until someone outside checks it. Either way it turns "inventory and vet before you connect" from good advice into a gate anyone on your team can run.

Sources: Insygna, AiThority

Try the Agent Report Card


A word from RapportScore

Insygna grades your agents this week. Who is grading how your reps show up on the call?

RapportScore measures how your team communicates in their real recorded calls and coaches them on it, the same way Insygna scores an agent before you trust it. It scores behavior in the conversation, and it is careful about its limits: it does not claim to read honesty or intent, and it tells you when the signal is thin instead of filling the gap. If you run a B2B sales or CS team and you have never seen your own rapport scored, that is the blind spot worth closing.

See how it works →

Worth a Click

  • 🏭 The systems integrators start selling agents into heavy industry. L&T Technology Services launched AgenticIQ, an end-to-end agent platform for engineering, manufacturing, and customer experience, with governance boundaries and IP retention built in. The interesting part is the channel. The seller is a systems integrator that already owns the delivery relationship inside automotive, semiconductor, and medical-device shops, ahead of the labs. It is a vendor launch with no third-party proof yet, so read the capability claims as framing. Business Wire via Yahoo
  • 🃏 Most agents flunk their own report card. Insygna says the median agent it tested scores under 50 out of 100. Grade that stat as the vendor's, then go run yours before the next demo talks you into a live login. Insygna

The Bottom Take

The through-line this week was machines more than model IQ. Grok Bot gets a cloud computer and signs into your apps. AWS will keep an agent's runtime alive for two weeks. Meanwhile the cheap middle keeps deploying with Solar Pro 4, and the tools worth your time are the ones that grade an agent before it logs in as you, like Insygna. The race was never really about which model is smartest. It is about what you let a logged-in agent do, and whether you checked it first.

— Ron