Three bets this week on where an agent’s real edge lives. ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏ ͏
The Agent Stack mascot
The Agent Stack _
Daily B2B AI automation brief · Saturday, August 29, 2026 · Issue #78

Hey there 👋

I read three of this week’s announcements back to back, hunting for the thread that ties them together. There isn’t one. There are three, and they flatly disagree with each other, which turned out to be the most useful thing about them.

Each is a bet on where an agent’s real edge lives. Salesforce, the company that spent years insisting it would stay neutral, just made one model the default across its whole stack. Perplexity and Nvidia went the opposite way and pulled the agent off the cloud onto a box that sits on your desk. And a cheap open model climbed to number two on a public leaderboard with no name attached, wagering that owning nothing beats signing anything. If I had to put money down I would split it three ways. But splitting it by accident is not the same as choosing, and most teams are about to do the first while telling themselves it was the second.


Tour de Headlines

🤝 Salesforce stops being model-agnostic. For years Salesforce told everyone it would stay neutral on which model you run. This week, in the middle of its earnings cycle, it made Claude the default reasoning model across Agentforce and Slack, and put a 37-skill Salesforce plugin inside Claude itself, in pilot now with open beta due in September. That second half is the part to sit with. Soon the assistant your reps already keep open all day will run CRM actions directly, so the model and the system of record stop being two tools you bolt together and start being one surface. They named it Claudeforce, which tells you how deep the coupling goes. This is the largest enterprise SaaS vendor trading optionality for a single deep bet, and the “stay agnostic” posture is quietly ending at the top of the market. If you build on Salesforce, the portability question you filed under “later” just moved to this quarter, because the platform’s own default is now a specific model with specific behavior.

Ship it? Watch the coupling. If your workflow now leans on one model’s specific behavior, price the switching cost before you depend on it, and keep an MCP-shaped seam between your logic and the model so a swap stays a config change instead of a rebuild.

Sources: Salesforce press release, “Salesforce and Anthropic announce Claudeforce” (Aug 26, 2026).

🖥️ Perplexity and Nvidia put a whole agent on your desk. Perplexity and Nvidia shipped an agent stack that runs entirely on Nvidia’s DGX Spark, on-device, with no per-token cloud bill. On the same open Qwen model they say it outruns the Pi harness and Nous Hermes. That comparison is theirs and nobody outside has rerun it yet, so believe the mechanism and hold the numbers loosely. The mechanism flips two constraints at once. Unit economics move from a metered bill to a fixed hardware cost you can amortize, and data residency becomes simple because nothing leaves the box. For regulated shops and high-volume pipelines, those were the two reasons local kept losing, and both just got answered. If your agent burns tokens on repetitive turns, model the amortized hardware cost against your cloud bill before your next renewal. Local is a real line item now, not a demo.

Sources: VentureBeat (Aug 25, 2026).

🧾 QueryStory wants your agent to show its receipts. QueryStory left stealth with a $6M seed and an agentic data platform that hangs a cybersecurity-style evidence chain on every answer. When the agent hands you a number, you can trace it back through the query, the source table, and the step that shaped it, one link at a time. Provenance is the trust layer that goes missing the moment an agent touches analytics or RevOps data, and the funding is a category signal worth reading: money is moving toward proving an agent’s answer, rather than only generating it. Before an agent’s output drives a real decision, ask what its evidence chain looks like. If the honest answer is “the model said so,” that gap is the entire pitch this class of tool is selling into.

Sources: TechCrunch (Aug 26, 2026).


Sponsor

Your agents can fake a log. Your reps can’t fake a call.

RapportScore measures how your reps communicate on real calls, then coaches the behavior that builds trust. Deterministic signals you can coach against, not vibes. See where your team stands.

See your team’s score →

Tool of the Day

🐂 GLM-5.3-Flash

A 320B open-weight multimodal model with a million-token context and MIT weights, priced near a tenth of frontier. The twist: it’s the anonymous “Ox Alpha” that climbed to number two on OpenRouter before anyone knew who built it.

Z.ai released GLM-5.3-Flash this week: a 320B-parameter, 18B-active mixture-of-experts, natively multimodal, with a 1M-token context window and weights under an MIT license. The reveal is the fun part. For a stretch it ran anonymously on OpenRouter as “Ox Alpha” and climbed to number two on reputation alone, before Z.ai stepped forward and claimed it. The benchmark table is Z.ai’s own, run on Z.ai’s setup, so read the percentages as a claim awaiting an outside referee. The way a model earns its rank tells you more than the score it posts, and this one earned real traffic while wearing a mask.

Two things make it worth an afternoon on your bench. MIT weights mean you self-host without a license fight, and a million tokens of context swallows a whole repo or a whole contract in one pass. Native multimodality helps too: images, tables, and screenshots go straight in without a separate vision bolt-on, which is what a document-heavy or screen-operating agent needs. The price is the other half, cheap enough to sit as the default for routine agent turns, with escalation to a frontier model saved for the tasks that earn it. That router pattern, cheap model first and expensive model on demand, is where the real savings live. Point it at one high-volume, low-stakes workflow, measure its output against your current default, and let the cost delta make the call.

Read the release →


Worth a Click

  • AccuKnox shipped Agentz, the governance layer sold as a product. A model-agnostic platform to build, run, and govern production agents, with zero-trust sandboxes, credentials injected at runtime, and an air-gapped deploy option. If you are standing up agents at scale and the control plane is your problem, this is the shape of the answer the market is converging on. (AccuKnox, Aug 27, 2026.)
  • The Ox Alpha reveal, in full. How an unnamed model reached OpenRouter number two on vibes before Z.ai claimed it. Read it for the mechanics of how open-weight distribution works now: reputation first, brand second, and the leaderboard as the audition.

Three teams placed three different bets this week on where an agent’s real edge lives. Salesforce went all-in on a partnership and one model. Perplexity and Nvidia bought the silicon and lifted the agent off the cloud completely. And a cheap open model climbed to number two with no name attached, wagering that depending on nobody outlasts any single deal. They can’t all be right. The tell is what each one agreed to depend on: a vendor, a box, or nobody at all. Pick yours on purpose. The default is depending on all three by accident and calling it a strategy after the fact.

More tomorrow,
— Ron

You’re receiving this because you subscribed to The Agent Stack. · Unsubscribe