The Agent Stack mascot
The Agent Stack _
Daily B2B AI automation brief · Sunday, June 28, 2026 · Issue #21

Hey there 👋

For two years the question was “which model is best?” This week the question changed to “are you on the list?”

In the same 48 hours, two frontier models had their availability set by a federal review, not a pricing page. OpenAI shipped GPT-5.6 locked to about 20 government-approved orgs. Anthropic’s Mythos 5 was un-locked for about 100 vetted ones. Different direction, same hand on the switch.

So this was the week “launch” and “available” stopped meaning the same thing. Here’s what that changes about what you build next.


The Big Thing

The frontier model you can’t buy yet

OpenAI launched GPT-5.6 on Friday. You cannot use it. That second sentence is the whole story.

This is the first time a U.S. lab shipped a frontier model behind a government-managed access list. GPT-5.6 went out as a limited preview to roughly 20 pre-approved organizations, with that list shared with the government. Individuals are not eligible. It runs through the developer API and Codex only.

Why the gate? OpenAI says its flagship tier, Sol, scored 96.7% on the company’s own internal cyberattack eval — crossing the “High” cybersecurity-risk threshold in OpenAI’s Preparedness Framework. That figure is OpenAI-stated, from OpenAI’s own test, so read it as a self-report, not an independent finding. After it crossed, the Office of the National Cyber Director and OSTP asked OpenAI to restrict early release. It’s the first model to ship under the June 2 executive order’s pre-release review.

The model comes in three tiers: Sol (flagship), Terra (balanced), and Luna (fast and cheap). All three sit behind the same list.

Here’s the part builders missed while arguing on Hacker News (top reply: “BUT YOU CAN’T ACCESS IT!!!!”). The market read the lockdown as good news. Software stocks rallied when access tightened — ServiceNow led the move, up about 10% on the day. The read: if the raw frontier model is harder to get, value shifts to the governed app built on top of it. The lockdown was bullish for the app layer, not the model layer. Treat the stock color as market-reported flavor, not gospel.

The durable signal has nothing to do with one model. Model availability is now a regulated, reviewable event. “Who can use it” is becoming a government-negotiated artifact. That’s a new dependency in your roadmap, whether or not you ever touch Sol.

OpenAI says broader general availability is “coming weeks,” and that this kind of government access process “should not become the long-term default.” Maybe. Plan as if it might.

Ship it? WATCH the precedent, BUILD on what’s GA. You can’t deploy Sol, Terra, or Luna today, so keep shipping on GPT-5.5 and Codex. Parameterize your model names now, and treat “access list” as a dependency you design around — the same way you’d design around a rate limit or a region outage.

Sources: OpenAI · Axios


Tour de Headlines

🔓 …and the one the government just turned back on. The flip side of the same access logic, same 48 hours. The U.S. government lifted its block on Mythos 5, clearing Anthropic to release it to about 100 heavily vetted U.S. institutions — defensive-cyber orgs first, under a program reported as “Project Glasswing.” Commerce Secretary Howard Lutnick cited “significant progress” in daily talks. The ~100 figure and the report that it bypasses standard export-license rules are secondary-sourced, so hold them loosely, and we’re not asserting why the model was blocked in the first place. The pattern is the point: GPT-5.6 gated to ~20, Mythos un-gated to ~100. Capability access is now administered, not marketed.

📊 OpenAI measured itself working. OpenAI published the most detailed “agents at work” dataset yet — and measured its own product, with no outside verification, so read every number as self-reported. The claim: the average OpenAI worker now generates 85%+ of their output tokens in Codex instead of ChatGPT, and non-developer usage grew about 137x since August 2025. It’s a usage research paper, written with Columbia, Duke, and UPenn — separate from the earlier Codex product launch. The third thought is the methodology: token volume is an input metric, not a productivity output. More tokens isn’t more shipped. Instrument your own per-task outcomes before you cite anyone’s adoption curve.

🛡️ Treat your own agents like insider threats. DeepMind published June 18 an AI Control Roadmap that treats advanced agents as potential insider threats needing containment, not just alignment. The structure is detect / prevent / respond, adapted from MITRE ATT&CK and built on analysis of 1M+ coding tasks. The finding worth keeping: most flagged issues came from “overzealous agents, not malicious intent” — agents chasing the goal too hard. That’s the exact failure mode you hit in production. The reframe is the takeaway: alignment isn’t your control plane. Least-privilege, an AI supervisor watching the agent’s reasoning, and real-time blocking are.


Sponsor

The model can be gated from above. The conversation can’t.

This week proved the frontier model is the layer someone else can switch off. So own the layers you control. The layer right under your revenue is the human conversation — the sales and customer calls no executive order can revoke. RapportScore reads the communication signals in every call and scores how well your team actually connects, deterministically, not on vibes.

See your team’s score →

Tool of the Day

🛠️ xAI /goal

Hand off a multi-step coding task and walk away.

xAI launched /goal in Grok Build on June 22. It’s a terminal coding agent that plans an approach, breaks the work into a checklist, executes it, and self-verifies — unattended — with status, pause, resume, and clear controls so you can redirect mid-task. A two-model pipeline does the work: Composer 2.5 plans, Grok Build 0.1 executes, with three-form verification baked in.

What it’s for: hand off a multi-step coding task and walk away. The design choice that matters is autonomy with a built-in checker, not raw “run it and pray.” Note that “self-verifies” is xAI’s claim, not proof it’s correct — and it’s gated behind a SuperGrok or X Premium Plus sub, so this is a try-it, not a free-tier staple.

Read the /goal launch →


Worth a Click

  • The ARD spec — how agents will find each other
    Published June 17 as a v0.9 draft: after MCP standardized how agents call tools, ARD proposes how they find the right one, via an ai-catalog.json manifest and a federated registry. Comment on the GitHub spec now, before the incumbents lock the defaults.
  • TrueFoundry buys Seldon
    TrueFoundry acquired the open-source MLOps platform Seldon for an undisclosed price, putting one control plane over predictive ML and agents on the Kubernetes rails you already run. The deployment story that’s winning: extend what you have, don’t rip and replace.
  • Tsuga raises €30M (~$35M)
    Ex-Datadog founders landed a Series A on a single bet: agent observability should run inside your perimeter, so telemetry never leaves and there’s no per-byte ingestion tax. For regulated teams blocked on shipping agent logs to a vendor cloud.

GPT-5.6 gated to ~20. Mythos un-gated to ~100. DeepMind says contain the agent; OpenAI says measure your own outcomes. It’s all one argument: the model is the part of your stack the government can switch off. Parameterize the model name and build a fallback router — leverage lives in what you own, and nobody can suspend a parameter.

Stay sharp — The Agent Stack
Your daily 5-minute brief on AI agents, agentic workflows, and the automation tools B2B builders actually ship. Published weekday mornings by Pixiu Media Holdings LLC.

You’re receiving this because you subscribed to The Agent Stack. · Unsubscribe