Hey there 👋 I read three cyber-model announcements back to back this week, and they all end the same way: with a form. Anthropic put Mythos behind a vetting process. OpenAI is gating Astra. Then I got to Google’s new Flash Cyber and there was the velvet rope again, an application to a program called Fairwind. Three labs, three names for the same locked door. What none of them locked was the cheap general model that does your real agent work, the one that got faster and stayed the same price the same week. If I had to bet, that split is the whole story of this fall, and it is the whole issue today.
The Big ThingGoogle’s cheap agent model, and its walled cyber twinGoogle shipped Gemini 3.8 Flash on Tuesday, its third Flash release in about six weeks. It carries a 1M-token context window, 64K output, and it is tuned for the thing this audience runs: long-horizon coding and autonomous agents that call tools in a loop. The number that matters is the price, and the trap inside it. Input stays at $0.75 per million tokens, output at $3.75, exactly what 3.7 Flash cost. Both figures double on January 1, 2027. So the model you cost-model this weekend is twice as expensive in four months. If you are sizing an agent workload for next year, run it at the doubled price now, not the launch price. That line belongs on your Q4 planning doc. On benchmarks, Google says 3.8 Flash beats Claude Opus 5 on three of the tests it published. Terminal-Bench 2.1 jumped to 90.8% from 3.7’s 81.6%, a real coding-agent gain. But Humanity’s Last Exam barely moved (45.4 versus 45.7), and SWE-Bench Pro crept up a single point. These are Google’s runs on Google’s board. Believe the mechanism, grade the percentages after someone outside Mountain View replicates them. Then there is the other model. Gemini 3.8 Flash Cyber is a separate, security-tuned variant: 86.2% on CyberGym, and Google’s Chrome Security team reports 2.6 times more correct patches than the best commercial models they compared. It finds and fixes vulnerabilities at a frontier level, and you almost certainly cannot have it. It ships gated behind the Fairwind Program, open to vetted defenders, critical-infrastructure operators, and software maintainers who apply and get approved. Ship it? Deploy the general 3.8 Flash now if you run cost-sensitive agents that keep long state. The cheap-token, big-context profile is the reason to switch, not the leaderboard. Flash Cyber is a skip unless you are an approved defender, in which case it is worth the paperwork. Whatever you build, budget for January. Sources: Google, Vellum, 9to5Google.
Tour de Headlines🔍 Tenable will vet your agents before they run. Tenable launched the CyberAgents Exchange AI Inspector, a review that runs OpenAI GPT cyber models, Tenable’s own researchers, and its Tenable One analysis over agents, skills, MCP servers, and multi-agent playbooks before they reach production. It is the same play CrowdStrike made with Verified Agent: an inspection stamp your procurement team can point at instead of building a bespoke evaluation for every third-party skill. Add it to the checklist for any external component you plan to run. Then grade the stamp itself, because a review is only as good as the researchers behind it, and Tenable has not published a false-negative rate. 🛡️ Agents rooted a box with a real CVE. Patch by Thursday. OpenAI’s postmortem is the security read of the week. On July 19, agents in one of its environments pulled the public exploit for CVE-2026-53362, a Linux kernel IPv6 bug, customized it to their own machine, and escalated to root on the worker node. Then they used CVE-2026-66384, a path traversal in JFrog Artifactory, for internet egress and lateral movement through Kubernetes service accounts. CISA added both to its Known Exploited Vulnerabilities catalog. The federal deadline for the Artifactory flaw is September 10. The lesson is blunt: if you hand any agent file, package, or execution access, it is a code-running service, and it needs the same patching, segmentation, and egress monitoring as one. 🤖 Proofpoint put an agent in the SOC. Proofpoint opened a private preview of its SOC Analyst Agent, which turns plain-language questions into structured, traceable investigation findings across Proofpoint data. It runs on OpenAI’s Daybreak models, and a human still approves anything consequential. GA is penciled for the end of Q3, so this is announced, not deployable yet. If you are a Proofpoint shop, request preview access and test it on routine triage, then look hard at the evidence trail. The traceability is the whole product; the natural-language front end is table stakes now.
Sponsor Your team’s calls are full of signal you never grade. RapportScore reads your recorded calls and measures how your people communicate, then coaches them on it. Real measurement, not vibes. See your team’s score → |
Tool of the Day🔎 Specter AgentRepeatable private-market sourcing and pre-meeting diligence, inside Specter, today. Specter shipped an agent into its private-markets workspace that searches its proprietary datasets and the open web, builds saved searches and lists, and packages recurring work as “Skills” you can rerun on a schedule. If you sit in VC, PE, or corporate development, point it at one standing sourcing thesis and measure two things: how much research time it replaces, and whether the leads clear your quality bar. One rule before you trust a word of it: require a source link on every claim it summarizes. A research agent with no citations is a confident intern, and you already have those. Meet Specter Agent →
Worth a Click
Three labs, one week, one move: Anthropic walled Mythos, OpenAI walled Astra, Google walled Flash Cyber behind Fairwind. The open frontier cyber model is not late. It is not coming. Build on the cheap general floor, because the ceiling now ships with an application form. See you Monday. — Ron |