|
Hey there 👋
Two numbers crossed the desk this morning that did not survive a second source. A funding round that shrank by sixty million dollars once I chased the primary filing, and a speed benchmark nobody outside the vendor has run. A normal morning, in other words. I cut both.
What held up is smaller and more useful. Anthropic let people reach into Claude Code and change how it behaves, with code, while it runs. My bias, up front: the agents worth betting on this year are the ones you can get your hands into. The most-used coding agent just became one of them, and if you build with agents, the floor shifted under you a little.
- Ron
The Big Thing
Anthropic let you rewrite Claude Code from the inside
Anthropic shipped mods for Claude Code yesterday: small TypeScript modules that can rewrite prompts, intercept tool calls, draw their own interface panels, register new commands and tools, or replace built-in features outright. They run in the CLI and the desktop app at no extra cost, and you install one through /plugin or the Claude directory.
The shape of a mod is middleware. You write a register(on, options) function and hook it to events the agent fires as it works: tool.call, prompt.submit, ui.render. On each one you can watch what is happening, rewrite the payload before it moves on, or stop it cold by returning { deny: ... }. Claude Code will even write a mod for you from a plain sentence of description and hot-reload it mid-session, so the agent can edit its own behavior on request.
Why this lands for anyone running agents: until now you shaped Claude Code from the outside, with settings files and system prompts. A mod drops you inside the loop. The examples people are already trading are down-to-earth: a confirmation gate before a destructive command, a live token-usage readout, a custom code-review pane, a drop-in replacement for the built-in /diff. The tool went from a product you accept as shipped to a surface you program to fit your own shop.
The catch, and it is the first thing most writeups skip: mods are unsandboxed. A mod runs with the same access to your machine as Claude Code itself, so one you pull from the directory deserves the same scrutiny you would give any program you install by hand. Anthropic's answer on Enterprise plans is a sec-default mod that loads before everything else, so a user mod cannot quietly override your permission rules. On a personal plan, that guard is you. Read the source before you run a stranger's mod.
Ship it? DEPLOY, with one habit attached. If you live in Claude Code, write a single mod this week and make it a confirmation gate on anything that deletes or overwrites. You will feel the difference the same day. Just treat third-party mods the way you treat any unvetted dependency, because under the hood that is what they are.
Sources: AlphaSignal, dev.ua, Anthropic release notes.
Tour de Headlines
🧩 Google turns Gemini into something that does the work, not just names it
Skills for Gemini rolled out this week: reusable, multi-step routines that run repetitive tasks across your apps, with the first push landing in the Chrome side panel as Gemini grows into a full browsing agent. Skills also replace Gems, Google's older custom-assistant feature, which retires in November. Existing Gems migrate across, so nothing you built is lost. For builders the read is clean: the assistant layer is settling on one primitive you define once and reuse everywhere, and Google just told you which one it picked. Verdict: WATCH, and go claim your Gems before the rename lands.
🎙️ Microsoft cut voice-agent transcription to ten cents an hour
Microsoft dropped MAI-Transcribe-2 to $0.10 per audio hour on Azure AI Foundry, down from $0.36 five months ago. For a voice agent that makes the listening layer close to free: a team running 100,000 audio hours a year goes from $36,000 to $10,000. The model covers 60 languages with speaker diarization, word-level timestamps, and code-switching for mixed-language audio, and posts a 5.2% word error rate on the FLEURS benchmark. The speed numbers (ten times faster than one rival, and so on) are Microsoft's own leaderboard figures, so clock it on your own audio before you believe them. Verdict: DEPLOY if you run voice.
🛡️ Google's new flagship model shipped to cyber defenders first
Google began rolling out Gemini 4 Argon this week, its larger frontier model, with early access held to trusted cyber defenders rather than a wide release. Reporting around the launch notes internal doubts about its coding, which is an unusual thing to hear out loud before a model reaches everyone. The gated rollout is the part worth watching: a frontier model whose first job is defending networks, with the demos held back. Treat any benchmark claim as provisional until someone outside Google runs it. Verdict: WATCH.
|
Sponsor
Your calls leave a signal. Read it.
RapportScore measures how your team communicates on real calls and coaches them on it. It scores what happened on the call, using human signals plus AI, so your people get better at the conversation itself, meeting after meeting.
See your team’s score →
|
Tool of the Day
🖥️ NVIDIA DGX Spark
Run your agents on a box that never phones home.
If per-token bills or data-egress rules are what keep your agents off the jobs you want them on, NVIDIA's local setup is worth an afternoon. NVIDIA Sync now clusters two to four DGX Spark units into a single pool with up to 512GB of unified memory, enough to serve roughly 400-billion-parameter models and multi-agent pipelines on hardware you own outright. On one box it reports up to 2.6x faster inference on Qwen3.6-35B using NVIDIA's NVFP4 quantized checkpoint and vLLM. The clustering software landed earlier this year; what is new is that the box itself opens for sale on October 15. The point is not the benchmark. It is where the data stays: on your machine, with no per-call meter ticking in the background.
Read the setup guide →
Worth a Click
Watch where the control keeps moving: outward, to the edges of the agent. For a month the fight was about where the agent sits and whose name was on its login. This week it moved to who gets to reshape the thing itself. Anthropic handed over the internals of Claude Code, Google turned its assistant into routines you define, and the price of the senses an agent needs fell through the floor. My bet, with a stake on it: the teams that win the next year are the ones who treat their agent as something they program. The vendors are, on purpose, handing you the screwdriver.
And the fun part that happens to be true: the first thing people built with that screwdriver was a brake. The headline example mod for Claude Code is a confirmation gate that stops the agent before it runs a destructive command. We spent two years trying to make these things more autonomous, and the opening move, once people could change one from the inside, was a pause button. That tells you where the real nerves are.
- Ron
|