|
Hey there 👋
The desk read two things back to back this morning, and both came from the same corner of the industry. One was Anthropic’s plan for research agents that propose their own safety fixes and then train them. The other was a letter, signed by 128 companies, asking for a coordinated cyber defense before agent-driven attacks scale up.
So here is my bias, for what it is worth. The most useful signal this week is a question about motive: who is asking for guardrails, and why now. The people closest to the capability spent the week building the safety net, not the demo reel. When the folks shipping the fastest agents start signing letters about fire departments, I stop skimming and start reading.
Five minutes. Let’s go.
— Ron
The Big Thing
Anthropic’s alignment researchers are agents now, and they beat the humans
Anthropic published a system it calls automated alignment researchers. These are agent systems that read the literature, propose fixes for alignment failures, and train those fixes with no human in the loop. Search, propose, train, score, repeat. The loop runs itself.
The number that will travel: Anthropic says the agents beat experienced human researchers on its benchmark tasks in about six hours. TechCrunch’s interview put the cost at roughly four dollars an hour of compute against about a hundred and fifty for the humans. On paper that reads like a researcher who works faster, costs less, and never sleeps.
Two things keep me from cheering yet. Anthropic wrote the tasks, and Anthropic ran the scoreboard. The result holds up on its own terms. It is also a home game, with the host writing the rules and keeping score. Grade the mechanism as real and the ranking as provisional, at least until a team with no stake in the answer runs the same setup on its own models.
The part worth your attention is what Anthropic released alongside the claim. It open-sourced the evaluation harness. For once you can run the thing instead of trusting the chart. If you own an alignment or evals function, that is an afternoon well spent: point it at a model you control and watch whether the loop holds up once it leaves the lab that built it.
One more caution, on the language. “Self-improving” is carrying a lot of weight in the headlines. The honest description is narrower and more interesting. Agents that search papers, draft a fix, train it, and score the result against a fixed set of tasks Anthropic chose. That is a genuine loop and a genuine result. It is not a machine rewriting itself toward superhuman science, and Anthropic says exactly that in its own limits section, which is the sentence most of the coverage skipped.
So what do you do about it on a Monday? Not much, and that is fine. This is a research artifact with a rare gift attached: the code to check it. Watch for the outside reruns. Kick the tires on the harness yourself. Hold the leaderboard claim loosely until somebody outside Anthropic reproduces it.
Ship it? Watch, then experiment. The open harness is worth an afternoon. The benchmark stays Anthropic’s number until an outside team reruns it.
Sources: Anthropic research · TechCrunch, Aug 28.
Tour de Headlines
🚨 128 companies want a cyber fire department. OpenAI, Anthropic, Google, Microsoft, AWS, CrowdStrike, Palo Alto Networks, Okta, Visa, Mastercard, Citi and roughly 120 more put their names on one call for collective cyber defense. The warning is blunt: AI-agent-enabled attacks are close to scaling, and no single vendor can hold the line alone. Read what it is before you read the letterhead. This is a coalition statement. It is not a standard, and no tool ships from it. What it buys you is cover. The people who build these agents are now on record saying the threat model is real, so use it to fund agent-threat defense this quarter instead of next.
🤝 Nvidia is reportedly circling Hugging Face. Hugging Face is where the open-weight hedge lives, the place teams go to self-host instead of renting a frontier API. The Information reports Nvidia agreed to buy it for about $12.9 billion, and Bloomberg, CNBC and TechCrunch all frame it the same careful way: reported talks, nothing signed. Neither company confirmed, and the outlets note the deal could still fall through. Treat every figure here as reporting, because that is what it is: nothing has been filed. If it closes, it would be Nvidia’s largest acquisition ever, and the dominant chip vendor would own the top home for open models and agent tooling. Nothing to do today. If your stack leans on HF for self-hosting, log the concentration risk and keep an eye on it.
⚖️ Aderant opens agents you can enroll in. The deployable pattern is hiding under the flashier news. Aderant opened Early Access to its Agent Center at ILTACON, with purpose-built agents (Appeals, Collections, Talent) wired straight into the billing and practice-management software law firms already run. Not a chatbot bolted on the side. The agents live inside the system of record. Enrollment is live now, which makes this a dated step past May’s intro announcement rather than another roadmap slide. Steal the shape even if you never touch legal software. An agent that sits inside the tool your team already opens every day tends to beat a smarter one parked in its own tab.
|
Sponsor
Your agents can fake a log. Your reps can’t fake a call.
RapportScore measures how your reps communicate on real calls, then coaches the behavior that builds trust. Deterministic signals you can coach against, not vibes. See where your team stands.
See your team’s score →
|
Tool of the Day
🛠️ Crescendo AI-Native CXP
What it is for: letting specialized agents run and keep improving a whole contact-center operation, well past simple ticket deflection.
Most CX AI is a single deflection bot with a polite greeting. Crescendo’s pitch is a crew of agents instead: Concierge on the front line, Assist backing up reps, Quality running QA, and Insights reading the outcomes, all operating the whole thing end to end and tuning from the results. The idea is sound and the role split is sensible.
The claim I would poke is the self-tuning one. A platform that grades its own quality has an obvious incentive to grade generously. So pilot it against your own resolution rate and CSAT before you let it mark its own homework. This is a single-vendor announcement, which means the capabilities are Crescendo’s word until a customer publishes numbers.
Ship it? Watch. Worth a real pilot for CX and support leaders, as long as you score it against your own baselines rather than the launch deck.
See the launch →
Worth a Click
- DBS put agentic AI in front of ~1,500 corporate bankers to draft credit memos. A regulated bank, a real production rollout, after a 150-user pilot. The agents handle 70-plus tasks to produce a review-ready first draft, and DBS is targeting at least a 30% cut in the time RMs and credit-risk managers spend on it. If you keep asking what pilot-to-production really looks like, this is the case study with numbers on it. dbs.com/newsroom
- Ambient.ai showed agentic “video walls” at GSX. An agent watches every camera feed and surfaces the one that matters, in plain language, about once a minute. This is off our core beat, it lives in physical security, but the shape is a clean agent-as-operator pattern for any always-on monitoring job. Confirm what is genuinely new before you bet on it. ambient.ai/press
Notice who is asking for the brakes this week. The lab shipping self-improving alignment agents. The 128 companies signing a cyber-defense letter. The chip vendor trying to buy the house the open models live in. All three moves come from the people closest to the capability, and every one of them lands on the safety net and the supply chain. They can see how fast the front end is moving, so they are spending their week on the brakes. When the builders themselves start asking for guardrails, that is the signal I would trust ahead of any launch deck.
— Ron
|