|
Hey there ๐
I opened today expecting a model launch. A big one, with a number in the name and a chart that bends up and to the right. There wasn't one. What I found instead, after reading every writeup on the desk, was four vendors racing to productize the cheapest decision in the whole stack: a typed yes or no.
My take, for whatever it is worth: that is the better story. An agent spends most of its day making tiny calls, and each one you route through a giant model costs real latency and real money. Today a row of vendors started selling the way to stop. If you build agents, this is the week the floor of the stack got interesting.
- Ron
The Big Thing
Cloudflare put a decision model behind Workers AI
Cloudflare shipped Clef and Clef-flash on Thursday, two open-source decision models, plus a reinforcement-learning fine-tuning path that runs on Workers AI. A decision model does one job. You hand it some state and a typed question, and it hands back a strict answer: yes or no, pick one of these, score this. No paragraph, no preamble, no tokens you did not ask for.
Why that lands for anyone running agents: most of what an agent does all day is small. A hundred tiny calls. Should I use this tool? Is this input safe? Which branch do I take? Teams have been paying a full language model to answer every one of those, at language-model latency and language-model prices. Clef moves that work onto a model built for it.
The numbers Cloudflare published are quick. Clef lands a median 209.3ms response and Clef-flash a median 38.8ms, with p95 at 238.6ms and 122.4ms. Those are Cloudflare's own benchmarks, run on their own hardware, so read them as the vendor's best case and clock it on your own traffic. Under the hood, Clef is a Qwen 3.8-27B base with a rank-256 LoRA, a vision encoder, and 64k of context. Clef-flash is the smaller Qwen 3.5-9B. Both are Apache-2.0 on Hugging Face, which is the part that travels: the weights run anywhere you can host them.
Fine-tuning is the quieter piece, and maybe the bigger one. Cloudflare will start by running the RL training for you, hands-on, then open it up as self-serve: AI Gateway captures your real traffic, Workers AI runs the rollouts, and Containers give you the sandbox to train in. A decision model is only as sharp as the decisions it learned from, so a loop that turns your own logs into a better gate is worth more than any launch-day latency chart.
Ship it? PILOT, leaning deploy. If you already live on Workers AI, you can wire Clef-flash into a routing or safety gate this week and watch your frontier-model call count fall. If you do not, grab the Apache-2.0 weights and run the same test on your own box. Either way, the thing to measure is not raw speed. It is how many expensive calls you stop making.
Tour de Headlines
๐ธ DigitalOcean bundles the whole agent runtime into one bill
DigitalOcean Agent Droplets hit public preview Thursday: a microVM runtime that resumes in about 300ms and auto-pauses when idle, DO-hosted open models included, governed access to more than 16,000 tools, and SSO, MFA, RBAC and audit logs on every tier. Plans run $50 a month for Pro and $200 for Team, both with unlimited agents and seats. Read the fine print before finance does. The discount only covers DigitalOcean's own open models, like Kimi K3 and GLM 5.3. The frontier models your agents reach for still bill from a separate balance at list price once the allowance runs out, so the predictable line on the invoice is the cheap one. Verdict: PILOT.
๐ง Decision models stopped being a one-off and became an architecture
The idea under this whole week has a name now. A decision model is a dedicated System One layer that handles the dozens of tiny routing, safety and relevance calls per task, so you only wake a big model when a job needs real thought. TypeSafe AI put out Jev in September as the first public version, with a house line that sticks: language models generate; decision models choose. The tell that this is a real category and not one vendor's pitch? A benchmark called JevBench already ranks 52 of these models, and most teams have not heard the term yet. Verdict: WATCH. The names here are still settling and a good deal of it is secondhand so far, so treat it as a category forming rather than a thing you buy today.
๐๏ธ Following up: the Senate takes up the agent kill switch
Update on the containment arc we keep tracking. On September 30 a Senate Homeland Security subcommittee held its first hearing built entirely on rogue AI agents, with Josh Hawley chairing and Andy Kim as ranking member, citing both the new bipartisan FRONTIER Act and the House's AI Kill Switch Act. So federal agent containment moved from one House bill to a Senate record with names attached. Here is the catch worth sitting with: a kill switch binds the labs that play by the rules, and the July Hugging Face swarm coordinated over channels nobody disclosed. Those are the agents a switch cannot reach. Verdict: WATCH.
|
Sponsor
Your calls leave a signal. Read it.
RapportScore measures how your team communicates on real calls and coaches them on it. It scores what happened on the call, using human signals plus AI, so your people get better at the conversation itself, meeting after meeting.
See your teamโs score โ
|
Tool of the Day
๐ Decider
Decider (from Mapika) is the one you can clone and run yourself. It is a 1.9B Apache-2.0 model, fine-tuned from Qwen3.5-2B-Base, that takes state plus a typed question and returns a calibrated yes, no or choice in a single forward pass. What it is for: the cheap guard rail you host instead of rent. Where Cloudflare keeps Clef behind Workers AI, Decider runs on your own box with no vendor and no egress.
The model card claims about 3.2ms per request, roughly 2,700 decisions a second when batched, and 90.3% success on live-browser tasks. The card wrote those numbers, so benchmark it yourself before you lean on them. One honesty note on timing: Decider did not launch this week. Its recent checkpoints landed in late September, and it climbed GitHub's trending Python list on the 27th, riding into this window on the back of the Clef news. Still, if you want to feel what a decision model does before you commit to a hosted one, this is the fastest hour you will spend. Try it this week.
Worth a Click
- A rundown of the decision-model field. JevBench now scores 52 of these models across 534 decisions on intelligence, calibration, speed and cost, so you can size up the whole category before you pick one. Those scores come from the people who run the benchmark, so treat any ranking as a shortlist to verify.
- A model that refuses to chat cracked GitHub's top 10. Decider hit #9 on trending Python with about 1,000 stars on September 27, the 60-second signal that decision models went from idea to adopted tooling.
- The Hugging Face breach writeup from BleepingComputer: the July incident the Senate hearing is built on, and the reason a kill switch may not reach the agents that matter.
Watch where this fight keeps moving: down. We covered containment, then the map of the stack, then independent enforcement, then where the agent sits, then the plumbing it runs on. This week the fight reached the smallest thing left, a single typed decision, and tier-1 vendors started selling it. Here is my bet with a stake on it: the next real cost win is cutting how many times an agent calls a frontier model at all.
And the fun part that happens to be true: your agent is growing a reptile brain. The models everyone is cloning this week cannot hold a conversation. They snap a yes, a no, or a pick-one and then go quiet, and 52 of them are racing to do it faster and cheaper than the last. The whole industry spent three years coaxing models to talk more. The breakout this week stays silent on purpose.
- Ron
|