|
Happy Saturday, there 👋
I spent an hour this morning trying to read Moonshot's own writeup of its new model, and the page would not load from anything on the desk. The version I ended up trusting came from Simon Willison, who skipped the launch post, ran the model himself on OpenRouter, and published the transcript and the bill. That is the right instinct with any lab's self-scored table: the launch post is marketing, and a hands-on run is the only review that counts.
This is where the usual story flips. The model sitting on top of the frontend-coding leaderboard this week is the open one, and the reflex about cheap Chinese models is dead on arrival, because this is the most expensive model Moonshot has ever priced. Open versus closed, budget versus premium: every prior you hold about model launches inverts at once.
The Big Thing
The open model just took the frontend crown
On Thursday, Chinese lab Moonshot announced Kimi K3, which it calls its "most capable model to date, with 2.8 trillion parameters" and the first "open 3T-class model," rounding 2.8 up to 3 and taking the size record from DeepSeek's 1.6T v4 Pro. The headline number is real, and the ranking under it needs a careful read, because this is where most coverage is already getting it wrong.
K3 is number one on Arena.ai's Frontend Code arena, ahead of Claude Fable 5. That is the specific board it leads. Widen the lens and the picture changes: on Artificial Analysis's long-horizon knowledge-work evaluation it lands second, behind only Fable 5, and on Moonshot's own full benchmark table it beats Claude Opus 4.8 and GPT-5.5 while losing to Fable 5 and GPT-5.6 Sol. So the accurate sentence is a narrow one: the open model took the frontend-coding crown. It did not win outright. That qualifier is the whole story.
Grade the source of each claim. The full table is Moonshot's own. The arena placements come from third parties, Artificial Analysis and Arena.ai, which run comparisons but do not audit them. The one fully independent data point is Willison pointing his own harness at it: a single SVG of a pelican, rendered correctly, vision and all, for 25 cents. Believe the mechanism first, and let the leaderboard settle over the next month.
The cost math is what makes an eng leader look twice. Artificial Analysis clocks K3 at $0.94 per task, roughly half of Opus 4.8's $1.80 and a shade under GPT-5.6 Sol's $1.04. Then the inversion lands in the price sheet. K3 runs $3 per million input tokens and $15 per million output, the same tier as Claude Sonnet and, by Willison's read, the priciest model any Chinese lab has shipped, up from Kimi K2.6 at $0.95 and $4. One catch worth knowing before you budget: K3 ships with a single "max" reasoning effort, and it is heavy. That 25-cent pelican burned 13,241 reasoning tokens for one picture.
Ship it? Test it this week on your own code; the arena is not your codebase. It is live now through an OpenAI-SDK-compatible API and on OpenRouter as moonshotai/kimi-k3, so wiring an eval to it is a config change you can do this afternoon. Open weights are promised "by July 27," which for now is a date on a blog, with the download still to come. If you are in a regulated shop, that self-host path, promised for July 27, matters more than the API does today.
Simon Willison's hands-on writeup · Moonshot's Kimi K3 blog
Tour de Headlines
🧭 Sierra started charging for the outcome, not the tokens. On Thursday, Bret Taylor's Sierra announced Horizon, a platform for agents that chase a goal across days, weeks, or months instead of a single chat: originating a loan, getting a prior authorization, closing a sale, scheduling a test drive. It pairs a "context engine" that stitches every interaction together with "long-horizon planning" that reasons between them. The line that will travel around your next planning meeting is the pricing one, verbatim: "you don't pay for tokens, you pay for business outcomes delivered," with Sierra eating the token spend. Agent OS already runs phone and chat for "almost half of the Fortune 50," by Sierra's own count. It is an announcement, with design partners so far and no self-serve yet. Ship it? Watch, and ask the obvious question the post does not answer: who absorbs the bill when a weeks-long agent loops a thousand times.
Sierra on Horizon
🔒 Pull the Claude Code update. This week's Claude Code releases read like a security patch cluster more than a feature drop. Anthropic closed a permission-check bypass in Windows PowerShell 5.1 sessions, made Bash commands over 10,000 characters always prompt instead of running on their own, stopped worktree subagents from touching the main repo checkout, and fixed approval previews so look-alike and zero-width characters can no longer dress up the message you are approving. Read plainly, it is the vendor hardening its own gaps, which is what a good changelog looks like. And it is a different beast from the config-file injections we keep flagging: this time the permission analyzer itself got stricter. Ship it? Deploy, run claude update before your next session.
Claude Code changelog
📞 A voice founder said the quiet part. Rime raised a $24M Series A and reports handling over 100 million calls a month for names like Mayo Clinic, Upstart, and Asurion, on voice models trained in its own recording studio rather than scraped from the web. The reason it is here is the founder's honesty. CEO Lily Clifford told TechCrunch the tech "is still not there to automate the vast majority of enterprise phone calls," and that talking to a voice agent today "is kinda like a new IVR, but with a better voice." Better speech-to-speech sits on the roadmap and has not shipped. Ship it? Watch, and keep that quote handy for the next vendor who promises to automate your whole call center.
TechCrunch on Rime
|
Sponsor
You benchmark the models. Benchmark the rapport.
This week you will swap a new model behind the same endpoint and move on by lunch. The highest-stakes workflow in your company still runs unmeasured: your live sales and CS conversations. The discovery call somebody misread. The deal where two people left with different understandings. The rapport that never formed, and closed-lost with no warning light. RapportScore reads the human signals in every call and email and scores how your team connects. B2B communication intelligence for teams that live on calls. You verify the agents. Verify the conversations too.
See your team’s score →
|
Tool of the Day
🧰 Emergent
Describe an internal app in plain language; it builds, deploys, hosts, tests, and debugs the thing. Self-serve at emergent.sh.
The pitch from CEO Mukund Jha is "an engineering team in a box," and this week Emergent crossed into unicorn territory a year after launch, which is a good excuse to look at what it is for. It targets the non-technical builder: the ops or RevOps lead with an internal-tool backlog and no eng ticket coming. The receipts are Emergent's own, told to TechCrunch, so price them accordingly: a $120M annualized run rate up 70% in four months, and more than 200,000 paying customers. What makes it read like real B2B work rather than demos is the customer list, trucking companies building shipment trackers, factories, construction firms standing up ERP systems, property managers building their own CRMs. Its closest comparison is Replit, and Jha concedes design is still a weak spot. The $130M round behind the headline is the footnote here.
Try Emergent →
Worth a Click
- Borrow OpenAI's scorecard, skip the sales pitch. CFO Sarah Friar published a way to score agent spend she calls "Useful Intelligence per Dollar": cost per successful task, plus a dependability split of ready to use, needs correction, needs escalation. The framework is genuinely usable this week. The benchmark figures wrapped around it are OpenAI's own, and the piece is quietly selling ChatGPT Work, so lean on the scoreboard and weigh those numbers with care. A scorecard for the AI age
- A bot that sits in your meetings watching for deepfakes. Polygraf announced Meeting Guard, which joins Zoom, Teams, and Meet as a visible participant and scans every face and voice for AI-generated content, cloned audio, and PII leaks, with cloud setup it says takes under 15 minutes. It is a single-source vendor announcement, so weigh the claims yourself; Polygraf cites Gartner that 62% of organizations have already hit a deepfake attack, 37% of those on video calls. The category is the real news. Polygraf Meeting Guard
For three issues the story was the attack surface. This week the capability surface moved. An open-weight model took the frontend-coding crown from the closed frontier and undercut Opus on cost per task, with the weights themselves promised for July 27. Sierra moved its pricing onto the results it delivers. A one-year-old shop became a unicorn selling an engineering team in a box to trucking companies and factories. That is the seam we flagged back in Issue #37 doing exactly what we said it would: above it, the model and the build both slide toward zero and toward portability.
The counterweight keeps this from being a victory lap. A voice founder calls her own product a nicer IVR, and OpenAI's CFO is handing out a scorecard precisely because plenty of these rollouts still cannot prove they paid for themselves. So spend the weekend on two things that cost nothing. Point an eval at K3 behind the same OpenAI SDK you already run, and start pricing your own agents by the outcome, not the token.
Ron
|