AI Weekly #9 — Sep 21–27, 2026
September 21–27, 2026
Anthropic cut the price of its flagship model while raising its scores, and three other stories this week went after the rest of an agent's bill.
Most weeks the AI news splits into capability on one side and money on the other. This week they were the same story told four times: what an agent costs to run, and who is changing that number.
Anthropic opened on Tuesday with Claude Opus 5.5, priced 20% below the Opus 5 it replaces, with cache reads 60% cheaper and a claimed 40% saving on typical workloads, while posting higher scores on every benchmark it chose to show. The same day Epoch AI published a measurement of how fast that sort of thing happens across the industry: the cost of reaching a fixed benchmark score has fallen about 47% a quarter since 2023, and faster still near the frontier. Opus 5.5 is one data point on Epoch's curve, and it lands roughly where the curve says it should.
Where the money goes instead
If the tokens are getting cheaper, the rest of the bill is what is left to argue about, and two items this week went after it from opposite ends.
Akamai signed an $11.6 billion, seven-year contract with Anthropic, with room to grow to about $20 billion, and it is explicitly for CPU workloads. That is the part of an agent that is not a model: running the tools, executing the code, fetching the pages, holding the sandbox open while the model thinks. A GPU deal would not have been news. A CPU deal of this size is a statement about where agent workloads now spend their time, and it came with warrants that give Anthropic a stake of up to about 5% of its supplier.
Stripe attacked the same cost from the application side. It turned on WebMCP across every Checkout page, so a shopping agent calls structured tools instead of reading and clicking a payment form. Stripe's own numbers are 42% fewer tokens and 38% fewer tool calls than browser automation, and the merchants did nothing to get there. The saving comes from the site describing itself properly, not from a cheaper model. That is the more durable of the two kinds of saving, and it is the argument for the protocols the MCP catalogue exists to track.
Firecrawl's $75 million round belongs in the same column. Its new Alexandria product is a bet that retrieval — the most token-hungry thing most agents do — will be bought separately from the model rather than bundled with it.
Cheap enough to run a thousand at once
The research item worth reading closely is Anthropic's report of about 950 Claude agents searching over 200,000 reverse transcriptases for 21 hours and turning up a family of phage enzyme systems nobody had characterised. Set the biology aside for a moment and look at what was published: a token count (210 million), a wall-clock time and an agent count. You can put a price on that run, and at this week's prices it is not a large one. Agent-science claims usually arrive without the inputs needed to check them; this one arrived with them, and a preprint.
Someone wants an off switch
California's governor named four advisers to carry out an executive order that includes designing an emergency shutoff for frontier models, with an independent organisation verifying on an ongoing basis that it works. There is no deadline in the announcement and no rule yet. But "prove you can turn it off, to someone you do not pay" is a different obligation from anything in current state law, and it is worth tracking from here rather than from the first draft regulation.
By the numbers
The directory took in 38 entries this week — 12 apps, 14 agent skills and 12 MCP servers — bringing the totals to 537 apps, 298 skills and 316 MCP servers.
When we opened the upstream sources, exactly two of those 38 were products launching for the first time inside the window: JetBrains Air and Dataiku Agent Management. Another eight shipped a real release that week — a new MCP write surface, a new SDK version, an app update. The remaining 28 either predate the week or could not be dated at all, and we did not guess. That includes all 14 skills: the skills column this week came from collections by Trail of Bits, OpenAI, SerpApi, Better Auth, AMD and others, with first commits running from December 2025 to July 2026, and none of them received more than maintenance commits between 21 and 27 September.
Two MCP servers we had dated to this week turned out to be something else: the dates on Typeform's and Canva's entries were when they appeared in the official MCP Registry, not when they launched. We corrected both. A registry listing is a real event, but it is not a release, and the difference is the one last week's issue was making: the categories have stopped launching and started accreting.
What we added
The eight picks below shipped something between 21 and 27 September. We checked each against a dated primary source: GitHub release tags, changelogs, launch posts and App Store version history. We did not rely on the date we catalogued them.
The week in AI
Anthropic ships Claude Opus 5.5 at a lower price than Opus 5
Anthropic released Claude Opus 5.5 on Tuesday, the first of a 5.5 family, priced at $4 per million input tokens and $20 per million output — 20% below Opus 5, with cache reads 60% cheaper. Anthropic claims roughly 40% lower cost on typical workloads, output over 30% faster, and a Terminal-Bench 4.0 score of 66.4% against Opus 5's 52.3%. Sonnet 5.5 and Haiku 5.5 are promised in the coming weeks.
Why it matters: A price cut on a flagship at the same time as a benchmark gain resets the default for agent and coding workloads that were being routed to cheaper models on cost alone.
Xiaomi puts a trillion-parameter MiMo model on Hugging Face under MIT
Xiaomi published open weights for MiMo-V2.6, including a Pro model that is a sparse mixture-of-experts with 1.02T total and 42B active parameters, taking text, image, video and audio with a 1M-token context. The licence is MIT. Flash and a 9B Qwen-based distillation arrived the same day, and the card lists SGLang and vLLM serving recipes.
Why it matters: Frontier-scale, multimodal and permissively licensed is still a rare combination; teams that need to self-host an agentic model now have one without licence carve-outs.
Akamai signs an $11.6bn, seven-year compute deal with Anthropic
Akamai announced the largest contract in its history: $11.6 billion over seven years to serve Anthropic's growing CPU workloads, with an option to add up to $9 billion more. Anthropic receives warrants for up to about 5% of Akamai's stock at $111.33 a share, vesting as the commitment grows, and Akamai expects around $5.5 billion of related capital spending, including memory it is pre-buying this year.
Why it matters: Agents spend much of their time running tools, not generating tokens, and that work lands on CPUs. This is the first contract of this size priced for that half of the bill.
950 Claude agents turn up an uncharacterised phage enzyme system
Anthropic reports that a run of about 950 parallel Claude agents screened more than 200,000 reverse transcriptases over 21 hours, flagged around 3,500 candidate systems and analysed the top 20 in depth. One family, which the team calls array-associated reverse transcriptases (ART), pairs the enzyme with a partner gene and CRISPR-like repeat arrays; lab follow-up was done at BSL-1/2 and a preprint is linked.
Why it matters: It is a search problem with a published token count (210M) and wall-clock time attached, which makes it one of the few agent-science claims you can cost out.
Stripe turns on WebMCP across every Checkout page
Stripe enabled WebMCP on all of its Checkout interfaces, so a browsing agent can call structured checkout tools instead of clicking through the form. Tools are disclosed progressively — a payment-submit tool only appears once the required fields are filled. In Stripe's tests agents used 42% fewer tokens and 38% fewer tool calls and finished 39% faster than with plain browser automation. Merchants do not need to change their integration.
Why it matters: Millions of merchants became agent-callable without doing anything, which makes WebMCP a deployed standard rather than a proposal.
Epoch AI: the price of a fixed capability falls about 47% a quarter
Epoch AI measured how fast the cost of reaching a given benchmark score falls, across AIME, FrontierMath, GPQA Diamond and two puzzle sets, and found a decline of about 47% per quarter — roughly 13x a year. Near-frontier performance gets cheaper faster (about 66% a quarter) than older performance levels (about 32%).
Why it matters: A usable number for budgeting: a capability that is too expensive for a product today is plausibly affordable within two or three quarters.
California starts designing a frontier-model kill switch
Governor Newsom named Jason Goldman, Gillian Hadfield, Alondra Nelson and Rob Reich to advise on his AI executive order, which tells the Government Operations Agency to speed up existing AI safety law and recommend stronger rules. Among them: requiring frontier developers to build an emergency shutoff for their models, with an independent verification organisation confirming on an ongoing basis that it works. The release sets no deadline.
Why it matters: Third-party verification of a shutdown capability would be a new kind of obligation for labs headquartered in the state.
Firecrawl raises $75M and launches a retrieval layer for agents
Firecrawl closed a $75M Series B led by Smash Capital and launched Alexandria, which lets an agent query the live web, official data providers, custom connectors and Firecrawl's own research, developer and government indexes through one interface. Firecrawl says agents using it scored 21% higher on answer quality than with built-in web tools across 845 tasks, and it plans to pay content contributors when agents use their material.
Why it matters: Retrieval is being unbundled from the model vendors' built-in search, and this is now a funded alternative with a free tier.
Google ships Gemini 3.8 Flash TTS with 30-second voice cloning
Google launched two speech models, Gemini 3.8 Flash TTS for directed, character-style voices and Flash-Lite TTS for cheap high-volume output. Together they cover over 100 languages and dialects and more than 2,000 voices; custom voices can be described in a prompt or cloned from a 30-second sample, and output is SynthID-watermarked. Both are in the Gemini API and AI Studio now, with no pricing in the announcement.
New on Onei this week
Released during this window and now in the catalogue.
- JetBrains Air App
Announced on Tuesday, and one of only two first launches among the 38 entries we indexed this week. What is worth checking is the review step rather than the agent list: diffs from a cloud run come back into the IDE you already use, which is where the pull-request bottleneck actually sits.
-
Announced at Dataiku's Succeed event on Thursday, with general availability in October. It is sold standalone, so you do not need the Dataiku platform, and it is priced per instance with monitoring metered per agent. It is aimed at the team that has to answer for agents it did not build.
- Basedash App
On Friday its MCP server gained write tools: create_chart, edit_chart, create_dashboard and edit_dashboard. A chart an agent builds now becomes a durable Basedash URL with a screenshot preview, not an answer that disappears when the chat ends.
- Arize AX App
Phoenix 20.16.0 shipped on Wednesday and already has cost tables and playground support for Claude Opus 5.5, released the day before. If you need to price this week's model changes against your own traces, that turnaround is the practical reason to look.
- Composio App
Three SDK minor versions in four days. The Thursday release (Python 0.24.0, TypeScript 0.21.0) adds experimental per-session policies for premium usage, which is the control you need before handing an agent a thousand tool integrations and a budget.
- Vantage MCP Server MCP server
v2.28.0 landed on Friday with tools for Kubernetes efficiency reports and saved filters, plus access-policy management gated behind an explicit switch. It was the only MCP server of the 12 we indexed this week that cut a real release inside the window.
- AI Applyd App
The app itself launched in April. This week's news is the MCP server, which went from 10 to 17 tools in v1.8.0 on Friday, including apply and review_application. That means a coding agent can now submit a job application. Read the review flow before you let it.
- Hemory App
First released in August. The Thursday update reorganises the app into Timeline, Activities, People and Agent tabs, and adds an alert when listening is interrupted. For an always-on recorder, knowing when it stopped listening matters more than any feature.