Skip to content

AI Weekly #6 — Aug 31–Sep 6, 2026

August 31 – September 6, 2026

Four frontier models in 72 hours — Fable 5.1, Gemini 3.8 Flash, Muse Spark 1.3 and GPT-6 Astra — and three of the four shipped the same weights twice, behind two different access gates.

Onei AI Editorial 10 min read How we make this

Four frontier models shipped inside 72 hours, and the tables of benchmark numbers are the least interesting thing about it. Three of the four vendors released the same model twice, under two different access policies.

Anthropic put out Fable 5.1 and Mythos 5.1 — by its own description the same underlying model, with Mythos restricted to vetted US organisations through cyber and life-sciences verification programmes. Google shipped Gemini 3.8 Flash generally and Gemini 3.8 Flash Cyber only through an application-gated Fairwind Program. OpenAI released GPT-6 Astra to a limited set of organisations first, then widened access, having rated it the first model to hit the Critical cybersecurity tier under its own Preparedness Framework.

That is a structural change worth naming. For two years a model launch was one artifact with one price. This week it was a capability and, separately, an access envelope — and the second is now a product decision the vendor makes at launch, not a policy it retrofits later. If you evaluate models, evaluate the tier you can actually buy, because the configuration in the blog post may not be the one your account can call.

The cost line moved further than the capability line

Read the four launches by price rather than by benchmark and a different pattern shows up.

Anthropic cut cache reads on Fable 5.1 to $0.25 per million tokens — a 75% reduction — while list input and output prices held at $10 and $50. Google priced Gemini 3.8 Flash at $0.75 and $3.75 per million through 31 December 2026, with both figures doubling on 1 January. Meta's pitch for Muse Spark 1.3 barely mentions capability: roughly 20% fewer tool calls and 25% fewer tokens than 1.2 on comparable coding work. And GitHub's HydraFusion preview reports a 67% cost reduction on TerminalBench 2.1 while scoring 4.9 points higher than a single frontier model.

Those are four different mechanisms — caching, promotional pricing, model efficiency, routing — all pointed at the same number, and it is not the price per token. It is the cost of one completed task. Per-token pricing is becoming a poor proxy for what an agent workload costs, because the token count is now something the vendor optimises too. Anyone still budgeting agents on a per-million-token rate card is measuring the wrong denominator.

Two of those mechanisms also carry a dated liability. An introductory price with a published expiry is a real line in next year's budget: anything built on Gemini 3.8 Flash costs twice as much on 1 January unless something changes.

A government took a position on training data

On 1 September the US Justice Department filed a statement of interest in the consolidated OpenAI copyright litigation, arguing that training a language model on copyrighted text is fair use and warning that a licensing requirement would concentrate model building among the few firms able to pay for it.

It decides nothing. A statement of interest is an opinion the court may ignore, and it does not touch the separate questions of how training data was acquired or whether outputs infringe. But it is the first time the federal government has taken a formal side in this fight, and it lands in the middle of every licensing negotiation currently running between a model vendor and a publisher. The negotiating posture on both sides of those tables changed on Tuesday, whatever the court eventually rules.

By the numbers

The directory took in 42 entries this week: 9 apps, 20 agent skills and 13 MCP servers. That brings the totals to 470 apps, 227 skills and 253 MCP servers.

The split is the part worth reading. Skills were almost half of everything added, and nearly all of them came from vendors publishing their own — .NET, AWS Labs, Snowflake, PlanetScale, Vercel, Replicate, Tinybird, Composio and MiniMax all shipped or expanded first-party skill repositories that reached us this week. A year ago agent skills were overwhelmingly community-written. They are turning into vendor documentation with an execution path attached, which is a better place for them: the company that changes the API is the company that should be updating the instructions for it.

The MCP catalogue tells a quieter version of the same story. Of the servers we added, the ones worth your time this week — Firebase, Backlog, Aave, Bird, Alpaca — are all official, published by the company whose API they wrap. Unofficial servers go stale the first time a field is renamed; official ones are maintained by the people doing the renaming.

What we added

Twelve entries below shipped upstream during the week. A note on the ones that did not: several vendor skill collections arrived in the directory this week but were published upstream months ago, so they are catalogued and searchable without being called new. Last week's issue covers what came before.

The week in AI

Model release

Anthropic ships Fable 5.1 and Mythos 5.1 — one model behind two access gates

Anthropic released Fable 5.1 into general availability and Mythos 5.1 as a restricted twin, limited to vetted US organisations through cyber and life-sciences verification programmes. List pricing holds at $10 per million input and $50 per million output tokens, but cache reads drop 75% to $0.25 per million, which Anthropic says works out roughly 25% cheaper than Fable 5 on typical workloads. Reported Terminal-Bench-Science 0.1 performance rises from 24.7% to 52.6%.

Why it matters: Capability and eligibility are now separate product decisions made at launch. Evaluate models on the tier your account can actually call, not the one in the announcement.

Anthropic

Model release

Google's Gemini 3.8 Flash arrives with a cyber twin behind an application form

Google published Gemini 3.8 Flash and Gemini 3.8 Flash Cyber on 2 September. The standard model runs at an introductory $0.75 per million input and $3.75 per million output tokens through 31 December 2026, after which both figures double. The Cyber variant is reachable only through the new Fairwind Program for government bodies, critical-infrastructure operators and software maintainers; Google's Chrome security team reports it produced 2.6 times more correct vulnerability patches than much larger commercial models.

Why it matters: An introductory rate with a published expiry date is a budget item, not a footnote. Anything built on 3.8 Flash doubles in unit cost on 1 January unless you move or renegotiate.

Google

Model release

Meta's Muse Spark 1.3 sells efficiency instead of a benchmark headline

Meta released Muse Spark 1.3 on 2 September through Muse Code and the Meta Model API. The claim it leads with is not capability but consumption: roughly 20% fewer tool calls and about 25% fewer tokens than Muse Spark 1.2 to finish comparable coding tasks. Weights remain closed, though Meta's stated roadmap includes a future open-weights release for the Muse Spark line.

Why it matters: Tool calls per completed task is the number that shows up on an agent bill and the one no public leaderboard tracks. Vendors publishing it themselves is a sign of where the competition has moved.

Meta AI Research

Model release

OpenAI launches GPT-6 Astra, its first model rated Critical for cyber capability

OpenAI released GPT-6 Astra on 3 September, first to a limited set of organisations and then to ChatGPT Plus, Pro, Business and Enterprise users, the API, Microsoft Azure and AWS Bedrock, at $10 per million input and $50 per million output tokens. OpenAI reports 72.6% on OSWorld 2.0 computer-use tasks against GPT-5.6 Sol's 65.7%, at roughly 40 minutes per task versus about 75, and says Astra is the first model to reach the Critical cybersecurity level under its Preparedness Framework.

Why it matters: Computer use has been the demo that never survived real workflows. A near-halving of time per task at a higher success rate is the first number that makes it worth retesting your own.

OpenAI

Policy

US Justice Department tells a federal court that training on copyrighted text is fair use

The Justice Department filed a statement of interest on 1 September in the consolidated OpenAI copyright litigation before Judge Sidney Stein in the Southern District of New York. It argues that training is "exceedingly transformative", that training alone causes no cognisable market harm, and that a licensing requirement would hand model building to the firms able to pay for licences. It is the federal government's first formal position on AI training and copyright, and it is not a ruling.

Why it matters: A statement of interest binds nobody, but it shifts the settlement maths for every publisher and author under the MDL and every licensing talk running alongside it.

US District Court, S.D.N.Y. (CourtListener) IPWatchdog

Ecosystem

GitHub puts multi-model orchestration into Copilot CLI as a research preview

Project HydraFusion, published 4 September, picks an execution pattern per request instead of a model per session: a single model, a cascade where a cheaper model drafts and a quality gate decides whether to escalate, or a critique loop where a model from another family reviews the draft before one revision. GitHub reports a 67% cut in estimated workflow cost on TerminalBench 2.1 alongside a 4.9-point quality gain against Claude Opus 5, plus 65% and 36% cost reductions on CheckpointBench and DeepSWE at near-parity quality. It runs on all Copilot plans via /experimental in the CLI.

Why it matters: If a vendor-run router beats a single frontier model on both price and quality, choosing one model per task stops being the sensible default.

The GitHub Blog

Funding

Gimlet Labs raises $300M to spread one inference workload across mixed silicon

Gimlet Labs announced a $300 million Series B on 4 September, led by Andreessen Horowitz and joined by Arm and Microsoft's M12 fund, six months after an $80 million round. Rather than pinning a model to one accelerator, the company splits a single inference workload across GPUs, CPUs, near-memory compute and dataflow architectures, and claims 5-10x speedups for the same power footprint. Reported valuation is around $3 billion.

Why it matters: Inference cost is the binding constraint on agent products right now, and capital is moving to the layer that makes a token cheaper without making the model worse.

Gimlet Labs

Product

Anthropic moves enterprise misuse monitoring into the customer's own cloud account

Enterprise Frontier Safeguards, announced 1 September, keeps zero data retention while letting activity-monitoring data sit in a customer's own AWS, Azure or Google Cloud account under their keys and access policies. When monitoring flags a pattern, the signal goes to the customer's reviewers instead of through Anthropic staff. Rollout is phased from autumn 2026 across Claude Code, Claude Enterprise, the Claude Platform, Amazon Bedrock, Google's Agent Platform and Microsoft Foundry, at no charge beyond the customer's own storage bill.

Why it matters: Zero retention and abuse monitoring have been sold as mutually exclusive. Moving the monitoring data to the customer's side of the line is how that trade-off gets unpicked in procurement.

Anthropic

New on Onei this week

Released during this window and now in the catalogue.

  • Kilo Code

    The JetBrains release is the part to look at: the same agent now runs in IntelliJ, VS Code and a terminal, so a team split across editors can standardise on one tool and one bill instead of three.

    Released Release notes Coding & Development
  • Monid
    Monid App

    Worth opening if you have ever abandoned an integration at its billing page rather than its docs. Per-call metering is a procurement argument as much as a technical one.

    Released Release notes Automation & Workflows
  • Articos

    Treat the output as a hypothesis generator, not evidence. The reason to try it is turnaround: half an hour is fast enough to run before a launch instead of arguing about it after.

    Released Release notes Marketing & Sales
  • Grove
    Grove App

    The design decision worth stealing is the single local socket. Your agent and your menu bar watch the same running process, so there is never a second answer to what is currently up.

    Released Release notes Coding & Development
  • dif
    dif App

    Keeping flags in the repo means a flag change goes through code review and rolls back with a revert. The generated context file is the piece that matters for agents: they read the flag set rather than guess at it.

    Released Release notes Coding & Development
  • claude-mem

    Most coding-agent setups paper over session memory with an ever-growing instructions file. Keeping the store local is the answer to the obvious objection about shipping your work into someone else's index.

    Released Release notes Development
  • Secure MCP Tunnel Client

    v0.0.14, out on 1 September, is the release that makes multi-replica deployments work — what you need if the tunnel is expected to survive a rolling restart rather than a demo.

    Released Release notes Developer Tools
  • Backlog MCP Server

    That Nulab publishes this themselves matters more than the tool count. An official server tracks the API as it changes; an unofficial one breaks the first time a field is renamed.

    Released Release notes Productivity
  • Firebase MCP Server

    It ships inside the Firebase CLI, so there is no second thing to keep current: the version you get is the version your CLI is on. That distribution choice is what most MCP servers get wrong.

    Released Release notes Databases & Data
  • Aave MCP
    Aave MCP MCP server

    Note what it declines to do: it hands back unsigned transactions. The agent reads positions and prepares a trade, the signing key stays with you, and that is the only sane shape for this category.

    Released Release notes Payments & Commerce
  • Bird MCP
    Bird MCP MCP server

    Messaging is the textbook irreversible action — you cannot unsend an SMS to forty thousand people. If you wire this in, wire in the approval step in the same sitting.

    Released Release notes Productivity
  • Alpaca MCP Server

    Paper-trading accounts make this one of the few finance servers you can genuinely experiment with. Start there and read the order-placement tools closely before you swap in live keys.

    Released Release notes Payments & Commerce

Browse the full catalogue →