AI Weekly #10 — Sep 28–Oct 4, 2026
September 28 – October 4, 2026
Google held its new frontier model, Gemini 4 Argon, back for cyber defenders while OpenAI made GPT-6.1 Sol near-flagship at a fifth of the price.
Two frontier models arrived this week, and the more telling of the two is the one most people cannot use yet.
Google announced Gemini 4 Argon on Wednesday and then held it back. The first users are security teams in its Fairwind Program, because Argon can find, validate and patch vulnerabilities on its own, and Google wants more time on the guardrails before anyone else gets it. OpenAI went the other way the day before: GPT-6.1 Sol, a week after GPT-6 Sol, priced at a fifth of GPT-6 Astra and available immediately. One lab is gating its top model; the other is making the tier just below it cheap enough that the top model matters less. Both are reasonable, and between them they describe the market better than either benchmark table does.
Defence first, by policy and by money
The Argon rollout was not the only place this week where the defensive side came first. Kevin Mandia's Armadin raised $255.5 million to run swarms of attacking agents against its own customers, which is offensive automation sold strictly as a defensive service. Google DeepMind published SynthID Bio, which watermarks AI-designed proteins so that a DNA-synthesis company can check whether an unfamiliar sequence came from a model with safeguards. All three assume the capability already exists and argue about who gets it first.
Washington produced its own version of that argument. Six labs signed a voluntary accord at the White House committing to internal controls, an independent auditor and a board committee to read the results. It has no deadline and no penalty. The executive order issued the same day is easier to mock — agencies must now say "Super Intelligence" instead of "AI" — but it carries the more consequential instruction: a proposed statutory definition within 60 days. Whatever that definition covers is what any future federal rule will reach.
The small models did the practical work
For anyone building agents, Cloudflare's Clef is the release to try this week. It is not a chat model. It takes an input and a fixed list of answers and returns one of them with a probability, which is what most agent steps actually are: route this ticket, pick the next tool, decide whether a page is relevant. A model that cannot answer outside the list cannot produce output your parser chokes on, and at 39ms median for the 9B variant it is cheap enough to call on every step. Weights are Apache 2.0, so it can run outside Cloudflare too.
By the numbers
The directory gained 12 entries this week — 3 apps, 7 agent skills and 2 MCP servers — and now lists 543 apps, 303 skills and 320 MCP servers.
None of the 12 launched for the first time this week. When we checked upstream, nine had no release inside the window at all. Three Hugging Face training skills went in at once, but they were added to Hugging Face's repository between November 2025 and May 2026. SAP's UI5 MCP server last shipped on 25 September, and Scrapling's last release was in August. All three apps were live before the window, as far as public records show, and one of them arrived with a release date of 4 October that no source supported, so we removed it rather than leave a guess on the page. That leaves the skills catalogue, not the MCP catalogue, as the part of the directory where things are still being built in public: two skills less than a fortnight old shipped real versions this week, and both are below.
What we added
The three picks below shipped a version between 28 September and 4 October. We dated each one against its GitHub release or version commit, not against the day we catalogued it. It is a short list because the week was short on releases, and we have not padded it.
The week in AI
Google announces Gemini 4 Argon, but starts with cyber defenders only
Google introduced Gemini 4 Argon on Wednesday as its new frontier model, aimed at long-horizon software engineering, legal and financial work, and vulnerability remediation. It raises the output limit to 1 million tokens from 64K and reports 77.9% on DeepSWE v1.1. Access opens first to vetted security teams in Google's Fairwind Program; developers, enterprises and consumers come later, at an introductory $2/$10 per million input/output tokens rising to $4/$20.
Why it matters: A frontier launch where the general release is deliberately deferred, and the first users are defenders, is now the pattern for models with strong offensive-security skills rather than an exception.
OpenAI's GPT-6.1 Sol undercuts its own flagship by a factor of five
At DevDay on Tuesday OpenAI shipped GPT-6.1 Sol, a week after GPT-6 Sol, positioning it close to GPT-6 Astra on agentic coding, computer use and professional work. The API lists it at $2 per million input tokens, $0.10 cached and $10 output, with a 1.05M-token context, 128K maximum output and five reasoning-effort levels from low to max. It is available in ChatGPT's paid work plans and Codex.
Why it matters: When the mid-tier model gets most of the way to the flagship at a fifth of the price, routing logic built around 'cheap model for easy steps' needs to be re-measured, not assumed.
Cloudflare open-sources Clef, small models that only pick from a list
Cloudflare released Clef (27B, built on Qwen 3.8) and Clef-flash (9B, on Qwen 3.5) under Apache 2.0 and put them on Workers AI. They are decision models: given an input and a fixed set of options, they return typed answers with probabilities instead of free text. Cloudflare reports median latency of 209ms for Clef and 39ms for Clef-flash against 524ms for Jev, a 64k context, image and video input, and Jev-API compatibility.
Why it matters: Most agent steps are classifications — is this urgent, which tool next — and a model that cannot answer outside the options removes a whole class of parsing failures from those steps.
Six AI labs sign a voluntary safety accord as a US order renames AI 'Super Intelligence'
On Tuesday the White House issued Executive Order 14434, which tells federal agencies to say 'Super Intelligence' instead of 'AI' in official communications and gives the President's science adviser 60 days to propose a statutory definition. The same day, leaders of Anthropic, OpenAI, Google, Meta, xAI and Nvidia signed a non-binding accord committing to four layers of oversight: internal controls, an internal team checking them, an independent external auditor, and an independent board committee.
Why it matters: The accord has no deadlines or penalties, but the 60-day definition is the piece with legal consequence: it decides what future federal rules will cover.
Kevin Mandia's Armadin raises $255.5M for agents that attack you first
Armadin, the offensive-security company founded by Mandiant's Kevin Mandia, raised a $255.5 million Series B co-led by Andreessen Horowitz and Accel, valuing it above $2.5 billion and taking total funding to $445 million seven months after leaving stealth. Its platform runs autonomous agents that chain individually low-severity weaknesses into validated attack paths, and it says it is in production with Fortune 500 and government customers.
Why it matters: Offensive automation is attracting the money in the same week a frontier lab gated its model to defenders; the two only make sense together if the defence side gets the same tools first.
DeepMind's SynthID Bio watermarks AI-designed proteins without breaking them
Google DeepMind published SynthID Bio in Nature: one method steers amino-acid choices to hide a signature in designed sequences, another fine-tunes part of AlphaFold 3's diffusion network to mark predicted structures. In wet-lab tests on binders for VEGF-A, the SARS-CoV-2 spike RBD and PD-L1, watermarked designs matched unmarked ones on hit rate, binding affinity and sequence diversity. Code, in vitro data and weights are being released to researchers.
Why it matters: It gives DNA-synthesis providers a way to check that an unfamiliar sequence came from a model with safeguards, which is the screening gap biosecurity people have been pointing at.
New on Onei this week
Released during this window and now in the catalogue.
- GoLive Skill
Alpha 8 on Saturday added read-only launch checks: one confirms a real live Stripe payment and a drained webhook queue without charging anything, another audits titles and Open Graph tags. The notes say they were tested against mocks, not live accounts, so treat it as alpha.
- Anidoodle Skill
Version 0.6 on Sunday, under two weeks after 0.1, gave scores real recorded instruments (piano, harp, marimba, timpani) as an opt-in download, and renders one timeline at 16:9, 1:1, 4:5 and 9:16 — the useful part if you cut the same clip for several feeds.
- Logo Design Skill
A small release, listed as one: 1.4.4 on Tuesday fixes presentation boards that printed handles like @marlow&finch for brand names with punctuation or accents. It came three days after 1.0, which says more about how fast this skill is moving than the fix itself.