Record — What changed, and when

Changelog

Entries here are updated in place: a model family moves to its next version on the same card, and an agent that gets renamed keeps its entry. That keeps the pages readable but erases its own history, so this is where the erased part is kept — particularly the replaced lines, which are the only record that a card used to say something else.

October 2026
8 Oct

MAI-Code-1.1-Flash parameter discrepancy noted

updatedModel Registry
Microsoft's Windows blog gives 137B total and 6.8B active for the on-device build, against the model card's 138B and 5B. The card keeps the model card's figures and records that the two disagree.
8 Oct

Step 5 Preview, weights pending

addedModel Registry
StepFun's new 600B / 27B-active flagship is API-only at $1 / $2.70, with open weights promised for 15 October and no licence named, so it carries the Weights pending licence. The Step 3.7 Flash card's link also now points at its repository instead of the Hugging Face home page.
8 Oct

DeepSeek V4 Flash → DeepSeek V4.1 Flash

replacedModel Registry
A new architecture, a causal encoder-decoder with an FP4 KV cache, in a 763B MIT checkpoint, against V4 Flash's 284B. It takes image input natively and is priced higher, at $0.30 / $1.20 peak. The old card also dated V4.1 Flash to June; DeepSeek and Hugging Face both say 10 September.
8 Oct

DeepSeek V4 Pro API status

updatedModel Registry
DeepSeek says V4 Pro is being phased out and has sent deepseek-v4-pro requests to V4.1 Flash since 14 September, while its pricing page still lists V4 Pro at $1.32 / $3.96 peak. The card now records both rather than picking one. The MIT weights are unaffected.
8 Oct

MAI-Code-1-Flash → MAI-Code-1.1-Flash

replacedModel Registry
138B / 5B active, adds image input, and costs $0.20 / $1.20 against $0.75 / $4.50, after GitHub deprecated 1.0 on 10 September. Microsoft says it now runs on-device too, but has published no weights or licence, so it stays proprietary here.
8 Oct

Foundation-Sec-8B → Foundation-Sec-1.1-8B-Instruct

replacedModel Registry
Cisco's version 1.1, an instruct model with context raised from 4K to 64K. Its repository states the dual licence, Llama 3.1 plus Apache 2.0 for Cisco's changes, that the original's Apache tag left out. The 1.1 base is not public.
8 Oct

Teuken-7B v0.4 → v0.6

replacedModel Registry
Re-pretrained on 6T tokens across the 24 EU languages, under CC-BY-NC-4.0. The commercial Apache 2.0 option exists only for v0.4, and the card says so.
8 Oct

Falcon 3 licence named as custom

fixedModel Registry
The Falcon3 repositories read falcon-llm-license, a custom agreement. The card now says that instead of implying a standard open licence.
8 Oct

GPT-Daybreak-Blue alias deprecated; Nova Premier wind-down

updatedModel Registry
OpenAI now marks gpt-daybreak-blue-latest deprecated, with no named replacement. On Nova, AWS is reported to be moving Premier to maintenance-only, but its lifecycle page gives Premier no end-of-life date, and the card records both.
8 Oct

GitHub Copilot picker now names MAI-Code-1.1-Flash

updatedCoding Agents
GitHub deprecated MAI-Code-1-Flash across Copilot on 10 September in favour of 1.1.
8 Oct

MiMo-V2.5-Pro → MiMo-V2.6-Pro

replacedModel Registry
The successor keeps V2.5-Pro's 1.02T/42B backbone, its 1M context and its $0.435/$0.87 price, and Xiaomi retires V2.5-Pro on 21 October. The licence is MIT, read from both repositories. The change is post-training: one mixed RL run that includes cybersecurity, where Xiaomi reports CyberGym rising from 40% to 94%.
8 Oct

MiMo-V2.5-Pro and Kimi K3 prices re-checked

updatedModel Registry
Both re-read against the vendors' own price sheets. Kimi K3 is unchanged at $3/$15 and gains its $0.30 cache-hit rate. MiMo-V2.5-Pro is now sourced first-party at $0.435/$0.87, and its pricing note records that Xiaomi deprecates it on 21 October.
8 Oct

Claude Haiku 5.5

addedModel Registry
Anthropic's first Haiku with an effort setting. It is also the first Claude model priced by prompt length: $0.10/$0.50 up to 100K tokens and $0.50/$2.50 above, against Haiku 4.5's flat $1/$5. The Fable 5.1 and Opus 5.5 cards now name Sonnet 5.5 and Haiku 5.5 as their current siblings.
8 Oct

Claude Sonnet 5.5

addedModel Registry
The registry's first Sonnet entry, added ten days after the 28 September launch. It keeps Sonnet 5's $2/$10 rate, and its cache reads were halved to $0.10 alongside Haiku 5.5. It is the first Sonnet to ship with Opus-class cyber safeguards, which hand higher-risk security tasks back to Sonnet 5.
7 Oct

Mistral Large 4 and Reflection Beam, weights pending

addedModel Registry
Both are announced as open weights due by the end of October and neither has published them, so both carry a new Weights pending licence: no licence recorded until there is a repository to read it from, and not counted as self-hostable. Large 4 is a priced public preview; Beam is a free, waitlisted beta.
5 Oct

Kolibri-1

addedModel Registry
Aleph Alpha's German-English MoE under Apache-2.0: 78.1B total, 3.46B active, trained from scratch in Germany and Finland. It needs Aleph Alpha's own vLLM plugin to serve, and its 1M context is an extension past a native 262K.
4 Oct

Jev, Clef and Clef-flash

addedModel Registry
The first decision models in the registry: they return a probability for each declared option instead of generating text. Jev is TypeSafe's hosted original; Clef and Clef-flash are Cloudflare's Apache-2.0, API-compatible answers on Qwen bases. A new decision model architecture value marks all three.
1 Oct

Gemini 4 Argon

addedModel Registry
Google's first frontier model since 3.1 Pro, carded as gated because Fairwind is today the only way in: a subset of partners get it without cyber guardrails, ahead of paid API and Ultra access Google has not dated. The announced introductory price is $2 in / $10 out, rising to $4 / $20, and the output ceiling moves to 1M tokens.
1 Oct

CodeMender's Fairwind model is now Gemini 4 Argon

updatedCoding Agents
The Fairwind page was rewritten around Argon and no longer names Gemini 3.8 Flash Cyber, so the CodeMender card now names Argon. The Flash Cyber card stays, with its deploy line saying the programme page has dropped it and Google has not said whether partners keep it.
September 2026
30 Sep

Claude Managed Agents and OpenAI Agents API

addedAgent Runtimes
The two labs' hosted agent platforms, carded together so they compare on the same fields. Both run the agent loop in the vendor's cloud and execution in a separate sandbox, both are filed as sandboxes, and neither discloses what isolates its hosted container. Both also default to open egress and exclude Zero Data Retention.
30 Sep

Cursor is owned by SpaceX

updatedCoding Agents
The card still named Anysphere alone, six weeks after SpaceX closed its $60 billion acquisition on 14 August 2026. Anysphere remains the operating entity, so the vendor now reads Anysphere (SpaceX), and Grok joins the model list, where the registry already recorded it on every Cursor plan.
30 Sep

OpenShell

addedAgent Runtimes
NVIDIA's open-source agent sandbox, the software half of the Open Agent Safety Platform announced on 28 September. It qualifies on mechanism: Landlock and seccomp inside a container or microVM, with socket calls brokered outside the workload so network policy cannot be bypassed by ignoring a proxy variable. Sentry, the platform's BlueField-4 watchdog, is a hardware reference design and gets no card.
30 Sep

Daybreak Blue did reach GPT-6 Sol and Luna

updatedModel Registry
The GPT-6 Sol card, the Blue card and the changelog entry for Sol's move to GPT-6 all said nothing extended Blue to the new models. The GPT-6 system card's Sol and Luna appendix, added the same day, says eligible Blue users get reduced cyber refusals on both. Luna now carries the Blue gate; Sol has since moved to 6.1, for which nothing names a tier, and the Blue alias itself still resolves to GPT-5.6 Sol.
30 Sep

dots

addedPersonal Agents
OpenAI's always-on personal agent, on GPT-6 Astra, launched at DevDay. It gets a card beside ChatGPT Work rather than inside it because the shape differs: its own cloud computer per dot, persistent learning, reachable from Slack, Teams and email, working unprompted in the background, and able to make purchases and use your own computer once you connect it.
30 Sep

GPT-6 Sol → GPT-6.1 Sol

replacedModel Registry
A week after GPT-6 Sol, OpenAI shipped 6.1 at the same $2/$10 with cached input halved to $0.10, and in place of the GPT-6.1 Astra it did not release. The change that matters here is the cyber rating: the system card addendum puts GPT-6.1 Sol at Critical, where GPT-6 Sol was not, making it the second model after Astra to cross that threshold.
22 Sep

Daybreak Blue is back, as its own card

addedModel Registry
Blue lived on the GPT-5.6 Sol card as an access field, and left the registry when that card became GPT-6 Sol this morning — the tier is still served, still behind GPT-5.6 Sol, and nothing extends it to the new model. OpenAI names, prices and documents the alias as a model of its own, so it now has a card as an access tier, beside GPT-5.6-Cyber for Red. The same pricing page gave GPT-5.6-Cyber a public rate of $12.50/$75, which its card had recorded as unpublished.
22 Sep

GPT-6 Astra's Daybreak gate opened, for Red only

updatedModel Registry
The card said reduced-safeguard access to Astra was coming through Daybreak in the following weeks. OpenAI's Daybreak help article now says it has arrived for Red customers and not for Blue, whose accounts get Astra with standard safeguards, and whose alias still points at GPT-5.6 Sol. The card's tier is now Red. The same article names an Astra Minor, which has no announcement, API page or system card yet, so it gets no card.
22 Sep

Two card names lose their parentheses

updatedModel Registry
Claude (Fable 5.1) is now Claude Fable 5.1, and Muse Spark (1.3) is now Muse Spark 1.3. The parentheses date from the first commit, when a card stood for a whole family and the bracket marked the part that got updated in place. Both families now have more than one card — Opus 5.5 beside Fable, Glimmer beside Spark — so the old form read as if the sibling were not part of the family. The anchors do not change, because the slug drops the punctuation either way.
22 Sep

Claude Opus 5.5

addedModel Registry
The registry had never carried an Opus card, and Opus 5.5 earns one: Anthropic's own docs now point most workloads at it rather than at Fable, at $4/$20 against Fable's $10/$50. It is also the first Opus to ship with Fable-class cyber and biology safeguards, so it carries the Life Sciences Verification gate, with the cyber programme recorded as announced rather than open. Fable 5.1's sibling list now names Opus 5.5 in place of Opus 5.
22 Sep

GPT-5.6 Sol → GPT-6 Sol

replacedModel Registry
Sol moves a generation and down a rung: GPT-6 Astra is the flagship now, and Sol is the cost tier beneath it, at $2/$10 with a 1.05M context and an April 2026 cutoff. The Daybreak Blue gate did not carry over, because nothing published extends it to the new model, and GPT-5.6-Cyber loses its base link — it was built on GPT-5.6 Sol, and pointing it at the card that replaced it would assert a lineage it does not have.
22 Sep

GPT-6 Luna

addedModel Registry
OpenAI's efficiency tier gets a card, at $0.10/$0.50 with the same 1.05M context as Sol; its GPT-5.6 predecessor never had one. It is also the GPT-6 model that reaches Free and Go users, through the desktop app.
22 Sep

The vendors are searchable, and the page stops being the exception

updatedSystem Cards
Twenty-four vendors across three sections with no way to find one: the page was the only one on the site with no search, which is also why the sitebar's ⌘K hint was missing here — the shortcut is only shown where there is a box to focus. It has the query line now, without the facet row, since the page has no facets to offer. The search blob is filled from each card's own text at runtime rather than written into the markup: these vendors are hand-written HTML rather than entries, so nothing builds a blob for them, and deriving it means a card edited in place is searchable by what it now says. Sections whose vendors all filter out hide their headings, and a search matching nothing says so instead of leaving the page blank.
22 Sep

All five layers are defined now, not the three that fitted in a row

updatedModel Runtimes
The layer definitions moved into the margin beside the intro, and the move exposed what the old row of three boxes had been hiding: the filter has always offered orchestrators and hosted providers — six of the twenty-three entries — while the page defined neither, and a numbered list that stops at three reads as complete. Layer 4 is packing, versioning and scheduling above the engine; layer 5 is open weights on someone else's metal, where the engine, the hardware and the retention policy are all theirs. Both are written from the entries filed under them rather than from a fresh claim. The weight-footprint paragraph moved up into the masthead, where it is a note about reading the registry's numbers rather than a lead-in to a list it does not depend on — and the list it introduced said “three things” above four of them, which is now four.
22 Sep

The two roles moved into the margin

updatedAgent Runtimes
Supervisors and sandboxes were a row of boxes under the intro; they are now annotations beside it, in the same format System Cards uses for its document types. Neither is in accent. Keeping the two apart and equal is this page's whole argument — “it runs my agents safely” is usually two products and people buy one — so marking either as the subject would answer the question the page exists to ask. The masthead gains two actions: into the entries, and across to Model Runtimes, which the intro has always had to name in its second sentence.
22 Sep

The document types moved into the margin

updatedSystem Cards
The page opened by defining its vocabulary in a boxed list, which put the index of vendors a screen and a half down — the definitions are reference a reader consults while reading the page, not steps they read in order. They are now annotations beside the intro: System card, Model card, Technical report and Retention policy, each keeping the claim that separates it from the other three. The examples that used to follow them did not survive the move; the format holds a claim, not a paragraph, which is the trade it makes for keeping them in view. The AI Act note stays prose rather than becoming a fifth annotation, because it is an argument rather than a definition — it sits in the spine, where it also answers the layout: four definitions run longer than three paragraphs, and the column beside them has to be given something true to say rather than padding. The masthead gains two actions: into the index, and out to the Commission's framework. The pattern is in the shared stylesheet, and Model Runtimes and Agent Runtimes took it the same day for their layer and role boxes.
22 Sep

The filters became one query line, and a filtered view became a link

updatedSite
Seven stacked rows of chips asked a reader to reconstruct the current query by scanning which chip in each row had gone dark, and they grew a row per facet — the registry was up to seven. There is now one panel: the query as a line of tokens, a row of controls that add to it, and what it currently matches. The controls are native selects sized to their own labels, so the dropdown, the keyboard handling and the mobile picker are the browser's rather than code here, and each one snaps back to its name after a pick because the tokens, not the controls, are where the query is written down. The query is also the URL — ?domain=cybersecurity&license=open-weights, spelled exactly as the tokens are so a filter you can see is one you can type, and with unknown values dropped rather than trusted — so a filtered view is a link, which is what the Copy query URL button hands you; queries can also be named and kept in the browser. The panel needs script and now says so by not being there without it, as the ⌘K hint does, and the full card list below it is unchanged either way. Pages declare facets rather than markup, so a new facet is a line in a list.
22 Sep

The nav reads as a catalogue, and the search box has a key

updatedSite
Seven sections in a row of identical uppercase labels gave a reader no way to say where they were except by reading all seven. Each now carries its number, and the current one sits in brackets with the number in accent alongside the label, so position is visible before the text is. The bar also ends in a ⌘K hint that focuses the filter box — rendered hidden and revealed by the same script that binds the key, because a shortcut advertised to someone with JavaScript off is chrome that lies, and System Cards has no filter box so it shows none. The bar's layout and the surface behind it moved into the shared stylesheet: both had been copied per page, and the surface had only ever been copied onto the registry, so the bar sat on fog on the other six. That also restores the layer boxes on both runtimes pages, which pick the fog token to contrast against a surface that was not there and had been fog on fog since those pages were built.
22 Sep

Antares was never purpose-built, and three cyber cards said so wrongly

updatedModel Registry
A sweep of the remaining cybersecurity entries against their repositories rather than their announcements. Cisco Antares was filed purpose-built with no base model; its own technical report says all three sizes are initialized from IBM Granite 4.0 and post-trained, so it is a fine-tune — and it gets no baseRef, because the registry carries Granite 4.2 and linking 4.0 to it would invent a lineage. Its licence is Apache 2.0, not the “vetted access” the card claimed: Hugging Face's gate is the automatic kind, a click-through nobody reviews, which 35,584 downloads a month makes plain. Foundation-Sec-8B cited a NOTICE file for its dual-licence reading and that file does not exist — the card says Apache 2.0 flatly, so the entry now states that and the unresolved Llama 3.1 question separately instead of attributing its own inference to the repo. Trend Cybertron had no link at all because the weights are published as Llama Primus, a name the product never uses. Two cards also pointed at huggingface.co, the website, which is not a citation.
Cisco Antares: purpose-built → fine-tune, base IBM Granite 4.0, context 128K (32K on the 350M), Apache 2.0, gate re-described, 3B recorded as announced-not-shippedFoundation-Sec-8B: link fixed to fdtn-ai/Foundation-Sec-8B, the phantom NOTICE citation removed, and Foundation-Sec-8B-Reasoning noted — built on this model and licensed `other`, not ApacheTrend Cybertron: link added, MIT named, base corrected to Llama 3.1 8B Instruct, the unshipped 70B recordedVerified unchanged: Sec-PaLM 2, Sec-Gemini v1 (still trusted-tester 17 months on), Claude Mythos 5.1, GPT-5.6-Cyber, Gemini 3.8 Flash Cyber, MAI-Cyber-1-Flash
22 Sep

The Fugu Cyber card listed a runtime it says it is not on

fixedModel Registry
Its deploy line has said “not on OpenRouter” since it was added, while a stored deployRef put an OpenRouter chip on the face of the same card — the two halves of one entry contradicting each other, and the chip is the half a reader sees first. Re-checked against OpenRouter's model list: Fugu Ultra, Fugu Ultra v2, Fugu Max and Namazu, still no Cyber tier, so the ref is gone rather than the sentence. The re-check also dated the rest: Sakana now sells four tiers, not three — Fugu Max arrived in September at $2 in / $6 out and is on the parent card with the subscription line that does not cover Cyber — and Fugu Cyber turns out to be gated by a sales conversation rather than a vetting programme, which on a page where every other cyber model sits behind Daybreak, Fairwind or a trusted-tester list is worth saying out loud. Sakana's own benchmark comparison is to GPT-5.5-Cyber and Mythos-Preview, both a version behind what the registry carries; the figures stay, with that attached.
22 Sep

Aikido Altar-1, a cyber model made by deleting experts

addedModel Registry
The first entry here whose domain specialization came from a prune rather than a training run: GLM-5.3 with 88 of its 256 routed experts per layer removed by REAP and the rest quantized to INT4, so it serves on four H200s inside a customer's own network. It earns a card because Aikido published the trade honestly — 60.4% recall on its 32-CVE benchmark against the unpruned model's 61.5% and full precision's 65.6% — and because “open-weight sovereign” is doing work in the announcement that the licence does not support: Hugging Face says `other` and the card inherits Z.ai's model-specific terms.
21 Sep

The filter contract is checked by the build

schemaSite
The filters broke for a month without anyone noticing, so the thing that catches it is a build failure rather than a line in the maintainer notes. After every build the emitted pages are read back: any page shipping chips must also ship the `[hidden]` override that lets a hidden card actually disappear, nothing else may claim `display` with `!important`, and filters.ts must still hide by the attribute those two checks are guarding. The last one is the point — a guard whose premise has moved passes quietly, which is how the original bug survived. All three were confirmed by reintroducing the bug and watching the build fail.
21 Sep

Every filter on the site was a no-op

fixedSite
Reported, not noticed: filtering the registry to cybersecurity changed the count and left all fifty cards on screen. The chips, the state and the counter were all correct — filters.ts hides a card by setting the `hidden` attribute, and every page then gave `.card` a `display` (`flex` on the card grids, `grid` on the changelog entries), which outranks the browser’s own `[hidden] { display: none }` and re-showed everything. It has been broken on all six filtered pages since the Astro rebuild in August, when the old code that re-rendered the grid was replaced by show-and-hide; the counter agreeing with the chips is what made it look like it worked. One rule in the shared stylesheet fixes every page.
21 Sep

Grok 4.6 → Grok 4.7

replacedModel Registry
The 4.6 card had flagged a successor Musk described in advance as 2.1T parameters with supplemental SpaceX engineering data. It shipped on 21 September 2026 and the launch post says neither thing — only that the base model is larger — so the new card carries no parameter count at all, which is the point of recording the claim as a claim. Price, context and modalities are unchanged; the knowledge cutoff moves to May 2026, the fast variant retreats to Cursor and Grok Build only, and Terminal-Bench nearly doubles to 38.0% while still trailing Fable 5.1 by twenty points. The vendor field changes with it: xAI joined SpaceX in February 2026 and now publishes as SpaceXAI, though the API and keys still say xAI.
21 Sep

CC, Google's household agent

addedPersonal Agents
Google Labs turned CC from a single-user briefing agent into a shared one on 17 September 2026, and the shared part is the finding. It holds its own verified Google account, so it is addressable — members forward it mail and Chat messages, and senders they nominate are auto-shared from then on, which is an untrusted-input path running straight from a school or a vet's mailbox into an agent with calendar and Drive write access. Its memory is the first on the page that pools several people, and it runs on Antigravity, the harness this site already lists as a coding agent. Contained by default on the strength of Google's own wording: an isolated cloud computer per instance, plus a distinct identity — which bounds the compute, not what the sharing settings let the account read.
21 Sep

ZCode, Z.ai's own harness

addedCoding Agents
Z.ai opened ZCode under Apache-2.0 on 20 September 2026, at desktop version 3.14.1 — a harness that already shipped, not a launch. It lands as byo-model rather than a vendor client: eight built-in providers span four competing model vendors and a router, and a custom provider takes any Anthropic- or OpenAI-compatible endpoint, which the docs say explicitly includes a self-hosted service on a private network. The card leads with what the project's own NOTICE.md admits — no default operating-system sandbox in the shared execution adapter, and `yolo` as the fallback mode for a non-interactive `--prompt` — because that is the fact a reader needs before pairing it with the WeChat, Feishu and Telegram remote-start surface.
16 Sep

Deploy lists checked against the runtimes, not the vendors

updatedModel Registry
Thirty-two of fifty entries now carry a deployAsOf, each one checked where the field says to check — the Ollama library's own tag list, OpenRouter's live model list, the Hugging Face repository — rather than against vendor availability copy. Three claims did not survive it: Sakana Fugu Cyber is not on OpenRouter at all, Muse Glimmer had a hosted route the card never mentioned, and Granite 4.2, Phi-4 and Falcon 3 each have a confirmed Ollama tag the card omitted. Two licences were wrong in the quiet way: Foundation-Sec-8B is dual-licensed and inherits Meta's Llama 3.1 Community terms, which "open weights" concealed, and Inkling is 952B by its own checkpoint against the 975B its announcement claimed.
16 Sep

Gemini 3 → Gemini 3.1 Pro

replacedModel Registry
Google's own model catalogue lists the Pro line as Gemini 3.1 Pro against endpoint gemini-3.1-pro-preview, so the card follows it. The prices are identical to Gemini 3's and the card keeps them; what changes is the date, the modality list (PDF input, 64K output ceiling) and the honest status — Google has left it marked Preview with no free tier since February. The contrast the entry now draws is the point: the Pro line has sat on a preview point release while Flash shipped three versions in six weeks.
16 Sep

Mistral Small 3.1 → Mistral Small 4

replacedModel Registry
Chasing the oldest price on the page found a model two versions behind. Small 4 is a different animal: 119B sparse against 3.1's 24B dense, with ~6B firing per token, moving the self-hosting floor from a 24GB card to an 80GB one even though the active count looks smaller, because memory needs the total — which is exactly what the card's derived footprint prints. It folds Instruct, Reasoning (formerly Magistral) and Devstral into one model with mode switching, takes 1M context, and stays genuinely Apache 2.0 with no added conditions, verified at the repository. Price moved $0.10/$0.30 to $0.15/$0.60.
16 Sep

Four stale prices re-sourced

updatedModel Registry
Every price stamp older than a quarter, checked against the vendor rather than a mirror. Three were unchanged and only the stamp had rotted — Gemini 3 at $2/$12 below 200K, Step 3.7 Flash at $0.20/$1.15, MAI-Code-1-Flash at $0.75/$4.50 — and all three gained the cache-read rates their vendors publish but the cards had as null. Two findings came with them: Google's Pro line now lists Gemini 3.1 Pro Preview, correcting a note that said no Pro model had shipped, and GitHub's plan docs list a MAI-Code-1.1-Flash that Microsoft has published no specification for, so that card says so rather than guessing.
16 Sep

Kimi K3 is not open-licensed, and its checkpoint did land

fixedModel Registry
The entry asked itself to verify whether the full checkpoint had shipped; it did, on 2 September — 96 shards, 2.78T parameters, ungated. Checking the repository also caught the registry's most repeated error one more time: the licence was recorded as a plain open-weight release and is actually a model-specific kimi-k3 agreement. It is MIT-shaped and free for almost everyone, but a Model-as-a-Service business past $20M revenue needs a separate agreement with Moonshot, and a product past 100M users must credit the model in its UI. Deploy list stamped after checking Hugging Face and OpenRouter.
16 Sep

Cross-page references, checked by the build

schemaSite
The relationships were already in the data as prose — 36 registry deploy lines naming a runtime, variants naming the model they came from, sandboxes naming the agents they run — and every one of them died at the page boundary. They are now explicit reference fields rendered as links, with an anchor on every card so there is something to link to. Stored rather than matched out of the text on purpose: "Gemini 3" is a substring of "Gemini 3.8 Flash Cyber", and a wrong link is worse than none. The build resolves each one against the other collection, so renaming an entry fails the build instead of leaving dead links; the reverse index on Model Runtimes is derived from the same data rather than stored twice.
11 Sep

Docker Sandboxes — a hypervisor boundary beside the kernel ones

addedAgent Runtimes
The page's first proprietary entry and its first microVM: a heavier wall around a machine the agent then completely owns, against nono's lighter per-tool policy on your real one. Its credential handling is the best recorded here — keys never enter the VM, a host proxy injects the headers — and its ceiling comes from Docker's own docs, which disclose that the default sbx run direct-mounts your working tree with no boundary at all, that hard links defeat workspace policy, and that the shared skills store is writable across sandboxes.
11 Sep

Agent Runtimes split out from the serving page

addedSite
Herdr and nono had been added as layers on the runtimes page, and between them proved the schema did not fit: formats, concurrency, governance and hardware were empty or reused on both. They now have a page whose fields answer their own question — enforcement (what does the containing, read from source rather than a README, with none as a real answer), contains, credentials and platforms. The old page becomes Model Runtimes; nothing about the serving layers changed. The two entries that prompted this stay filed under Model Runtimes, because that is where they landed — this log records where a change happened, not where the thing ended up.
11 Sep

nono, and a sandbox layer beside the supervisors

addedModel Runtimes
The second layer today that never serves a model, added because the Herdr entry had just finished saying containment arrives from somewhere else and could not name where. nono sandboxes the agent and, separately, each tool the agent shells out to — gh receiving a scoped token through a proxy rather than the agent's own credentials. The ceiling comes from reading the source rather than the pitch: enforcement is Landlock on Linux and Seatbelt on macOS, so it weakens with kernel ABI instead of refusing, and Windows means WSL2.
11 Sep

Herdr, and a supervisor layer above the runtimes

addedModel Runtimes
Herdr runs coding agents rather than models, so neither existing page could hold it honestly — the agents page would have called it an agent, and the five serving layers all assume a model somewhere. It gets a sixth layer instead, and the page premise widens from what serves the models to what runs the stack. The ceiling is the reason it is worth a card: it owns terminals, not behaviour, so nothing about it contains an unattended agent, and the persistence is narrower than the pitch — a restart returns the layout, not the processes. The notes record where containment does come from, since it is not Herdr: sandbox wrappers it accommodates, and an E2B plugin that mirrors a working tree into a cloud box per agent — borrowing local harness credentials, and running each agent with approval prompts skipped because nobody is watching the pane.
9 Sep

Nova pricing re-sourced from Bedrock itself

updatedModel Registry
The figures now come from AWS rather than an OpenRouter route, which corrected Micro to $0.035 and added latency-optimised Pro. It also surfaced something no other entry has: Bedrock prices the same model by region and inference profile, so Nova 2 Lite runs $0.15 to $0.5775 and Premier reaches $21.875 output. Nova 2 Omni, in preview with audio input, appeared in the same tables.
9 Sep

Amazon Nova

addedModel Registry
The last major cloud vendor missing from the registry, listed for its understanding tier only — Canvas, Reel, Sonic and the embeddings model are not language models. Two limits are recorded that the marketing does not lead with: a 10K output cap on the Nova 1 tiers against 64K–131K elsewhere, and Bedrock Guardrails applying to text only on the multimodal tiers.
9 Sep

architecture — recording models that are not ordinary transformers

schemaModel Registry
Mercury 2.5 exposed a gap: nothing on a card said it generates text a fundamentally different way. The field marks departures only, ordered by size — diffusion, hybrid Mamba-transformer, then sparse-attention variants — so the paradigm difference is not flattened into an attention-operator detail. Omitted means standard, so 44 entries needed no edit.
9 Sep

Mercury 2.5 — the registry's first diffusion model

addedModel Registry
Every other entry generates tokens sequentially; Mercury refines many in parallel, and sells on latency rather than intelligence. At $0.04 / $0.15 on OpenRouter it also undercuts the previous floor, though Inception's own site still lists $0.20 / $0.75 — the note records both.
9 Sep

Muse, Meta's personal agent

addedPersonal Agents
The first entry whose reach leaves software entirely — smart home devices and vehicles alongside mail, browsing and purchases, so both were added to the hot reach set. It is also the first isolation: default that reads like engineering rather than policy: nspawn containment, a single egress authority, credentials the agent never sees, and a browser sub-agent given accessibility trees rather than DOM. Meta prices the page's core risk at up to $130,000 for a prompt injection affecting one user.
8 Sep

GLM-5.3-Flash is MIT, and two local-model entries sharpened

updatedModel Registry
A practitioner's CTI comparison caught an error: GLM-5.3-Flash's licence had been inferred from GLM-5.3 and recorded as model-specific, when Hugging Face says MIT. Qwen3.8-Flash-Next gains its real licence name, ~6B active, and the fact that ~51B of its total is an n-gram table meant to sit off the GPU — so the derived footprint overstates VRAM. Both it and DeepSeek V4 Flash now carry an independent throughput measurement rather than only vendor benchmarks.
5 Sep

MiniMax-M3

addedModel Registry
A gap found while checking a Saudi sovereign-AI story: MiniMax was missing entirely. ~428B total with ~23B active, natively multimodal, 1M context on sparse attention claiming 9x prefill and 15x decode speedups over M2. Its custom community licence is not Apache or MIT. humain-m3, the Arabic model built on it, gets no entry — there is no public repository to verify.
4 Sep

GPT-6 Astra — the first Critical cyber model

addedModel Registry
Shipped broadly, but its advanced cyber capability did not: as launched it refuses proof-of-concept exploit work, with less restrictive safeguards promised through Daybreak in the coming weeks. Measured without production safeguards it scores 100% on ExploitBench and found two previously unknown zero-days during evaluation. The GPT-5.6-Cyber card, which had carried Astra as forthcoming, now points at its entry.
3 Sep

Muse Spark 1.2 → 1.3

replacedModel Registry
An efficiency and judgement release: roughly 20% fewer tool calls and 25% fewer tokens for the same work, plus stronger prompt-injection resistance and better calibration of irreversible actions. Max reasoning is withheld pending further safety testing, and Meta lists a Muse Spark open-weights release on its roadmap — which would partly reverse the closure the original release represented.
3 Sep

Crush

addedCoding Agents
Charm's Go terminal agent: reads your language servers for context, switches model mid-session without losing it, and takes any OpenAI- or Anthropic-compatible endpoint including local Ollama. Tagged open with a caveat in the notes — FSL-1.1-MIT is source-available rather than open source, converting to MIT two years after each release.
3 Sep

CodeMender is only half-gated now

updatedCoding Agents
Google's Fairwind page splits its availability: inside the programme it pairs with the gated Gemini 3.8 Flash Cyber for governments, critical infrastructure and core platform companies, but any Google Cloud customer can run CodeMender against publicly available models. The harness is no longer the gated part — the model is.
3 Sep

Fairwind's terms recorded in full

updatedModel Registry
More than 650 partners, and admission carries operational conditions rather than vetting alone: access confined to internal security, incident response or penetration testing teams, with protections such as MFA required. Membership grants the CodeMender harness alongside the model.
2 Sep

Gemini 3.7 Flash → 3.8 Flash, and 3.5 Flash Cyber → 3.8 Flash Cyber

replacedModel Registry
Google's third Flash release in six weeks, at the same introductory $0.75/$3.75 running to 31 December. The cyber variant's gate changed with it: the CodeMender pilot is replaced by the new Fairwind Program, which widens trusted-defender access from governments and partners to critical infrastructure operators and software maintainers.
2 Sep

Claude siblings corrected, Astra and Grok 4.7 noted

updatedModel Registry
Checking a landscape summary against Anthropic's docs found the Fable card still naming Opus 4.8 and Sonnet 4.6 as siblings, where the lineup is now Opus 5, Sonnet 5 and Haiku 4.5; output limit and model ID added too. GPT-5.6-Cyber records that OpenAI has classified the unreleased Astra as its first Critical cyber model, and Grok 4.6 records its announced successor.
1 Sep

Claude Fable 5 → 5.1, Mythos 5 → 5.1

replacedModel Registry
Headline token prices are unchanged, but cached reads fell 75% to $0.25, so the saving is in reuse rather than per token. Mythos gets a published price for the first time. Its gate also changed name: this announcement describes the Cyber and Life Sciences Verification Programmes and never mentions Project Glasswing, and cyber access to Mythos-class is stated as coming rather than open.
August 2026
31 Aug

deployAsOf — when the deployment list was last verified

schemaModel Registry
Deployment lists go stale in a way nothing else on a card does: a runtime library gains or drops a model and no vendor document changes, so the entry stays wrong and looks fresh. This is the same discipline the pricing asOf already imposes, applied to the other field that decays silently. Absent renders nothing, so only entries actually checked against the runtimes carry a stamp — no existing entry needed touching.
31 Aug

Nemotron 3.5 no longer claims Ollama

fixedModel Registry
The deploy field listed Ollama among the self-hosting runtimes, but the Ollama library has no Nemotron 3.5 tag — ollama.com/library/nemotron still carries only the Llama-3.1-era 70B. Ollama is dropped from the list and the notes now say what the local path actually is: a manual Modelfile import of a third-party GGUF. The rest of the list survives, because ggml-org and Unsloth both publish those GGUFs, so llama.cpp and LM Studio were never in question.
30 Aug

OpenClaw 2.0

updatedPersonal Agents
Shipped 30 August by 933 contributors across 16,000+ pull requests, roughly half the project's history. Shared cloud sessions make it multiplayer, which the card records as an exposure change rather than a feature: the same gateway serves one operator or a team whose members trust each other, and joining a session reaches the host. Isolation is unchanged — tools still run on the host unless sandboxing is configured.
30 Aug

Hy4 preview — Tencent's first entry

addedModel Registry
A 770B MoE with 49B active under Apache-2.0, 1M context, published two days ago. Notable for the licence at that size, and for a model card that names its own faults: an early version that over-thinks complex tasks and over-verifies its own work, shipped deliberately to find out what breaks.
30 Aug

The Qwen 3.8 middle: Flash, Flash-Next and 27B

addedModel Registry
The registry had Qwen's top and bottom but not the tiers most people would actually run. Flash is hosted-only at $0.15 in and $0.47 out with a 1M context; Flash-Next is open at 180B across 512 experts, but on a Qwen4 experimental architecture and under a custom licence rather than Apache-2.0. Qwen3.6-35B-A3B is replaced by the dense Apache-2.0 Qwen3.8-27B.
30 Aug

Granite 4 → Granite 4.2

replacedModel Registry
Dense 3B, 8B and 30B with native chain-of-thought and three switchable thinking modes, 128K context extending to 512K on the 30B. The architecture change is the part that matters for self-hosting: dense rather than Granite 4's hybrid Mamba-2 MoE, so the active-parameter saving is gone.
30 Aug

DeepSeek Harness

addedCoding Agents
DeepSeek ships its own MIT-licensed agent, and it is model-agnostic rather than a client for its own API — catalog providers plus custom endpoints for a gateway or self-hosted server. Recorded with its developer-preview status and its own safety notice, which says it is unaudited and that approval prompts do not guarantee isolation.
30 Aug

Goose, plus three in-house agents: Inspect, Minions and River

addedCoding Agents
Ramp's Inspect, Stripe's Minions and Shopify's River are doing serious volume behind closed doors — 75% of Ramp's merged PRs, ~1,300 a week at Stripe, one in eight at Shopify. They are marked in-house so nobody reads them as products, and each records why the team rejected the off-the-shelf option. Goose is the adoptable one, now under the Agentic AI Foundation.
30 Aug

availability and rationale

schemaCoding Agents
An in-house chip separates agents you can adopt from agents that merely exist, and a why-they-built-it field records the clearest public statement of where off-the-shelf harnesses stop. Entries without availability default to public, so nothing else needed touching.
30 Aug

Gemma 3 → Gemma 4

replacedModel Registry
E2B and E4B for edge plus 12B, 26B and 31B, at 256K context. The licence is the change that matters: Apache-2.0, where the family ran on the custom Gemma terms through Gemma 3.
30 Aug

GLM-5.3-Flash

addedModel Registry
The first natively multimodal model in the GLM-5 series, 320B total against 18B active, cutting attention compute 3.0x and KV cache 4.4x against GLM-5.3 — the figure that decides whether a 1M context is affordable to serve. Tested anonymously as ox-alpha, and served on Chinese AI chips.
30 Aug

gpt-oss-120b

addedModel Registry
OpenAI's one open-weight release at this scale, Apache-2.0, 117B total and 5.1B active, MXFP4-quantised to fit a single 80GB accelerator. Included because the rest of that catalogue is closed.
30 Aug

GLM-5.3 and Qwen3.8-Max weights have landed

updatedModel Registry
Both cards flip to open and on-prem, and both promises arrived altered: GLM-5.3 under a model-specific glm-5.3 licence rather than the MIT announced, and Qwen3.8-Max as Qwen3.8-2.4T-A95B — text-only and always thinking, where the hosted Max adds vision and optional thinking. Self-hosting Qwen means running the base, not the product.
30 Aug

Nemotron tiers renamed to NVIDIA's own

updatedModel Registry
Lightning 30B A3B, Nano 30B A3B, Super 120B A12B, Ultra 550B A55B — matching the vendor's model list, with Ultra's positioning for planning, code generation and deep research recorded.
18 Aug

Derived weight footprint and paramsActive

schemaModel Registry
Every self-hostable model now shows what its weights actually occupy at Q4 and FP8, plus the hardware class that implies — computed from parameters, so it cannot go stale. Active-parameter counts moved out of prose into a field and print beside the total, because reading 49B active as the memory requirement for a 1.6T model is the usual mistake.
14 Aug

Runtimes page

addedSite
What actually serves an open-weight model, split by layer: engines, wrappers over them, gateways in front, orchestrators above, and hosted providers where the engine stops being yours. The site leaned on this layer in three places — byoModel, localModels, onPrem — without documenting it. Notes TGI's move to maintenance, archived 21 March 2026.
14 Aug

Nemotron 3 → Nemotron 3.5, adding Lightning

updatedModel Registry
NVIDIA released Nemotron 3.5 Lightning on 11 August, a 30B MoE with 3B active built as the execution layer under Ultra's planning, alongside the NeMo Switchyard router. It joins the family card as a tier — NVIDIA calls it the smallest member of the Nemotron 3 family — rather than getting a card of its own.
14 Aug

Project Perception split out of the MDASH entry

addedCoding Agents
Perception has its own product documentation now, with named red, blue and green team agents, so folding it into MDASH understated it. MDASH keeps its own entry for the model-escalation harness.
14 Aug

GLM-5.2 → GLM-5.3

replacedModel Registry
Same base model, post-training only, and the card flips to proprietary while the MIT weights are held back two weeks for safety hardening. Z.ai reports emergent cyber capability: best-in-class on CyberGym, and 2,436 vulnerabilities found across 269 open-source projects, the oldest introduced in 1981.
14 Aug

Sakana Fugu and Fugu Cyber

addedModel Registry
A multi-agent system sold as a single model, with routing and constituent models undisclosed — so parameters, context and lineage are recorded as inapplicable rather than unknown. Not available in the EU/EEA pending GDPR compliance.
14 Aug

Gemini 3.7 Flash

addedModel Registry
Google's workhorse tier, three weeks after 3.6 Flash. Priced at an introductory $0.75 in / $3.75 out that doubles on 1 January 2027 — the card says so, because the rate on it is guaranteed to expire.
14 Aug

Gemini Spark's model named

updatedPersonal Agents
Google's 3.7 Flash announcement states Spark runs on it for AI Pro and Ultra subscribers, filling a field that had read "version not stated". The Spark product page still does not say.
13 Aug

Qwen3.8-Max

addedModel Registry
2.4T parameters, 95B active. Recorded as proprietary with onPrem false: Alibaba calls it the first Max-class Qwen to be open-weighted, but the weights are a week out and the licence is unnamed, so the fields describe today and not the promise.
13 Aug

DeepSeek V4 line refreshed, V4.1 Pro renamed V4 Pro

updatedModel Registry
Prices had drifted badly since April: input $1.74 to $0.435, output $3.48 to $0.87. Flash's cached-input rate was wrong by a factor of ten. Both now carry the peak/off-peak change that landed 16 August.
13 Aug

Grok 4.5 → Grok 4.6

replacedModel Registry
The family finally has a verified price, $2 in / $6 out, replacing a "not verified" placeholder. Parameters went back to undisclosed: 1.5T described 4.5, and neither the announcement nor the docs give a figure for 4.6.
12 Aug

Grok Bot

addedPersonal Agents
The first entry with no containment documented. Bots sign into your tools with your credentials and drive the interfaces, so permissions and audit trails see a user session rather than an agent; they also instruct each other with no human in between.
12 Aug

Memory and retention recorded per vendor

addedSystem Cards
System cards cover evaluations and safeguards, not what a service keeps about you. Each frontier card now names where that document actually lives — and xAI, which has none, is styled as a finding rather than left blank.
12 Aug

Personal Agents page

addedSite
Assistants that hold your accounts, indexed by blast radius rather than features: what they can reach, what contains them, who may address them, and what has already gone wrong. Seeded with seven entries.
11 Aug

Muse Glimmer

addedModel Registry
A 30B Apache-2.0 distillation of Muse Spark, so a new card rather than a version bump. It also made the Spark note wrong: Meta closed the open-weight era at the frontier, not outright.
11 Aug

localModels — which vendor-model agents can run local models

schemaCoding Agents
Kept separate from byoModel, because being model-agnostic by design and being redirectable through a runtime are not the same promise. Codex ships its own --oss flag; Claude Code works but Anthropic documents it as unsupported; Copilot manages chat only.
11 Aug

Factory Droid

addedCoding Agents
Verified while checking Ollama's integration list, and it earned an entry of its own: BYOK against any OpenAI- or Anthropic-compatible endpoint plus local Ollama and LM Studio, with keys never leaving the machine.
11 Aug

access — the programme a model is gated behind

schemaModel Registry
Daybreak has no home in a schema that only knows licences, and the same was true of Project Glasswing, the CodeMender pilot and Cisco's vetted release. Gating is now structured and filterable, and independent of licence: weights can be open yet vetted.
11 Aug

GPT-5.5-Cyber → GPT-5.6-Cyber

replacedModel Registry
Also reclassified from access-tier to fine-tune: OpenAI trained this one on top of GPT-5.6 Sol rather than only relaxing its safeguards. Daybreak Blue, which does only that, sits on the Sol card instead.
11 Aug

The price row reads consistently on every card

fixedModel Registry
Priced models labelled the row "$ / 1M tok" while unpriced ones labelled it "Price", so the field looked absent on exactly the cards that had a number.
11 Aug

Codex Security, Mistral Vibe, Trae, Qwen Code, Kimi Code CLI

addedCoding Agents
Closing three gaps at once: the site's security focus, and the first non-US entries on a page that had tracked country all along but listed only American and global-open projects.
11 Aug

Windsurf → Devin Desktop

replacedCoding Agents
Cognition retired the Windsurf brand on 2 June 2026 in an over-the-air update; windsurf.com redirects to devin.ai. Cascade became Devin Local.
11 Aug

Gemini CLI retired, succeeded by Antigravity CLI

replacedCoding Agents
Google stopped serving Google AI Pro, Ultra and free individual accounts on 18 June 2026; API-key and Gemini Code Assist enterprise auth were unaffected. The Antigravity entry became Antigravity IDE, the brand now spanning several surfaces.
11 Aug

Kiro: international launch, CLI and ACP

updatedCoding Agents
Launched internationally in May 2026 as a replacement for Amazon Q Developer, and no longer tied to its editor.
11 Aug

Site published

addedSite
Three pages at launch — Model Registry, Coding Agents and System Cards — with 36 models, 9 of them cyber-specialized, and 21 agents. The models below are the starting set; anything since is in the entries above.
Cisco AntaresFoundation-Sec-8BSec-PaLM 2Sec-Gemini v1Claude Mythos 5GPT-5.5-CyberGemini 3.5 Flash CyberMAI-Cyber-1-FlashTrend CybertronGPT-5.6 (Sol)Claude (Fable 5)Muse Spark (1.2)InklingGemini 3Grok 4.5Llama 4DeepSeek V4.1 ProQwen3.6-35B-A3BGLM-5.2DeepSeek V4 FlashStep 3.7 FlashMiMo-V2.5-ProKimi K3Mistral Small 3.1Phi-4MAI-Thinking-1MAI-Code-1-FlashGranite 4Nemotron 3Gemma 3Falcon 3SmolLM3GPT-SW3Teuken-7BBloombergGPTMedGemma