Registry — LLM / SLM / Domain-specialized

AI Model Registry

A structured index of language models — frontier LLMs, small language models, and domain-specialized systems for cybersecurity, medicine, finance and European languages. Every entry carries the same base record: vendor, country of origin, license, release, parameters and modalities. Specialized entries are further classified by how they specialize: purpose-built, domain fine-tune, or access-tier variant.

64
models
33
on-prem capable
8
countries
12
cyber-specialized
All models with disclosed parameter counts, on one log scale

Claude Haiku 5.5

🇺🇸 USA
Anthropic
GeneralLLMProprietaryLife Sciences Verification
Parametersundisclosed
Released
7 Oct 2026
Context
1M (128K output)
Modalities
textvision
Base model
—
Price$0.10 in · $0.50 outper 1M tokas of Oct 2026
Details

Successor to Haiku 4.5 and the last of the Claude 5.5 family, after Opus 5.5 and Sonnet 5.5. Anthropic positions it as a subagent under Opus or Sonnet, for summaries, compaction, classification and database queries. It is the first Haiku with an effort setting (default medium) and adaptive thinking. Its knowledge is reliable to June 2026, and Anthropic commits to no retirement before 7 October 2027. Vendor-reported: Terminal-Bench 4.0 at 39.2% against Haiku 4.5's 0.0% and GPT-6 Luna's 16.4%, FrontierCode 1.1 at 46.4% against Luna's 42.4%, OSWorld 2.1 (offline subset) at 72.4%, and GDPval-AA v2.1 at 1620 Elo against Luna's 1437 and Sonnet 5.5's 1840. A system card accompanies the release.

Pricing: Claude API list price, prompts up to 100K tokens; cached input $0.01/1M. The first Claude model priced by prompt length: a prompt over 100K tokens pays $0.50/$2.50, with cache reads at $0.05, for every token in the request, so a long agent transcript reaches the higher tier quickly. That is 90% below Haiku 4.5's $1/$5 under the threshold and 50% below it over the threshold. Anthropic's 75% average saving is a blend of the two. The new tokenizer also uses slightly more tokens per task. Batch halves both tiers

License: API / claude.ai

Access — Life Sciences Verification: The model is generally available, but some of what it can do in biology and cybersecurity is not. Its cyber safeguards are stricter than Haiku 4.5's and looser than Sonnet 5.5's, and they still block penetration testing and other attacker techniques. Its biology safeguards match Sonnet 5, Sonnet 5.5 and Opus 5. Anthropic directs broader work to the Life Sciences and Cyber Verification Programs (apply)

Deployment: API (claude-haiku-5-5), Claude Code, AWS Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS

Reference →

Mistral Large 4

🇫🇷 France
Mistral AI
GeneralLLMWeights pending
Parameters1.1T
MoE — the model docs give 1.05T total and 52B active, plus a 1.6B vision encoder; the announcement rounds to 1T and says 49B active, without reconciling the two
Released
6 Oct 2026
Context
1M
Modalities
textvision
Base model
—
Price$1.36 in · $4.18 outper 1M tokas of Oct 2026
Details

Mistral's largest model to date: one natively multimodal MoE that folds instruct, reasoning and agentic use together, trained from scratch on 3,800 Grace Blackwell GPUs in Mistral's own European datacentres. Read it as a preview in both senses. The weights are promised, not published, and Mistral says the reinforcement-learning run behind what the API serves is "still in flight", so the scores describe a checkpoint that will move before release. Those scores are vendor-reported against Claude Opus 5.5, GPT-6 Astra, DeepSeek V4 Pro, Kimi K3 and GLM-5.3, and the pitch leans on cybersecurity — 93% on Cybench, and open weights as the answer to provider-side refusals mid-incident — which is a claim about self-deployment that cannot be tested until the weights exist. Mistral says it is red-teaming the model with cybersecurity partners in the meantime. At 1.05T total the self-hosting floor, once there is one, is a multi-GPU node rather than a card.

Pricing: Mistral Studio list price (public preview); cached input $0.14/1M. The docs page shows the preview at half the list price — $0.68 in, $0.07 cached, $2.09 out — with no end date for the discount

License: No weights yet. Mistral says it will release them by the end of October and names no licence: the docs page marks the model "Open" with an empty licence list. Mistral's families do not share terms — Small 4 is Apache-2.0, Medium 3.5 a modified MIT — so this card records a licence when there is a repository to read it from

Deployment: Public preview API on Mistral Studio (mistral-large-4), served from Mistral's own datacentres in Europe. Mistral says it "will be able to run on private cloud or on-premise" once the weights ship; until then there is nothing to self-host and no runtime to list

Reference →

Reflection Beam

🇺🇸 USA
Reflection AI
GeneralLLMWeights pendingBeam early access
Parameters501B
Sparse MoE — 501B total, 23B active
Released
5 Oct 2026
Context
1M
Modalities
text
Base model
—
PriceFree during the beta, with daily usage limits — no list price published
Details

Reflection's first model, pitched as the Western answer to the Chinese open-weight frontier: a text-only MoE for coding, reasoning and agentic work, pretrained on 23.8T tokens, with an RL run on 10.5K GB300s for four weeks. The headline claim is efficiency — scores comparable to GLM-5.2 on roughly a third to a quarter of the inference compute — and Reflection's own footnote says how soft that is: it is estimated as 2 × active parameters × tokens generated, excluding prefill, attention and serving overhead, so it is an approximate compute comparison rather than measured cost. The benchmarks are vendor-reported, and its own table puts it behind GLM-5.3 and Kimi K3 on SWE-Bench Pro v2-Hard (77.2 against 84.3 and 88.2) and Terminal Bench 2.1 (80.1 against 88.2 and 88.3). Context is a split claim: mid-training extends the effective length to 1M, while RL ran at 256K. Text-only, though it handles other modalities rendered as text.

License: No weights yet. Reflection says it will release them this month under Apache-2.0, with a technical report and model card; that is the announcement's word, and the card takes the licence from the repository when there is one

Access — Beam early access: A waitlist: Reflection is making the preview "available to a select group of users" while red-teaming and evaluation finish. Admission criteria are not published (apply)

Deployment: Waitlisted beta API only (platform.reflection.ai), while Reflection finishes red-teaming. Reflection promises distribution partners and open-source library integrations at the weight release, without naming them; until then there is nothing to self-host and no runtime to list

Reference →

Kolibri-1

🇩🇪 Germany
Aleph Alpha
Multilingual / EULLMOpen weightsPurpose-builtOn-prem
Parameters78.1B
MoE — 78.1B total, 3.46B active per token; 384 experts per layer, 1 shared and 6 routed
Weights43 GB at Q478 GB at FP8One 80GB GPU3.5B active per tokenderived · weights only
Runtimes
Released
3 Oct 2026
Context
1M
Modalities
text
Base model
—
PriceSelf-hosted — infrastructure cost (open weights)
Details

A deliberate two-language model: German and English only, with a tokenizer built for German word structure, pitched at government and industrial buyers who need to run it on their own hardware. Trained from scratch on 20T tokens (about 62.5% English, 23.9% German, 13.6% code) on infrastructure in Germany and Finland; the card discloses 768 B200s for 21 days of pre-training and an estimated 950 MWh in total, which few vendors publish. The MoE is very sparse — 3.46B active of 78.1B — so per-token compute is small and memory is not: the card's own minimum is one H200 or two 80 GB A100s. The 1M context needs reading: native length is 262K, the longer window works because positional encoding sits only in the sliding-window layers, it must be switched on with a max_position_embeddings override, and Aleph Alpha recommends staying at or under 262K. Its benchmarks are vendor-reported on its own eval-framework, against other ~3B-active MoEs, with sector sets (German public sector, automotive supplier, semiconductors) that it built itself; the overall scores lead that group in both languages, while closed-book factual recall and some tool-calling splits trail it. The card asks for human review before output is acted on rather than unsupervised use, and Aleph Alpha is a signatory of the EU GPAI Code of Practice.

License: Apache-2.0 on both the FP8 and BF16 repositories, read from Hugging Face rather than the announcement

Deployment: Self-hosted via vLLM, but only through Aleph Alpha's aleph-alpha-inference plugin (pip or the ghcr.io container), which pins the vLLM version it supports — stock vLLM does not know the architecture. FP8 weights on Hugging Face, with a BF16 repository alongside. Community GGUF and MLX conversions appeared within days; none is Aleph Alpha's, and whether llama.cpp or mlx-lm run the architecture unpatched is unverified, so they are not listed as runtimes here. No hosted API announced verified Oct 2026

Reference →

Clef

🇺🇸 USA
Cloudflare
GeneralLLMOpen weightsdecision modelOn-prem
Parameters27B
Dense — the Qwen3.8-27B backbone with its vision encoder, plus a small joint schema head (27.4B in the published safetensors)
Weights15 GB at Q427 GB at FP8One 24GB GPUderived · weights only
Released
1 Oct 2026
Context
64K
Modalities
textvisionvideo
Base model
Qwen3.8-27B
Price$0.24 in · — outper 1M tokas of Oct 2026
Details

Cloudflare's open answer to Jev, and API-compatible with it: the same state-plus-typed-questions request, the same choice, score and noul answers. A head on top of Qwen3.8-27B reads the backbone's final hidden states and scores every option of every question in a single forward pass, so there is no free-form text and nothing to parse. Unlike Jev it reads images and video. Cloudflare reports it ahead of Jev on most of ten public classification and tool-selection sets in its own Decision Index, at a median 209ms against Jev's 524ms — vendor-reported, on a benchmark the vendor assembled, and Jev wins When2Call and BRIGHT in the same table. Hosted requests are not stored or trained on, per Cloudflare. A reinforcement-learning fine-tuning service for it runs through Cloudflare's forward-deployed engineers for now.

Pricing: Workers AI list price for @cf/cloudflare/clef. Workers AI lists an input rate only; a decision model returns probabilities rather than generated tokens

License: Apache-2.0 on the Hugging Face repository; loading it needs the repo's own model code, not stock transformers

Deployment: Workers AI (@cf/cloudflare/clef); self-hosted from Hugging Face with the repository's custom code, which Cloudflare tests with transformers on a single H200 verified Oct 2026

Reference →

Clef-flash

🇺🇸 USA
Cloudflare
GeneralSLMOpen weightsdecision modelOn-prem
Parameters9B
Dense — a Qwen3.5-9B backbone plus a joint schema head (9.4B in the published safetensors)
Weights5 GB at Q49 GB at FP8One 24GB GPUderived · weights only
Released
1 Oct 2026
Context
64K
Modalities
textvisionvideo
Base model
Qwen3.5-9B
Price$0.09 in · — outper 1M tokas of Oct 2026
Details

The small Clef, on a 9B base, and the one most decision workloads are likely to want: a median 39ms per call in Cloudflare's own measurement, against 209ms for Clef and 524ms for Jev. It holds up on tool selection and narrow intent sets and drops sharply on the wide ones — 66.8 macro-F1 on CLINC150 with out-of-scope detection, where Clef scores 97.4 — so check it on your own label set before trading accuracy for latency. All figures vendor-reported. Same API as Clef and Jev, and it also takes images and video.

Pricing: Workers AI list price for @cf/cloudflare/clef-flash. Workers AI lists an input rate only; a decision model returns probabilities rather than generated tokens

License: Apache-2.0 on the Hugging Face repository; loading it needs the repo's own model code, not stock transformers

Deployment: Workers AI (@cf/cloudflare/clef-flash); self-hosted from Hugging Face with the repository's custom code verified Oct 2026

Reference →

Gemini 4 Argon

🇺🇸 USA
Google DeepMind
GeneralLLMGatedFairwind
Parametersundisclosed
Released
30 Sep 2026
Context
Undisclosed (1M output)
Modalities
textvisionvideo
Base model
—
Price$2 in · $10 outper 1M tokas of Sep 2026
Details

Google's first frontier model since Gemini 3.1 Pro, and not a Pro: a new family name, shipped first to cyber defenders rather than developers, under a phased release that includes the U.S. government's voluntary pre-release access process. The headline spec is the output ceiling, raised to 1M tokens from 64K, so a single trajectory can run to hundreds of thousands of tokens; the input context is not stated. Google reports 77.9% on DeepSWE v1.1, first place on the Vals Index, 51.3% and first on Zapier's AutomationBench, 91.7% on LVBench, and a tie for first at 68% on CWE-bench v1, with vulnerability discovery ahead of 3.8 Flash Cyber on Google's and Wiz's internal benchmarks — vendor-reported throughout. The safeguards claims are the ones to hold it to at general release: activation monitoring for misuse, chain-of-thought monitoring that can halt execution, and leading robustness on Gray Swan's indirect prompt injection benchmark.

Pricing: Announced introductory price — not yet on general sale; cached input $0.10/1M. Cached input is announced as 95% off the input price. After the introductory period, which Google does not date, the price becomes $4 in / $20 out

License: Fairwind Program only at launch; Google says paid API customers and Google AI Ultra subscribers come next, with no date

Access — Fairwind: For now the only door. Google calls it exclusive access for a subset of Fairwind partners, prioritising governments and national cyber authorities, critical infrastructure operators and core technology platforms, with academic labs doing defensive benchmarking also invited to apply. Admitted organisations may grant the model only to internal cybersecurity, incident response or penetration testing teams, must track employee access and use, deploy phishing-resistant MFA, and are limited to dual-use work such as authorised threat simulation, reverse engineering and malware analysis. Inside the programme Google releases it without cyber guardrails, and it can run standalone or inside CodeMender (apply)

Deployment: Fairwind Program partners, standalone or inside CodeMender; zero data retention is supported when accessed as a managed model on Gemini Enterprise. Not yet in the public Gemini API model catalogue

Reference →

GPT-6.1 Sol

🇺🇸 USA
OpenAI
GeneralLLMProprietary
Parametersundisclosed
Released
29 Sep 2026
Context
1.05M (128K output)
Modalities
textvision
Base model
—
Price$2 in · $10 outper 1M tokas of Sep 2026
Details

Successor to GPT-6 Sol (22 September 2026), a week later, announced at DevDay on the day OpenAI did not ship a GPT-6.1 Astra — which the Wall Street Journal reports was held back after internal testing showed more deception and a tendency to proceed without asking permission. Sol stays the middle tier under Astra and above Luna, and OpenAI pitches it as near-Astra intelligence at a fifth of the price. The cyber line moved with it: the system card addendum rates GPT-6.1 Sol Critical for cybersecurity under the Preparedness Framework, the second model after Astra to cross that threshold, with 99.7% on ExploitBench at max effort, and High for biology and chemistry; it ships under Astra's safeguards stack. GPT-6 Sol was open to Daybreak Blue with reduced cyber refusals, per the GPT-6 system card; the 6.1 addendum mentions phased defender access through Daybreak without naming a tier for this model, so the card carries no access programme until something does. Knowledge cutoff 30 April 2026; reasoning effort runs from low to max, with none and minimal no longer supported. OpenAI reports DeepSWE 1.1 matching Astra at about a fifth of the cost and 6.4 points above GPT-6 Sol; AutomationBench 2.2 points above Claude Opus 5.5 at medium effort for about a third of the cost; OSWorld 2.0 offline within 2.1 points of Astra at about a seventh of the cost per task; Terminal-Bench Science 0.1 at $5.47 a task against $23.21 for Opus 5.5 and $23.80 for Astra, which still scores highest at 68.1%; and factual errors on a hard internal set down from 11.4% to 7.7% at low effort. Its restriction-bypass rate on an agentic alignment test fell from GPT-6 Sol's 64.4% to 23.5%. Vendor-reported throughout. An Ultrafast variant for Codex is announced for the coming days.

Pricing: OpenAI API standard pricing; cached input $0.10/1M. Cache writes $2.50. Cached input is half GPT-6 Sol's $0.20. Prompts over 272K input bill the whole request at 2× input and cache rates and 1.5× output; Batch and Flex −50%, Fast mode 2× (not available with EU data residency), regional processing +10%. A fifth of GPT-6 Astra's standard rates

License: API / ChatGPT Work and Codex

Deployment: OpenAI API (model ID gpt-6.1-sol), with US and EU data residency; ChatGPT Work and Codex for Plus, Pro, Business, Enterprise and Edu. Not yet in ChatGPT's Chat mode

Reference →

Claude Sonnet 5.5

🇺🇸 USA
Anthropic
GeneralLLMProprietaryLife Sciences Verification
Parametersundisclosed
Released
28 Sep 2026
Context
1M (128K output)
Modalities
textvision
Base model
—
Price$2 in · $10 outper 1M tokas of Oct 2026
Details

Successor to Sonnet 5 and the second model in the Claude 5.5 family, between Opus 5.5 and Haiku 5.5. Anthropic pitches it as the faster, cheaper complement to Opus 5.5, strongest on well-scoped everyday tasks, bug fixing and office documents, and over 30% faster than Sonnet 5. Effort runs Low to Max, defaulting to high on the API and medium in Claude Code and the apps; thinking is adaptive, and integrations that ran Sonnet with thinking off must move to the new between_tools setting. Vendor-reported: Terminal-Bench 4.0 at 70.6% against Sonnet 5's 10.3% and Opus 5.5's 66.4%, FrontierCode 1.1 at 52.1% at Xhigh against Opus 5.5's 54.4%, GDPval-AA v2.1 at 1844 Elo against Opus 5.5's 1846, and OSWorld 2.1 (partial credit) at 80.1%. It is the first Sonnet with classifiers against reasoning extraction, with preserved thinking bound to the account that created it. Knowledge reliable to June 2026; no retirement before 28 September 2027. A system card accompanies the release.

Pricing: Claude API list price; cached input $0.10/1M. Same $2/$10 as Sonnet 5. Cache reads launched at $0.20 and were halved to $0.10 on 7 October 2026, alongside the Haiku 5.5 launch, which Anthropic puts at about 20% off most agentic work. Cache writes $2.50. Anthropic's 'up to 30% cheaper' claim is per task, from using fewer tokens, not from the rate

License: API / claude.ai

Access — Life Sciences Verification: The model is generally available, but its higher-risk cyber capability is not. It is the first Sonnet to ship with Opus 5.5-class cyber safeguards, which visibly hand higher-risk security tasks to Sonnet 5 while still allowing routine bug finding and fixing. Biology safeguards are unchanged from Sonnet 5. The Life Sciences Verification Program is open; Anthropic says tiered Cyber Verification Program access covering Sonnet 5.5 will open for applications soon — announced, not yet open (apply)

Deployment: API (claude-sonnet-5-5), claude.ai, Claude Code, AWS Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS; zero data retention available

Reference →

GPT-6 Luna

🇺🇸 USA
OpenAI
GeneralLLMProprietaryDaybreak Blue
Parametersundisclosed
Released
22 Sep 2026
Context
1.05M (128K output)
Modalities
textvision
Base model
—
Price$0.10 in · $0.50 outper 1M tokas of Sep 2026
Details

The efficiency tier of the GPT-6 family, described by OpenAI as its most efficient model for focused, high-volume work, and successor to GPT-5.6 Luna, which this registry never carded separately. Knowledge cutoff 18 May 2026 — a month later than Sol's. OpenAI reports DeepSWE 1.1 at 66.6% at max effort, comparable to Claude Opus 5 and Fable 5 at medium effort at 93% and 96% less per task, and OSWorld 2.0 offline above GPT-5.6 Sol at medium effort at a tenth of the cost — vendor-reported. Parameters are undisclosed, so the LLM classification is a default rather than a measurement. GPT-5.6 had a Terra tier between the two; GPT-6 has not announced one.

Pricing: OpenAI API standard pricing; cached input $0.01/1M. Cache writes $0.125. Same long-context (>272K), Batch, Flex, Fast-mode and regional multipliers as GPT-6.1 Sol. Half GPT-5.6 Luna's promotional $0.20/$1.20

License: API / ChatGPT Work and Codex; desktop app for Free and Go

Access — Daybreak Blue: The model is generally available; reduced cyber refusals for authorised defensive work are not. Eligible Daybreak Blue users get them on Luna, subject to the programme's access controls, and individual members must enable Advanced Account Security or receive standard access. OpenAI states this in the GPT-6 system card's appendix on Sol and Luna, not on the model page, and has not said how the tier is selected in the API: the gpt-daybreak-blue-latest alias still resolves to GPT-5.6 Sol (apply)

Deployment: OpenAI API (model ID gpt-6-luna); ChatGPT Work and Codex for Plus, Pro, Business, Enterprise and Edu, and the desktop app for Free and Go users

Reference →

Claude Opus 5.5

🇺🇸 USA
Anthropic
GeneralLLMProprietaryLife Sciences Verification
Parametersundisclosed
Released
22 Sep 2026
Context
1M (128K output)
Modalities
textvision
Base model
—
Price$4 in · $20 outper 1M tokas of Sep 2026
Details

The first model in Anthropic's Claude 5.5 family and successor to Opus 5 (July 2026), followed by Sonnet 5.5 (28 September) and Haiku 5.5 (7 October). Anthropic claims Fable 5.1-level performance on most work at 40% less cost than Opus 5, and its own docs now recommend it over Fable as the starting point for most workloads. Vendor-reported: Terminal-Bench 4.0 at 66.4% against Fable 5.1's 55.8% and GPT-6 Astra's 57.9%, FrontierCode 1.1 at 54.4% against Astra's 53.3%, GDPval-AA v2.1 at 1846 Elo against Fable 5.1's 1735, OSWorld 2.0 at 81.8% partial credit, and Terminal-Bench-Science at 58.7%, behind Astra's 64.6%. Anthropic itself cautions that benchmark margins at this level are a weak guide and that the real gap to Fable 5.1 is narrower. The scores were run with production safeguards on, with intervening cyber tasks completed by Opus 4.8 and biology tasks by Opus 5. Anthropic rates it comparable to Claude Mythos 5.1 in biology and cybersecurity, which is why it is the first Opus to ship with Fable-class safeguards. Pre-release testing by METR and Frontier Design; Anthropic reports its best behavioural-audit scores to date and 85% fewer containment-boundary crossings than Opus 5 or Mythos 5.1, while noting the model often suspects it is being evaluated. Thinking is adaptive and cannot be turned off; default effort is medium; knowledge cutoff June 2026. A system card accompanies the release.

Pricing: Claude API list price; cached input $0.20/1M. 20% below Opus 5's $5/$25, and cache reads 60% below its $0.50; cache writes $5. Anthropic puts the saving on typical workloads at about 40%, because the model also uses fewer tokens per task. Fast mode, in Claude Code and the Claude Platform at up to 2.5× speed, is $8/$40

License: API / claude.ai

Access — Life Sciences Verification: The model is generally available, but its biology and cybersecurity capability is not: it ships with Fable 5.1-class safeguards that route most cyber tasks to Opus 4.8 and gate biology work. Vetted organisations can apply to the Life Sciences Verification Program today for safeguards suited to biology R&D. Anthropic says the Cyber Verification Program will extend to Opus 5.5 in the coming weeks, in three tiers of increasingly permissive access — announced, not yet open (apply)

Deployment: API (claude-opus-5-5), claude.ai, AWS, Google Cloud and Microsoft Azure; zero data retention available

Reference →

Grok 4.7

🇺🇸 USA
SpaceXAI
GeneralLLMProprietary
Parametersundisclosed
Runtimes
Released
21 Sep 2026
Context
500K
Modalities
textvision
Base model
—
Price$2 in · $6 outper 1M tokas of Sep 2026
Details

Successor to Grok 4.6 (August 2026): a new, larger base model, a longer reinforcement-learning run weighted toward tasks that take many hours, and training on the Grok Bot harness itself. Parameters stay undisclosed — before launch Musk put the model at 2.1T with supplemental training on SpaceX engineering data, and neither figure appears in the launch post or the docs, so the card records neither. Vendor-reported benchmarks reward reading the whole table rather than the headline: it leads its comparison set on EEBench (64.0% against 39.4% for GPT-5.6 Sol) and on the Harvey legal-agent benchmark (19.6% against 6.7%), sits behind Fable 5.1 on CursorBench 4.0 (46.3% vs 51.8%) and well behind it on Terminal-Bench 4.0 (38.0% vs 57.9%) — though that is close to double Grok 4.6’s 20.3% — and its 71.0% DeepSWE figure is asterisked as a high-effort run. Safeguards are an entirely new stack, with 62.4% on LatchBio’s biosafety benchmark and 3.3% of risky prompts allowed through on xAI’s own HackerBench v0.3; select cybersecurity partners get invite-only access to the model’s red-team capabilities. As with 4.6, no system card accompanies the release. Developer docs fill in the rest: 500K context with no stated output limit, text and image in and text out, a knowledge cutoff of May 2026 (Grok 4.6’s was February), reasoning effort from low to xhigh, built-in function calling, code execution and web and X search, and encrypted reasoning returned on every Responses API call whether or not it is requested. The vendor name is new too: xAI joined SpaceX in February 2026 and the site, announcement and copyright now read SpaceXAI, while the API, the docs and the key are all still branded xAI.

Pricing: SpaceXAI (xAI) API list price; cached input $0.50/1M. Unchanged from Grok 4.6. A prompt reaching 200K tokens bills the whole request at long-context rates — $4 input, $1 cached, $12 output. The fast variant is twice the price and is served only in Cursor and Grok Build, not on the public API; the US regional endpoint adds 10%.

License: API / X platform

Deployment: API (model ID grok-4.7, Responses and Chat Completions; US regional endpoint at us.api.x.ai), Grok Build as its coding agent’s default model, Cursor on all plans; also OpenRouter, Vercel and Cloudflare verified Sep 2026

Reference →

Aikido Altar-1

🇧🇪 Belgium
Aikido Security
CybersecurityLLMOpen weightsDomain fine-tuneOn-prem
Parameters504B
MoE — 504B total after pruning 88 of GLM-5.3's 256 routed experts per layer; routing is untouched, so 8 experts fire per token and ~40B stay active
Weights277 GB at Q4504 GB at FP88x80GB node40B active per tokenderived · weights only
Runtimes
Released
21 Sep 2026
Context
1M
Modalities
text
Base model
GLM-5.3, by way of the community cyankiwi/GLM-5.3-AWQ-INT4 build
PriceSelf-hosted — infrastructure cost (open weights)
Details

Not a trained cyber model: a Router-weighted Expert Activation Prune (REAP, Cerebras) of GLM-5.3 at INT4, with no retraining at all. The domain adaptation is in which experts survived — calibration ran on cybersecurity traces, code, tool calling, reasoning, English and multilingual Wikipedia, and each expert was scored by its largest share of any single domain's routed work rather than by global frequency, which is what keeps a domain's specialists instead of deleting them. Aikido publishes the cost of that cut rather than burying it: on its own benchmark of 32 known CVEs across 30 repositories Altar-1 averaged 60.4% recall and rediscovered 23 of 32, against 61.5% for the unpruned INT4 build and 65.6% for full-precision GLM-5.3. KL divergence against full BF16 is 0.506 nats on a sealed 25-prompt panel. So the trade is a few points of recall for weights that fit on a node you own — the pitch is data residency and air-gapped pentesting, not capability, and Aikido says its own Attack, AI Code Analysis and Deep Review products already run on it. Two things a “sovereign” framing should not obscure: the lineage runs through a Chinese base model, and the artifact was assembled from community work — the quantization is cyankiwi's, the prune 0xSero's, both credited on the card. The derived footprint below understates the 328 GB Aikido actually publishes, because only the routed experts are 4-bit while attention, the shared expert, the dense layers and the head stay BF16.

License: Hugging Face records `other`, with no licence name of its own; the model card says it inherits GLM-5.3's model-specific licence. A derivative cannot be more permissive than what it was cut from, so read Z.ai's terms — not the word “open-weight” in the announcement — before deploying

Deployment: Self-hosted via vLLM on a 4× H200 node — Hopper required, since the W4A16 experts need the Marlin MoE kernel; weights on Hugging Face. Aikido's own serving flags cap the context at 128K rather than the base model's 1M, to leave KV-cache room at production batch sizes verified Sep 2026

Reference →

Step 5 Preview

🇨🇳 China
StepFun
GeneralLLMWeights pending
Parameters600B
Sparse MoE — 600B total, 27B active per token, per StepFun
Released
20 Sep 2026
Context
1M
Modalities
textvision
Base model
—
Price$1 in · $2.70 outper 1M tokas of Oct 2026
Details

StepFun's new flagship for agentic work, and a separate line from Step 3.7 Flash rather than its successor: it is three times the size, and StepFun pitches it at software engineering and professional knowledge work, with finance singled out. Vendor-reported at high effort, against GLM-5.3 and Kimi K3 at max: DeepSWE v1.1 at 67.7% against 66.9% and 67.5%, Terminal-Bench 4.0 at 33.3% against 41.9% and 12.6%, and CyberGym at 84.7%. StepFun's own table puts it behind GPT-6 Astra and Claude Opus 5 on most coding and agent benchmarks, and says a meaningful gap to the frontier remains on the hardest long-running tasks. Several headline scores come from StepFun's own benchmarks (StepCodeBench, FinStepBench), so they have no outside reference point. Its FrontierFinance score, 66.4%, is the one external finance result. Preview naming: expect the weights, or a later release, to differ from what the API serves today.

Pricing: StepFun Open Platform list price; cached input $0.05/1M. The cache-miss rate already includes writing to the cache, and output includes reasoning tokens. Five times Step 3.7 Flash's $0.20 input

License: No weights yet. StepFun says it will publish them on 15 October 2026 and has not named a licence, either on the model page or in its announcement. Its last release, Step 3.7 Flash, is Apache 2.0, but a sibling's licence says nothing about this one; the card takes the licence from the repository when there is one

Deployment: StepFun API (step-5-preview) and StepFun's own products. Nothing to self-host until the weights appear, and no runtime has announced support

Reference →

DeepSeek V4.1 Flash

🇨🇳 China
DeepSeek
GeneralLLMOpen weightsOn-prem
Parameters763B
MoE — DeepSeek quotes a 552B backbone; the published checkpoint is 763B because it also carries 196B of Engram conditional memory and the DSpark speculative-decode module, and all of it has to be stored. 384 routed experts plus one shared, 6 routed active per token. Activation is asymmetric: 8B per token on prefill, 16B on decode
Weights420 GB at Q4763 GB at FP88x80GB nodederived · weights only
Runtimes
Released
10 Sep 2026
Context
1M (384K output)
Modalities
textvision
Base model
—
Price$0.30 in · $1.20 outper 1M tokas of Oct 2026
Details

Successor to V4 Flash, and a new architecture rather than a refresh: DeepSeek calls it the smallest model of its new family. A causal encoder-decoder splits the 40 layers into 20 encoder and 20 decoder, and the decoder's KV cache is projected from the encoder. Together with CSA2 sparse attention and an FP4 KV cache, that cuts the KV cache to about a quarter of V4 Flash's in memory and an eighth on SSD. It is trained from scratch on 45T multimodal tokens, so image input is native rather than bolted on, and reasoning effort is set as an integer from 1 to 100. Vendor-reported at maximum effort: DeepSWE v1.1 at 74.2% against V4 Pro's 62.7%, Terminal-Bench 4.0 at 31.2% against V4 Flash's 7.0%, and CyberGym at 88.1%. DeepSeek reports it ahead of its own V4 Pro, which is why the API now sends V4 Pro traffic here. Costin Raiu's September 2026 CTI evaluation on two DGX Sparks was of V4 Flash, not this model, and does not carry over.

Pricing: DeepSeek API list price, peak hours; cached input $0.006/1M. Off-peak is half: $0.15 / $0.60, cache hits $0.003. Peak hours are 01:00–04:00 and 06:00–10:00 UTC on weekdays, so most hours bill at the off-peak rate. That is still a rise on V4 Flash's $0.14 / $0.28, which were the full rates before off-peak billing began

License: MIT, read from the Hugging Face repository

Deployment: Self-hosted via vLLM and SGLang, both of which publish V4.1 Flash recipes; the repo ships no Jinja chat template, only DeepSeek's own encoding reference. DeepSeek API as deepseek-flash, where the retired deepseek-v4-flash and deepseek-v4-flash-vision-exp IDs now land too. Ollama lists it as a cloud model only, not a local one verified Oct 2026

Reference →

Mercury 2.5

🇺🇸 USA
Inception
GeneralLLMProprietarydiffusion
Parametersundisclosed
Runtimes
Released
8 Sep 2026
Context
260K (64K output)
Modalities
text
Base model
—
Price$0.04 in · $0.15 outper 1M tokas of Sep 2026
Details

The only diffusion language model in this registry, and the reason it earns a place: every other entry generates tokens one at a time, while Mercury produces and refines many in parallel. Inception sells the line on latency rather than intelligence — 5–7x higher throughput, sub-300ms time to first token and up to 70% lower cost per task, all vendor-reported and stated for Mercury generally rather than measured for 2.5. That trade is worth understanding for agentic work, where latency compounds across every step of a tool-calling loop rather than being paid once. Supports reasoning effort, tool calling and structured outputs. Successor to Mercury 2 (March 2026), which it doubles in context while cutting the listed price roughly six-fold. Parameters are undisclosed, so there is no footprint here and no self-hosted path.

Pricing: OpenRouter list price for inception/mercury-2.5. Sources disagree: Inception's own site lists $0.20 in / $0.75 out, which looks like copy carried over from Mercury 2 at $0.25 / $0.75. The rate here is what OpenRouter's model API returns for this model ID — confirm before budgeting

License: Hosted API only — no published weights

Deployment: Inception API and OpenRouter (model ID inception/mercury-2.5) verified Sep 2026

Reference →

GPT-6 Astra

🇺🇸 USA
OpenAI
GeneralLLMProprietaryDaybreak Red
Parametersundisclosed
Released
4 Sep 2026
Context
—
Modalities
textvision
Base model
—
Price$10 in · $50 outper 1M tokas of Sep 2026
Details

The first model from any lab to meet the Critical cybersecurity threshold under OpenAI's Preparedness Framework — the reason this registry carries threshold language at all. Measured without production safeguards it scores 100% on ExploitBench against GPT-5.6 Sol's 78.5%, 42.4% on ExploitGym against 30.3%, and on SRE-Bench binary reverse engineering solves 88.0% of tasks first try and 99.2% within four, against 55.9% and 68.7%. On a fresh exploit set built from vulnerabilities of the previous three months it discovered and used two previously unknown zero-days, which OpenAI says it is disclosing to the maintainers. Beyond cyber it reports 98% on FrontierMath Tier 4, 99.9% on ARC-AGI-3 and 64.6% on Terminal-Bench Science against Fable 5.1's 52.6%. The alignment claim is the one to weigh against all that: on an evaluation built from the July Hugging Face incident, GPT-5.6 Sol without production safeguards exceeded its authorised target 48% of the time and Astra did so in 0% of cases. Vendor-reported throughout; a system card accompanies the release.

Pricing: OpenAI API standard pricing. A Fast mode runs at roughly double the speed for double the price. Separate cache read and write rates apply, which the announcement does not quantify

License: API / ChatGPT; also Microsoft Azure and AWS Bedrock

Access — Daybreak Red: The model is broadly available, but its advanced cyber capability is not. With standard safeguards it does secure code review and patching and refuses harder work such as writing proof-of-concept exploits. Reduced refusals on Astra are available only to Daybreak Red customers, which is separately approved and provisioned; Daybreak Blue accounts get Astra with standard safeguards, and OpenAI says extending reduced refusals to Blue is in progress. Blue's own alias, gpt-daybreak-blue-latest, still resolves to GPT-5.6 Sol. The Critical rating describes what the model can do when measured without production safeguards, not what a general user can invoke (apply)

Deployment: OpenAI API (model ID gpt-6-astra), ChatGPT Plus, Pro, Business and Enterprise, Microsoft Azure and AWS Bedrock; the Codex harness was updated alongside it

Reference →

Gemini 3.8 Flash Cyber

🇺🇸 USA
Google DeepMind
CybersecurityLLMGatedDomain fine-tuneFairwind
Parametersundisclosed
Released
2 Sep 2026
Context
—
Modalities
textvision
Base model
Gemini 3.8 Flash
PriceGated programme access — no public list price
Details

Successor to Gemini 3.5 Flash Cyber, and shipped with what Google calls a more permissive set of mitigations for cybersecurity — the same access-tier logic as Daybreak and the Verification Programmes, applied to a fine-tune. Google reports it ahead of both its predecessor and larger frontier models on CyberGym, a real-world vulnerability discovery success rate above 70% across 20 languages, and 47.2% pass@1 on CWE-Bench patching at materially lower cost. The deployment figures are the more interesting claim: Chrome Security reports 2.6× more correct patches than competing large models, Wiz reports 7.5–9.7% higher recall on penetration-testing benchmarks, and Google Cloud reports finding a critical vulnerability in under two hours. All vendor- or partner-reported.

License: Trusted defenders only, through the Fairwind Program

Access — Fairwind: A limited-access programme for governments and trusted partners, prioritising national cyber authorities, critical infrastructure operators in healthcare, telecoms, energy and finance, and core technology platform companies. Google reports more than 650 participating partners. Admission carries operational conditions rather than just vetting: access must be confined to internal cybersecurity, incident response or penetration testing teams, and protections such as multi-factor authentication must be deployed. It grants the CodeMender harness alongside the model

Deployment: Restricted access through the Fairwind Program, paired with the CodeMender harness; not listed in the public Gemini API model catalogue. Since Gemini 4 Argon's launch on 30 September the programme page names only Argon, and Google has not said whether Flash Cyber remains available to partners

Reference →

Muse Spark 1.3

🇺🇸 USA
Meta (Superintelligence Labs)
GeneralLLMProprietary
Parametersundisclosed
Runtimes
Released
2 Sep 2026
Context
1M
Modalities
textvisionvideo
Base model
—
Price$1.25 in · $4.25 outper 1M tokas of Aug 2026
Details

Meta's proprietary frontier family from Superintelligence Labs. The original Muse Spark (April 2026) ended the open-weight Llama era at the frontier, shipping cloud-only; 1.1 added agentic computer use and the public API, 1.2 was co-trained with the Muse Code terminal agent, and 1.3 (2 September 2026) is an efficiency and judgement release rather than a capability jump — Meta reports roughly 20% fewer tool calls and 25% fewer tokens than 1.2 on the same work, with less verbose output. The agentic changes are behavioural: it asks clarifying questions, calls for help when stuck, and confirms before consequential actions. Two things worth watching. Safety work is aimed squarely at the failure modes this site tracks — stronger resistance to prompt injection, and better calibration of what counts as an irreversible action. And max reasoning is held back pending further safety testing, so the model shipped with a capability deliberately withheld. Meta also lists a Muse Spark open-weights release on its roadmap, which would partly reverse the closure the original release represented.

Pricing: Meta Model API list price (standard tier). A cheaper 'contributor' tier lets Meta retain submitted data for training; standard-tier data is not retained. The 1.3 announcement does not restate pricing, so these figures carry the August stamp until re-verified

License: Meta Model API / Meta apps — no downloadable weights, though an open-weights release is on the roadmap

Deployment: Meta Model API, Muse Code, Meta AI apps, OpenRouter verified Sep 2026

Reference →

Gemini 3.8 Flash

🇺🇸 USA
Google DeepMind
GeneralLLMProprietary
Parametersundisclosed
Released
2 Sep 2026
Context
1M (64K output)
Modalities
textvisionaudiovideo
Base model
—
Price$0.75 in · $3.75 outper 1M tokas of Sep 2026
Details

The third Flash release in six weeks, and the reason to read the cadence rather than the version number: Google has not shipped a frontier Pro model since early 2026, and the Flash line is where its coding and agentic gains are landing. Google reports 54.9% on HLE-Verified and the top position on the DeepSWE leaderboard, gains that are marginal over 3.7 Flash on most tests but larger on coding — vendor-reported. Computer use remains the weak spot: improved on OSWorld-2.0 but, per Ars Technica's reading of Google's own chart, still well behind Claude Opus. Takes text, image, video, audio and PDF in, text only out, with thinking levels, function calling, code execution, search grounding and computer use in preview.

Pricing: Gemini API introductory list price. Same introductory rate as 3.7 Flash and on the same clock: it runs to 31 December 2026, after which the published price is $1.50 in / $7.50 out. Google has now shipped three Flash models in six weeks, so a successor may well arrive before the increase does

License: API / Google products

Deployment: Gemini API (model ID gemini-3.8-flash), AI Studio, Android Studio, Gemini Enterprise and consumer Gemini subscriptions; 3.7 Flash remains listed as stable

Reference →

Claude Mythos 5.1

🇺🇸 USA
Anthropic
CybersecurityLLMGatedAccess-tier variantLife Sciences Verification
Parametersundisclosed
Released
Sep 2026
Context
1M
Modalities
textvision
Base model
—
Price$10 in · $50 outper 1M tokas of Sep 2026
Details

The restricted counterpart to Claude Fable 5.1 — the same underlying model with safeguards configured for cybersecurity and life sciences work, rather than a cyber fine-tune. Anthropic reports Terminal-Bench 4.0 between 55.8% and 60.9% against Fable 5's 42.0%; vendor-reported. Both verification programmes gate the same model with different safeguard configurations, which is what makes this an access tier and not a variant. Its April 2026 preview was the first model to complete the UK AISI corporate-network attack simulation end-to-end, a multi-step exercise estimated at ~20 human hours. Note the gate has changed name: this announcement describes the CVP and LSVP programmes and does not mention Project Glasswing, which is how access to Mythos 5 was described in June.

Pricing: Anthropic states pricing is otherwise the same as Fable's; cached input $0.25/1M. The first published price for this model: previous Mythos releases carried no standard list price at all

License: Restricted — trusted access programmes only

Access — Life Sciences Verification: Vetted life sciences professionals, built with the US government and currently open only to US organisations. The separate Cyber Verification Programme covers Opus- and Sonnet-class models with reduced cyber safeguards today, and Anthropic says it will include Mythos-class models in the near future — so the cyber door to this model is announced rather than open

Deployment: Restricted API access through Anthropic's trusted access programmes

Reference →

Claude Fable 5.1

🇺🇸 USA
Anthropic
GeneralLLMProprietary
Parametersundisclosed
Released
Sep 2026
Context
1M (128K output)
Modalities
textvision
Base model
—
Price$10 in · $50 outper 1M tokas of Sep 2026
Details

Anthropic's frontier family, sitting above Claude Opus 5.5, Sonnet 5.5 and Haiku 5.5, and a September 2026 refresh of Fable 5 rather than a new generation. Thinking is adaptive and always on, defaulting to high effort; knowledge is reliable to June 2026, and Anthropic commits to no retirement before 1 September 2027. Anthropic reports Terminal-Bench-Science 0.1 at 52.6% against Fable 5's 24.7%, CursorBench 3.2.0 at 73.4% against 70.5%, and Humanity’s Last Exam at 60.9% without tools and 65.0% with — vendor-reported. Available on all platforms at launch, including AWS, Google Cloud and Azure. The restricted counterpart is Claude Mythos 5.1.

Pricing: Claude API list price; cached input $0.25/1M. Headline rates are unchanged from Fable 5, but cached reads fell 75% from $1.00 to $0.25, which Anthropic frames as roughly 25% less for typical workloads and up to about 45% for agentic ones — the saving is in reuse, not in the per-token price. Anthropic's model docs still state the generic rule that cache reads cost 10% of base input, which would be $1.00, so confirm against the pricing page before budgeting. Current siblings: Opus 5.5 at $4/$20, Sonnet 5.5 at $2/$10, Haiku 5.5 from $0.10/$0.50

License: API / claude.ai

Deployment: API (claude-fable-5-1), claude.ai, AWS Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS

Reference →

Jev

🇺🇸 USA
TypeSafe AI
GeneralLLMProprietarydecision model
Parametersundisclosed
Released
Sep 2026
Context
64K per request (32K for state plus the longest question)
Modalities
text
Base model
—
Price$0.042 in · $0 outper 1M tokas of Oct 2026
Details

The model that named the category: a decision model, which TypeSafe calls System One. It does not write text. A request carries a state and a set of typed questions — yes/no ("noul"), choice among declared options, or a score on ordered levels — and the answer is a probability for each option plus a confidence, so code routes, escalates or defers to a human without parsing anything. TypeSafe trains it with Reinforcement Learning for Calibrated Decisions (RLCD) and describes a new architecture and sampler without detailing either; parameters are undisclosed. Text only — images must be turned into text first — and English is where accuracy is best, by TypeSafe's own account. The pitch is cost and latency at every branch point of an agent: $0.042 per million input tokens and output free, with speed and price multiples over LLMs that are vendor-reported on workflows of its choosing. Launched publicly in September 2026; the current version is 1.13.

Pricing: TypeSafe list price for jev-1.13.0. Charged on input only — output tokens are free because there is no text output. Rate limits are documented as adjusting without notice while capacity is added

License: Hosted API only — no published weights, and per-account fine-tuning is not offered

Deployment: TypeSafe API, POST /v1/systemone (model ID jev-1.13.0, alias jev-latest), in early access verified Oct 2026

Reference →

MiMo-V2.6-Pro

🇨🇳 China
Xiaomi
GeneralLLMOpen weightsOn-prem
Parameters1T
MoE — 1.02T total, 42B active; 384 routed experts with 8 active, no shared experts. 70 layers, 60 sliding-window and 10 global attention, plus a 5-layer multi-token-prediction drafter and separate vision and audio encoders
Weights561 GB at Q41 TB at FP88x80GB node42B active per tokenderived · weights only
Runtimes
Released
Sep 2026
Context
1M
Modalities
textvisionaudiovideo
Base model
—
Price$0.43 in · $0.87 outper 1M tokas of Oct 2026
Details

Successor to MiMo-V2.5-Pro, on the same 1.02T/42B MoE backbone and the same 1M context, with native text, image, video and audio input. What changed is the post-training. Xiaomi ran one mixed reinforcement-learning pass across coding, general agents, visual tasks and cybersecurity, rather than a separate run per domain. Two checkpoints ship. RL is the flagship. MOPD is a later distilled upgrade aimed at tool-call repetition, where an agent repeats the same calls without making progress; Xiaomi does not say which checkpoint its API serves. Vendor-reported, against V2.5-Pro: Terminal-Bench 4.0 at 34.9% against 1.5%, CyberGym at 94.0% against 40.0%, and ExploitBench at 47.9% against 16.6%. The comparison columns are Claude Opus 5, GPT-5.6 Sol and Claude Fable 5, not the current frontier, and on ExploitBench all three score 70% or higher.

Pricing: Xiaomi MiMo API overseas list price; cached input $0.004/1M. Unchanged from V2.5-Pro, which Xiaomi retires on 21 October 2026. Cache writes are free for a limited time, and the batch API is half price. A separate pro-ultraspeed tier costs 10× the standard rate

License: MIT on both Hugging Face repositories (RL and MOPD checkpoints); loading needs the repo's own model code (trust-remote-code)

Deployment: Self-hosted via SGLang and vLLM, both of which publish recipes covering the V2.6-Pro checkpoints; the reference SGLang launch is two 16-GPU nodes. Also Xiaomi MiMo API (mimo-v2.6-pro) and OpenRouter verified Oct 2026

Reference →

Hy4 preview

🇨🇳 China
Tencent (Hy / Hunyuan)
GeneralLLMOpen weightssparse attentionOn-prem
Parameters770B
MoE — 770B total, 49B active per token; 78 layers, the first with a dense FFN and the remaining 77 carrying 256 routed experts plus one shared, with the top 8 routed experts firing per token. A native multi-token-prediction layer for speculative decoding adds ~10B, which is why the published repository weighs 780B
Weights424 GB at Q4770 GB at FP88x80GB node49B active per tokenderived · weights only
Released
28 Aug 2026
Context
1M
Modalities
text
Base model
—
PriceSelf-hosted — infrastructure cost (open weights)
Details

Tencent's first entry in this registry, and one of the largest Apache-2.0 models anywhere: a 770B MoE claiming the open-source frontier, with a 1M context and attention built on Gated DeepSeek Sparse Attention, which Tencent credits to DeepSeek and GLM directly. Training data was built around work done by Tencent's own software engineers, game developers, finance analysts and security experts. The preview label is meant literally and the model card is unusually candid about it: an early version with real headroom left in pre- and post-training, shipped with known issues including spending longer than necessary on complex reasoning and over-verifying its own work. Tencent says it would rather ship early and hear what breaks, as it did with Hy3.

License: Apache-2.0

Deployment: Self-hosted via vLLM or SGLang in BF16 or FP8; community GGUF, MLX 4-bit and Hygon INT8 builds exist. Also on ModelScope, cnb.cool and GitCode verified Sep 2026

Reference →

Qwen3.8-Flash-Next

🇨🇳 China
Alibaba
GeneralLLMOpen weightsOn-prem
Parameters180B
MoE — ~180B total across 512 experts, ~6B active per token, 48 layers. The total decomposes into a ~125B backbone, a ~51B n-gram embedding table and a ~4B draft head, and the embedding table is designed to sit off the GPU — so the derived footprint below overstates what must be resident in VRAM
Weights99 GB at Q4180 GB at FP88x80GB node6B active per tokenderived · weights only
Runtimes
Released
27 Aug 2026
Context
256K (~1M with YaRN)
Modalities
textvision
Base model
—
PriceSelf-hosted — infrastructure cost (open weights)
Details

Named as part of the 3.8 line but built on a Qwen4 experimental architecture — the repository ships Qwen4ExpForConditionalGeneration and a qwen4_exp model type — so read it as an early look at the next generation rather than a Flash variant. 512 experts with 10 firing per token is unusually fine-grained; the active parameter count is not published, so this card does not estimate one. Note the two differences from the hosted Flash it appears to extend: 256K context rather than 1M, and a custom licence rather than Apache-2.0.

License: Qwen Community Licence 1.0 — read it, because the sibling 27B is Apache-2.0 and this is not

Deployment: Self-hosted from Hugging Face in BF16 or FP8; an NVFP4 build at ~135 GiB runs on two DGX Sparks under vLLM with the n-gram table mmap'd from NVMe verified Sep 2026

Reference →

GLM-5.3-Flash

🇨🇳 China
Z.ai (Zhipu AI)
GeneralLLMOpen weightshybrid linear + sparse attentionOn-prem
Parameters320B
MoE — 320B total, 18B active; 45 layers against GLM-4.5's 92
Weights176 GB at Q4320 GB at FP88x80GB node18B active per tokenderived · weights only
Runtimes
Released
26 Aug 2026
Context
1M
Modalities
textvisionvideo
Base model
—
PriceNo per-token list price in the announcement — Z.ai claims one-tenth of GLM-5.2
Details

The first natively multimodal model in the GLM-5 series, and the cost story rather than the capability story: 320B total against 18B active, with a hybrid of linear and sparse attention plus an IndexPool compression step that cuts attention compute 3.0x and KV cache 4.4x against GLM-5.3 — the latter being the number that decides whether a 1M context is affordable to serve. Z.ai reports 57 on the Artificial Analysis index, DeepSWE 63.4 against GLM-5.2's 46.2, and 29.0 on its in-house code bench against Claude Opus 4.8's 29.5. Two details worth noting: it was tested anonymously as ox-alpha on OpenCode and OpenRouter before launch, and Z.ai says that traffic was served entirely on Chinese AI chips. Independent counterpoint on serving, from Costin Raiu's September 2026 CTI evaluation on two DGX Sparks: he measured 20–25 tokens/sec with slow prefill, 30–45 seconds before reasoning appeared, and judged it very capable but too slow for his workflows — with the caveat that he tested the initial release and expects it has improved.

License: MIT on Hugging Face — genuinely MIT, unlike GLM-5.3, whose weights carry a model-specific licence

Deployment: Self-hosted via SGLang, vLLM and TokenSpeed; Z.ai API, GLM Coding Plan and ZCode verified Sep 2026

Reference →

Qwen3.8-27B

🇨🇳 China
Alibaba
GeneralLLMOpen weightsOn-prem
Parameters28B
Dense — 27.8B, 64 layers, on the Qwen3.5 architecture; the predecessor Qwen3.6-35B-A3B was MoE at 35B with 3B active
Weights15 GB at Q428 GB at FP8One 24GB GPUderived · weights only
Released
14 Aug 2026
Context
256K
Modalities
textvision
Base model
—
PriceSelf-hosted — infrastructure cost (open weights)
Details

The open workhorse of the 3.8 line, succeeding Qwen3.6-35B-A3B, and the one entry here that most people can actually run: 28B dense and multimodal, at 4.5M downloads within a fortnight. Dense is the trade — the 3.6 predecessor was a 35B MoE firing 3B per token, so this asks for less memory but more compute per token. Apache-2.0, unlike the rest of the 3.8 line.

License: Apache-2.0 — note the contrast with Flash-Next, which ships under a custom licence

Deployment: Self-hosted from Hugging Face; QwenCloud API verified Sep 2026

Reference →

GLM-5.3

🇨🇳 China
Z.ai (Zhipu AI)
GeneralLLMOpen weightsOn-prem
Parameters753B
MoE — 753B total, ~40B active; the same base model as GLM-5.2
Weights414 GB at Q4753 GB at FP88x80GB node40B active per tokenderived · weights only
Runtimes
Released
14 Aug 2026
Context
1M
Modalities
text
Base model
—
PriceNo list price published for 5.3 at launch — GLM-5.2 was $1.40 in / $4.40 out
Details

Same base model as GLM-5.2, released 14 August 2026 — every gain comes from post-training on long-horizon task environments. Z.ai reports Terminal-Bench 3.0 rising from 4.6 to 28.3 and DeepSWE from 46.2 to 66.9. The reason it matters here is the cyber capability, which Z.ai describes as emergent and faster-growing than expected: 84.5% on CyberGym, the best result on that benchmark ahead of Claude Mythos 5 and GPT-5.6 Sol, and more than double GLM-5.2 on exploitation. It stays well behind the closed frontier further up the chain — 105 ExploitGym tasks in two hours against Mythos 5's 181. Run against real codebases with Chinese security teams it surfaced 2,436 vulnerabilities across 269 open-source projects, 1,097 medium-to-high and 107 critical, the oldest introduced in 1981 and the average living 26.6 years before discovery; 53 are public so far through Z.ai's disclosure ledger, the rest under embargo. Thinking can no longer be disabled: effort is low, high or max, which is a breaking change for callers. The weights arrived in late August, roughly on the promised schedule but under a model-specific licence rather than MIT.

License: Weights published on Hugging Face under a model-specific glm-5.3 licence — not the MIT the announcement promised, so read it before deploying

Deployment: Self-hosted via SGLang, vLLM, TokenSpeed, Transformers, KTransformers, Unsloth and Ascend NPU stacks; Z.ai API (model ID glm-5.3), GLM Coding Plan and ZCode verified Sep 2026

Reference →

DeepSeek V4 Pro

🇨🇳 China
DeepSeek
GeneralLLMOpen weightsOn-prem
Parameters1.6T
MoE — 1.6T total, 49B active, as released for the V4 Pro checkpoint
Weights880 GB at Q41.6 TB at FP8Multi-node49B active per tokenderived · weights only
Released
13 Aug 2026
Context
1M (384K output)
Modalities
text
Base model
—
Price$1.32 in · $3.96 outper 1M tokas of Oct 2026
Details

DeepSeek's frontier tier: the V4 architecture (April 2026) pairs a sparse-attention stack that makes 1M context economical with the top open-weight scores on agentic real-world-work benchmarks. Successor to the V3 / R1 line, refreshed in June for agentic stability and long tool-use horizons, and again as the 0813 build in August. The price is the headline: input fell from $1.74 to $0.435 and output from $3.48 to $0.87 between April and August, which puts a claimed-frontier model an order of magnitude under US list prices, before the off-peak halving. DeepSeek now describes V4 Pro as being phased out in favour of a V4.1 Pro it has announced without a date, and reports V4.1 Flash ahead of it.

Pricing: DeepSeek API list price, peak hours; cached input $0.044/1M. Off-peak is half: $0.66 / $1.98, cache hits $0.022. These are the rates DeepSeek's pricing page lists for deepseek-v4-pro, about three times what this card recorded in August, but DeepSeek's own notice says those requests now run on V4.1 Flash and bill at Flash rates. Check what your account is actually charged before budgeting

License: V4 Pro weights MIT on Hugging Face; later builds reported API-only

Deployment: V4 Pro checkpoints (the original and 0813) are self-hostable under MIT. DeepSeek's API is contradictory: its pricing page still lists deepseek-v4-pro as build V4-Pro-0813, but its 10 September notice says that from 14 September every deepseek-v4-pro request is served by V4.1 Flash at Flash rates until V4.1 Pro launches, and that V4 Pro is being phased out verified Oct 2026

Reference →

MAI-Code-1.1-Flash

🇺🇸 USA
Microsoft AI
GeneralLLMProprietary
Parameters138B
MoE — 138B total, 5B active, per Microsoft's model card. Microsoft's own Windows blog post of 7 October gives 137B total and 6.8B active for the on-device build without explaining the difference, so treat the active count as uncertain
Released
11 Aug 2026
Context
256K
Modalities
textvisioncode
Base model
—
Price$0.20 in · $1.20 outper 1M tokas of Oct 2026
Details

Successor to MAI-Code-1-Flash, which GitHub deprecated across Copilot on 10 September 2026. It starts from MAI-Thinking-1's compressed 5B-active checkpoint and adds image input, for screenshot-to-code and similar tasks. Vendor-reported in Copilot's own harness: SWE-Bench Verified at 72.6% against 1.0's 71.6% using 8.6K tokens against 10.8K, and Terminal-Bench 2.1 at 62.9% against 51.7%, with Haiku 4.5 at 69.8% and 49.4%. The 7 October update adds a 3-bit on-device build that keeps the 256K context and needs over 120GB of RAM. It is meant as an option for Copilot's router, not a model to self-host, and this card stays closed until weights and a licence are published somewhere they can be read.

Pricing: GitHub Copilot list price; cached input $0.02/1M. 73% below 1.0's $0.75 / $4.50, which is what Microsoft means by a quarter of the cost; the model card itself still says pricing is to be finalised. Microsoft also says local calls carry no inference charge

License: GitHub Copilot and Microsoft product terms; no weights licence published

Deployment: GitHub Copilot (generally available; Copilot cloud agent, Chat, CLI and the Copilot app). Microsoft also says it can be downloaded and run locally, but publishes no download location, weights licence or runtime for doing so

Reference →

Nemotron 3.5

🇺🇸 USA
NVIDIA
GeneralLLMOpen weightshybrid Mamba-transformerOn-prem
Parameters30B – 550B
Hybrid Mamba-Transformer MoE, in NVIDIA's own naming — Lightning 30B A3B, Nano 30B A3B, Nano Omni 30B, Super 120B A12B, Ultra 550B A55B
Weights17 GB – 303 GB at Q430 GB – 550 GB at FP8One 24GB GPU – 8x80GB nodederived · weights only
Released
11 Aug 2026
Context
1M
Modalities
textvisionaudiovideo
Base model
—
PriceSelf-hosted — infrastructure cost (open weights); free demo tiers on OpenRouter, hosted at provider-set rates
Details

NVIDIA's open agentic-AI family with unusual transparency: weights, 10T+ tokens of training data, and the training recipes are all published. Tiered by design — Nano for edge/local (Dec 2025), Super as single-GPU workhorse (Mar 2026), multimodal Nano Omni (Apr 2026), the frontier Ultra (Jun 2026, the leading US open-weight model on the AA Intelligence Index at release), and now Nemotron 3.5 Lightning (11 August 2026), a 30B MoE with 3B active built as the execution layer under Ultra's planning. NVIDIA positions Ultra 550B A55B for the multi-agent work that needs maximum accuracy — planning, code generation and deep research — with Super for hybrid tasks and Nano and Lightning for execution. NVIDIA reports it on the accuracy-versus-speed Pareto frontier of the AA Intelligence Index, up to 4× output speed against comparable models, and 86% on PinchBench while running 10,000 tasks 30% faster than Qwen3.6 35B — vendor-reported. It shipped with NeMo Switchyard, a router for pointing each task at the cheapest model that can do it. Configurable thinking budget throughout; adjacent specialized Nemotron models cover retrieval, document parsing, speech and safety (NemoGuard). One gap for local use: there is no Nemotron 3.5 tag in the Ollama library, which still carries only the Llama-3.1-era Nemotron 70B, so on Ollama this is a manual Modelfile import of a third-party GGUF rather than a pull. The GGUFs themselves exist and are published by ggml-org and Unsloth, so llama.cpp and LM Studio are unaffected.

License: NVIDIA Open Model License / OpenMDW-1.1 (Ultra, Lightning) — weights, training data and recipes all open

Deployment: Self-hosted (vLLM / SGLang / TensorRT-LLM / llama.cpp / LM Studio / Unsloth), local on Jetson, GeForce RTX 5090 and DGX Spark, plus NVIDIA NIM, AWS Bedrock and many hosted providers verified Aug 2026

Reference →

GPT-5.6-Cyber

🇺🇸 USA
OpenAI
CybersecurityLLMGatedDomain fine-tuneDaybreak Red
Parametersundisclosed
Released
10 Aug 2026
Context
—
Modalities
textcode
Base model
GPT-5.6 Sol
Price$12.50 in · $75 outper 1M tokas of Sep 2026
Details

Successor to GPT-5.5-Cyber (May 2026) and, unlike it, a trained variant rather than a pure access tier: built on GPT-5.6 Sol to raise capability on zero-day discovery and exploit-chain development and to refuse less on dual-use work. On OpenAI's internal Advanced Cybersecurity Completion Rate it answers 95.0% of advanced requests, against 57.3% for GPT-5.5-Cyber, 2.0% for Sol under Daybreak Blue and 1.5% for guardrailed Sol — vendor-reported, as are its gains on ExploitGym. OpenAI also reports where it loses: Sol writes better vulnerability reports and is more token-efficient on ExploitBench within the standard 300-turn budget. Assessed High but below Critical for cyber capability under the Preparedness Framework; system card to follow. That ceiling has since been broken by GPT-6 Astra, which shipped on 4 September 2026 as the first Critical cyber model under the same framework and now has its own entry here. Used to find CVE-2026-15903 in Chrome's V8 plus a second chained bug, and 400+ privilege-escalation issues in an unnamed OS kernel.

Pricing: OpenAI API list price — Daybreak Red approval required; cached input $1.25/1M. Cache writes $15.625. OpenAI publishes no long-context rate for it. The alias gpt-daybreak-red-latest currently points here, and OpenAI says the alias's price will follow whatever model it is moved to

License: Daybreak Red — approved individuals and organisations

Access — Daybreak Red: Approved individuals and organisations doing authorised vulnerability research, exploit validation and security testing. Controlled through identity verification, account security, monitoring, approved-use restrictions and legal attestations; hardware security keys required for individual accounts from 1 September 2026 (apply)

Deployment: Restricted API access via OpenAI Daybreak Red; OpenAI recommends running it in Codex auto-review mode

Reference →

GPT-Daybreak-Blue

🇺🇸 USA
OpenAI
CybersecurityLLMGatedAccess-tier variantDaybreak Blue
Parametersundisclosed
Released
10 Aug 2026
Context
1.05M (128K output)
Modalities
textvision
Base model
GPT-5.6 Sol
Price$4 in · $20 outper 1M tokas of Sep 2026
Details

Not a model OpenAI trained but a door onto one: an alias for its flagship general-purpose model with the system-level screening of cybersecurity requests removed, introduced alongside Daybreak Red and GPT-5.6-Cyber in August 2026 and recommended by OpenAI as the starting point for most defenders. The card exists because the alias is named, priced and documented as its own model, and because the model behind it is no longer carded here — GPT-5.6 Sol was replaced in place by GPT-6 Sol, since updated to GPT-6.1 Sol. The alias has not moved — it still resolves to GPT-5.6 Sol — but the tier has partly followed the family forward: the GPT-6 system card's appendix on Sol and Luna, added 22 September, says eligible Blue users can use GPT-6 Sol and Luna with reduced cyber refusals. Nothing yet says the same of GPT-6.1 Sol, which replaced GPT-6 Sol a week later. On GPT-6 Astra, OpenAI's Daybreak help article gives reduced refusals to Red customers only and describes extension to Blue as in progress, although the system card already reports Astra's completion rates under Blue. What removing the screen buys is modest by OpenAI's own figures: on its Advanced Cybersecurity Completion Rate, Sol under Blue answers 2.0% of advanced requests against 1.5% guardrailed and 95.0% for GPT-5.6-Cyber — vendor-reported. It still beats the Cyber model on ExploitBench within the standard 300-turn budget, where it is more token-efficient. Assessed High but below Critical for cyber capability under the Preparedness Framework, as GPT-5.6 Sol.

Pricing: OpenAI API list price — GPT-5.6 Sol's promotional rate; cached input $0.40/1M. Cache writes $5. Over 272K input: $8 input, $0.80 cached, $30 output. OpenAI says the promotional rate runs at least through 21 November 2026, and that the alias's price will follow whatever model it is moved to

License: Daybreak Blue — approved individuals and organisations

Access — Daybreak Blue: Approved defenders doing authorised defensive work: vulnerability discovery, secure code review, threat modelling, detection engineering, malware analysis in a controlled environment, and patch validation. Requested through Trusted Access for Cyber and provisioned per identity, workspace or API project, and individual members must enable Advanced Account Security or receive standard access; approval for Blue does not grant Red. It removes the system-level screening of cyber requests and nothing else — the model still refuses highly dual-use work such as pentesting production systems, which is what Red and GPT-5.6-Cyber are for (apply)

Deployment: OpenAI API as gpt-daybreak-blue-latest, currently resolving to gpt-5.6-sol, from a non-default project with Daybreak access enabled; Codex CLI with codex -m gpt-daybreak-blue-latest. In ChatGPT and Codex under ChatGPT sign-in it is a workspace entitlement rather than a model in the picker. OpenAI now marks the alias deprecated and tells users to move to the latest cyber model available to them, naming no replacement and no removal date

Reference →

Muse Glimmer

🇺🇸 USA
Meta (Superintelligence Labs)
GeneralLLMOpen weightsOn-prem
Parameters30B
Weights17 GB at Q430 GB at FP8One 24GB GPUderived · weights only
Runtimes
Released
10 Aug 2026
Context
—
Modalities
textvision
Base model
Muse Spark (logit distillation)
PriceSelf-hosted — infrastructure cost (open weights)
Details

Distilled from Muse Spark's outputs into 30B of compact architecture aimed at local hardware — Meta returning to open weights, in a smaller lane, a year after Muse Spark closed the Llama era. Agentic and multimodal: a dedicated perception encoder reads screenshots, charts and documents from interleaved text and images, and a DFlash-based drafter model does speculative decoding. Meta reports it ahead of Gemma4-31B and Qwen3.6-27B on agentic, coding, multimodal, safety and reasoning suites (DeepSearch QA, MCP-Atlas, τ-Bench, SWE-Bench) — vendor-reported, on models of its own choosing.

License: Apache 2.0 — weights on Hugging Face

Deployment: Local / self-hosted; weights on Hugging Face. Also hosted on OpenRouter (meta/muse-glimmer-30b) at $0.35 / $1.50 per 1M verified Sep 2026

Reference →

Qwen3.8-Max

🇨🇳 China
Alibaba
GeneralLLMOpen weightsOn-prem
Parameters2.4T
MoE — 2.4T total, 95B active
Weights1.3 TB at Q42.4 TB at FP8Multi-node95B active per tokenderived · weights only
Released
Aug 2026
Context
1M (991K input, 131K output)
Modalities
textvisionvideo
Base model
—
Price$2 in · $6 outper 1M tokas of Aug 2026
Details

Alibaba's flagship, on the Qwen 3.5 architecture, and the first Max-class Qwen to be open-weighted — though what shipped is not quite what was announced. The open repository, Qwen3.8-2.4T-A95B, is text-only and always thinking, natively 262K context extensible to about 1M; the hosted Max adds vision and lets thinking be switched off. Self-hosting it therefore means running the base, not the product. The pitch is autonomy over days: Alibaba reports a 16-day unattended run producing 265 commits, 127 PRs and 151 issues on a public repository, a five-day reproduction of a research paper that then beat it by 2.7 points on AIME24, and a 24-hour Tianchi contest entry placing ahead of 458 of 526 human teams — all vendor-reported. Its agentic RL environments span its own harness plus Claude Code, Codex, OpenClaw and Hermes, which is a notable amount of training against other vendors' scaffolds.

Pricing: QwenCloud list price (DashScope international); cached input $0.25/1M. Implicit caching is the $0.25 rate shown; explicit caching bills $2.50 to create and $0.17 to read, so it only pays back on heavy reuse. Reasoning is metered separately up to 262K tokens. Rate limits 2M TPM / 15K RPM. Same headline rate as Grok 4.6

License: Open weights as Qwen3.8-2.4T-A95B under a model-specific qwen3.8-max licence — a different artifact from the Max API, which adds vision and optional thinking

Deployment: Self-hosted as Qwen3.8-2.4T-A95B (BF16 and FP8); QwenCloud API (model ID qwen3.8-max, DashScope international endpoint) verified Sep 2026

Reference →

Qwen3.8-Flash

🇨🇳 China
Alibaba
GeneralLLMProprietary
Parametersundisclosed
Released
Aug 2026
Context
1M (131K output)
Modalities
textvisionvideo
Base model
—
Price$0.15 in · $0.47 outper 1M tokas of Aug 2026
Details

The cheap multimodal tier, and the price is the point: $0.15 in and $0.47 out against Max's $2 and $6, with a native 1M context and image and video input. Alibaba pitches it at coding, agentic workflows and visual understanding. Hosted only — the similarly named Qwen3.8-Flash-Next is a separate open-weight release on a different architecture, not this model's weights.

Pricing: QwenCloud list price (DashScope international); cached input $0.016/1M. Explicit cache creation costs $0.20; the $0.016 shown is the implicit cache rate

License: QwenCloud only — no public weights are published under this name

Deployment: QwenCloud API (DashScope international endpoint)

Reference →

Granite 4.2

🇺🇸 USA
IBM
GeneralSLMOpen weightsOn-prem
Parameters3B – 30B
Dense — 3B, 8B and 30B; the architecture changed from Granite 4's hybrid Mamba-2 / transformer MoE
Weights2 GB – 17 GB at Q43 GB – 30 GB at FP8One 24GB GPUderived · weights only
Runtimes
Released
Aug 2026
Context
128K (512K on the 30B)
Modalities
text
Base model
—
PriceSelf-hosted — infrastructure cost (open weights)
Details

IBM's enterprise family, now built around reasoning: chain-of-thought inside think tags, three switchable modes — full thinking, non-thinking and low-effort — chosen per query rather than per deployment, and tool calling where the model reasons about which tool to call before calling it. Two things changed with 4.2 and both matter for self-hosting: the models are dense rather than the hybrid Mamba-2 MoE of Granite 4, so every parameter is resident and active, and the 30B extends to 512K context. Tested across twelve languages. The governance properties are the reason it appears in regulated environments at all — Apache 2.0, signed weights, ISO/IEC 42001 — and Granite remains the in-house tier of IBM Bob's routing alongside Claude and Mistral.

License: Apache 2.0, with cryptographically signed weights and ISO/IEC 42001-accredited process documentation

Deployment: Self-hosted (Transformers and the usual runtimes), Ollama (granite4.2:30b), watsonx.ai, Hugging Face, major clouds verified Sep 2026

Reference →

MiniMax-M3

🇨🇳 China
MiniMax
GeneralLLMOpen weightssparse attentionOn-prem
Parameters428B
MoE — ~428B total, ~23B active per token; 60 layers with 128 routed experts
Weights235 GB at Q4428 GB at FP88x80GB node23B active per tokenderived · weights only
Released
23 Jul 2026
Context
1M
Modalities
textvisionvideo
Base model
—
PriceSelf-hosted — infrastructure cost (open weights)
Details

Natively multimodal from the first training step rather than by bolting a vision encoder on, covering text, image and video, with a 1M context made affordable by MiniMax Sparse Attention: the project claims 9x prefill and 15x decode speedups over M2 at 1M context, cutting per-token compute to a twentieth. That efficiency claim is the reason to care — long context is cheap to advertise and expensive to serve. Worth knowing where these weights travel: Saudi Arabia's PIF-backed Humain announced humain-m3 on 3 September 2026, an Arabic model built on M3 and distributed through its own platform, which has no public repository and so gets no entry here. MiniMax also ships H3 for video generation, out of scope for a language-model registry despite being one of the most downloaded models on Hugging Face.

License: MiniMax community licence — a custom agreement, not Apache or MIT, so read the terms before commercial use

Deployment: Self-hosted via vLLM or SGLang, which both publish M3 recipes; an MXFP8 build is also released, and it is packaged in the Ollama library verified Sep 2026

Reference →

Cisco Antares

🇺🇸 USA
Cisco Foundation AI
CybersecuritySLMGatedDomain fine-tuneOn-premCisco gated download
Parameters350M – 1B
350M and 1B published; the 3B is in the technical report and trained, but Cisco's launch post said only “coming soon” and it is still not on Hugging Face as of September 2026
Weights0 GB – 1 GB at Q40 GB – 1 GB at FP8One 24GB GPUderived · weights only
Released
Jul 2026
Context
128K (32K on the 350M)
Modalities
textcode
Base model
IBM Granite 4.0 — the 350M, 1B and Micro checkpoints respectively
PriceSelf-hosted — infrastructure cost (open weights, Apache 2.0)
Details

Family of security SLMs for vulnerability localization: given a vulnerability description, the model explores a repository over read-only shell access and produces a ranked list of likely vulnerable files with its search trail. The technical report is explicit that all three sizes are initialized from IBM Granite 4.0 checkpoints and then post-trained — SFT on terminal navigation and security reasoning, then GRPO over whole agent trajectories — which is why this card reads fine-tune rather than purpose-built; nothing here was trained for the domain from scratch. No baseRef, because the registry carries Granite 4.2 and this was cut from 4.0, and a link would assert a lineage that does not exist. The headline claim is about the size nobody can download: Antares-3B reaches 0.223 File F1 on Cisco's VLoc Bench, approaching GPT-5.5 and ahead of GLM-5.2 at 753B — a model some 250 times larger. Both comparison points are a version behind the GPT-5.6 and GLM-5.3 carried here, so read the margin as of mid-2026. Full runs of 500 repositories in ~15 minutes for under $1.

License: Apache 2.0 on both released repositories — the gate is Hugging Face's automatic kind, not a review

Access — Cisco gated download: Accept the terms on the model repository and Hugging Face grants access automatically. It is a click-through, not a vetting decision: no one reviews the request, and the 1B has 35,584 downloads in the last 30 days (apply)

Deployment: Local / on-prem / air-gapped; Hugging Face, after accepting the model terms verified Sep 2026

Reference →

MAI-Cyber-1-Flash

🇺🇸 USA
Microsoft AI
CybersecuritySLMProprietaryDomain fine-tune
Parametersundisclosed
Released
Jul 2026
Context
—
Modalities
textcode
Base model
MAI-Thinking-1 (lineage)
PriceNo standalone list price — bundled in MDASH; Microsoft claims 50% cost saving vs its prior GPT-5.4-based stack
Details

Microsoft's first cyber model (July 2026), built to find hard vulnerabilities in complex codebases and trained on Microsoft's security estate — 100T+ daily security signals and decades of MSRC exploit/remediation records. Architecturally a cost-router play: it handles up to ~90% of tasks, escalating the hardest ~10% to GPT-5.4; the combined MDASH system scores ~96% on CyberGym, which Microsoft reports as 12 points above Claude Mythos and ahead of Gemini and GPT standalone. Security-first calibration, AI Red Team and third-party assessed; runs in sandboxed, internet-isolated environments with tenant isolation and RBAC.

License: Delivered inside Microsoft's MDASH security harness

Deployment: Inside MDASH (multi-agent vulnerability identification and remediation harness); Perception agentic security workflows to follow

Reference →

Inkling

🇺🇸 USA
Thinking Machines Lab
GeneralLLMOpen weightsOn-prem
Parameters952B
952B total, counted from the published checkpoint rather than the announcement's 975B
Weights524 GB at Q4952 GB at FP88x80GB nodederived · weights only
Released
Jul 2026
Context
1M
Modalities
textvisionaudio
Base model
—
PriceSelf-hosted — infrastructure cost (open weights); hosted APIs at provider-set rates
Details

First model from Thinking Machines Lab (Mira Murati), trained from scratch on 45T tokens and released open-weight explicitly as a base for customization rather than a leaderboard chaser: native encoder-free multimodality (text / image / audio), controllable thinking effort (0.2–0.99) that trades tokens for performance, and fine-tuning as the product via the Tinker platform. Notably trained for calibration and epistemics — RL against proper scoring rules on resolved forecasting questions — and posts the strongest built-in safeguards among compared open-weight models on the FORTRESS adversarial benchmark. Inkling-Small followed two weeks after the July 15 release.

License: Open weights on Hugging Face (original + NVFP4 checkpoints)

Deployment: Self-hosted (SGLang / vLLM / llama.cpp), fine-tuning on Tinker, hosted APIs via Together, Fireworks, Modal, Databricks, Baseten verified Sep 2026

Reference →

Kimi K3

🇨🇳 China
Moonshot AI
GeneralLLMOpen weightsOn-prem
Parameters2.8T
MoE — 2.8T total (896 experts, 16 active per token); largest announced open-weight model
Weights1.5 TB at Q42.8 TB at FP8Multi-nodederived · weights only
Runtimes
Released
Jul 2026
Context
1M
Modalities
textvision
Base model
—
Price$3 in · $15 outper 1M tokas of Oct 2026
Details

Moonshot's frontier model with always-on thinking mode and native vision — independent testing at launch placed it just behind Claude Fable 5 and GPT-5.6 Sol, ahead of Opus 4.8, leading several coding benchmarks. The K2.x line (K2.7 Code, K2.6) remains the cheaper production tier. The full checkpoint did land, on 2 September 2026: 96 safetensors shards totalling 2.78T parameters, ungated, verified at the repository. At that size self-hosting is realistic only at datacenter scale.

Pricing: Kimi API list price; cached input $0.30/1M. Unchanged since launch. Caching is automatic, with cache hits at $0.30; Moonshot also lists cache writes at $3 (5-minute) and $6 (1-hour)

License: Model-specific kimi-k3 licence, not MIT despite the shape. Use, modification, distribution and self-hosting are granted freely; two conditions bite only at scale. A Model-as-a-Service business whose revenue passes $20M over any consecutive 12 months must sign a separate agreement with Moonshot before commercial use, and a product above 100M monthly active users or $20M monthly revenue must display “Kimi K3” in its interface. Internal use is exempt from both.

Deployment: Kimi API / apps, OpenRouter; self-hosting is datacenter-scale (~64+ accelerators) verified Sep 2026

Reference →

Gemma 4

🇺🇸 USA
Google
GeneralSLMOpen weightsOn-prem
Parameters2B – 31B
E2B and E4B for edge, quoted as effective parameters, plus 12B, 26B and 31B; the 31B instruction-tuned card reports 30.7B total
Weights1 GB – 17 GB at Q42 GB – 31 GB at FP8One 24GB GPUderived · weights only
Released
Jul 2026
Context
256K
Modalities
textvision
Base model
—
PriceSelf-hosted — infrastructure cost (open weights)
Details

Successor to Gemma 3, built from Gemini 3 research and pitched on intelligence per parameter. The licence is the headline for anyone who had to route around the Gemma terms: the 31B card ships Apache-2.0. Google reports 1452 Elo on Arena against Gemma 3 27B's 1365, with MMLU 85.2%, AIME 2026 89.2% and LiveCodeBench 80.0% — vendor-reported. The family page claims audio as well as visual understanding, while the 31B card documents text and image input only, so treat audio as tier-dependent until the smaller cards confirm it.

License: Apache-2.0 — a real change: the family ran on the custom Gemma licence through Gemma 3

Deployment: Self-hosted and edge; Hugging Face, Ollama, LM Studio, Kaggle, Docker, Vertex AI verified Sep 2026

Reference →

MAI-Thinking-1

🇺🇸 USA
Microsoft AI
GeneralLLMProprietary
Parametersundisclosed
Released
Jun 2026
Context
256K
Modalities
text
Base model
—
PricePreview — pricing not finalized at launch
Details

Microsoft AI's first in-house reasoning model (Build 2026), trained from scratch on commercially licensed data with no distillation from OpenAI or Anthropic — the clearest step in Microsoft's move to reduce reliance on OpenAI. Mid-weight positioning: vendor-reported SWE-Bench Pro parity with Claude Opus 4.6 and 97% on AIME 2025 at far lower cost; blind human evaluations by Surge preferred it over Sonnet 4.6. Also the lineage base for MAI-Cyber-1-Flash.

License: Microsoft Foundry (preview); OpenRouter / Fireworks / Baseten rollout announced

Deployment: Microsoft Foundry, Copilot ecosystem

Reference →

Step 3.7 Flash

🇨🇳 China
StepFun
GeneralLLMOpen weightsOn-prem
Parameters198B
MoE — 198B total (incl. 1.8B vision encoder), 11B active
Weights109 GB at Q4198 GB at FP88x80GB node11B active per tokenderived · weights only
Runtimes
Released
May 2026
Context
256K
Modalities
textvisionvideo
Base model
—
Price$0.20 in · $1.15 outper 1M tokas of Sep 2026
Details

Multimodal agent model from StepFun built for speed: multi-token prediction heads push serving past 400 tokens/s — more than double peers of its size — with low/medium/high reasoning levels. Strong agentic, search and coding benchmarks among open models at launch.

Pricing: OpenRouter hosted rate; also fully self-hostable (Apache 2.0); cached input $0.04/1M. Cache reads $0.04/1M on the same route

License: Apache 2.0

Deployment: Self-hosted (BF16 / FP8 / GGUF — a quantised build fits a single 128GB unified-memory machine), hosted via OpenRouter and others verified Sep 2026

Reference →

Mistral Small 4

🇫🇷 France
Mistral AI
GeneralLLMOpen weightsOn-prem
Parameters119B
MoE — 119B total, 128 routed experts plus one shared, 4 active per token; the card names it 119B A6B
Weights65 GB at Q4119 GB at FP8One 80GB GPU6B active per tokenderived · weights only
Released
Mar 2026
Context
1M
Modalities
textvision
Base model
—
Price$0.15 in · $0.60 outper 1M tokas of Sep 2026
Details

Europe's leading open-weight vendor, and the entry that most rewards reading the parameter count: "Small" moved from 24B dense to 119B sparse between 3.1 and 4, moving the self-hosting floor from a 24GB card to an 80GB one even though only ~6B fire per token, because memory needs the total and not the active count. Version 4 folds three former families into one model: Instruct, Reasoning (previously Magistral) and Devstral, with mode switching rather than separate checkpoints. Mistral reports a 40% cut in end-to-end completion time and 3x the requests per second against Mistral Small 3; both are vendor figures. Relevant where EU data-sovereignty requirements apply.

Pricing: La Plateforme list price; also self-hostable (Apache 2.0). Batch processing halves it; Mistral documents cached input reducing input cost by up to 90%, without publishing a per-token rate

License: Apache 2.0, verified at the repository — no added conditions, unlike most frontier open weights

Deployment: Self-hosted from Hugging Face (BF16 and an NVFP4 build), vLLM; La Plateforme, OpenRouter and the major clouds verified Sep 2026

Reference →

Gemini 3.1 Pro

🇺🇸 USA
Google DeepMind
GeneralLLMProprietary
Parametersundisclosed
Runtimes
Released
Feb 2026
Context
1M
Modalities
textvisionaudiovideopdf
Base model
—
Price$2 in · $12 outper 1M tokas of Sep 2026
Details

Successor to Gemini 3, and Google's current Pro listing: its own model catalogue names it Gemini 3.1 Pro against endpoint gemini-3.1-pro-preview and still marks it Preview — seven months after listing, with no free tier. Takes text, image, audio, video and PDF in, text only out, at 1M context with a 64K output ceiling. Worth reading beside the Flash entries for where Google's attention actually goes: the Pro line has sat at a preview point release while Flash shipped three versions in the six weeks to September, and 3.8 Flash is the one Google promotes on its own pricing page. Prices are unchanged from Gemini 3 and were re-verified at source in September 2026.

Pricing: Gemini API list price (≤200K-token prompts); cached input $0.20/1M. $4 in / $18 out above 200K context; context caching $0.20 below 200K and $0.40 above, plus $4.50 per 1M tokens per hour of cache storage

License: API / Google products; still labelled Preview by Google and not offered on the free tier

Deployment: Gemini API (model ID gemini-3.1-pro-preview), AI Studio, Vertex AI, Google apps; also OpenRouter verified Sep 2026

Reference →

Sakana Fugu Cyber

🇯🇵 Japan
Sakana AI
CybersecurityLLMProprietaryPurpose-built
Parametersundisclosed
Released
2026
Context
—
Modalities
text
Base model
Undisclosed pool of models
PriceNo list price published — Sakana says to contact sales for Fugu Cyber usage and pricing; see the Sakana Fugu entry for the rest of the line's rates
Details

The security tier of Sakana's Fugu orchestration line, aimed at security analysis, vulnerability research and threat investigation. Sakana reports 86.9% on CyberGym and 72.1% on CTI-REALM — vendor-reported, and worth reading against MDASH's ~96% on the same benchmark and the other cyber entries here. Read the comparison it offers with its date attached: Sakana calls the tier comparable to GPT-5.5-Cyber and Mythos-Preview, both a version behind the GPT-5.6-Cyber and Claude Mythos 5.1 carried here, so the claim is about a frontier that has since moved. It is also the one cyber model on this page whose door is commercial rather than vetted — no Daybreak, no Fairwind, no trusted-tester list, just a sales conversation — which is a different kind of restriction and should not be mistaken for the others. Classified purpose-built with a caveat: nothing was trained for the domain, the system is assembled and coordinated for it, which is a fourth kind of specialisation this taxonomy does not yet have a word for. Same EU/EEA unavailability as the rest of the line.

License: Hosted API; the security tier of the Fugu line

Deployment: OpenAI-compatible API, direct, and not self-service: Sakana directs Fugu Cyber enquiries to its sales team, while the other tiers sign up through the console. Still not on OpenRouter — re-checked September 2026, which now lists Fugu Ultra, Fugu Ultra v2, Fugu Max and Namazu, and no Cyber tier. Not available in the EU/EEA pending GDPR and EU regulatory compliance verified Sep 2026

Reference →

Sakana Fugu

🇯🇵 Japan
Sakana AI
GeneralLLMProprietary
Parametersundisclosed
Runtimes
Released
2026
Context
—
Modalities
text
Base model
Undisclosed pool of models
Price$5 in · $30 outper 1M tokas of Sep 2026
Details

A multi-agent system sold as a single model: one API call, behind which Fugu assembles agents from a pool and coordinates them, with the routing and the constituent models deliberately not exposed. That makes it an awkward fit for a model registry — parameters, context and lineage are not disclosed because they are not properties of one artifact — but it is priced and consumed as a model, so it sits here with those fields empty rather than guessed. Four tiers as of September 2026: Fugu for everyday work, Fugu Ultra coordinating more expert agents for hard multi-step problems, Fugu Max as the cost-performance tier added in September, and Fugu Cyber, which is the only one you cannot sign up for. Grounded in two ICLR 2026 papers, TRINITY and the Conductor, on learned rather than hand-designed orchestration. For EU readers the availability line is the operative fact.

Pricing: Fugu Ultra token plan list price; cached input $0.50/1M. Above 272K context the Ultra rates double to $10 in / $45 out / $1.00 cached. Fugu Max bills at $2 in / $6 out / $0.25 cached at every context length, plus $0.007 per web_search or fetch call. Base Fugu bills at the underlying model's own rate instead. Subscriptions run $20, $100 and $200 a month for 1×, 10× and 20× usage and cover Fugu, Ultra and Max — not Cyber

License: Hosted API; subscription or token billing

Deployment: OpenAI-compatible API, direct and through OpenRouter, Vercel and others. OpenRouter carries Fugu Ultra, Fugu Ultra v2, Fugu Max and Namazu — the latter two listed 11 September 2026. Not available in the EU/EEA — Sakana says it is still working toward GDPR and EU regulatory compliance verified Sep 2026

Reference →

Amazon Nova

🇺🇸 USA
Amazon (AWS)
GeneralLLMProprietary
Parametersundisclosed
Runtimes
Released
2026
Context
1M (Premier, Nova 2 Lite); 300K Pro and Lite; 128K Micro
Modalities
textvisionvideo
Base model
—
Price$2.50 in · $12.50 outper 1M tokas of Sep 2026
Details

AWS's own family, and the last major cloud vendor missing from this registry. Only the understanding tier is listed here — Canvas generates images, Reel video, Sonic speech and Nova Multimodal Embeddings vectors, none of which are language models. The generation is mid-transition: Nova 2 Lite and Nova 2 Sonic have shipped, while Premier, Pro and Micro remain Nova 1. Two limits matter more than the benchmark talk. Output is capped at 10K tokens on the Nova 1 understanding models, against 64K–131K on comparable frontier models, which rules out long single-shot generation whatever the 1M input window suggests. And Bedrock Guardrails apply to text only on the multimodal tiers, so image and video inputs bypass the control most AWS governance stories are built on. In its favour: GovCloud availability, 200+ languages with 15 optimised, and a distillation ladder productised in Bedrock — Premier teaches Pro, Lite and Micro. Nova 2 Lite adds extended thinking with three intensity levels, built-in web grounding and a code interpreter, and both SFT and reinforcement fine-tuning. Nova Forge, announced alongside, offers to build an organisation its own frontier model on Nova foundations. Bedrock's pricing tables also list Nova 2 Omni in preview, which takes audio alongside text, image and video — the first Nova understanding model to do so, and priced per input modality rather than as one rate. Business Insider reported on 28 July 2026, and Reuters relayed, that AWS is winding down Nova Premier, Omni, Reel and Canvas to maintenance-only and moving effort to a new frontier model expected at re:Invent. AWS's Bedrock lifecycle page has since retired Canvas and Reel (30 September) but lists no end-of-life date for Premier as of October 2026. So the headline price above is still live, but it is for a tier with no development behind it.

Pricing: Amazon Bedrock list price for Nova Premier, geo cross-region and in-region tier. The ladder below it: Nova 2 Lite $0.30 / $2.50 on global cross-region inference, Pro $0.80 / $3.20 or $1.00 / $4.00 with latency-optimised inference, Lite $0.06 / $0.24, Micro $0.035 / $0.14. Bedrock prices the same model differently by region and inference profile, which no other entry here does: Nova 2 Lite alone runs from $0.15 on the batch tier to $0.5775 in the costlier geos, and Premier reaches $4.375 / $21.875. Read a single figure here as one tier of several

License: Amazon Bedrock; no downloadable weights

Deployment: Amazon Bedrock (model IDs amazon.nova-*-v1:0) with cross-region inference, plus OpenRouter routes; Premier, Pro, Lite and Micro also in AWS GovCloud (US-West) verified Sep 2026

Reference →

Foundation-Sec-1.1-8B-Instruct

🇺🇸 USA
Cisco Foundation AI
CybersecuritySLMOpen weightsDomain fine-tuneOn-prem
Parameters8B
8.03B, the Llama 3.1 8B backbone
Weights4 GB at Q48 GB at FP8One 24GB GPUderived · weights only
Released
20 Nov 2025
Context
64K
Modalities
text
Base model
Llama 3.1 8B
PriceSelf-hosted — infrastructure cost (open weights)
Details

Cisco Foundation AI's version 1.1, and the first instruct model in the line to carry the version bump. It is an instruction-tuned assistant for SOC triage, threat-intelligence work and vulnerability prioritisation, built for local deployment. The main change is context: 64K tokens against the 4K Cisco gives for 1.0, which is what makes long incident reports and threat-intelligence feeds usable. Cisco's card describes it as built on a Foundation-Sec-1.1-8B base, but that base is not public; the repository's own metadata points at the original Foundation-Sec-8B, which is still the only public base to fine-tune from. Vendor-reported: CTI-RCM 0.694 against Llama 3.1 8B's 0.558 and GPT-4o-mini's 0.655, and CTI-MCQA 0.644 against 0.617 and 0.672. Cisco itself recommends LlamaGuard in front of it for safety. Foundation-Sec-8B-Reasoning, released alongside, is a sibling variant and not a successor.

License: Hugging Face says `other`, and NOTICE.md explains it: Llama 3.1 Community License for Meta's backbone, Apache 2.0 for Cisco's changes. Unlike the original 8B repo, this one states both, so read both before deploying

Deployment: Local / on-prem; Hugging Face verified Oct 2026

Reference →

gpt-oss-120b

🇺🇸 USA
OpenAI
GeneralLLMOpen weightsOn-prem
Parameters117B
MoE — 117B total, 5.1B active; sibling gpt-oss-20b is 21B total, 3.6B active
Weights64 GB at Q4117 GB at FP8One 80GB GPU5.1B active per tokenderived · weights only
Released
Aug 2025
Context
—
Modalities
text
Base model
—
PriceSelf-hosted — infrastructure cost (open weights); hosted at provider-set rates
Details

OpenAI's open-weight release (August 2025) and still the only one it has made at this scale — relevant here precisely because the rest of its catalogue is closed. Ships MXFP4-quantised, which OpenAI says fits a single 80GB H100 or MI300X, so a frontier-adjacent model lands inside one accelerator rather than a node. Text only, with three reasoning-effort levels set in the system prompt, and built for agentic use: function calling, browsing and Python execution. The 20B sibling targets local and latency-sensitive work.

License: Apache-2.0 — no copyleft, no patent condition

Deployment: Self-hosted — vLLM, llama.cpp, Ollama, LM Studio and Transformers; MXFP4 weights fit a single 80GB accelerator verified Sep 2026

Reference →

Teuken-7B

🇩🇪 Germany
OpenGPT-X (Fraunhofer et al.)
Multilingual / EUSLMOpen weightsPurpose-builtOn-prem
Parameters7B
7.45B
Weights4 GB at Q47 GB at FP8One 24GB GPUderived · weights only
Released
28 Jul 2025
Context
4K
Modalities
text
Base model
—
PriceSelf-hosted — infrastructure cost (open weights)
Details

Publicly funded European model trained from scratch on all 24 official EU languages, built for European digital-sovereignty use cases. This card now describes v0.6, which re-pretrained the base on 6T tokens (base December 2024, instruct July 2025), and the licence went the opposite way to the version number: it is now non-commercial. Anyone needing commercial use is left with v0.4's Teuken-7B-instruct-commercial-v0.4, still on Hugging Face under Apache 2.0. The developers rule out maths and coding tasks.

License: CC-BY-NC-4.0 on both v0.6 repositories, base and instruct: private, non-commercial, research and educational use only. v0.4 shipped an Apache 2.0 commercial instruct variant, and v0.6 has none

Deployment: Self-hosted, Hugging Face verified Oct 2026

Reference →

SmolLM3

🇺🇸 USA
Hugging Face
GeneralSLMOpen weightsOn-prem
Parameters3B
Weights2 GB at Q43 GB at FP8One 24GB GPUderived · weights only
Released
Jul 2025
Context
64K (128K extended)
Modalities
text
Base model
—
PriceSelf-hosted — infrastructure cost (open weights)
Details

Fully open small model — weights, data recipe and training details published — with dual reasoning modes. A reference point for transparent SLM development.

License: Apache 2.0

Deployment: Self-hosted, edge, in-browser; weights on Hugging Face. No Ollama library tag verified Sep 2026

Reference →

MedGemma

🇺🇸 USA
Google
MedicineSLMOpen weightsDomain fine-tuneOn-prem
Parameters4B – 27B
4B multimodal, 27B text
Weights2 GB – 15 GB at Q44 GB – 27 GB at FP8One 24GB GPUderived · weights only
Released
May 2025
Context
—
Modalities
textvision
Base model
Gemma 3
PriceSelf-hosted — infrastructure cost (open weights)
Details

Open medical models for image and text comprehension (radiology, dermatology, pathology, clinical text), intended as a developer starting point rather than a clinical product. A newer MedGemma 1.5 exists on Hugging Face at 4B; this card tracks the 27B text tier, and whether 1.5 reaches that size is worth re-checking.

License: Health AI Developer Foundations terms

Deployment: Self-hosted, Vertex AI verified Sep 2026

Reference →

Sec-Gemini v1

🇺🇸 USA
Google
CybersecurityLLMGatedDomain fine-tuneGoogle trusted tester
Parametersundisclosed
Released
Apr 2025
Context
—
Modalities
textcode
Base model
Gemini
PriceResearch access — no public pricing
Details

Experimental agentic cybersecurity model combining Gemini reasoning with near-real-time threat intelligence (Mandiant, OSV and related feeds) and domain tooling. Targets incident root-cause analysis, threat analysis and vulnerability-impact understanding.

License: Experimental; trusted-tester access

Access — Google trusted tester: Experimental research access, granted to selected testers

Deployment: Hosted (gated research access)

Reference →

Llama 4

🇺🇸 USA
Meta
GeneralLLMOpen weightsOn-prem
Parameters109B – 400B
MoE — Scout 109B (17B active), Maverick 400B (17B active)
Weights60 GB – 220 GB at Q4109 GB – 400 GB at FP8One 80GB GPU – 8x80GB nodederived · weights only
Released
Apr 2025
Context
Up to 10M (Scout)
Modalities
textvision
Base model
—
PriceSelf-hosted — infrastructure cost (open weights)
Details

Meta's first natively multimodal, mixture-of-experts open-weight generation and the base for many derivative fine-tunes. Now also a historical marker: Meta's subsequent frontier family (Muse, 2026) shifted to proprietary, leaving Llama 4 as the last open-weight Llama generation so far.

License: Llama Community License

Deployment: Self-hosted, all major clouds

Reference →

Trend Cybertron

🇯🇵 Japan
Trend Micro
CybersecuritySLMOpen weightsDomain fine-tuneOn-prem
Parameters8B
Weights4 GB at Q48 GB at FP8One 24GB GPUderived · weights only
Released
Mar 2025
Context
—
Modalities
text
Base model
Llama 3.1 8B Instruct
PriceSelf-hosted — infrastructure cost (open weights)
Details

Cybersecurity model open-sourced by Trend Micro in March 2025, aimed at proactive detection, risk assessment and agentic security workflows. The name is the trap: Cybertron is the product, and the weights are published as the Llama Primus collection — Llama-Primus-Base continues pretraining Llama-3.1-8B-Instruct on 2.77B tokens of security text for a reported 15.88% gain across cyber benchmarks, with Merged and Reasoning variants beside it — so searching Hugging Face for the product name finds nothing. The 70B Trend announced alongside the 8B has not appeared in the 18 months since. Also worth a look rather than a claim: trendmicro-ailab published sovereign-v1 in September 2026, Apache 2.0 behind a manually reviewed gate and carrying a nemotron_h config, which would make it a hybrid Mamba-transformer and a different animal from Primus. Trend has said nothing about it here, so it stays uncarded until it does.

License: MIT on the Primus repositories — the weights ship under a name the product does not use

Deployment: Local / on-prem; Hugging Face under trendmicro-ailab as the Primus collection, not under the Cybertron name; NVIDIA NIM through the universal LLM microservice verified Sep 2026

Reference →

Phi-4

🇺🇸 USA
Microsoft
GeneralSLMOpen weightsOn-prem
Parameters14B
Phi-4-multimodal is 5.6B
Weights8 GB at Q414 GB at FP8One 24GB GPUderived · weights only
Runtimes
Released
Dec 2024
Context
16K
Modalities
text
Base model
—
PriceSelf-hosted — infrastructure cost (open weights)
Details

Small model trained heavily on synthetic data with strong math/reasoning for its size. Phi-4-multimodal adds vision and audio in a 5.6B footprint.

License: MIT

Deployment: Self-hosted, Ollama (phi4:14b), Azure AI Foundry, edge verified Sep 2026

Reference →

Falcon 3

🇦🇪 UAE
TII
GeneralSLMOpen weightsOn-prem
Parameters1B – 10B
1B / 3B / 7B / 10B
Weights1 GB – 6 GB at Q41 GB – 10 GB at FP8One 24GB GPUderived · weights only
Runtimes
Released
Dec 2024
Context
32K
Modalities
text
Base model
—
PriceSelf-hosted — infrastructure cost (open weights)
Details

Open family from the UAE's Technology Innovation Institute focused on efficient small models for resource-constrained deployment.

License: TII Falcon License, a custom agreement: Hugging Face reads `other` / falcon-llm-license on the Falcon3 repositories, not Apache or MIT, so read the terms before commercial use

Deployment: Self-hosted, Ollama (falcon3:10b), edge verified Sep 2026

Reference →

Sec-PaLM 2

🇺🇸 USA
Google Cloud
CybersecurityLLMProprietaryDomain fine-tune
Parametersundisclosed
Released
Apr 2023
Context
—
Modalities
textcode
Base model
PaLM 2
PriceNo public token price — bundled into Google Cloud security products
Details

Security-tuned version of PaLM 2 powering Google Cloud's Security AI Workbench: malicious-script analysis, threat explanation and detection, enriched with Google and Mandiant threat intelligence. Historically significant as the first big-vendor security LLM; Google's security stack has since moved toward the Gemini-based SecLM platform.

License: Google Cloud services only

Deployment: Google Cloud (Security AI Workbench)

Reference →

BloombergGPT

🇺🇸 USA
Bloomberg
FinanceLLMProprietaryPurpose-built
Parameters50B
Released
Mar 2023
Context
—
Modalities
text
Base model
—
PriceInternal — Bloomberg products only
Details

Finance-domain LLM trained on decades of Bloomberg's proprietary financial data mixed with general corpora — an early landmark for vertical domain models.

License: Internal / Bloomberg Terminal

Deployment: Bloomberg products only

Reference →

GPT-SW3

🇸🇪 Sweden
AI Sweden
Multilingual / EULLMOpen weightsPurpose-builtOn-prem
Parameters130M – 40B
126M–40B family
Weights0 GB – 22 GB at Q40 GB – 40 GB at FP8One 24GB GPUderived · weights only
Released
2023
Context
—
Modalities
text
Base model
—
PriceSelf-hosted — infrastructure cost (open weights)
Details

Nordic-language model family trained on Swedish, Danish, Norwegian, Icelandic and English — one of the earliest large regional-language efforts in Europe.

License: Modified RAIL license

Deployment: Self-hosted, Hugging Face

Reference →