Reference — Documentation & evaluations

System Cards & Model Documentation

Where to find the primary documentation for the models in the registry. These are the documents that answer governance questions — what a model was trained on, how it was evaluated, what it refuses, and what the vendor claims versus what independents measured.

For EU deployments these documents map directly onto obligations: the AI Act's general-purpose AI rules (applicable since August 2025) require providers to maintain technical documentation and training-data summaries, and downstream deployers inherit the need to understand them.

Frontier labs — system cards per release

Proprietary vendors publish safety and system documentation per model release, at varying depth.

Anthropic 🇺🇸 USA

System cards for each Claude release, plus a transparency hub covering safety framework, evaluations and government commitments. Among the most detailed system cards published.

Memory & retentionDocumented separately from the system cards, in the privacy centre. Prompts and outputs for covered models are retained 30 days from 9 June 2026, including for organisations that previously held zero-retention agreements through Console, Bedrock, Google Cloud or Azure Foundry; flagged or legally held content is kept longer. Claude memory stores stated preferences and context rather than transcripts. Retention practices →

OpenAI 🇺🇸 USA

System cards per release (GPT-5 onward under the unified safety framework) plus preparedness framework reports and the safety hub.

Memory & retentionNot in the system cards, which cover safety evaluations, preparedness and safeguards only. The commitments sit in the enterprise privacy pages: business data is not trained on by default, retention is customer-controlled on Enterprise, Healthcare and Edu, and connected apps inherit the workspace's existing permissions. Consumer ChatGPT memory is documented separately again, in the help centre. Enterprise privacy →

Google DeepMind 🇺🇸 USA

Model pages with technical reports and model/system cards for the Gemini family, plus the Frontier Safety Framework documentation.

Memory & retentionSplit away from the model documentation entirely, into the Gemini Apps privacy hub. Activity is kept 18 months by default, adjustable to 3 or 36; with activity off, conversations are still held up to 72 hours to run the service; conversations selected for human review are kept up to three years, disconnected from the account. Workspace tenants get the same 18-month default under admin control. Gemini Apps privacy hub →

Microsoft AI 🇺🇸 USA

Model pages for the MAI family; the MAI-Thinking-1 technical report also covers the lineage behind MAI-Cyber-1-Flash.

Memory & retentionThe model pages don't carry it; the product documentation does, and in more operational detail than most. Prompts, responses and Graph-grounded data are stored as Copilot activity history, encrypted, not used to train the foundation models, discoverable through Purview with admin-set retention policies, and deletable by the user. EU traffic stays inside the EU Data Boundary — with Anthropic subprocessor models currently excluded from it. Copilot data & privacy →

Meta 🇺🇸 USA

Developer documentation for the Muse family and legacy Llama model cards. Note the documentation split after the shift from open-weight Llama to proprietary Muse.

Memory & retentionSurfaces as a pricing decision rather than a policy document: the Meta Model API's cheaper contributor tier lets Meta retain submitted data for training, while standard-tier data is not retained — so on this platform the retention posture is the tier you picked. Check the API terms for the tier you are actually on.

xAI 🇺🇸 USA

Model announcements and documentation for the Grok family; safety documentation is comparatively thin — worth noting in assessments.

Memory & retention — not foundNo retention or memory documentation accompanies the Grok model announcements. The gap matches the thin safety documentation noted above: for an assessment you are left with the general consumer privacy policy rather than anything specific to the deployed model.

Thinking Machines Lab 🇺🇸 USA

The Inkling announcement doubles as a technical report — architecture, training data scale, safety benchmarks (FORTRESS) and calibration methodology.

Memory & retentionNot applicable in the same way: the report documents a model with open weights rather than a hosted product holding user history, so retention becomes a question about whoever serves it — which, if you self-host, is you.

Open-weight vendors — model cards with the weights

For open models the primary documentation lives on the Hugging Face organisation pages, next to the checkpoints. These cards carry no retention line by design: weights do not remember anything. Memory and retention are properties of whatever serves the model, so on a self-hosted deployment the answer is whatever your own stack does — which is the strongest reason to write an internal card for it, since no vendor will.

DeepSeek 🇨🇳 China

Model cards and technical reports for the V4 family and predecessors, published with MIT-licensed weights.

Alibaba Qwen 🇨🇳 China

Cards for the Qwen 3.x open-weight line, including the 3.6-generation MoE models.

Z.ai (Zhipu) 🇨🇳 China

GLM-5.x model cards and agentic evaluation detail, MIT-licensed.

Moonshot AI 🇨🇳 China

Kimi model cards; check here for the K3 open-weight checkpoint status.

Xiaomi MiMo 🇨🇳 China

MiMo-V2.6-Pro cards and technical report, with training methodology and agentic and cybersecurity benchmark detail.

StepFun 🇨🇳 China

Step 3.7 Flash cards including quantised deployment recipes for local use.

NVIDIA 🇺🇸 USA

Nemotron 3 model cards — unusual for including the training datasets themselves, not just their description.

Mistral AI 🇫🇷 France

Model cards and API documentation for the open and commercial lines — the EU-jurisdiction reference vendor.

IBM Granite 🇺🇸 USA

Granite 4 cards with the governance extras: signed weights and ISO/IEC 42001-accredited process documentation.

Cisco Foundation AI 🇺🇸 USA

Cards for the security models — the Foundation-Sec line and the Antares family (gated access).

Model-card standard 🌐

Hugging Face's documentation of the model-card format itself — useful when writing internal cards for fine-tuned models.

Independent evaluation & tracking

Vendor documentation states claims; these sources test or track them independently — always pair the two.

UK AI Security Institute 🇬🇧 UK

Government pre-deployment evaluations of frontier models, including the cyber-range exercises referenced for Mythos-class models.

Artificial Analysis 🌐

Independent intelligence, speed and price benchmarking across hosted models — the fastest way to sanity-check a vendor benchmark table.

Stanford HELM 🇺🇸 USA

Academic holistic evaluation framework with transparent, reproducible scenario coverage.

LMArena 🌐

Blind human-preference rankings — noisy but hard to game in the ways static benchmarks are.

Epoch AI 🌐

Data and analysis on compute, parameters and training trends — the source for scale context behind the registry's log axis.

models.dev & OpenRouter 🌐

Machine-readable model specification and pricing databases — practical sync sources for keeping the registry's pricing fields current.