Overview — Agents & harnesses

Coding Agents

An agent is a model plus a harness — the scaffolding of tools, permissions, context management and workflow around it. The frontier models inside these tools have largely converged, so the harness now does much of the differentiating work. This page indexes the coding agents worth knowing, plus the security-focused agent systems built on the same pattern. Agents tagged bring your own model can run against local or self-hosted models from the registry — the privacy-preserving combination.

Claude Code

🇺🇸 USA
Anthropic
CodingTerminalProprietaryLocal models — unsupported
Models
Claude frontier models (Fable/Opus/Sonnet tiers)
Local models
Ollama exposes an Anthropic-compatible API, and `ollama launch claude-code` sets ANTHROPIC_BASE_URL to localhost:11434 so the harness talks to a local model. It works, but Anthropic's own gateway documentation states it does not support routing Claude Code to non-Claude models — treat it as your own configuration, not a supported path.
Terminal-first agent, also in VS Code, JetBrains, a desktop app, web and mobile. The deepest harness in the field: lifecycle hooks, Skills, plugins, subagents and MCP, with workflows that orchestrate many parallel subagents in one session. Enterprise routes exist via AWS Bedrock and Google Vertex.

OpenAI Codex

🇺🇸 USA
OpenAI
CodingCloud agentProprietaryLocal models — first-party
Models
GPT-6 Astra and the GPT-5.6 family, plus codex-tuned variants
Local models
The CLI ships an --oss flag for local open-weight models, and config.toml takes a model_providers block pointing at any OpenAI-compatible endpoint; ollama and lmstudio are reserved provider names. The cloud sandboxes remain OpenAI-hosted — only the local CLI runs offline.
Cloud sandboxes running parallel asynchronous jobs, plus a CLI and IDE integration. Built for throughput: delegate several tasks, review the pull requests. OpenAI reports millions of weekly users and near-universal internal adoption. The harness was updated alongside GPT-6 Astra on 4 September 2026 for faster computer use, which OpenAI puts at 1.9x quicker task completion than the previous Sol experience on Mind2Web.

GitHub Copilot

🇺🇸 USA
GitHub / Microsoft
CodingExtensionProprietaryLocal models — partial
Models
Multi-model picker: MAI-Code-1.1-Flash, Claude, GPT, Gemini
Local models
VS Code takes bring-your-own-key providers and a custom endpoint, with local models through the official Ollama extension (the built-in provider is deprecated). Chat and utility tasks only: inline completions, semantic search and anything using embeddings still require a GitHub account, so the stack is never fully offline.
The default enterprise path: completions, chat, an agent mode with issue-to-PR flow, and a CLI. Moved to usage-based AI Credits billing in June 2026, with per-model token rates; basic completions stay unmetered on paid plans. BYO-key support in VS Code can point chat at other providers.

Cursor

🇺🇸 USA
Anysphere (SpaceX)
CodingIDEProprietary
Models
Claude, GPT, Gemini and Grok, plus in-house Composer models
The AI-native IDE: fast tab completion, multi-file agent edits and cloud agents, with its in-house Composer line tuned for fast agentic editing. Strong editor UX; large refactors still reward close review. Owned by SpaceX since 14 August 2026, when a $60 billion all-stock acquisition of Anysphere closed; Anysphere remains the operating entity and the brand is unchanged. That makes Cursor the one multi-vendor editor here whose owner also sells a competing model family — Grok is on every Cursor plan, and its fast variant is served only in Cursor and SpaceXAI's own Grok Build. Cursor's announcement names no change to the other vendors' models or to data terms.

Amp

🇺🇸 USA
Amp Inc. (spun out of Sourcegraph)
CodingTerminalProprietary
Models
Multi-model routing: GPT-5.6, Claude Fable 5 and fast models, picked per task and agent mode
CLI-first agent with VS Code, Zed and Neovim extensions, a web thread surface and a Slack integration. Built by the Sourcegraph team around big-codebase context, spun out as an independent company in December 2025. Pay-as-you-go credits at zero markup for individuals and teams, plus a limited free tier; you can link a ChatGPT subscription, but not arbitrary endpoints.

Antigravity CLI

🇺🇸 USA
Google
CodingTerminalProprietary
Models
Gemini 3 family, auto-routed
Successor to the Apache-2.0 Gemini CLI, which stopped serving Google AI Pro, Ultra and free individual accounts on 18 June 2026 (API-key auth and Gemini Code Assist enterprise licences were unaffected). A single Go binary (agy) that migrates skills, MCP servers, agents and memory from the old CLI on install; closed-source, unlike its predecessor.

Antigravity IDE

🇺🇸 USA
Google
CodingIDEProprietary
Models
Gemini 3 family plus selected third-party models
Google's agent-first IDE: an agent manager surface for supervising multiple agents across editor, terminal and browser, with artifact-based reporting of what agents did. Now one of three Antigravity surfaces alongside the CLI and the Antigravity 2.0 command centre.

Jules

🇺🇸 USA
Google
CodingCloud agentProprietary
Models
Gemini
Asynchronous cloud coding agent: point it at a GitHub repo and an issue, get a reviewed diff back. The fire-and-forget lane of Google's lineup.

Muse Code

🇺🇸 USA
Meta (Superintelligence Labs)
CodingTerminalProprietary
Models
Muse Spark 1.3 (the 1.2 generation was co-trained with the agent)
Meta's terminal agent, released August 2026 alongside Muse Spark 1.2 — notable because model and harness were co-trained, the pattern Microsoft also used for MAI-Code-1-Flash inside Copilot. Muse Spark 1.3 shipped into it on 2 September, tuned on months of Muse Code usage: Meta reports roughly 20% fewer tool calls and 25% fewer tokens for the same work, which is a harness-level efficiency claim as much as a model one.

IBM Bob

🇺🇸 USA
IBM
CodingPlatformProprietary
Models
Routes per task: Anthropic Claude, Mistral, IBM Granite, specialized fine-tunes
Agentic SDLC platform (GA April 2026) whose pitch is the router: each task goes to the cheapest model that can do it well, with security and next-edit models in the mix. IBM claims 55–65% cost reduction versus direct-to-frontier. An on-prem variant is announced.

Devin Desktop

🇺🇸 USA
Cognition
CodingIDEProprietary
Models
In-house SWE models plus frontier options
Formerly Windsurf: Cognition retired the brand on 2 June 2026 in an over-the-air update, and windsurf.com now redirects here. Cascade became Devin Local, rewritten in Rust with subagents, and the agent manager is the default surface. Speaks the open Agent Client Protocol, so other agents can drive the editor. The local half of the same lineup as Devin's cloud agent.

Devin

🇺🇸 USA
Cognition
CodingCloud agentProprietary
Models
Proprietary stack over frontier models
The most hands-off tier: an autonomous engineer that takes a ticket through to a PR in its own environment. Works best on tightly scoped tasks with a high review bar; loop risk on fuzzy ones.

Kiro

🇺🇸 USA
AWS
CodingIDEProprietary
Models
Claude models
AWS's spec-driven agentic IDE: requirements and design documents first, then agent execution against them — a governance-friendly workflow for teams that want auditable intent. Launched internationally in May 2026 as a replacement for Amazon Q Developer, and now also ships a CLI and speaks the Agent Client Protocol, so the agent is no longer tied to the editor.

Replit Agent

🇺🇸 USA
Replit
CodingPlatformProprietary
Models
Frontier models (Claude, GPT)
Prompt-to-deployed-app in a hosted workspace — the low-floor end of the spectrum, strongest for prototypes and internal tools rather than existing large codebases.

Factory Droid

🇺🇸 USA
Factory
CodingTerminalProprietaryBring your own model
Models
Factory-managed models, or your own: Anthropic Messages API, OpenAI Responses API, any OpenAI-compatible endpoint (OpenRouter, Fireworks, DeepInfra, Groq, Baseten, Hugging Face, Gemini) and local Ollama or LM Studio
Delegate a task, review the diff, merge — from the terminal, a desktop app or cloud sessions, customised through AGENTS.md, connectors, MCP and plugins. The BYOK story is unusually complete for a proprietary harness: custom models go in ~/.factory/settings.json and the keys stay on your machine rather than on Factory's servers. Two caveats worth knowing before planning a private stack around it: custom models work in the CLI only, not the web or mobile surfaces, and Factory tests only the official Anthropic and OpenAI APIs, warning that models under ~30B do markedly worse at agentic coding.

Mistral Vibe

🇫🇷 France
Mistral AI
CodingTerminalProprietary
Models
Mistral Medium for multi-step work, Devstral 2 for agent tasks, Codestral for completion, Codestral Embed for code search
Successor to Mistral Code: a terminal-native agent with VS Code, JetBrains and Zed extensions. The only European entry here, and the strongest fit for teams that cannot send code to US-hosted APIs — the whole stack can be self-hosted, up to air-gapped on-premise GPUs. Model-locked to Mistral's own line, so it misses the BYO-model filter, but the deployment story is what earns it a place.

Trae

🇨🇳 China
ByteDance
CodingIDEProprietary
Models
International build offers Claude, GPT, Gemini and Grok; the China build runs Doubao, DeepSeek and Kimi
ByteDance's AI-native IDE, a VS Code fork with MCP support, priced aggressively against Cursor. Ships as two distinct products: a global edition and Trae CN, which is free and routes only to domestic models. Now paired with a TraeWork surface for non-coding agent tasks.

Goose

🇺🇸 USA
Block, now Agentic AI Foundation (Linux Foundation)
CodingTerminalOpen sourceBring your own model
Models
Any — 15+ providers including Anthropic, OpenAI, Google, Azure, Bedrock, OpenRouter and local Ollama; existing Claude, ChatGPT or Gemini subscriptions can be used through ACP
A general-purpose agent that runs on your machine rather than a coding-only harness, in a native desktop app, a CLI and an embeddable API. Around 70 extensions connect through MCP. Started at Block and has since moved under the Agentic AI Foundation at the Linux Foundation, which is a governance answer few agents on this page have: the harness outlives its sponsor's product decisions.

Inspect

🇺🇸 USA
Ramp
CodingPlatformProprietaryIn-house — not available
Models
Not disclosed; model agnosticism was one of the reasons OpenCode was chosen as the harness
Why they built it
Off-the-shelf agents could not run many sessions at once on a laptop, gave designers nothing for frontend work, and could not reach the remote environments needed to debug interacting systems at scale. Ramp built on OpenCode specifically because it exposed an HTTP API and was model-agnostic.
Built on OpenCode with React and Vite, Cloudflare Durable Objects, SQLite, Modal sandboxes, VS Code Server and Chromium. The distinguishing feature is verification rather than generation: backend changes checked against tests, telemetry and feature flags, frontend changes against screenshots and live previews, with sessions on remote sandboxes rather than a developer's machine. Ramp reports Inspect authoring 75% of merged pull requests, built by a 5.5-person team with 150+ engineers contributing to it, and more than 80% of Inspect itself written in Inspect sessions.

Minions

🇺🇸 USA
Stripe
CodingCloud agentProprietaryIn-house — not available
Models
Not disclosed
Why they built it
Built for unattended work rather than pair programming: a task goes in and a reviewable pull request comes out, with nobody steering in between — the opposite end of the spectrum from Cursor or Copilot.
Work starts from a Google Doc, a ticket, a Slack thread or a single emoji reaction, and each agent gets its own isolated code and services so runs do not contend for a machine. What Stripe calls blueprints combine deterministic code with agent loops. Reported at roughly 1,300 pull requests merged a week containing no human-written code — figures from Stripe engineers speaking publicly rather than from a company document, so treat the number as reported.

River

🇺🇸 USA
Shopify
CodingCloud agentProprietaryIn-house — not available
Models
Not disclosed
Why they built it
Deliberately refuses to work in private: River runs only in public Slack channels, so that prompt patterns and debugging technique spread by being watched. The constraint is pedagogical, not technical.
Slack-native, handling code questions, codebase navigation, review, tests, data queries and pull requests, on an internal platform called Aquifer for durable multi-participant agent sessions, against a single repository with reproducible Nix environments. Reported at 5,938 employees across 4,450 channels in 30 days, roughly one in eight merged pull requests, and a merge rate climbing from 36% to 77% in two months — reported figures rather than a company publication.

Aider

🌐 —
Open source
CodingTerminalOpen sourceBring your own model
Models
Any — API keys or local models via Ollama and compatible endpoints
The original open-source terminal pair programmer: git-native, maps the repo, commits as it goes. Fully offline-capable when paired with local models — a clean privacy stack.

Cline

🌐 —
Open source
CodingExtensionOpen sourceBring your own model
Models
Any — OpenAI-compatible endpoints, incl. local and gateway (LiteLLM) setups
Open-source VS Code agent with plan/act modes, MCP support and per-step approval. Model-agnostic by design, so it pairs naturally with a self-hosted gateway.

Roo Code

🌐 —
Open source
CodingExtensionOpen sourceBring your own model
Models
Any — same endpoint flexibility as Cline
Cline fork grown into its own project: multiple configurable agent modes (architect, coder, reviewer) with fine-grained permissions. Popular in teams standardizing on internal model gateways.

OpenHands

🌐 —
All Hands AI
CodingPlatformOpen sourceBring your own model
Models
Any — configurable per deployment
Open-source autonomous development platform (formerly OpenDevin): sandboxed agent runtime you can self-host end to end, from UI to model. The reference choice for fully private autonomous agents.

Crush

🌐 —
Charm
CodingTerminalOpen sourceBring your own model
Models
Any — a built-in catalogue plus anything behind an OpenAI- or Anthropic-compatible API, including local Ollama by base URL; Charm also sells its own subscription provider, Hyper
A Go terminal agent from the Charm ecosystem, and unusually polished for the tier. Two things set it apart: it reads your language servers, using LSP output as context the way you would, which most agents do not do; and it switches model mid-session while preserving context, so a cheap model can carry the routine turns and a frontier one the hard ones. Sessions are per project, MCP extends it over http, stdio and sse, and it runs first-class on macOS, Linux, Windows, Android and the BSDs. The licence deserves attention: FSL-1.1-MIT is source-available rather than open source — you may read, modify and self-host it, but not compete with it — and each release converts to MIT two years on. GitHub reports it as NOASSERTION because FSL is not a recognised SPDX licence.

OpenCode

🌐 —
Open source
CodingTerminalOpen sourceBring your own model
Models
Any — bring your own key or local endpoint
Fast-growing open terminal agent, model-agnostic with a polished TUI — the open-source counterpart to Claude Code and Codex CLIs.

DeepSeek Harness

🇨🇳 China
DeepSeek
CodingPlatformOpen sourceBring your own model
Models
Any — DeepSeek by default, a catalog covering Anthropic, OpenAI, Bedrock, Vertex and Azure, and custom providers configured by base URL and protocol for a company gateway or a self-hosted server
DeepSeek's own harness (dsh), which matters mainly because the vendor of one of the largest open-weight families now ships a model-agnostic agent rather than a client for its own API. Everything is a plugin, on the Cordis composition framework; the agent works inside a selected workspace, reading and editing files, running commands, delegating and keeping a plan, through a local web UI, CLI modes or a Python SDK. API keys are write-only and stored outside the settings file. Two caveats belong together: it is explicitly a developer preview promising breaking changes, and its own safety notice says it is unaudited, executes model-generated code, loads third-party plugins, and that approval prompts and sandboxing do not guarantee isolation — DeepSeek recommends a disposable VM or container. With a third-party plugin ecosystem forming around a topic tag, that is the same supply-chain shape that turned into a problem for OpenClaw's skill registry.

Qwen Code

🇨🇳 China
Alibaba (Qwen team)
CodingTerminalOpen sourceBring your own model
Models
Any — OpenAI, Anthropic, Gemini and Qwen APIs, plus local Ollama and vLLM endpoints, switchable at runtime
Apache-2.0 terminal agent that began as a fork of Gemini CLI v0.8.2 and stopped syncing upstream to become its own multi-protocol framework, with headless, IDE-plugin and daemon modes. Outlived the CLI it forked from precisely because it is endpoint-agnostic: nothing about it depends on a vendor's free tier surviving.

Kimi Code CLI

🇨🇳 China
Moonshot AI
CodingTerminalOpen source
Models
Moonshot's Kimi models
Apache-2.0 terminal agent for code and shell work, successor to Kimi CLI with config migrated on install. Moonshot's open-weight K2 coders are the draw: the same family can be self-hosted, though this client is pointed at Moonshot's own endpoints.

ZCode

🇨🇳 China
Z.ai (Zhipu AI)
CodingPlatformOpen sourceBring your own model
Models
Any — GLM-5.3 by default and the tuning target, with GLM-5.3-Flash built in for screenshot and image work; built-in providers for Anthropic, OpenAI, OpenRouter, Moonshot, MiniMax and Xiaomi MiMo, and a custom provider takes any Anthropic- or OpenAI-compatible endpoint, including a self-hosted one on a private network
Z.ai's own harness, its own codebase rather than a fork of an existing CLI, opened under Apache-2.0 on 20 September 2026 — at desktop version 3.14.1, so the product long predates the repository. One repository ships three surfaces: an Electron workbench with terminal and Git panel, a local web UI, and a TUI, with `zcode` entering the terminal and `zcode --web` the browser; remote workspaces run over SSH or WSL, and skills, plugins and MCP come with an official marketplace. Goals carry long-running work through plan-execute-verify cycles across parallel agents, which can also be started and steered from WeChat, Feishu or Telegram. Read NOTICE.md before wiring that up: Z.ai states there that the shared agent execution adapter provides no default operating-system sandbox, and that the standalone CLI given a non-interactive `--prompt` without `--mode` falls back to `yolo`. Documenting your own defaults that precisely is the exception and not the rule — but what is documented is an unsandboxed agent with a remote-start surface.

Codex Security

🇺🇸 USA
OpenAI
SecurityCloud agentProprietary
Models
Codex models (GPT-5 family)
Announced as Aardvark in October 2025, renamed and opened as a research preview on 6 March 2026 to ChatGPT Pro, Business, Enterprise and Edu, through the Codex web surface. Builds a threat model of a repository, then finds vulnerabilities, rates them by real-world impact and proposes patches. OpenAI reports 1.2M commits scanned in 30 days, turning up ~800 critical and 10,000+ high-severity issues, including in Chromium, OpenSSL, PHP and GnuTLS.

CodeMender

🇺🇸 USA
Google DeepMind
SecurityCloud agentProprietary
Models
Gemini 4 Argon for a subset of Fairwind Program partners (multiple cooperating agents); publicly available Gemini models otherwise
Code-security agent that finds, validates and fixes vulnerabilities, upstreaming patches to open-source projects, running several agents in parallel on one report. Its availability now splits in two, which is the thing to read: inside Google's Fairwind Program it is paired with the gated Gemini 4 Argon, which replaced 3.8 Flash Cyber on the programme page at Argon's launch, for governments, critical infrastructure operators and core platform companies, under conditions including restricting use to internal security teams; separately, any Google Cloud customer can run CodeMender against publicly available models. The harness is no longer the gated part — the model is.

MDASH

🇺🇸 USA
Microsoft
SecurityPlatformProprietary
Models
MAI-Cyber-1-Flash (~90% of tasks) escalating to GPT-5.4
Microsoft's multi-agent vulnerability identification and remediation harness — 100+ expert-tuned agents — and the clearest production example of specialized-model-plus-frontier-escalation economics: the small cyber model handles the bulk, the frontier model takes the hardest tenth. Its agents also feed Project Perception, listed separately.

Project Perception

🇺🇸 USA
Microsoft
SecurityPlatformProprietary
Models
Not named in the documentation; Microsoft's May 2026 announcement describes MDASH's agents extending into Perception
A security-operations agent fleet inside the Defender portal, in invitation-only Limited Public Preview. Three categories run as one loop: red team (Recon, mapping attack paths, choke points and excess permissions), blue team (Triage, Threat Intelligence, Attack Investigation, Detection Authoring) and green team (Posture Prioritization). Playbooks assign a team objectives; a chat surface picks the playbook for you; sessions can be watched, redirected or stopped. Read it as an agent with standing access to your estate rather than as a detection product: green-team agents change posture, red-team agents consume attacker-shaped input by design, and the real control is the approval gate on high-impact actions plus per-agent identity and execute-versus-view permissions. "Agentic security" here is a product name, not a standard — there is nothing to audit against.

XBOW

🇺🇸 USA
XBOW
SecurityCloud agentProprietary
Models
Frontier models with a proprietary offensive harness
Autonomous penetration testing: continuously attacks your applications the way a human red team would, at machine scale. The offensive counterpart to the defensive agents above — useful context for threat modeling what attackers now automate.