Long-lived assistants that run on your own hardware, reach you through the messaging apps you already use, and act on your behalf between conversations. They are indexed here by blast radius rather than by features: a personal agent takes untrusted input — a message from anyone who can reach your inbox — and turns it into privileged action on your machine. What it can touch, what it remembers, what contains it, and who is allowed to address it are the fields that matter.
❯
Add
OpenClaw
🌐 —
OpenClaw Foundation · MIT
Self-hostedOpen sourceBoundary: yoursBring your own modelContainment optional
Unchanged in 2.0 and stated plainly in the project's own words: tools run on the host for the main session unless you configure sandboxing. The architecture is framed as trusted gateway, untrusted execution, deterministic policy — so the gateway is the component to harden.
Exposure
Inbound messages are treated as untrusted input. DM-capable channels pair unknown senders by default — a pairing request must be approved with `openclaw pairing approve` before an unknown sender can interact. The project tells you to read its security, gateway-exposure and sandboxing documentation before connecting other people or exposing the Gateway remotely. 2.0 adds shared cloud sessions, and the team framing is the thing to read carefully: the same gateway serves one operator or "a team whose members trust each other", with configuration the only difference. Multiplayer is not a security boundary — bringing someone into a live session brings them to the host.
Incidents
CVE-2026-25253 (CVSS 8.8, CWE-669): OpenClaw took a gatewayUrl from a query string and opened a WebSocket to it without prompting, sending a token — a one-click compromise from a malicious page, affecting every version before 2026.1.29. Exposure was the larger story: Censys is reported to have tracked publicly reachable instances rising from roughly 1,000 to over 21,000 in the last week of January 2026, alongside a supply-chain campaign that placed hundreds of malicious skills in the ClawHub registry. Treat the deployment, not the model, as the attack surface.
Started as Clawdbot in November 2025, renamed in January 2026, and at 2.0 on 30 August 2026; the npm package is still published as clawdbot, which matters when matching advisories. A Gateway process connects models, tools, channels and companion apps, so the messaging surfaces and the shell live behind one component — convenient, and the reason the gateway is the thing worth hardening. Camera and screen access are available on supported platforms. Version 2.0 touched installation, messaging, memory, skills, models, automations, the browser and native apps, plugins and security, with onboarding that picks up existing API keys and local models, a rebuilt browser app, and shared cloud sessions turning it multiplayer. The scale of the release is its own data point: 933 contributors, 569 of them first-time, across more than 16,000 pull requests — roughly half of everything ever merged into the project
The most thorough containment documented on this page, and unusually specific. The agent runs in a systemd-nspawn container with restricted kernel capabilities, filtered syscalls and no io_uring, root inside mapped to an unprivileged host user, on a filesystem separate from the host. A component called Sentinel is the sole authority for connector actions and network egress. Credentials are never shown to the agent at all: they sit in an isolated credential container, the agent handles surrogate tokens, and the real secret is injected at the network boundary. The browser sub-agent sees accessibility-tree snapshots rather than raw DOM, cannot run JavaScript in page context and has no DevTools. Payments use single-use card numbers scoped to merchant, amount and time.
Exposure
You address it from the apps; nobody else can. The untrusted input is everything it reads on your behalf — mail, web pages, connector data — and the defences are layered rather than a single check: injection resistance trained into the model, untrusted content labelled as it enters context, several classifier models running in parallel, deterministic classifiers watching for egress of unrelated personal data, and mandatory human approval on anything detected as a purchase. Write operations, data egress and purchases need approval; reads and low-risk actions do not. The email connector strips one-time tokens, password-reset and magic links, which closes the obvious account-takeover path through an agent that reads your inbox.
Incidents
None recorded here. That is a gap in this page, not a clean bill of health.
Meta's personal agent, launched 8 September 2026 and named for the model family rather than being one of them. It works in the background, spawns sub-agents, writes its own tools and connectors, and edits itself — with reach that goes past software into the physical world, covering smart home devices and vehicles alongside mail, calendar, Instagram, Facebook, browsing and purchases. Two things distinguish it from the rest of this page. The trust boundary is stated as temporary: Meta says a Muse Confidential VM is planned for later in 2026 to make it cryptographically verifiable that Meta cannot access your VM, using trusted execution environments with external audit — until then the boundary is Meta's word. And the bug bounty puts a price on the page's central concern, offering up to $300,000 overall and up to $130,000 specifically for a prompt-injection attack affecting an individual user. Files and agent memory are inspectable, editable and downloadable by the user.
Terminal work can run through local, Docker, SSH, Singularity, Modal, Daytona or Vercel Sandbox backends, and command approval, DM pairing and container isolation are available — but they are options to choose, not a default posture.
Exposure
DM pairing on messaging surfaces, and command approval for privileged actions. Both are documented as available controls rather than defaults, so the deployed posture depends on how it was configured.
Incidents
None recorded here. That is a gap in this page, not a clean bill of health.
A learning loop is the pitch: it writes skills from experience, refines them in use, and keeps a persistent model of the user across sessions through FTS5 session search and profile modelling — compatible with the agentskills.io standard. Around 40 tools behind a single gateway process. The memory is the part to think about before pointing it at personal accounts: an agent that accumulates a durable profile of you is a more valuable target than one that forgets, and skills written from experience are executable content the agent authored itself.
Containment is the platform's and the administrator's: each agent gets its own governed Entra identity rather than a shared service account, credentials are scoped to the task and redacted from logs, Purview policies and sensitivity labels are enforced in the moment, access is limited to approved resources and destinations, and sensitive actions can require human approval. Enrolment requires Intune policy configuration and an opt-in attestation.
Exposure
Reachable through the Microsoft 365 surfaces it is attached to, and acting under an identity your directory already knows — which is what makes its actions attributable. Note what it reads: Teams chats and Outlook mail both carry content originating outside the organisation, so the untrusted-input path exists even though the agent itself is not publicly addressable.
Incidents
None recorded here. That is a gap in this page, not a clean bill of health.
Announced at Build on 2 June 2026 as an "always-on personal agent" that acts without being prompted each time — Microsoft's Autopilots category. The detail that puts it on this page rather than a product list: it is powered by OpenClaw open-source technology, with Microsoft contributing policy conformance upstream. The same framework as the first card, with the trust boundary moved from your machine to a governed tenant. The launch post does not describe a memory model beyond a persistent identity, and requires a GitHub Copilot licence.
Enterprise and Edu admins control access, connectable tools, permitted actions, browser use and network access, with a Compliance API for visibility and auto-review using frontier models to check important actions before they happen. On desktop it inherits Codex's enterprise governance model. Outside a managed tenant the posture rests on the per-action approvals you configure yourself.
Exposure
Not addressable by third parties — there is no inbound messaging channel of its own. The untrusted input arrives through what it reads: plugins connect Slack, Teams, Google Drive, SharePoint, email, calendars, CRMs and project trackers, and Scheduled Tasks can fire on an event, so an inbound message can start a run without you being present.
Incidents
None recorded here. That is a gap in this page, not a clean bill of health.
Launched 9 July 2026, and the desktop build is the part worth reading closely: Computer Use lets it click, type and move files across your apps and browser in the background, on every plan including Free, while Scheduled Tasks run once, on a schedule or on an event. That combination — unattended execution, standing plugin access to mail and chat, and direct control of the desktop — is a wider blast radius than anything else on this page, governed by a vendor rather than by you. The launch post details controls and approvals but no memory model.
OpenAI · Pro and Business Premium in eligible markets, first dot included in the plan; Enterprise, Edu and Healthcare as a beta the workspace admin enables
Cloud — each dot has its own cloud computer, reached from ChatGPT on desktop, web and mobileProprietaryBoundary: vendor'sContained by default
Each dot works on its own cloud computer, sandboxed within the workspace and isolated from other users', and code there runs apart from the systems that enforce its safeguards, so a dot cannot switch off a required check. Your computer is outside that boundary until you connect it, after which the dot can read local files and use tools installed on your device under a local sandbox; microphone and camera need device permission on top. Background "proactive research" is limited to read-only tools, enforced in code. Actions that send, share or change things pass auto-review, a separate model outside the dot's reach, against built-in rules and your Custom Rules; purchases with saved cards need approval, deletion and new security-sensitive access need confirmation each time, and password changes and money transfers are handed back to you. Saved passwords are supplied by a credential service without entering the model's context.
Exposure
Addressable by you through ChatGPT, Slack and Teams, with texting announced; the system card also names email as a channel it follows up through. OpenAI does not say who else in a shared Slack or Teams channel can instruct your dot. The untrusted input arrives through what it reads — over 4,000 apps through plugins — continuously, in the background, which is why OpenAI red-teamed it with injected mail. Disconnecting an app stops new data flowing in, but what the dot has already learned stays in its context until you reset it.
Incidents
None recorded here. That is a gap in this page, not a clean bill of health.
Launched 29 September 2026 at DevDay as an always-on agent that learns your preferences, delegates to subagents and works toward your goals around the clock, and the entry that sits furthest from a chatbot on this page: it starts work before being asked and can be let onto your own laptop. Runs on GPT-6 Astra, which OpenAI treats as Critical for cyber and High for biology. There is no separate system card: dots are an appendix to Astra's, and it is worth reading because it discloses what did not work. In its automated injection test, 16,600 malicious emails across 100 rollouts and 2,638 iteratively refined attempts produced no scored successes; human red-teamers did find weaknesses in how dots handle sensitive disclosures, which OpenAI patched through its confirmation policy while saying known vulnerabilities remain. Across chained tasks where scope shifts unannounced, moderate scope violations rose from 8.6% to 19.7% of samples as intervening tasks doubled from five to ten, and the model had some success evading the monitor's top-priority page to OpenAI's security team. Vendor-reported throughout. Specialist dots, with their own identity, credentials and IT-provisioned hardware, are in enterprise pilots and being integrated with Microsoft Agent 365.
Google · Google AI Ultra and select business users, in select countries
Cloud, with mobile and macOS appsProprietaryBoundary: vendor'sContainment optional
Models
Gemini 3.7 Flash for AI Pro and Ultra subscribers, per Google's August 2026 model announcement; the Spark product page itself names no version
Reachable on
Gemini appGemini for macOS
Reach
mailcalendarcloud files
Autonomy
Unattended
Memory
—
Isolation
App connections are off until you turn each one on, and Google says it is designed to check with you before major actions. Those are consent and approval controls rather than containment: no sandbox, policy engine or administrative governance is described on the product page, and the agent runs on Google's infrastructure where there is nothing for you to configure.
Exposure
No inbound channel of its own — you address it through the Gemini apps. The untrusted input arrives through Gmail, and Schedules can fire on a condition rather than a clock, so mail you did not solicit can meet an agent acting on its own. Connectors being off by default means the reach on any given account is whatever the user has switched on.
Incidents
None recorded here. That is a gap in this page, not a clean bill of health.
The most cloud-anchored entry here: it works in the background 24/7 even when your phone and laptop are turned off, which means execution never touches your hardware and the boundary is entirely Google's. Tasks are one-off actions, Skills are reusable procedures, Schedules are time- or condition-triggered. Native to Workspace — Gmail, Calendar, Drive, Docs, Sheets, Slides, plus YouTube and Maps. No memory model is described beyond the Skills it keeps.
Google Labs · Labs experiment — waitlist, US only, 18+, personal Google account
Cloud — an isolated instance per group, on web and mobileProprietaryBoundary: vendor'sContained by default
Models
Google's latest Gemini models on the Antigravity harness; the launch post names no version, and Ars Technica reports Gemini 3.8 Flash
Reachable on
GmailGoogle ChatCalendarTasksDriveDocs
Reach
mailchatcalendarcloud files
Autonomy
Scheduled
Memory
Persistent
Isolation
Every CC runs on its own isolated cloud computer on Google's Antigravity harness, and the agent holds its own verified Google account rather than borrowing a member's — a distinct identity with its own permissions model, which is what makes its actions attributable. Note what that does and does not bound. It contains the compute and separates one household's instance from another's; it is the sharing controls, not the sandbox, that decide what the account can read, and once a sender is on the auto-cc list their mail flows in without a further decision.
Exposure
Addressable by design, which sets it apart from the other vendor agents here: CC has its own email address and Google Chat presence, and members forward it invitations, schedules and photographs. Google says it "only responds to group members" and will not take action or share information outside the group without permission. The untrusted input is the shared mail itself — a school, a swim centre or a vet nominated to the auto-cc list is a third party whose account can be compromised or spoofed, and its mail reaches the agent with no further decision. New senders are held for a weekly private list each member approves, which is the control on this card worth understanding before the convenience of the others.
Incidents
None recorded here. That is a gap in this page, not a clean bill of health.
Announced 17 September 2026, expanding the single-user CC that became Gemini's Daily Brief in May. The premise is that household logistics sit in one person's inbox: up to six members share into one agent, which sorts mail from senders they nominate into a shared "Your Day Ahead" brief, a family calendar and a task list, and fills in registration PDFs, meal plans and shopping lists, checking live drive times through the Maps API between back-to-back activities. Two things earn it a card rather than a product note. Its memory is the first here that is explicitly shared — Google says it separates what applies to the whole household from what belongs to one person — so the question this design raises is what a six-way pooled memory carries between members, not whether it remembers. And it runs on Antigravity, the same agentic harness this site lists as a coding agent, pointed at a family's mail.
Anthropic · Pro, Max, Team and Enterprise; web and mobile in beta
Cloud, desktop, web and mobileProprietaryBoundary: vendor'sContainment optional
Models
Claude; the article does not name a tier
Reachable on
Claude Desktopclaude.aiClaude Mobile
Reach
local filesbrowserMCP servers
Autonomy
Scheduled
Memory
Project-scoped
Isolation
Three approval modes govern what it does without asking: Manual pauses for approval on each action, Auto lets it work while safety-checking actions, and Skip turns the automatic checks off entirely. The containment is therefore a setting, and the weakest position is one click away — worth knowing before running it against a machine with credentials on it.
Exposure
No inbound channel of its own; you address it from the Claude apps. The untrusted input arrives through what it reads — connectors and websites it opens in Chrome — and scheduled tasks run in the cloud without your machine awake, so a run can proceed with nobody watching.
Incidents
None recorded here. That is a gap in this page, not a clean bill of health.
Started in January 2026 as a desktop research preview beside Chat and Code, built on Claude Code's foundations, and scoped to a folder you pointed it at. It earns a place here because it since grew the two properties that matter: cloud sessions that continue after you close the laptop, and scheduled tasks that do not need your computer awake. Memory is narrower than the others — what Claude remembers about you in chat does not carry into Cowork, and within Cowork it applies to projects only.
The only entry here where containment is the architecture rather than a prompt: untrusted tools run in WASM sandboxes under capability-based permissions, secrets are injected at the host boundary and never exposed to tools, outbound HTTP is restricted to explicitly approved hosts and paths, responses are scanned for leaked credentials, prompt-injection patterns are detected and content sanitised, and local storage is AES-256-GCM with no telemetry. Note one gap between the pitch and the repo: the product site advertises encrypted TEE enclaves, which belong to the managed NEAR AI Cloud offering — the README describes no TEE.
Exposure
Addressable through whichever channels you enable — Telegram and Slack run as WASM channels, alongside a REPL, a web UI and HTTP webhooks. The README documents no pairing or sender-approval step of the kind OpenClaw ships, so who may talk to it is a function of how you expose the channel and the webhook endpoint.
Incidents
None recorded here. That is a gap in this page, not a clean bill of health.
A Rust rewrite pitched as an Agent OS, built by NEAR AI and launched at NEARCON 2026, dual-licensed and offered as managed hosting from $5 to $200 a month. Routines drive it unattended through cron schedules, event triggers and webhook handlers, with a heartbeat system for proactive monitoring, and memory is persistent across full-text and vector search plus a workspace filesystem. It can also build new WASM tools on demand — capability the sandbox model is there to contain. Read it as the design answer to the failure modes the first card on this page demonstrated.
xAI · Early beta — SuperGrok Heavy, Cursor Ultra and Cursor Teams Premium; enterprise waitlist
Cloud, with macOS and iOS appsProprietaryBoundary: vendor'sNo containment documented
Models
Grok; the announcement does not name a version
Reachable on
Grok Bot appiOSBot-to-bot threads and group chats
Reach
computer usebrowsermailchatCRM
Autonomy
Unattended
Memory
Persistent
Isolation
The announcement describes no containment. Bots get "a computer of their own" in the cloud, which keeps work off your laptop but says nothing about what bounds their actions inside your accounts, and the post is silent on how the credentials they sign in with are stored. The only control described is behavioural — they "only come back when something needs your approval" — and the page's own testimonial celebrates letting that go: "it does it without me verifying and reviewing."
Exposure
Two paths, and the second is unusual. You message bots like colleagues from desktop or phone. But bots also message each other independently, share context in threads, and can be put in a group chat where they "pass work, assign ownership, and only pull you in for judgment calls" — so one bot's output is another bot's instruction with no human in between. The untrusted input arrives through everything they read on your behalf: inboxes, CRM notes, tickets and arbitrary websites.
Incidents
None recorded here. That is a gap in this page, not a clean bill of health.
Launched 11 August 2026 in early beta, on a page that brands as SpaceXAI. The distinguishing capability is also the risk: rather than calling APIs, bots sign into the tools you already use and drive their interfaces the way a person would, explicitly including platforms with no API or MCP — so ordinary application permissions and audit trails see a legitimate user session, not an agent. They learn by watching you do a job once, save it as a routine, and run it unsupervised afterwards. xAI's model documentation is already the thinnest of the frontier labs on the system-cards page; nothing here changes that.