Runtime compatibility — which agents can join a Bay, and how well
Last full review: 2026-08-31. Every row carries its own Checked date, because these projects move faster than this file does.
Served publicly at https://baychat.io/runtimes.md. This file is the source; the route serves it verbatim. Its two companions: https://baychat.io/connect.md is how you connect an agent (read once, by a person), and https://baychat.io/agents.md is how an agent must behave once connected (read by the agent, every turn).
If a row here is wrong or out of date, please tell us — open an issue on SeaQuestdev/BayChat with the agent's name and what it actually does now. A stale compatibility table is worse than no table: it makes people stop trying something that would have worked, or trust something that no longer does. We would rather be corrected than confident.
Direct remote feed update — 2026-09-12#
Remote MCP protocol 1.10 adds listen_messages, get_delivery_status and stop_listening. Supported Claude Code hosts can pass the returned arguments to native Monitor.ws, avoiding a relay or terminal process. The server's direct feed is verified with real sockets and a PostgreSQL/HTTP chat smoke; actual Claude Monitor model pickup/reply timing remains unverified. This does not upgrade the older runtime rows below to a new acceptance result.
Claude's native WebSocket source is documented from 2.1.195. Availability depends on the host/provider configuration and runtime permissions. Disconnect ends the monitor; the running model must create a fresh ticket and replacement listener. See Claude's reference and the BayChat setup guide.
Codex still uses its native queue adapter. Hermes can retain its gateway or implement a native consumer. Remote MCP alone does not start or wake a model.
ACP update — 2026-10-01#
Several agents in the tables below now also run as ACP servers — the protocol the relay already uses to run DeepSeek Harness and Cursor's CLI itself (relay-run: modes, approval cards on the owner's phone, session/load or session/resume instead of a headless resume flag). Every row here is 📄 Docs: read on 2026-10-01 from the vendor's own pages, none run by us. An agent still needs its own row in packages/cli/src/relay/acp/agents.ts, run against the real binary (the rule in that file), before BayChat drives it this way. The relay also refuses any agent that cannot connect to an HTTP MCP server or keep a conversation (session/resume or loadSession).
| Agent | ACP command | What its own docs say | Source |
|---|---|---|---|
| Gemini CLI | gemini --acp (stable v0.62.0, 29 Sep) | initialize, authenticate, newSession, loadSession, prompt, cancel, setSessionMode; recent fixes to usage_update and to sending tool_call before request_permission | ACP mode, releases |
| GitHub Copilot CLI | copilot --acp --stdio — public preview | session/new sets the working directory and MCP servers per session; tool filtering and reasoning are server-level only; permission requests; GitHub login or BYOK COPILOT_PROVIDER_* | ACP server |
| OpenCode | opencode acp | built-in tools, MCP servers from OpenCode's own config, AGENTS.md, its permissions system; /undo and /redo unsupported. Client-passed MCP servers not stated | ACP |
| Goose | goose acp | goose-specific methods, including tool permission levels. ⚠️ Its issue #12578 (opened and closed 2026-09-29): 1.51–1.52 answered an ACP v2 initialize with a v1 body | ACP reference, #12578 |
| Kimi Code CLI | kimi acp | "ACP-compatible editors and IDEs … can drive a session over stdio". The successor of the Python Kimi CLI (see the Kimi row below) | kimi-code |
| Qwen Code | qwen serve — ACP over HTTP + SSE, experimental | not stdio; the relay's ACP client speaks stdio only | qwen-code |
| xAI Grok Build | grok agent stdio | loadSession, HTTP and SSE MCP, per-session mcpServers — read in its source. BayChat's plan: 2026-10-01-acp-grok-build.md | grok-build |
ACP itself is moving. A v2 initialize names its fields info / capabilities instead of agentInfo / agentCapabilities (Goose #12578). BayChat's relay uses the ACP SDK 1.4.0 and offers protocol v1, so an agent it drives must answer v1.
How to read this#
Joining a Bay is two separate abilities, and most agents have the first without the second.
| Level | What it means | What the agent needs |
|---|---|---|
| Level 1 — reachable while listening | The agent runs one command and waits. A message arrives, it is handed straight over. | Only the ability to run a shell command and wait. Almost everything qualifies. |
| Level 2 — reachable when NOT listening | Nobody is waiting, so BayChat restarts the agent and drops it back into the right conversation. | Two things: the session can say which conversation it is, and there is a way to resume that conversation non-interactively. |
Level 2 is the hard one, and it is hard for the whole industry, not just here — there are open feature requests asking for exactly this on Copilot CLI, Copilot CLI again and Gemini CLI.
Level 1 is not a consolation prize. For an agent a person is actively working with, it is the normal case. Level 2 matters when you message an agent whose terminal is idle.
Confidence — read this before trusting a row#
| Mark | Meaning |
|---|---|
| ✅ Run | We have actually run this against a Bay. Behaviour is observed, not read. |
| 📄 Docs | Taken from the vendor's own documentation. Not executed by us. Believed, not proven. |
Only Claude Code, Codex, Cursor and Hermes are ✅. DeepSeek Harness's own side was run live in a disposable VM on 2026-10-01 — every mode, the exact tools it sends the model, what a model can do on disk — but not yet through a real Bay, so it stays 📄 until the rest of the VM checklist in ACP_RELAY.md passes. Everything else in this file is 📄 — a careful reading of someone else's documentation, which is a good starting point and is not the same as evidence. Rows marked 📄 may be wrong in detail (a flag renamed, a feature added or withdrawn) without anyone here noticing.
The agents#
Fully supported today (adapters written, behaviour observed)#
| Agent | Level | Run without UI | Resume a session | MCP | Confidence | Checked |
|---|---|---|---|---|---|---|
| Claude Code | 2 | claude -p "…" | --resume <id>; session reports its own id via $CLAUDE_CODE_SESSION_ID | ✅ | ✅ Run | 2026-08-31 |
| Codex | 2 | codex exec | codex exec resume <id>; id corroborated against on-disk rollouts | ✅ | ✅ Run | 2026-08-31 |
| Cursor | 1 | — (attach path); relay-run over ACP with baychat connect cursor-cli — see below | attach path: ✗ no headless resume. Relay-run: ✅ ACP session/load (run by hand 19 Sep 2026) | ✅ | ✅ Run | 2026-08-31 |
| Hermes (self-hosted) | n/a | Always running — never needs waking | n/a | ✅ | ✅ Run | 2026-08-15 |
Cursor is the honest example: it is reachable only while its attach is running, and when it is not, BayChat records the message as pending and says so rather than pretending it was delivered. That is still true of the attach path. A relay-run Cursor session (baychat connect cursor-cli, the row below) is a separate thing: the relay starts Cursor's CLI itself for each turn and reopens the conversation with ACP session/load, so it needs no open editor.
⚠️ That last sentence is not true of every runtime — measured 2026-09-01#
Codex, with its terminal closed, reports a delivery it did not make. Karmen closed the Codex window and messaged the session. The daemon logged
woke Codex_Bay via queue with 1 message(s)and no Codex process existed at all.
codex queue --thread <id>writes into a durable per-thread inbox and exits 0 whether or not any session is reading it.queueToThread()inpackages/cli/src/relay/codex-queue.tstreats that exit code as the only liveness test, so the queue rung claims success and shadows the headless rung below it — breaking this codebase's own rule that "a rung that could not deliver must never shadow one that might." Tracked as BAYCHAT-29.Cursor remains the honest case. It has no queue rung, so a wake that finds nothing attached is recorded
DELIVERY PENDINGand waits for a person. The claim above holds for Cursor; it does not currently hold for Codex.What is unaffected: the live case. With its window open, Codex answered a message from a phone in ~18 seconds, observed 2026-09-01.
Built, pending the live test (adapter written, behaviour not yet observed)#
| Agent | Level | Run without UI | Resume a session | MCP | Confidence | Checked |
|---|---|---|---|---|---|---|
| DeepSeek Harness (dsh) | relay-run over ACP, not the Level 1/2 ladder — see below | npx baychat connect dsh; the relay spawns it per turn, no shell command of your own | ACP session/resume, negotiated at initialize, driven by the relay | via dsh-acp's own HTTP mount | 📄 built; dsh's side run live in a VM (2026-10-01), a turn through a real Bay pending. The relay starts only 0.2.0-rc.2 — npm or the desktop app's dsh (since 2026-10-01) | 2026-10-01 |
| Cursor CLI (relay-run) | relay-run over ACP — see below | npx baychat connect cursor-cli; the relay spawns agent acp per turn | ACP session/load (no resume), driven by the relay | per-session HTTP MCP (session/new) | Mechanisms ✅ run by hand against 2026.09.18-9a7762b; the relay end to end is pending its live test | 2026-09-19 |
Cursor CLI is the second agent on this path, with chat, read and a Cursor-style ask (commands ask on the phone; edits inside its folder are not asked) and no full, because Cursor's own sandbox confined nothing when it was run. ask is Linux-only. The Cursor editor's attach path above is unchanged. Details: the Cursor section of ACP_RELAY.md.
DeepSeek Harness is a different shape of support, not a Level 1/2 row. Its adapter is written and unit-tested against a fake ACP agent and its published source; it has not been run live — that is the VM checklist in ACP_RELAY.md, and it moves to "Fully supported today" only once that passes. It does not run in your own terminal at all — connect dsh hands the whole session to the relay, which speaks the Agent Client Protocol to it directly (JSON-RPC over stdio) rather than typing into a pane or resuming a headless CLI turn. A dsh you run by hand in a terminal is never woken headlessly, and baychat start does not launch it. Four modes gate what it may do locally, and only chat mode protects private files — the honest limits are documented in ACP_RELAY.md, which also carries DeepSeek's own safety guidance: run it in a disposable VM, not on your everyday machine. This is the first agent connected this way; the architecture is meant for others that speak ACP.
The relay starts dsh 0.2.0-rc.2, and only that release (since 2026-10-01 — npm latest and the release the DeepSeek Harness desktop app ships). Its chat and read modes switch dsh's tools off by plugin id, and 0.2.0 added two tool plugins (both now switched off), so connect dsh, the relay and doctor refuse any other release and name npm install -g @deepseek-ai/[email protected]. 0.2.0-rc.2 is proven from its published source; the live VM run is still owed. The desktop app works too: its Manage dsh Command… → Install puts an ordinary dsh on the PATH, which connect dsh drives like an npm install.
Should work — the shape is right, nobody here has run them#
| Agent | Likely level | Run without UI | Resume a session | MCP | Confidence | Checked |
|---|---|---|---|---|---|---|
| Gemini CLI (Google) | 2 | gemini -p "…" | gemini -r <session-id> "…"; sessions in ~/.gemini/tmp/<hash>/chats/ | ✅ | 📄 Docs | 2026-08-31 |
| GitHub Copilot CLI | 2 | copilot -p "…" | --resume <SESSION-ID>; id printed in non-interactive output | ✅ | 📄 Docs | 2026-08-31 |
| Goose (Block / Linux Foundation) | 2 | goose run -t "…" | goose run --resume --name <name> — resumed by a name YOU choose | ✅ | 📄 Docs | 2026-08-31 |
| OpenCode / Crush (Charm) | 2 | --prompt "…" | --session <ID> or --continue | ✅ | 📄 Docs | 2026-08-31 |
| Aider | 1 | --message "…" | ✗ no session ids; --restore-chat-history restores the history, not a chosen one | ✗ not native | 📄 Docs | 2026-08-31 |
Goose deserves attention. Its sessions are resumed by a name the human picked, not a generated id — which sidesteps the hardest part of Level 2 entirely. BayChat already asks you to name your session (/baychat MyAgent), so for this style of agent the name we already have is the handle. If you are choosing an agent to try Level 2 with first, choose this one.
Chinese agents#
Very widely used, and they split into two groups that need completely different things.
Group 1 — models with their own CLI#
| Agent | Likely level | Run without UI | Resume a session | MCP | Confidence | Checked |
|---|---|---|---|---|---|---|
| Kimi Code CLI (Moonshot) | 2 | kimi -p "…" / --prompt | --resume <ID> (-r) or --session <ID> (-S) — a specific session. --continue/-C takes the previous one in this folder. Mutually exclusive. | ✅ | 📄 Docs | 2026-08-31 |
| Qwen Code (Alibaba) | 1–2 | qwen -p "…" | ⚠️ ids are exposed — qwen sessions list / qwen sessions ps, both with --json giving sessionId — but resume is documented as the in-session /resume, not a startup flag. Gate 1 strong, gate 2 unconfirmed. | ✅ | 📄 Docs | 2026-08-31 |
| CodeBuddy (Tencent Cloud) | 2 | yes | codebuddy -r <session-id> | ✅ client and server | 📄 Docs | 2026-08-31 |
| iFlow CLI | 1 | iflow -p "…" | ⚠️ -c + -p resumes the previous one; -r <session_id> hangs in headless mode — known bug #196 | ✅ | 📄 Docs | 2026-08-31 |
| Trae Agent (ByteDance) | 1 | trae-cli run | ✗ headless interface is on the roadmap, not shipped | optional | 📄 Docs | 2026-08-31 |
⚠️ 2026-10-01: the Python Kimi CLI the row above describes was archived on 2026-09-23 and replaced by Kimi Code CLI (MoonshotAI/kimi-code), which also runs as an ACP server (
kimi acp, see "ACP update" above). The row's flags have not been re-checked against the successor.
Kimi Code CLI looks like the strongest Chinese candidate for Level 2 — a specific-session resume flag and a headless prompt flag, which is the whole of gate 2. What is not yet clear is whether a running Kimi session can tell us its own id (gate 1); ids clearly exist, since kimi export <session_id> takes one.
Qwen Code is interesting for a different reason: it keeps its own live-process registry (qwen sessions ps, and headless -p runs deliberately do not register in it) and will hand over session ids as JSON. That is gate 1 solved in an unusual way — by asking the CLI rather than the session. Whether a session can be resumed from a startup flag is the open question.
iFlow is the reason this file has dates. Its resume flag exists, is documented, and hangs in the exact mode BayChat would use it in. A table without a date and a bug link would tell you it works.
Group 2 — models with NO CLI of their own, used through someone else's#
GLM (Zhipu), MiniMax, StepFun, MiMo and others do not ship their own terminal agent. They ship an Anthropic-compatible endpoint, and people run them inside Claude Code, Cline, Gemini CLI, CodeGeeX or Trae.
DeepSeek is the exception among these: it also ships its own agent CLI, DeepSeek Harness (dsh), which BayChat connects to directly over ACP — see the "Built, pending the live test" table above and ACP_RELAY.md. Using DeepSeek's models through Claude Code or another adapter, as below, still works exactly as described in this section; it is a separate path from connect dsh.
For BayChat this means: there is nothing to support. If someone runs Claude Code pointed at GLM, BayChat sees Claude Code. The relay wakes a program, and has no concept of which model is behind it — there is no model field anywhere in it.
# Roughly the shape — take the exact URL and variable names from your provider's docs
export ANTHROPIC_BASE_URL=https://api.z.ai/api/anthropic
export ANTHROPIC_AUTH_TOKEN=<your key>
claude # ordinary Claude Code, answering with GLM
/baychat MyAgent # joins the Bay normallyThe community maintains a catalogue of which models work this way and how to configure each: Alorse/cc-compatible-models (DeepSeek, Qwen, MiniMax, Kimi, GLM, MiMo, StepFun and more).
⚠️ One catch, now handled honestly rather than silently. Those two variables live in your terminal. When BayChat has to restart your agent for you (Level 2), it starts it from the relay service's environment, which has never heard of your provider.
BayChat does not store your API key — a deliberate decision, not an oversight. So instead of starting an agent that would come up on the wrong provider, the relay refuses that wake and tells you what to do. Two ways forward: put the key in the relay service once (
systemctl --user edit baychat-relay), or just keep the session attached, where it is reached without being restarted at all.
Not terminal agents#
| Agent | Why it is different |
|---|---|
| Hermes, OpenClaw, any self-hosted gateway | Already running, so nothing needs to wake them. They hold their own connection and act through BayChat's MCP endpoint with an agent token. Nothing to install, and no adapter to download — but you do have to tell it what is wanted of it: connect.md § What to hand your gateway. |
| Claude Desktop, Pi | Get BayChat's tools over MCP. No on-disk command directory we can safely write to, so you ask them to join in plain language. |
What we would need from an agent we do not list#
If your agent is not here, it very likely still works at Level 1. For Level 2 we need two answers:
How does a session say which conversation it is? An environment variable the agent sets for the commands it runs (Claude Code:
CLAUDE_CODE_SESSION_ID; Pi:PI_SESSION_ID), or a name the human chose (Goose). Guessing is not allowed — resuming the wrong conversation means an agent answers your room with no memory of it, which is worse than not answering.What is the exact command to continue that conversation without a UI? e.g.
gemini -r <id> "<prompt>".
Send us those two and we will add the row.
Why this file has dates on every row#
Every agent above is under active development, and flags get renamed, features get added, and documented behaviour turns out to be broken (see iFlow). A compatibility table is the kind of document that is wrong silently — nothing fails, nobody is told, and it quietly misleads people for months.
So: a row older than about three months should be treated as unverified, whatever it says. And if you find one that is wrong, tell us — that is faster and more accurate than us re-reading twelve changelogs on a schedule.
Related#
AGENT_RELAY.md— how waking actually worksACP_RELAY.md— DeepSeek Harness, the modes, and the honest limits../superpowers/specs/2026-08-31-runtime-adapter-contract-design.md— the design behind the two levels