Rove

Engines

An engine is the AI coding CLI a task runs on: claude, codex, copilot, kimi, pi, omp, or one you register yourself. Rove runs the real interactive CLI inside the task's terminal session.

Managed task = one git worktree + one branch + one or more terminal tabs

A Task can own several terminal tabs. Each engine tab has its own PTY process and conversation, but opening another engine or shell tab does not create another worktree: ordinary sibling tabs use the same Task directory. Two agents in sibling tabs can therefore edit each other's work; create another task when you need git-level isolation and a separate branch.

Which engines are supported

EngineIdAccount detectActivity badgeHistoryEffort levelsModel flag
Claude Codeclaude✓✓✓—--model (aliases fable/opus/sonnet, or a full id)
Codexcodex✓✓ (after you trust hooks)✓none/low/medium/high/xhigh/max--model (a slug)
GitHub Copilotcopilot✓✓ (screen-based)✓——
Kimi Codekimi✓✓handoff only—--model (an alias from its config)
Pipi—✓✓off/minimal/low/medium/high/xhigh/max--model (pattern or provider/id; listed by pi --list-models)
OMPomp—✓✓off/minimal/low/medium/high/xhigh/max--model= (pattern or provider/id; listed by omp models --json)
IBM Bobbobsigned in / out✓ (screen-based)✓——
Cursor Agentcursorbinary only✓ (screen-based, plus a session hook)———
Droid, Devin, Qoder CLIcontribbinary only✓ (screen-based, plus a session hook)———
Gemini CLI, OpenCode, Grok CLI, Amp, Cline, Kiro CLI, Maki, Antigravitycontribbinary only✓ (screen-based)———
Anything you registercustombinary only————

A model is pinned per task in the engine's own spelling (see Model below); an engine with no model flag refuses one rather than dropping it.

Claude Code is the default and the most complete: its quota probe drives rate-limit auto-resume and the Settings usage dashboard.

Codex has a quota probe too, and it works differently. There is no endpoint to call — the Codex CLI writes the server's rate_limits block into its rollout JSONL, so Rove reads the newest rollouts off disk. That makes it a snapshot of the last response Codex received: a window whose reset time has already passed is dropped, and an account that hasn't run Codex recently publishes nothing at all rather than a stale number.

Contrib engines are launch + badge only. Rove ships a catalog of well-known coding CLIs (gemini, opencode, cursor, grok, droid, amp, devin, qodercli, cline, kiro, maki, antigravity) so they appear in the engine selector whenever the binary is on your PATH, with a proper name, a launch command, and screen-based activity badges. A catalog entry also declares how its CLI takes a first message: OpenCode's positional argument is a project directory, so Rove pastes the prompt after launch instead of appending it to the command line.

Settings → Engines lists them (and your own registered engines) with their binary discovery, and that is all detection can answer for them. No login state, history, or model picker; those need a real adapter, which is what promotes an engine to built-in.

IBM Bob is built in, for its history and account reads rather than for hooks: Bob ships Claude's nested hook schema but nothing fires from it on 2.0.5, so session identity comes from its history store keyed by worktree, but Rove does not resume Bob conversations across tab restarts. Restarting a Bob tab launches a new conversation. Rove launches it as bob chat --trust — chat is Bob's TUI subcommand, and the first-run folder dialog would otherwise stop every task spawned into a fresh worktree, which is what breaks a parallel round. Its first message is pasted rather than appended: bob chat declares no positional and discards a stray one without an error. Account detection reports whether you are signed in, not who — Bob keeps only an opaque token, and Rove does not decode credential material.

Plugins can contribute engines too: a plugin manifest's [[engines]] entries register engines with a display name, launch command, screen manifest, and identity. Unlike the contrib catalog, plugin engines are offered without a binary check — installing the plugin is the opt-in. See Plugin authoring.

Kimi is partial. Rove finds the binary, reads its login state, and can locate each session's transcript, enough to watch it for activity and to hand the conversation to another engine. It still doesn't parse that transcript (the wire format is unverified), so auto-title keeps the placeholder and rove api read-output reports engine_unsupported rather than guessing.

Picking an engine

Per task, at creation time, or with v in the sidebar. For new tasks the per-project last-used engine wins; Settings → Engines (defaultVendor) is the fallback, then claude. From a script it is rove api add --command <engine-or-command-line>; see engine presets and protocols.

Engines whose CLI isn't installed are hidden from the new-task dialog, and so are engines you switched off in Settings → Engines (Settings still lists them so you can switch them back on). Custom engines always show; you added it, so Rove assumes you meant it.

Reasoning effort

Codex accepts none, low, medium, high, xhigh, max, passed as -c model_reasoning_effort=<level>. Pi and OMP take off, minimal, low, medium, high, xhigh, max as --thinking <level> — minimal in place of Codex's none, plus off. The remaining engines have no effort flag Rove can drive; a selected effort is ignored there rather than passed through.

Three places select one:

  • rove api add --command codex --effort LEVEL, which is the only one that reaches the task's first session — the other two rebuild it;
  • the sidebar row menu's Change engine entry, whose second row lists the engine's levels (←→ picks one, and "engine default" clears it). Engines that declare no levels show no row;
  • rove api update --task-id ID --effort LEVEL from a shell.

Wherever Rove shows a level it puts a fill glyph in front of it, ○ for the engine's lowest through ◔ ◑ ◕ ● to ◉ for its highest, placed by the level's position in that engine's own list (Codex's high reads ◕ high). With two-cell tab rows (sidebar.tabRowHeight), an agent tab's second line shows the pinned level after the engine name.

The board's start-a-task-from-an-issue picker chooses an engine but not a level, so a task started that way runs its first session on the engine's default until you set one.

Model

A task can pin a model the same way it pins an effort. The value is passed to the engine verbatim, in that engine's own spelling — a claude alias or full id, a codex slug, a pi/omp fuzzy pattern or provider/id — so the list an engine can print is a set of suggestions, never a closed list. rove api engine-list shows each engine's models: what pi --list-models or omp models --json answered, a short alias list for claude and codex, and null for an engine Rove cannot list (kimi's aliases live in its own config; copilot has no model flag).

The same three places select one:

  • rove api add --command pi --model cliproxy/claude-fable-5, which reaches the task's first session;
  • Change engine and the new-task dialog, whose model row is a free-text input with the engine's list as suggestions underneath (tab reaches the row in the change-engine picker; empty = the engine's default);
  • rove api update --task-id ID --model MODEL from a shell.

An engine that declares no model flag (copilot, contrib, custom) refuses a model up front (BAD_MODEL) instead of dropping it at launch.

Auto routing

Rather than picking the three fields by hand, a new task can be started at a depth — swift, standard or deep — and Rove fills the engine, model and effort from the table in Settings → Auto routing (autoRouting.<tier>.* in state.json, see CONFIGURATION.md). The three fields stay on screen and editable; the tier is recorded on the task (.task.tier). A tier whose target cannot start — engine not listed, not logged in, a model or effort its engine cannot carry — says so in Settings and is refused by rove api add --tier (TIER_UNAVAILABLE), not at launch.

--tier auto goes one step further and picks the depth from the prompt itself, through a classifier that is off until you configure it — see the tier classifier, which also spells out where the prompt text goes.

Workspace trust

Five of the six builtin engines gate a first launch in a never-seen directory behind a trust dialog, and every task worktree is such a directory, so a hosted session can't answer it — nobody is at the pane to press a key, so the launch sits on the dialog instead of starting the turn (and with Kimi, whether a stray Enter accepts or exits the process depends on the Kimi version; Copilot's cursor sits on a session-only "Yes", so it returns every launch). Before spawning an engine into a Rove-created worktree, Rove writes that vendor's own trust record for the path, merging into existing entries, never clobbering:

EngineTrust record
Claude~/.claude.json → projects[<path>].hasTrustDialogAccepted
Codex~/.codex/config.toml → [projects."<path>"] trust_level = "trusted"
Copilot~/.copilot/config.json → trustedFolders
Kimi~/.kimi-code/workspace-trust/<record>
Pi~/.pi/agent/trust.json → { "<canonical path>": true }
OMPnone — OMP has no project-trust gate

Claude, Codex, and Kimi trust writes follow CLAUDE_CONFIG_DIR, CODEX_HOME, and KIMI_CODE_HOME, respectively; Pi's follows PI_CODING_AGENT_DIR (the directory itself — ~/.pi/agent by default). With CLAUDE_CONFIG_DIR set, the Claude trust file is <CLAUDE_CONFIG_DIR>/.claude.json. Blank overrides use the default paths above. Each write resolves the current profile again.

Claude and Codex leave unreadable, non-regular, oversized (over 8 MiB), or invalid configuration unchanged; launch continues with the engine's own trust prompt. Codex can repair up to five duplicate standalone trust tables that contain only the same trust_level = "trusted" entry, after validating the complete repaired TOML. Existing Kimi records are never replaced. JSON rewrites and new Codex config files use owner-only read/write permissions (0600); updates to existing Codex files retain their permissions.

This only ever fires for worktrees Rove itself created from a repo you already work in; your own directories are untouched.

Contrib, plugin, and custom engines have no trust record Rove knows how to write, so whether a task worktree launches cleanly is up to the CLI. OpenCode does not gate at all (checked against 0.6.3 in an unseen directory, including with an empty config, so it needs nothing). Cursor Agent does gate — its --trust flag only applies to --print/headless runs, so an interactive task worktree stops at its trust prompt and you have to answer it once in the pane. The rest are unmeasured; if a task sits on a dialog at launch, that is what you are looking at.

Custom launch commands

Override any engine's launch command in Settings → Engines, or by hand in state.json:

{ "engineCommand.claude": "claude --model opus" }

Quotes are honored, so claude --append-system-prompt "be terse" works. For Claude, Rove appends its own --session-id so the tab stays resumable. If your override already pins the conversation (--session-id, --resume, --continue, --from-pr), Rove leaves it alone.

If an engine exits non-zero, the terminal stays open with a banner pointing at Settings → Engines, and drops you into a shell.

Activity badges

The sidebar shows what each session is doing: working, done, or needs input. There's nothing to configure. Rove reads the engine's own hook events, falling back to its transcript when hooks aren't available.

One thing worth knowing: the depth of the badge depends on the engine. Claude, kimi, and omp report the full working / done / needs-input vocabulary through hooks, sub-second. Codex reports working and done through hooks, but not needs-input: its only "waiting" event is a permission decision hook, and Rove will not install an observer on a hook that gates approvals. Codex ships no screen rules either, so it has no needs-input source at all today — a Codex session waiting on you reads as whatever its last hook said. Engines without hooks or a readable transcript (copilot today) rely on screen reading for everything: Rove classifies the visible terminal against engine-declared rules, which still distinguishes working from waiting-on-you but can't see a completed turn the way a transcript marker can. Rove labels the gap honestly rather than guessing.

Screen detection runs only for tabs attached to an open TUI. A matched blocked rule reaches the sidebar and attention inbox within one poll (up to six seconds), unless a hook claim takes precedence. The claim clears when the dialog disappears or polling detaches. Unopened tabs still need hooks to report approval requests.

Programs can report their own state with OSC 7501. Any engine (Claude Code 2.1.295+ does) may write ESC ] 7501 ; state=working|blocked|done|idle|error|clear [:msg=…] ST into its terminal. The PTY host parses it across output chunks, for every hosted tab whether or not a TUI is attached, and the daemon writes it to the tab's activity. The sidebar badge, the attention inbox and the phone bridge then read the new state from there. Precedence: a hook report at the same time or later wins over OSC 7501, and OSC 7501 wins over screen and title inference for that tab; an exited engine still reads as exited. blocked means waiting on you (kind=permission marks an approval). clear or a terminal reset drops the claim. Statuses carrying an id= describe child work and do not move the badge. ESC ] 7501 ; ? ST is answered so a program can detect support. OSC 3008 context records are parsed and kept but drive no UI yet.

Pi and OMP differ from each other on waiting. They share one adapter and one hook file — OMP is Stencil Labs' fork of the pi coding agent, so both load a TypeScript extension from <agent dir>/extensions/ and dispatch the same pi.on(...) events. Rove writes rove-activity.ts there, and it is the only hook install that shells out through an engine API (pi.exec) rather than editing a settings file. OMP emits tool_approval_requested when it blocks on its native approval prompt — and its question tool reports the same way — so its needs-input badge is exact. pi has no tool-approval prompt at all, so there its needs-input source is the screen rules (its trust dialog, and any other modal). pi and OMP also report a user interrupt from the assistant message itself (stopReason: "aborted"), which is the one native interrupt signal Rove sees anywhere — on claude and codex an interrupt has to be inferred from the terminal title.

Codex won't run Rove's hooks until you trust them once via /hooks, so Codex badges stay dark until you approve. That's by design: Rove writes the hook definition but never bypasses the trust prompt for you.

Pi and OMP hook installation writes rove-activity.ts into <agent dir>/extensions/, and cleanup removes that file. The agent directory is PI_CODING_AGENT_DIR when set, else ~/.pi/agent or ~/.omp/agent. If that directory does not exist, nothing is written: there is no CLI to read it.

Cursor installs one hook, sessionStart, into hooks.json under CURSOR_CONFIG_DIR (unset or blank uses ~/.cursor). That hook reports which cursor session is live in a worktree; the badge itself keeps coming from the screen rules, because cursor's other hook events — beforeSubmitPrompt, beforeShellExecution, beforeMCPExecution, stop, sessionEnd — gate the agent's own actions, and Rove does not install an observer on a hook that gates approvals. If ~/.cursor does not exist, nothing is written and no directory is created: there is no CLI there to read it. Your own entries in hooks.json, and every other event, are left alone.

Droid, Devin, and Qoder CLI get the same single observer, SessionStart, written into their Claude-shaped settings file: ~/.factory/settings.json for Droid, $XDG_CONFIG_HOME/devin/config.json (else ~/.config/devin/config.json) for Devin, and $QODERCLI_CONFIG_DIR/settings.json (else ~/.qoder/settings.json) for Qoder CLI. As with Cursor, the hook only reports which session is live; the badge comes from the screen rules. If the settings file's directory does not exist, nothing is written.

Every hook Rove installs carries the version of the shape that wrote it, so Rove can tell its own current entry from one an older version left behind. Settings → Engines reads that back per engine as installed, outdated or not installed; an entry written before versions existed reads as outdated. The install runs on every launch and replaces its own entries rather than stacking a second copy, and the row at the bottom of that section runs the same install on demand.

Claude and Codex hook installation and cleanup use settings.json under CLAUDE_CONFIG_DIR and hooks.json under CODEX_HOME. Unset or blank overrides use ~/.claude and ~/.codex. Invalid JSON or hook structure, unreadable files, non-regular files, and files over 8 MiB are left unchanged. Other user settings and commands in a shared hook group survive cleanup. Cleanup recognizes literal rove/rove invocations, including absolute executables and Bun/Node source or bundle entry paths. Commands behind shell wrappers or compound shell commands are left for manual review.

Adding a hook to another engine

Hooks are no longer a built-in-only privilege. A shipped catalog entry in packages/rove/src/engine/contrib-engines.ts may declare a createHookAdapter alongside its screenManifest, and cursor is the worked example (engine/cursor-local/hook-adapter.ts). Declaring one costs the other catalog entries nothing — an engine without it keeps the no-op adapter, installs nothing, and warns about nothing.

Two things are worth getting right. A hook adapter does not REPLACE the screen rules: hooks, transcript markers and screen reading cover different gaps, so an engine that gains a hook keeps its manifest. And the set of hook-capable engines is derived in one place, activityHookAdapters() in engine/hook-adapter.ts — both the launch-time installer and rove doctor read it, so a new adapter reaches both without touching either.

Mechanics: design/engine-internals.md.

Resuming and forking

Resume. Whether a restarted tab comes back into its old conversation depends on the engine. Claude Code accepts a caller-set session id, so Rove pins one at launch and a tab whose process is gone (a reboot, say) relaunches into the same conversation instead of a blank one. Codex and Kimi mint their own ids — Codex announces its in the terminal title, Kimi's is discovered from its session store after the fact — and each reopens the last conversation with its own resume verb (codex resume <id>, kimi -S <id>). Pi pins one too (--session-id, reopened with --session <id>); OMP mints its own and Rove learns it from the hook payload or the session store, then reopens it with -r <id prefix>. Copilot and custom engines have no resume verb Rove knows, so their tabs relaunch fresh. A tab that never sent a first message isn't resumed; there's no transcript yet.

Continue in a new tab. ctrl+a c opens the continuation flow in the same worktree. What happens depends on the source and destination engines:

Same engine → a native fork only when that CLI can actually branch a conversation. The two resulting tabs keep the source context and then diverge:

EngineNative fork?How
claude✓--resume <src> --fork-session
codex✓codex fork <src>
pi✓--fork <src>
omp—no fork verb — transcript handoff
copilot—starts a fresh Copilot session with a transcript handoff
kimi—has a kimi fork verb, but it exits instead of opening the session — transcript handoff instead
custom—refused, unless the preset declares a built-in protocol (then it forks like that engine)

Copilot's --resume reopens rather than branches, which would put two live processes on one transcript. Kimi 0.40.1 does ship a fork [sessionId] subcommand, but it is a one-shot that prints the new session id and exits, so it cannot BE a tab's launch command — branching there would mean forking and then resuming, two launches. Rove therefore uses the same transcript handoff as a cross-engine continuation: the new tab is a fresh conversation that reads where the previous one stopped. A custom engine without a known session store is refused instead of silently opening a blank continuation.

Each engine declares its own fork verb, so a custom preset with an engineProtocol of claude or codex forks like that engine — the same resolution that gives it their session-id flags.

Different built-in engine → a handoff. This is the move that saves you when you hit a usage limit mid-task. The new engine starts fresh with a first prompt that points it at the old session's transcript and asks it to state where the previous one stopped. That sentence is how you check the handoff landed.

Handoffs work in every direction between the four built-ins. The receiving agent is handed the previous transcript's path and reads it in whatever format it finds, so no format ever has to be converted. Only a custom engine can't be a handoff source (Rove doesn't know its session store); that case is refused with a reason. A handoff to a custom engine works fine.

Custom engines

Any CLI can be an engine. A custom engine is a named preset: an id, a launch command, an optional display name, and, declared once at registration, the protocol Rove speaks to it. Settings → Engines → + Add engine asks for each. Or by hand:

{
  "customEngineIds": ["aider"],
  "engineCommand.aider": "aider --model sonnet",
  "engineName.aider": "Aider",
  "engineProtocol.aider": "claude"
}

Being in customEngineIds is the registration.

engineProtocol.<id> is optional and says "talk to my binary the way you talk to claude". Set it when your CLI is a wrapper around a built-in (a different binary name, a fixed set of flags, a company shim) so the transcript reader, workspace-trust pre-answer, and first-message delivery all apply. Leave it out and the preset gets the generic protocol: it launches and runs fine with settle-then-paste delivery, but Rove reads no history, pre-answers no trust dialog, and shows no working badge (the activity observer only recognises engines with an adapter or screen manifest, so a generic engine reads as at-rest). If the running session later reveals a built-in engine behind the command, Rove upgrades the task's protocol automatically — see engine presets and protocols.

Press x on an engine row in Settings to reset a built-in's overrides, or remove a custom engine entirely.

Engine presets and protocols

Rove separates two things a "vendor" used to conflate:

  • the command. What actually runs. A preset id (claude) or a full command line (codex --search). Nothing validates its flags; probe an unfamiliar CLI with <cmd> --help first.
  • the protocol. Which adapter Rove uses for that command: whose transcripts to read, whose trust store to pre-answer, whether the first message may ride the launch argv.

The protocol is derived from the command, never declared beside it:

  1. argv[0] names a preset — a built-in, a contrib engine from the table above, or one of yours → that preset's protocol. Deterministic, and it answers before anything spawns. This is the normal path.

  2. Otherwise Rove can recognise a known engine binary through wrappers (env FOO=1 claude, node …/codex.js), the same walk the process probe uses at runtime.

  3. Neither → generic, described above — until the live session says more. While a generic task's engine tab runs, the daemon watches for a built-in engine's fingerprint in it: the engine's process in the tab's process tree, or a status glyph only one engine writes into its terminal title. When exactly one engine is identified, the task's protocol is upgraded in place (the command it launches never changes), so history, trust pre-answer, and delivery start applying mid-session. The sniff is deliberately conservative: ambiguous or absent evidence leaves the task generic, and a task whose protocol is already known is never flipped.

    When that engine tab was launched from one of your registered presets and the walk found a built-in's process in it, Rove also writes engineProtocol.<id> for the preset itself, so the next task on it starts named instead of sniffing again. Only the process walk mints that key — a title glyph names one session, and this outlives every session on the preset — and a protocol you declared is never overwritten.

rove api engine-list prints every entry with its raw command and resolved protocol; copy one into rove api add --command verbatim, or edit a flag first. See API.md for the dispatch verbs.

Where conversations are stored

Engines own their own history. Rove reads it, never writes it.

EngineTranscripts
claude~/.claude/projects/<encoded-cwd>/<sessionId>.jsonl
codex~/.codex/sessions/<YYYY>/<MM>/<DD>/rollout-*.jsonl
copilot~/.copilot/session-state/<id>/events.jsonl
kimi~/.kimi-code/session_index.jsonl maps each session to its dir; the stream is <sessionDir>/agents/main/wire.jsonl
bobone SQLite database, ~/.bob/db/bob.db — tasks keyed by workspace, messages per task. No per-session file, so nothing to hand another engine in a handoff
pi, omp<agent dir>/sessions/<encoded-cwd>/<timestamp>_<session id>.jsonl — ~/.pi/agent and ~/.omp/agent unless PI_CODING_AGENT_DIR says otherwise. pi encodes the cwd as an absolute path, OMP as one relative to your home (or the temp root); Rove reads both spellings

That's why a crash never loses a conversation, and why history survives rove reset and a machine reboot.

On this page