SkillAtlasSkill 详情

codeman

Mission control for AI coding agents

审核状态:已审核Quality 72Security 52

复制安装命令

用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。

复制前请先查看来源、License 和安全提示。

项目 README

来源文件:README.md

抓取于 2026年9月14日

Codeman

Mission control for AI coding agents

Claude Code • OpenCode • Codex • Antigravity • Gemini • Pi • Grok • OMP • Terminal - One Dashboard • Any Device

License: MIT Node.js 22+ TypeScript 5.9 Fastify npm version GitHub stars Contributors Total commits

English • 简体中文

Codeman — parallel subagent visualization

Codeman is a self-hosted mission control for AI coding agents. It spawns Claude Code, OpenCode, Codex, Antigravity, Gemini, Pi, Grok, or OMP inside persistent tmux sessions, streams the real terminal to any browser, and keeps agents productive after you walk away: it re-prompts on idle, resumes when a usage limit resets, runs scheduled jobs, and shows every background agent working in real time.

Get started in one line (macOS & Linux, Windows via WSL):

curl -fsSL https://getcodeman.com/install | bash
codeman web
# Open http://localhost:3000 and start your first session

The installer asks before every system change, and re-running the same line updates in place. Full details: Quick Start - Installation.

Codeman dashboard tour: session tabs per case, one-click Run for new agents, live plan usage in the header


Quick Start - Installation

curl -fsSL https://getcodeman.com/install | bash

This installs Node.js, tmux and a build toolchain if missing (node-pty ships no Linux prebuilds, so it compiles from source), clones Codeman to ~/.codeman/app, and builds it. A few things worth knowing:

  • It asks first. Every system change (package installs, AI CLI download) is prompted, and a menu at the end lets you choose: run Codeman in this terminal, install it as a background service (systemd/launchd, auto-start on boot), or don't start yet. Nothing runs in the background unless you pick it.
  • How it's reachable, your choice. The installer offers three ways to reach the dashboard: Tailscale (loopback bind fronted by tailscale serve, so you get https://<machine>.<tailnet>.ts.net with a real certificate and your tailnet as the login, no password needed), any device on your network (0.0.0.0, with a strongly recommended password prompt), or this machine only (127.0.0.1, safest). Skipping the password on a network bind requires an explicit confirmation and ends with a loud warning. The highlighted default reflects what is already on the machine (Tailscale when it is already in use, your existing binding on a re-run), and a bare Enter never pulls in new software. A bare codeman web started by hand still defaults to loopback.
  • Re-run to update. The same one-liner updates a finished install in place: local changes in ~/.codeman/app are stashed (never discarded), and a running service is restarted and verified. If a first install was interrupted, re-running resumes the full setup instead. install.sh update and install.sh uninstall also exist.
  • CI / headless: without a terminal attached, steps that would change your system abort with instructions instead of running silently. Set CODEMAN_NONINTERACTIVE=1 to approve them for automation.

You'll need at least one AI coding CLI installed — Claude Code, OpenCode, Codex, Antigravity, Gemini CLI, Pi, Grok Build, DeepSeek Harness, or OMP (any combination works; Gemini CLI is enterprise-only since Google's consumer cutover, and Antigravity is its successor). The installer detects whichever of the nine is present; if none is found, it offers to install Claude Code or OpenCode, or you can skip and install one yourself later. After install:

codeman web
# Open http://localhost:3000 and start your first session

Sharing with a small team? Start it in multi-user mode instead: each person gets their own login and workspace.

codeman users add alice --admin      # create the first admin account
codeman web --multiuser              # named logins + per-user case spaces

Prefer Docker Compose? A local-image Compose deployment ships in docker/: copy docker/.env.example to docker/.env, set CODEMAN_PASSWORD, then run bash docker/Start-Codeman.sh on Linux. Codeman runs in a container and spawns Docker cases as sibling containers through the host socket. See the Docker deployment guide for direct Compose commands, storage and networking options.

Details in Multi-User Mode below.

Keep it running in the background

To outlive the shell you started it in, without setting anything up:

codeman web -d          # detach; logs to ~/.codeman/web.log
codeman web --status    # is it up, and on which pid
codeman web --stop      # graceful SIGTERM; agents keep running in tmux

-d waits until the server actually answers before reporting success, and refuses to start a second one on the same data dir (two servers sharing a tmux socket attach to each other's sessions).

To have it come back after a reboot, install it as a service instead. The installer's final menu does this for you (option 2); codeman service is the equivalent for an npm i -g aicodeman install:

codeman service install     # systemd user unit (Linux) or LaunchAgent (macOS)
codeman service status
codeman service uninstall

service install writes the unit with your current PATH baked in, which matters more than it sounds: launchd hands a job /usr/bin:/bin:/usr/sbin:/sbin, so a Homebrew or nvm node, tmux or claude is invisible to a hand-written plist. It never copies CODEMAN_PASSWORD into the unit file; add that yourself if the service needs auth.

To write the unit by hand instead:

Linux (systemd):

mkdir -p ~/.config/systemd/user
cat > ~/.config/systemd/user/codeman-web.service << EOF
[Unit]
Description=Codeman Web Server
After=network.target

[Service]
Type=simple
ExecStart=$(which node) $HOME/.codeman/app/dist/index.js web
Restart=always
RestartSec=10

[Install]
WantedBy=default.target
EOF
systemctl --user daemon-reload
systemctl --user enable --now codeman-web
loginctl enable-linger $USER

macOS (launchd):

mkdir -p ~/Library/LaunchAgents
cat > ~/Library/LaunchAgents/com.codeman.web.plist << EOF
<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN"
  "http://www.apple.com/DTDs/PropertyList-1.0.dtd">
<plist version="1.0">
<dict>
  <key>Label</key>
  <string>com.codeman.web</string>
  <key>ProgramArguments</key>
  <array>
    <string>$(which node)</string>
    <string>$HOME/.codeman/app/dist/index.js</string>
    <string>web</string>
  </array>
  <key>RunAtLoad</key><true/>
  <key>KeepAlive</key><true/>
  <key>StandardOutPath</key>
  <string>/tmp/codeman.log</string>
  <key>StandardErrorPath</key>
  <string>/tmp/codeman.log</string>
</dict>
</plist>
EOF
launchctl bootstrap gui/$(id -u) ~/Library/LaunchAgents/com.codeman.web.plist
Windows (WSL)
wsl bash -c "curl -fsSL https://getcodeman.com/install | bash"

Codeman requires tmux, so Windows users need WSL. If you don't have WSL yet: run wsl --install in an admin PowerShell, reboot, open Ubuntu, then install your preferred AI coding CLI inside WSL (Claude Code, OpenCode, Codex, Antigravity, Gemini CLI, Pi, Grok Build, DeepSeek Harness, or OMP). After installing, http://localhost:3000 is accessible from your Windows browser.


Mobile-Optimized Web UI

The most responsive AI coding agent experience on any phone. Full xterm.js terminal with local echo, swipe navigation, and a touch-optimized interface designed for real remote work — not a desktop UI crammed onto a small screen.

Mobile — answering an agent's plan prompt with the keyboard accessory bar and Enter buttonMobile toolbar: accessory bar with /init, /clear, clipboard and Esc above the Run, case, stop, Enter, voice and settings controls
Answering prompts by touchAccessory bar + dedicated Enter button
Terminal AppsCodeman Mobile
200-300ms input lag over remoteLocal echo — instant feedback
Tiny text, no contextFull xterm.js terminal
No session managementSwipe between sessions
No notificationsPush alerts for approvals and idle
Manual reconnecttmux persistence
No agent visibilityBackground agents in real-time
Copy-paste slash commandsOne-tap /init, /clear, /compact
Password typing on phoneQR code scan — instant auth
  • Keyboard accessory bar — /init, /clear, /compact quick-action buttons above the virtual keyboard; destructive commands require a double-press to confirm, so you never fire one by accident
  • Dedicated Enter button — replays the keypress through the terminal, so text buffered by local echo is flushed first rather than stranded
  • Swipe navigation & smart keyboard handling — swipe left/right to switch sessions; toolbar and terminal shift up when the keyboard opens (visualViewport API)
  • Built for phones — safe-area insets for notch and home indicator, 44px touch targets, bottom-sheet case picker, native momentum scrolling
codeman web --https
# Open on your phone: https://<your-ip>:3000

localhost works over plain HTTP. Use --https when accessing from another device, or use Tailscale (recommended): the installer can set it up for you (choose Tailscale at the network-access prompt, or run bash ~/.codeman/app/install.sh tailscale on an existing install). That gives you https://<your-machine>.<tailnet>.ts.net with a real certificate: private to your tailnet, no password required, and PWA install + push notifications work on your phone.

Secure QR Code Authentication

Typing passwords on a phone keyboard is miserable. Codeman replaces it with cryptographically secure single-use QR tokens — scan the code displayed on your desktop and your phone is authenticated instantly.

Each QR encodes a URL containing a 6-character short code that maps to a 256-bit secret (crypto.randomBytes(32)) on the server. Tokens auto-rotate every 60 seconds, are atomically consumed on first scan (replays always fail), and use hash-based Map.get() lookup that leaks nothing through response timing. The short code is an opaque pointer — the real secret never appears in browser history, Referer headers, or Cloudflare edge logs.

The security design addresses all 6 critical QR auth flaws identified in "Demystifying the (In)Security of QR Code-based Login" (USENIX Security 2025, which found 47 of the top-100 websites vulnerable): single-use enforcement, short TTL, cryptographic randomness, server-side generation, real-time desktop notification on scan (QRLjacking detection), and IP + User-Agent session binding with manual revocation. Dual-layer rate limiting (per-IP + global) makes brute force infeasible across 62^6 = 56.8 billion possible codes. Full security analysis: docs/qr-auth-plan.md


Using Codeman — A Human's Guide

A start-to-finish walkthrough for driving Codeman from the browser. If you just installed, this is where to begin.

1. Launch the server

codeman web                       # localhost:3000 (loopback only — safe default)
codeman web --port 8080           # custom port (or set CODEMAN_PORT)
codeman web --https               # self-signed TLS (only needed for remote access)
codeman web -H 0.0.0.0            # bind LAN — REQUIRES CODEMAN_PASSWORD (see Security)
codeman web -d                    # detach: survives closing the shell (--status, --stop)
codeman service install           # systemd/launchd service: comes back after reboots

Open the printed URL. The page is a single dashboard; everything below happens there.

2. Create your first session

Click + New Session (or Quick Start). A session is one AI CLI running in its own tmux-backed terminal. You choose:

FieldWhat it does
Working directory / caseThe folder the agent operates in. A "case" is just a named working dir Codeman remembers. Add Case creates one from scratch, links an existing folder, or clones a GitHub repo straight into one (Clone Repo).
CLI / run modeClaude (default), OpenCode, Codex, Antigravity, Gemini, Pi, Grok, OMP, or Terminal (plain shell).
ModelPer-session model (App Settings → Models → New Claude sessions). A soft default — /model still works in-session.
Effort / UltracodeReasoning effort (low–max) or ultracode for dynamic multi-agent workflows. Switchable anytime with /effort.

Hit start — Codeman spawns the CLI via a real PTY and streams it to your browser over SSE.

3. Read the dashboard

  • Tabs (top) — one per session. Alt+1-9 to jump, Ctrl+Tab for next, drag to reorder (tab order syncs across your devices).
  • Terminal (center) — a real xterm.js terminal; full TUIs render correctly. Type directly and press Enter to send. Shift+Enter inserts a newline.
  • Side panels — Respawn, Orchestrator, Cron, Subagents, Settings (toggled from the toolbar).

4. Talk to the agent

  • Type prompts straight into the terminal — input is delivered exactly-once even across reconnects (a dropped link never loses or double-sends a prompt).
  • Paste or drag-and-drop images directly into the session.
  • Voice input — Ctrl+Shift+V (Deepgram Nova-3, with auto-silence stop).
  • Attachments — register external files/docs and preview Office/PDF inline.

5. Make it autonomous

ModeUse it forWhere
RespawnLong unattended runs — auto-restarts the CLI on idle/limit, with adaptive timing. Presets: solo-work, overnight-autonomous, …Respawn tab
OrchestratorTurn one goal into a phased plan and drive it to completion across agents.Orchestrator panel
CronSaved, named jobs on a schedule (once/interval/daily/weekly) that spawn a session and send a prompt when due.⏰ Cron button (opt-in: App Settings → Header & Panels → Scheduling)
Auto-resumeAutomatically continue after a subscription rate-limit resets.Respawn tab (top)

6. Reach it from anywhere

  • Phone/tablet — the UI is fully touch-optimized; scan the desktop QR code to log in without typing a password.
  • Outside your network — ./scripts/tunnel.sh start opens a Cloudflare tunnel (set CODEMAN_PASSWORD first).
  • SSH — codeman tui is a full-screen dashboard in the terminal (codeman tui --list to list, codeman tui 2 to attach straight to one).

7. Operate & maintain

  • App Settings — model, effort, permission startup mode, theme/skin, notifications, display toggles, per-CLI options, a synced custom display name, and per-device English/Simplified Chinese UI language.
  • Run it in the background — codeman web -d detaches from your shell (--status, --stop); codeman service install makes it a systemd user unit / macOS LaunchAgent that survives reboots. Both verify the server actually answers before reporting success, and both refuse to start a second server on one data dir. See Keep it running in the background.
  • Self-update — git-clone installs update in place from App Settings → System → Updates.
  • Deploy your own changes — see Development.

⚠️ Safety: if you're working inside a Codeman-managed session (echo $CODEMAN_MUX → 1), never run tmux kill-session / pkill claude directly — use the web UI or ./scripts/tmux-manager.sh.


Zero-Lag Input Overlay

Zerolag demo: instant local echo next to 600ms-2.7s server echo, side by side on two phones

When accessing your coding agent remotely (VPN, Tailscale, SSH tunnel), every keystroke normally takes 200-300ms to round-trip. Codeman implements a Mosh-inspired local echo system that makes typing feel instant regardless of latency.

A pixel-perfect DOM overlay inside xterm.js renders keystrokes at 0ms. Background forwarding silently sends every character to the PTY in 50ms debounced batches, so Tab completion, Ctrl+R history search, and all shell features work normally. When the server echo arrives 200-300ms later, the overlay seamlessly disappears and the real terminal text takes over — the transition is invisible.

  • Ink-proof architecture — lives as a <span> at z-index 7 inside .xterm-screen, completely immune to Ink's constant screen redraws (two previous attempts using terminal.write() failed because Ink corrupts injected buffer content)
  • Font-matched rendering — reads fontFamily, fontSize, fontWeight, and letterSpacing from xterm.js computed styles so overlay text is visually indistinguishable from real terminal output
  • Full editing — backspace, retype, paste (multi-char), cursor tracking, multi-line wrap when input exceeds terminal width
  • Persistent across reconnects — unsent input survives page reloads via localStorage
  • Enabled by default — works on both desktop and mobile, during idle and busy sessions

Extracted as a standalone library: xterm-zerolag-input — see Published Packages.


Live Agent Visualization

Watch background agents work in real-time. Codeman monitors agent activity and displays each agent in a draggable floating window with animated Matrix-style connection lines back to the parent session.

Subagent Visualization: three parallel Explore agents as floating windows with live tool-call feeds

  • Floating terminal windows — draggable, resizable panels for each agent with a live activity log showing every tool call, file read, and progress update as it happens
  • Connection lines — animated green lines linking parent sessions to their child agents, updating in real-time as agents spawn and complete
  • Status & model badges — green (active), yellow (idle), blue (completed) indicators with Haiku/Sonnet/Opus model color coding
  • Auto-behavior — windows auto-open on spawn, auto-minimize on completion, tab badge shows "AGENT" or "AGENTS (n)" count
  • Nested agents — supports 3-level hierarchies (lead session -> teammate agents -> sub-subagents)

Multi-agent Workflow runs ("ultracode") get the same treatment: a floating run window tracks the whole workflow live, with phases, per-agent token counts, and the current tool of every agent:

Ultracode workflow visualization: a live run window with per-agent tokens and phases

Agent Teams — first-class support for Claude Code's native multi-agent teams (CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1). TeamWatcher polls ~/.claude/teams/, matches teammates to their lead session, and surfaces them as live subagent windows with team-aware idle detection — so the Respawn Controller won't fire while teammates are still working. See docs/agent-teams/.


Respawn Controller

The core of autonomous work. When the agent goes idle, the Respawn Controller detects it, sends a continue prompt, cycles context management commands for fresh context, and resumes — running 24+ hours completely unattended.

WATCHING → IDLE DETECTED → SEND UPDATE → /clear → /init → CONTINUE → WATCHING
  • Multi-layer idle detection — completion messages, AI-powered idle check, output silence, token stability
  • Auto-resume on usage limit (opt-in, off by default) — when Claude halts on a subscription limit ("You've hit your limit · resets 3pm"), Codeman parses the reset time, waits it out plus a 2-minute safety buffer, then dismisses the rate-limit dialog and sends continue — so an overnight run survives the 5-hour window instead of stalling until morning. Recognizes every Claude Code limit-message format, retries if still limited, survives Codeman restarts, and holds respawn cycles while paused so /clear can't wipe the waiting conversation. Enable per session at the top of the Respawn tab
  • Circuit breaker — prevents respawn thrashing when Claude is stuck (CLOSED -> HALF_OPEN -> OPEN states, tracks consecutive no-progress and repeated errors)
  • Health scoring — 0-100 health score with component scores for cycle success, circuit breaker state, iteration progress, and stuck recovery
  • Built-in presets — solo-work (3s idle, 60min), subagent-workflow (45s, 240min), team-lead (90s, 480min), ralph-todo (8s, 480min), overnight-autonomous (10s, 480min)

Orchestrator Loop

Beyond single-session respawn, the Orchestrator turns a high-level goal into a phased plan and drives it to completion across multiple agents — a state machine that runs idle → planning → approval → executing → verifying → (replanning) → completed.

  • Plan, then execute — generates a phased plan from your goal and pauses for approval before touching anything; reject with feedback to regenerate
  • Per-phase verification gates — each phase is verified before the next begins; on failure the orchestrator replans instead of barreling ahead
  • Multi-agent execution — fans phases out to team agents / a task queue, coordinating work too big for one session
  • Crash-safe — full state persists under the orchestrator key in state.json, so it survives restarts
  • Driven from the UI or API — the Orchestrator panel, or POST /api/orchestrator/start → /approve → /status (10 endpoints)

Full design: docs/orchestrator-loop-architecture.md.


Multi-Session Dashboard

Run 20 parallel sessions with full visibility — real-time xterm.js terminals at 60fps, per-session token and cost tracking, tab-based navigation, and one-click management.

Persistent Sessions

Every session runs inside tmux — sessions survive server restarts, network drops, and machine sleep. Auto-recovery on startup with dual redundancy. Ghost session discovery finds orphaned tmux sessions. Managed sessions are environment-tagged so the agent won't kill its own session.

Session Manager & Command Palette

Ctrl/Cmd/Alt+K opens a fuzzy session palette; Browse all sessions opens the Session Manager: one deduped list of everything Codeman knows about (live sessions, past sessions from state and lifecycle history, and Claude transcripts), each row showing its first and most recent prompt.

  • Pinning: pin a session to float it to the top of the list. Pinned sessions even survive kill (they demote to a lightweight stopped entry that stays visible and resumable).
  • Name retention: resuming a past session keeps its original name instead of minting a new one.
  • Cross-device tab order: drag-reordered tabs persist server-side, so your ordering follows you from desktop to phone.

Hostname-Aware Window Title

Running Codeman on multiple hosts (laptop, dev box, NAS)? The browser tab title is codeman:<hostname> so you can tell which backend each tab points at without clicking in:

codeman web                                # codeman:<os.hostname()>
codeman web --title-hostname dev-box       # codeman:dev-box (manual override for noisy hostnames)

The title is templated into the served HTML on first byte, so it's correct from the very first paint and works without JavaScript. The same hostname prefix is applied to the tab-flash format (⚠️ (N) codeman:<host>) and to OS-level desktop notifications (codeman:<host>: <event>), so cross-host alerts in the system notification center are also unambiguous.

Smart Token Management

ThresholdActionResult
110k tokensAuto /compactContext summarized, work continues
140k tokensAuto /clearFresh start with /init

Tab Alerts

Session tabs: a regular active tab beside a yellow waiting-for-input tab and a red needs-decision tab, both with a breathing glow

Every tab tells you its state at a glance. A running session keeps its green status dot. When a session stops and waits for input, its tab turns yellow: steady ring, tinted background, yellow dot, with a slow breathing glow on top. When a permission prompt or question is blocking the agent, the tab turns red with a faster pulse. The base tint never blinks off, so even a split-second glance (or a screenshot) reads the true state; the ring stays visible while the tab is selected, and a page reload re-arms pending alerts from the server, so a blocked session can never hide behind a fresh-looking tab.

Notifications

Real-time desktop alerts when sessions need attention — permission_prompt and elicitation_dialog trigger critical red tab blinks, idle_prompt triggers yellow blinks. Click any notification to jump directly to the affected session. Hooks auto-configured per case directory.

Run Summary

Click the chart icon on any session tab to see a timeline of everything that happened — respawn cycles, token milestones, auto-compact triggers, idle/working transitions, hook events, errors, and more.

Zero-Flicker Terminal

Terminal-based AI agents (Claude Code's Ink, OpenCode's Bubble Tea) redraw the screen on every state change. Codeman implements a 6-layer anti-flicker pipeline for smooth 60fps output across all sessions:

PTY Output → 16ms Server Batch → DEC 2026 Wrap → SSE → Client rAF → xterm.js (60fps)

More Features

  • Background daemon & service install — codeman web -d runs the server detached with a pidfile, ~/.codeman/web.log, and verified startup (it polls the server until it answers, so a port clash never reads as success); codeman service install writes a systemd user unit (Linux) or LaunchAgent (macOS) with your shell's PATH baked in, so an nvm or Homebrew node, tmux and claude are actually found. Secrets are never written into unit files
  • Self-update — git-clone installs under systemd/launchd update in place from App Settings → System → Updates: it detects the latest release, auto-stashes a dirty tree, and streams build progress across the service restart (npm installs report as non-updatable)
  • Clone a GitHub repo as a case — paste a repository URL into Add Case → Clone Repo and Codeman clones it into ~/codeman-cases/<name> and registers it as a normal case, ready to run an agent in. It preflights the URL while you type (tells you whether it can be cloned anonymously and offers the repo's real branches and tags for the optional branch/tag field), fills the case name in from the URL, and lets you pick which CLI the Run button should use. Public repositories over https://; Codeman never collects or stores credentials
  • Multi-CLI — run Claude Code, OpenCode, Codex, Antigravity, Gemini, Pi, Grok, or OMP per session; env-var prefixes auto-gate (CLAUDE_CODE_* vs OPENCODE_* vs CODEX_* vs ANTIGRAVITY_* vs GEMINI_*/GOOGLE_* vs PI_* vs GROK_*/XAI_* vs OMP_*). See docs/opencode-integration.md, docs/pi-integration.md, docs/grok-integration.md and docs/omp-integration.md
  • Docker sessions — run a case inside an isolated, hardened container. One checkbox on Create New spins up a container with sensible defaults and starts the agent inside it; multiple sessions share one per-case container; export a container + its workspace to a portable .tar.gz to move it to another machine. See docs/docker-cases.md
  • Remote SSH sessions — point a case at another machine and run the agent there inside a durable remote tmux: survives SSH drops, auto-reconnects, and can discover + attach sessions already running on the host. See docs/remote-sessions.md
  • Effort & Ultracode — set a per-session default effort (low–max) or enable ultracode (dynamic multi-agent workflows). Soft defaults only — switchable anytime with /effort in-session. Extended-thinking budget is configurable too
  • Voice input — dictate prompts with Deepgram Nova-3 (Web Speech API fallback): toggle recording, auto-silence stop, live level meter (Ctrl+Shift+V)
  • Image input — paste or drag-and-drop images straight into a session
  • Gesture control (opt-in) — a MediaPipe hand-tracking overlay to grab/drag session windows and pinch buttons, hands-free. Enable with CODEMAN_GESTURE=1 + App Settings → Terminal & Input
  • Multi-monitor span (macOS) — one click opens a browser window maximized across all displays, so floating agent/gesture panels can cross the physical seam
  • File Viewer button (opt-in) — a header button that toggles the built-in file browser panel with one tap; enable under App Settings → Header & Panels → Header buttons
  • CJK / IME input — full composition support for Chinese / Japanese / Korean
  • OS notifications & hostname-aware titles — desktop alerts and tab titles are prefixed codeman:<host> so multi-host setups stay unambiguous

Isolated Docker Sessions

Run a case inside its own hardened Docker container instead of directly on your host — for security isolation, reproducible toolchains, and one-click portability.

  • One click — on New Case → Create New, tick 🐳 Run in an isolated Docker container. Codeman creates the case folder, spins up a container with default settings, and starts the agent inside it. No host/image/network fields to fill in.
  • Resource templates — expand the checkbox for a Small / Medium / Large / GPU preset (memory, CPUs, GPU), or set your own. Disk is elastic — storage grows as data flows in, no fixed cap.
  • Shared per-case container — many sessions can docker exec into the same container; killing one session never tears the container out from under the others.
  • Hardened by default — non-root, --cap-drop ALL, no-new-privileges, PID/memory caps, never --privileged or the docker socket; a sealed profile (no host credentials, network off) is one toggle away.
  • Seamless auth, isolated credentials — your host Claude / Codex / Antigravity / Gemini / OpenCode / Pi logins work inside the container out of the box: credentials are seeded (copied) in at launch and onboarding/trust prompts are pre-answered, so no login wizard appears. The container keeps its own copies and never writes back to your host credential stores; only conversation transcripts are shared, and exports never capture secrets.
  • Seamless auth, isolated credentials — your host Claude / Codex / Antigravity / Gemini / OpenCode / OMP logins work inside the container out of the box: credentials are seeded (copied) in at launch and onboarding/trust prompts are pre-answered, so no login wizard appears. The container keeps its own copies and never writes back to your host credential stores; only conversation transcripts are shared, and exports never capture secrets.- Move it to another machine — export a container's whole environment (toolchain + workspace) to a portable .tar.gz, docker load it on the other side, and import it into a fresh case.
  • Durable — reconnect after a restart lands back in the same live agent; a container stop/reboot resumes the conversation from the bind-mounted transcript.

Prerequisite: just Docker (or Podman). The agent base image builds itself automatically on first use, with progress streamed to the UI (or pre-build it with node scripts/build-agent-image.mjs). Full guide: docs/docker-cases.md.


Remote SSH Sessions

Point a case at another machine and run the agent there, over SSH, with the same dashboard, mobile UI, and autonomy features. Your laptop is just a window onto a session that lives on the remote host.

  • Durable by design: the agent runs inside a dedicated tmux session on the remote host, so a dropped SSH connection, network change, or laptop sleep never kills the run. Reconnecting lands back in the same live conversation.
  • Auto-reconnect: a bounded-backoff watcher notices a dead SSH pane and silently reattaches to the still-running remote session (kill-switch in settings; intentional kills are never revived).
  • Discover & attach: list the codeman-* sessions already running on a host (started by that machine's own Codeman, or by another operator) and attach to one. Attached sessions you don't own detach on tab close, never kill.
  • Shared sessions: several clients can attach the same remote session at different window sizes without clamping each other; discovery shows a "shared" badge with the client count.
  • Injection-safe: every ssh command line flows through a single shell-escaping builder, and host/path/identity fields are schema-guarded.

Set it up under New Case → Remote (host, user, identity file, optional jump host). Full design: docs/remote-sessions.md.


Multi-User Mode (opt-in)

Share one Codeman with a small trusted team, each person getting their own login and workspace. Off by default — without the flag, nothing changes.

Enable with codeman web --multiuser (or CODEMAN_MULTIUSER=1). Create the first admin, then manage users from the CLI or the Users tab in App Settings:

codeman users add alice --admin      # prompts for a password (or --password-stdin)
codeman users add bob                # a regular user
codeman users list
  • Per-user spaces — each user's cases live under ~/codeman-users/<name>/cases; sessions, cases, search, and real-time events are scoped to their owner. Admins see everything.
  • Individually revocable logins — named users with scrypt-hashed passwords in ~/.codeman/users.json; disable, reset (one-time password), or delete an account at any time. Admin actions are audited to ~/.codeman/admin-audit.jsonl.
  • Safer defaults for regular users — non-admins run Claude in --permission-mode auto (Anthropic's classifier-guarded mode); raw shell sessions, cron launchCommand, and skip-permissions require an explicit per-user grant.

⚠️ This separates workspaces; it does not sandbox users from each other. Every session runs as the same OS account, so a determined user's agent can still reach another user's files. For real isolation, pair users with Docker cases or run separate instances under separate OS accounts. See docs/multi-user-plan.md and the multi-user section of docs/security-architecture.md.


Remote Access — Cloudflare Tunnel

Access Codeman from your phone or any device outside your local network using a free Cloudflare quick tunnel — no port forwarding, no DNS, no static IP required.

Browser (phone/tablet) → Cloudflare Edge (HTTPS) → cloudflared → localhost:3000

Prerequisites: Install cloudflared and set CODEMAN_PASSWORD in your environment.

# Quick start
./scripts/tunnel.sh start      # Start tunnel, prints public URL
./scripts/tunnel.sh url        # Show current URL
./scripts/tunnel.sh stop       # Stop tunnel
./scripts/tunnel.sh status     # Service status + URL

The script auto-installs a systemd user service on first run. The tunnel URL is a randomly generated *.trycloudflare.com address that changes each time the tunnel restarts.

Persistent tunnel (survives reboots)
# Enable as a persistent service
systemctl --user enable codeman-tunnel
loginctl enable-linger $USER

# Or via the Codeman web UI: App Settings → System → Remote access → Cloudflare Tunnel
Authentication
  1. First request → browser shows Basic Auth prompt (username: admin or CODEMAN_USERNAME)
  2. On success → server issues a codeman_session cookie (24h TTL, auto-extends on activity)
  3. Subsequent requests authenticate silently via cookie
  4. 10 failed attempts per IP → 429 rate limit (15-minute decay)

Always set CODEMAN_PASSWORD before exposing via tunnel — without it, anyone with the URL has full access to your sessions.

QR Code Authentication

Typing a password on a phone keyboard is terrible. Codeman solves this with ephemeral single-use QR tokens — scan the code on your desktop, and your phone is instantly authenticated. No password prompt, no typing, no clipboard.

Desktop displays QR  →  Phone scans  →  GET /q/Xk9mQ3  →  Server validates
→  Token atomically consumed (single-use)  →  Session cookie issued  →  302 to /
→  Desktop notified: "Device authenticated via QR"  →  New QR auto-generated

Someone who only has the bare tunnel URL (without the QR) still hits the standard password prompt. The QR is the fast path; the password is the fallback.

How It Works

The server maintains a rotating pool of short-lived, single-use tokens. Each token consists of a 256-bit secret (crypto.randomBytes(32)) paired with a 6-character base62 short code used as an opaque lookup key in the URL path. The QR code encodes a URL like https://abc-xyz.trycloudflare.com/q/Xk9mQ3 — the short code is a pointer, not the secret itself, so it never leaks through browser history, Referer headers, or Cloudflare edge logs.

Every 60 seconds, the server automatically rotates to a fresh token. The previous token remains valid for a 90-second grace period to handle the race where you scan right as rotation happens — after that, it's dead. Each token is single-use: the moment a phone successfully scans it, the token is atomically consumed and a new one is immediately generated for the desktop display.

Security Design

The design is informed by "Demystifying the (In)Security of QR Code-based Login" (USENIX Security 2025), which found 47 of the top-100 websites vulnerable to QR auth attacks due to 6 critical design flaws across 42 CVEs. Codeman addresses all six:

USENIX FlawMitigation
Flaw-1: Missing single-use enforcementToken atomically consumed on first scan — replays always fail
Flaw-2: Long-lived tokens60s TTL with 90s grace, auto-rotation via timer
Flaw-3: Predictable token generationcrypto.randomBytes(32) — 256-bit entropy. Short codes use rejection sampling to eliminate modulo bias
Flaw-4: Client-side token generationServer-side only — tokens never leave the server until embedded in the QR
Flaw-5: Missing status notificationDesktop toast: "Device [IP] authenticated via QR (Safari). Not you? [Revoke]" — real-time QRLjacking detection
Flaw-6: Inadequate session bindingIP + User-Agent stored for audit. Manual session revocation via API. HttpOnly + Secure + SameSite=lax cookies

Timing-Safe Lookup

Short codes are stored in a Map<shortCode, TokenRecord>. Validation uses Map.get() — a hash-based O(1) lookup that reveals nothing about the target string through response timing. There is no character-by-character string comparison anywhere in the hot path, eliminating timing side-channel attacks entirely.

Rate Limiting (Dual Layer)

QR auth has its own rate limiting, completely independent from password auth:

  • Per-IP: 10 failed QR attempts per IP trigger a 429 block (15-minute decay window) — separate counter from Basic Auth failures, so a fat-fingered password doesn't burn your QR budget
  • Global: 30 QR attempts per minute across all IPs combined — defends against distributed brute force. With 62^6 = 56.8 billion possible short codes and only ~2 valid at any time, brute force is computationally infeasible regardless

QR Code Size Optimization

The URL is kept deliberately short (/q/ path + 6-char code = ~53-56 total characters) to target QR Version 4 (33x33 modules) instead of Version 5 (37x37). Smaller QR codes scan faster on budget phones — modern devices read Version 4 in 100-300ms. The /q/ prefix saves 7 bytes compared to /qr-auth/, which alone is the difference between QR versions.

Desktop Experience

The QR display auto-refreshes every 60 seconds via SSE with the SVG embedded directly in the event payload (~2-5KB) — no extra HTTP fetch, sub-50ms refresh. A countdown timer shows time remaining. A "Regenerate" button instantly invalidates all existing tokens and creates a fresh one (useful if you suspect the QR was photographed).

When someone authenticates via QR, the desktop shows a notification toast with the device's IP and browser — if it wasn't you, one click revokes all sessions.

Threat Coverage

ThreatWhy it doesn't work
QR screenshot sharedSingle-use: consumed on first scan. 60s TTL: expired before the attacker can act. Desktop notification alerts you immediately.
Replay attackAtomic single-use consumption + 60s TTL. Old URLs always return 401.
Cloudflare edge logsShort code is an opaque 6-char lookup key, not the real 256-bit token. Single-use means replaying from logs always fails.
Brute force56.8 billion combinations, ~2 valid at any time, dual-layer rate limiting blocks well before statistical feasibility.
QRLjacking60s rotation forces real-time relay. Desktop toast provides instant detection. Self-hosted single-user context makes phishing implausible.
Timing attackHash-based Map lookup — no string comparison timing leak.
Session cookie theftHttpOnly + Secure + SameSite=lax + 24h TTL. Manual revocation at POST /api/auth/revoke.

How It Compares

PlatformModelComparison
DiscordLong-lived token, no confirmation, repeatedly exploitedCodeman: single-use + TTL + notification
WhatsApp WebPhone confirms "Link device?", ~60s rotationComparable rotation; WhatsApp adds explicit confirmation (acceptable tradeoff for single-user)
SignalEphemeral public key, E2E encrypted channelStronger crypto, but exploited by Russian state actors in 2025 via social engineering despite it

Full design rationale, security analysis, and implementation details: docs/qr-auth-plan.md


Security

By default Codeman launches sessions with --dangerously-skip-permissions, so the web UI is by design a remote-code-execution surface for whoever can reach it — the whole security model exists to control who that is. (The startup permission mode is configurable; see below.) Recent hardening (v0.9.0 + v0.9.5) closes the browser-driven attack paths that bite self-hosted dev tools. Full model: docs/security-architecture.md. Found a vulnerability? See SECURITY.md for private disclosure and the list of known limitations.

Network & access

  • Loopback by default — the server binary binds 127.0.0.1, reachable only from the same machine, so the no-password default is safe out of the box (the guided installer asks about network access and configures the binding + password for you). Binding a non-loopback host without CODEMAN_PASSWORD starts but prints a loud warning with three concrete fixes (set a password, loopback + an authenticated tunnel, or explicitly acknowledge with --allow-unauthenticated-network)
  • Optional auth, real sessions — HTTP Basic via CODEMAN_USERNAME (default admin) / CODEMAN_PASSWORD. Success issues an opaque 256-bit codeman_session cookie (randomBytes(32)) — validated server-side, not client-signed, so it can't be forged offline (24h TTL, auto-extend, device-context audit log)
  • Per-IP rate limiting — 10 failed attempts → 429 with Retry-After (15-min decay). A valid cookie or correct password recovers immediately even while an attacker hammers the same IP — important because all tunnel traffic shares one loopback IP. QR auth has its own separate limiter
  • Configurable permission mode - --dangerously-skip-permissions is only the default. App Settings → Agents & CLIs → Claude → Startup Mode can switch new sessions to Anthropic's classifier-guarded auto mode (low-prompt, needs Claude Code 2.1.207+), normal prompting, or an explicit allowed-tools list. In multi-user mode, non-granted users are forced to auto, and shell sessions / skip-permissions require an explicit per-user grant

Always-on browser hardening (v0.9.5)

These run for every request — before auth, even on the default no-password loopback install:

  • Host-header allowlist → blocks DNS rebinding. A custom domain rebound to 127.0.0.1 is rejected with 403 host not allowed before any handler runs. Allowed: localhost, any IP literal, the bind host, .ts.net / .trycloudflare.com / .cfargotunnel.com, the active managed tunnel, and CODEMAN_ALLOWED_HOSTS (add custom reverse-proxy domains here — comma-separated; exact host or leading-dot .suffix for subdomains)
  • Cross-site Origin / CSRF guard. On state-changing methods (POST/PUT/PATCH/DELETE) the Origin must pass the same allowlist, else 403 cross-site request blocked. A missing Origin is allowed (so curl, the CLI, and Claude Code hooks keep working); only a present-but-foreign or opaque null origin is rejected
  • Raw text/plain bodies. The global parser no longer JSON-parses text/plain, closing the CORS "simple request" CSRF vector where a cross-site fetch could smuggle JSON into a write route with no preflight
  • WebSocket origin validation. The terminal WS upgrade runs the same Host + Origin check and closes with code 4003 on failure (anti-CSWSH)
  • XSS-escaped agent output. AI-derived strings (tool names, command arguments, subagent descriptions) are HTML-escaped at every injection site before rendering in the subagent / activity panels

Input, files & headers

  • Schema-validated inputs — every API body is checked with Zod v4 schemas; a CLAUDE_CODE_* / OPENCODE_* / CODEX_* / ANTIGRAVITY_* / GEMINI_* / GOOGLE_* / PI_* env-prefix allowlist gates which settings each CLI can receive
  • Path containment — file routes realpath before boundary checks (no TOCTOU); .., absolute paths, and symlinks resolving outside the working dir are rejected. Caps: 10 MB text preview / 50 MB raw & download; /api/download blocklists sensitive paths (.env, *credentials*, ~/.ssh/, .aws/credentials). SVG/HTML is served octet-stream + nosniff + attachment so it downloads rather than executes
  • Security headers — Content-Security-Policy (default-src 'self', every exception enumerated), X-Content-Type-Options: nosniff, X-Frame-Options: SAMEORIGIN, HSTS over HTTPS, and CORS reflected only for localhost / 127.0.0.1 / ::1

Supply chain & isolation

  • Pinned & verified deps — security-sensitive transitive deps are forced to patched versions via npm overrides; lockfile integrity is checked on every commit/PR (all entries resolve to registry.npmjs.org with sha512 hashes). Public assets are NUL-byte-scanned and node --check-validated in CI
  • Multi-instance isolation — CODEMAN_INSTANCE scopes both the tmux socket (-L codeman-<name>) and data dir (~/.codeman-<name>) so two instances never attach each other's live sessions

Mobile login uses single-use, 60-second QR tokens — see QR Code Authentication above for the full design (it addresses all 6 flaws from USENIX Security 2025's QR-login study).


Terminal UI (codeman tui)

A full-screen dashboard for your sessions, in the terminal. Same states as the web UI, because it is a client of the same server:

codeman tui              # the dashboard
codeman tui --list       # numbered session list, then exit (scriptable)
codeman tui 2            # attach straight to session 2 of that list

Sessions are grouped NEEDS YOU → WORKING → IDLE → RECENT, longest-waiting first. ↑↓/j/k select, 1-9 and [/] switch between sessions, Enter attaches into the tmux pane (F1 to come back). Inside a pane the bar across the top keeps the session strip visible and Alt+1-Alt+9 switch without leaving. y/n/digit answer a pending permission dialog right from the list, p sends a one-line prompt, n starts a session and opens straight into it, x kills one (y confirms), / searches, g shows the away digest, ? is help, q quits. Below 72 columns it drops the preview pane and becomes a single-column list, so it stays usable in Termius on a phone. With no server running it still starts in attach-only degraded mode.

The web UI remains the primary surface; see docs/tui.md for the full guide.


Keyboard Shortcuts

Ctrl bindings also accept Cmd on macOS.

ShortcutAction
Ctrl/Cmd+WKill active session
Ctrl/Cmd/Option+KFind open session or start a new one
Ctrl/Cmd+TabNext session
Alt/Option+[ / Alt/Option+]Previous / next session
Alt/Option+1-Alt/Option+9Switch to tab N (physical keys, so macOS Option layouts work)
Alt/Option+BCollapse / expand the session sidebar (sidebar layout only)
Ctrl+Shift+{ / Ctrl+Shift+}Move active tab left / right
Ctrl/Cmd+CCopy selection, or interrupt when nothing is selected
Ctrl+Shift+CCopy selection (never interrupts)
Ctrl/Cmd+VPaste, or upload a clipboard image and paste its path
Ctrl/Cmd+LClear terminal
Ctrl+Shift+RRestore terminal size
Ctrl+Shift+VToggle voice input
Ctrl/Cmd + / -Font size
Ctrl/Cmd+?Keyboard help
Shift+EnterInsert newline (sent to terminal)
EscapeClose panels & modals

Driving Codeman from an Agent — Programmatic Guide

For AI agents and automation that control Codeman without a browser: an agent that spins up worker sessions, a CI bot, or Claude Code running inside a Codeman session orchestrating other sessions. Everything the UI does is HTTP + a CLI, so an agent can do it too.

The agent skill (start here)

Everything in this section also ships as a Claude Code skill in skills/codeman. Install it once and you never paste API docs into a prompt again. You ask for what you want in plain English, and the agent already sitting inside a Codeman session loads the recipes and drives the API itself.

Step 1: install it

HowCommandScope
Skills CLInpx skills add Ark0N/Codeman --skill codeman -gGlobal, works for any skills-aware agent
Claude Code plugin/plugin marketplace add Ark0N/Codeman then /plugin install codeman@codemanGlobal, through Claude Code's plugin manager; /plugin update codeman follows releases. Pick this OR a codeman skill install, not both: a Claude Code with both lists the skill twice (codeman and codeman:codeman)
Bundled CLIcodeman skill installGlobal (~/.claude/skills/codeman), for npm installs that never cloned the repo
Bundled CLIcodeman skill install --case <name>One case only
Web UIApp Settings → Agents & CLIs → Claude → Agent SkillAuto-injects into each case on Claude session create (agentSkillEnabled, SYNCED, default off)

codeman skill uninstall [--case <name>] reverses the CLI installs, and never touches a skills/codeman you wrote yourself.

Step 2: ask for things

That is the entire interface. No curl, no endpoint names, no session ids. These prompts work as written:

You sayThe skill does
"What sessions are running right now?"Lists them with name, mode and status. Read-only, safe to ask anytime.
"Start a shell worker on the myapp case, run the test suite, tell me if it passes."Spawns, waits on a split completion marker, reads back the exit code, cleans up.
"Spin up 3 workers for lint, typecheck and tests. Run them in parallel, report failures."The fan-out flow: one session per task, all started first, then gathered as each finishes.
"Have a claude worker on refactor-auth summarize src/session.ts, then close it."Spawns, runs the readiness ladder (first-run trust dialog included), send-and-wait, reads the clean transcript answer, deletes.
"Watch session w4 and tell me if it gets stuck on a permission prompt."Blocks on the blocked signal and surfaces the question to you. It never answers another session's prompt itself.

Step 3: nothing

The agent deletes every session it started. Watch the tabs appear and disappear in the dashboard while it works.

A real run, start to finish

You: spin up 3 shell workers, run lint / typecheck / the frontend syntax check in parallel, and tell me which failed.

lint       -> 9f2d8e5f   dispatched
typecheck  -> aff9c691   dispatched     3 tabs appear in the dashboard
syntax     -> be9f1f15   dispatched

lint         DONE_lint_17909      rc=0
typecheck    DONE_typecheck_3409  rc=0   gathered as each one finishes
syntax       DONE_syntax_18501    rc=0

deleted 9f2d8e5f, aff9c691, be9f1f15    tabs disappear

Those DONE_<task>_<random> strings are the skill's split marker trick, and they are why the fan-out is reliable on hook-less shell sessions: the typed line contains ${M}_17909, so only the command's real output ever contains DONE_17909. An unsplit marker would match the echo of your own keystrokes before the command had even run.

What's in the box

FileContents
SKILL.mdSafety rules, the ready-made fast path (spawn N workers, task them, collect), and the verb index. Always loaded.
reference/verbs.mdThe 14 verbs in detail: readiness, send-and-wait, markers, interrupts, cleanup. On demand.
reference/recipes.md6 worked multi-worker flows (fan-out, blocked-worker watch, messaging fan-out). On demand.
reference/endpoints.mdFull endpoint tables, error codes, per-mode signal table, capacity limits. On demand.
reference/messaging.mdTalking to claude workers directly via Claude Code cross-session messaging. On demand.

Every recipe in there was verified against a live server, and the comments record the failure modes that were measured rather than guessed.

Two things worth knowing

  • It self-gates. Outside a Codeman session (CODEMAN_MUX unset) the skill refuses to act and does not guess an API URL, so a global install costs an unrelated Claude Code session nothing.
  • It is deliberately conservative. Unprompted, it may only spawn sessions, prompt them, and delete ones it created in that same conversation, by exact id, through a fail-closed guard that refuses to delete the agent's own session. Deleting a case (which erases a real directory of your code), bulk kills, respawn/ralph/cron/orchestrator changes and settings writes all require you to ask, naming the target.

⚠️ Turning agentSkillEnabled back off does not remove already-injected copies (a create-time sweep would yank the skill out from under other live sessions sharing that .claude/ dir). Remove them per case with codeman skill uninstall --case <name>.


The rest of this section is the manual path: the same operations as raw HTTP, for a CI bot, a shell script, or any agent without skill support.

Detect that you're inside Codeman

When a CLI runs in a Codeman-managed session, these environment variables are set — read them instead of hardcoding anything:

VariableMeaning
CODEMAN_MUX=1You're in a managed tmux session. Never tmux kill-session / pkill claude / pkill tmux — you'll kill yourself or a sibling.
CODEMAN_API_URLBase URL of the API (e.g. https://127.0.0.1:3000). Use it for every call below.
CODEMAN_SESSION_IDYour own session id. Use it to avoid acting on yourself.
CODEMAN_HOOK_SECRET_FILEPath to the hook secret (required on /api/hook-event while a managed tunnel is up).

Rules of the road (read before you POST)

  1. Single-line input, ending in \r. Programmatic input is sent as literal text, and Enter fires only when the input contains a carriage return: {"input":"run tests\r"}. Without the \r the text sits on the session's prompt unsubmitted (and a combined wait runs its full timeout on a turn that never started). Embedded newlines are stripped rather than rejected, so "echo A\necho B\r" runs the joined command echo Aecho B: send one line per call.
  2. Make input idempotent. Include a stable clientId and a monotonic per-session seq on POST …/input. The server de-duplicates, so a retry after a dropped connection can't double-deliver a prompt.
  3. Auth. If CODEMAN_PASSWORD is set, send HTTP Basic auth (user admin or CODEMAN_USERNAME) or a codeman_session cookie. The default loopback install is passwordless. A missing Origin header is allowed, so plain curl works; cross-site browser origins are rejected (CSRF guard). ⚠️ A 401 replies with the bare string Unauthorized, not the JSON envelope, so piping it into jq throws a parse error instead of showing the failure: check the status before parsing.
  4. Response envelope. Most endpoints return { "success": true, "data": … } (errors: { "success": false, "error", "errorCode" }). A few legacy GETs return bare bodies — handle both (body.data ?? body).
  5. /api/v1/* is a stable alias of /api/*.
  6. Wait instead of polling, and don't treat a timeout as an error. The wait endpoints answer with HTTP 200 and wait.timedOut: true when nothing happened in time, so loop over short waits (60s is the default) rather than issuing one long call, because tunnels cut idle connections. wait.timeoutMs tells you the timeout the server actually applied after clamping (600s ceiling).
  7. Only claude sessions emit stop and blocked. Those two come from Claude Code hooks; shell and the external CLIs (opencode/codex/gemini/antigravity/pi) accept only idle, working and exit. Asking for stop explicitly on those is a 400; omitting until is always safe. ⚠️ On a shell session idle fires once, at startup, and never again, so send-and-wait there can only time out; synchronize hook-less sessions with a wait-output marker.
  8. Only claude sessions emit stop and blocked. Those two come from Claude Code hooks; shell and the external CLIs (opencode/codex/gemini/antigravity/omp) accept only idle, working and exit. Asking for stop explicitly on those is a 400; omitting until is always safe. ⚠️ On a shell session idle fires once, at startup, and never again, so send-and-wait there can only time out; synchronize hook-less sessions with a wait-output marker.8. Nothing reports "ready", so wait for it explicitly. A new session answers {"signal":"exit","immediate":true} (that means not started, not crashed) until its PID exists, and a claude worker in a fresh case then sits on the CLI's trust dialog. Prompt it there and the wait resolves on idle in ~2s looking exactly like a finished turn, while the text sits stuck in the dialog. Recipe 2b below is the sequence that avoids it.

Recipes

# CODEMAN_API_URL is auto-set inside every Codeman session, correct scheme included.
# The fallback below fits a stock install; on a --https install set the https:// URL
# yourself and add -k to each curl (self-signed cert).
API="${CODEMAN_API_URL:-http://127.0.0.1:3000}"
# (add  -u admin:"$CODEMAN_PASSWORD"  to each call if a password is set)

# 1. See what's running
curl -s "$API/api/sessions" | jq '.data // .'

# 2. Spin up a worker session (a "case" = named working dir)
curl -s -X POST "$API/api/quick-start" \
  -H 'Content-Type: application/json' \
  -d '{"caseName":"refactor-auth","mode":"claude","effort":"high"}' | jq

# 2b. Wait until that worker is actually READY (see rule 8): composer marker first,
#     first-run trust dialog only as the fallback. (Probing trust first and sending
#     a blind Enter misfires on re-runs: the dialog text stays in the buffer forever,
#     so the probe matches stale text and the Enter lands in a ready composer.)
#     Match single tokens: TUI text can reach the matcher without its spaces.
until [ "$(curl -s "$API/api/sessions/$SID" | jq '.data.pid')" != null ]; do sleep 1; done
R=$(curl -sG "$API/api/sessions/$SID/wait-output" --data-urlencode 'match=bypass' \
      --data-urlencode 'from=buffer' --data-urlencode 'timeout=5000')
if ! jq -e '.data.wait.matched' <<<"$R" >/dev/null; then
  T=$(curl -sG "$API/api/sessions/$SID/wait-output" --data-urlencode 'match=trust' \
        --data-urlencode 'from=buffer' --data-urlencode 'timeout=2000')
  jq -e '.data.wait.matched' <<<"$T" >/dev/null && \
    curl -s -X POST "$API/api/sessions/$SID/input" -H 'Content-Type: application/json' \
      -d '{"input":"\r","useMux":true}'        # accept the first-run trust dialog
  curl -sG "$API/api/sessions/$SID/wait-output" --data-urlencode 'match=bypass' \
    --data-urlencode 'from=buffer' --data-urlencode 'timeout=45000' >/dev/null
fi

# 3. Send a prompt into a session (exactly-once: clientId + seq)
curl -s -X POST "$API/api/sessions/$SID/input" \
  -H 'Content-Type: application/json' \
  -d '{"input":"Run the test suite and summarize failures\r","useMux":true,"clientId":"agent-1","seq":1}'

# 4. Send a prompt and BLOCK until that turn is done (registers the wait before
#    writing, so it can't answer with the previous turn's idle state)
curl -s -X POST "$API/api/sessions/$SID/input" \
  -H 'Content-Type: application/json' \
  -d '{"input":"Run the test suite and summarize failures\r","useMux":true,
       "clientId":"agent-1","seq":2,"wait":"stop,exit","waitTimeout":60000}' \
  | jq '.data.wait'      # -> {"signal":"stop","timedOut":false,"waitedMs":41230,...}
#    (`stop` is the definitive end-of-turn hook. Adding `idle` makes it resolve on a
#     spinner pause too, and on anything that redraws a ❯ prompt — like a dialog.)

# 4b. Timed out? That's a 200, not a failure. Loop over short waits.
curl -s "$API/api/sessions/$SID/wait?until=stop,exit&timeout=60000" | jq '.data.wait'

# 4c. Or wait for a marker in the output (works for shell sessions too).
#     ⚠️ Unique per call (tmux repaints replay old screen text), and SPLIT so the
#     typed line never contains it: your own keystrokes echo into the output
#     stream, so an unsplit marker matches before the command has run. from=buffer
#     catches a marker that printed before the wait landed.
N=$RANDOM
curl -s -X POST "$API/api/sessions/$SID/input" -H 'Content-Type: application/json' \
  -d "{\"input\":\"M=DONE; npm test; echo \${M}_$N rc=\$?\r\",\"useMux\":true}"
curl -sG "$API/api/sessions/$SID/wait-output" \
  --data-urlencode "match=DONE_$N" --data-urlencode 'from=buffer' \
  --data-urlencode 'timeout=60000' | jq '.data.wait'

# 5. Read the terminal back. ⚠️ Use terminal?tail=, NOT /output: the latter's
#    textOutput is empty for every tmux-backed (i.e. every interactive) session.
#    tail counts BYTES, and what comes back is terminal data, ANSI included.
curl -s "$API/api/sessions/$SID/terminal?tail=8000" | jq -r '.data.terminalBuffer'

# 6. Stream live events (session output, agent activity, status)
curl -sN "$API/api/events"          # Server-Sent Events

# 7. Schedule recurring work (cron-style job)
curl -s -X POST "$API/api/cron/jobs" \
  -H 'Content-Type: application/json' \
  -d '{"name":"nightly-deps","agentType":"claude","workingDir":"/home/me/proj",
       "promptMode":"inline_text","promptText":"Update dependencies and open a PR",
       "inputMode":"typed","scheduleType":"daily","dailyTime":"03:00",
       "enabled":true,"concurrencyPolicy":"warn_only"}' | jq

# 8. Inspect background sub-agents and their transcripts
curl -s "$API/api/subagents" | jq '.data // .'
curl -s "$API/api/subagents/$AID/transcript" | jq -r '.data // .'

# 9. Whole-system snapshot (sessions, settings, respawn, stats)
curl -s "$API/api/status" | jq

Or use the bundled CLI

The same operations are available as commands (codeman <cmd>, aliases in parentheses) — handy from a shell tool inside a session:

codeman session start -d /path/to/repo   # (s)  start a session
codeman session list                     #      list sessions
codeman session logs <id>                #      tail output
codeman task add "fix the failing test"  # (t)  queue a task
codeman attach <path>                    #      show an attachment card for a local file
codeman tui --list                       #      numbered session list (plain text when piped)
codeman tui 3                            #      attach to session 3 of that list

Hooks (events flowing back to Codeman)

Codeman registers Claude Code hooks that POST /api/hook-event (permission_prompt, idle_prompt, stop, task_completed, …) so the dashboard reacts in real time. This endpoint is auth-exempt on loopback but, under a managed tunnel, requires the X-Codeman-Hook-Secret header (read it from $CODEMAN_HOOK_SECRET_FILE). You normally don't call this by hand — Codeman wires it up — but it's how the autonomy layers "see" what the agent is doing.

Full endpoint list and request/response shapes follow.


API

REST over Fastify — ~200 handlers across 21 route modules, plus an SSE stream and a WebSocket terminal channel. All responses use the ApiResponse<T> envelope ({success, data} / {success, error, errorCode}); /api/v1/* is a stable alias. A representative subset:

Sessions

MethodEndpointDescription
GET/api/sessionsList all
POST/api/quick-startCreate case + start session ({caseName?, mode?, effort?, envOverrides?})
POST/api/sessions/:id/inputSend input ({input, useMux?, clientId?, seq?, wait?, waitTimeout?}: clientId+seq = exactly-once; wait blocks until the turn ends)
GET/api/sessions/:id/terminalRead terminal output (?tail=<bytes>, ?full=1); the read path for interactive sessions
GET/api/sessions/:id/outputParsed one-shot output (textOutput is empty for tmux-backed sessions)
GET/api/sessions/:id/waitBlock until a signal fires (?until=stop,idle,exit&timeout=&fresh=); a timeout is a 200
GET/api/sessions/:id/wait-outputBlock until a literal string appears (?match=&nocase=&from=now|buffer&timeout=)
GET/api/sessions/unifiedUnified live + history list (Session Manager) — ?q=&limit=
POST/api/sessions/:id/pinPin/unpin in the Session Manager ({pinned})
PUT/api/session-orderSync tab order across devices ({order: [ids]})
DELETE/api/sessions/:idDelete session

Respawn

MethodEndpointDescription
POST/api/sessions/:id/respawn/enableEnable with config + timer
POST/api/sessions/:id/respawn/stopStop controller
PUT/api/sessions/:id/respawn/configUpdate config

Orchestrator

MethodEndpointDescription
POST/api/orchestrator/startStart orchestration from a goal
POST/api/orchestrator/approveApprove the generated plan
GET/api/orchestrator/statusCurrent phase + progress
POST/api/orchestrator/stopStop and clean up

Cron (scheduled jobs)

MethodEndpointDescription
GET / POST/api/cron/jobsList / create cron jobs
PUT / DELETE/api/cron/jobs/:idUpdate / delete a job
PUT/api/cron/jobs/:id/enabledEnable / disable
POST/api/cron/jobs/:id/runRun now
GET/api/cron/jobs/:id/runsRun history

Subagents

MethodEndpointDescription
GET/api/subagentsList all background agents
GET/api/subagents/:idAgent info and status
GET/api/subagents/:id/transcriptFull activity transcript
DELETE/api/subagents/:idKill agent process

System

MethodEndpointDescription
GET/api/eventsSSE stream
GET/api/statusFull app state
POST/api/hook-eventHook callbacks
GET/api/system/update/checkCheck for a new release
POST/api/system/updateSelf-update (git-clone installs)
POST/api/clipboardPush text to all connected browsers ({text})
GET/api/sessions/:id/run-summaryTimeline + stats

Building something on top of Codeman? docs/extending-codeman.md is the integration guide: render your own UI as a tab, subscribe to the SSE event stream to react when an agent needs you, drive Codeman from a script, and the traps worth knowing before you start. Codeman has no plugin runtime on purpose, so an integration is just your own process talking HTTP.


Architecture

flowchart TB
    subgraph Codeman["CODEMAN"]
        subgraph Frontend["Frontend Layer"]
            UI["Web UI<br/><small>xterm.js + Agent Windows</small>"]
            API["REST API<br/><small>Fastify</small>"]
            SSE["SSE Events<br/><small>/api/events</small>"]
        end

        subgraph Core["Core Layer"]
            SM["Session Manager"]
            S1["Session (PTY)"]
            S2["Session (PTY)"]
            RC["Respawn Controller"]
            ORC["Orchestrator Loop"]
        end

        subgraph Detection["Detection Layer"]
            SW["Subagent Watcher<br/><small>~/.claude/projects/*/subagents</small>"]
            TW["Team Watcher<br/><small>~/.claude/teams/*</small>"]
        end

        subgraph Persistence["Persistence Layer"]
            SCR["Mux Manager<br/><small>(tmux)</small>"]
            SS["State Store<br/><small>state.json</small>"]
        end

        subgraph External["External"]
            CLI["AI CLI<br/><small>Claude Code / OpenCode / Codex / Antigravity / Gemini / Pi</small>"]
            CLI["AI CLI<br/><small>Claude Code / OpenCode / Codex / Antigravity / Gemini / OMP</small>"]            BG["Background Agents<br/><small>(Task tool)</small>"]
        end
    end

    UI <--> API
    API <--> SSE
    API --> SM
    SM --> S1
    SM --> S2
    SM --> RC
    SM --> ORC
    SM --> SS
    S1 --> SCR
    S2 --> SCR
    RC --> SCR
    ORC --> SCR
    SCR --> CLI
    SW --> BG
    SW --> SSE
    TW --> SSE

Development

npm install
npx tsx src/index.ts web    # Dev mode
npm run build               # Production build
npm test                    # Run tests (same suite CI runs; browser/mobile/perf suites have their own commands)

See CLAUDE.md for full documentation.


Community

Questions, setup help, and ideas live in GitHub Discussions: the Q&A section answers the most common ones (phone access, overnight runs, updating), and the roadmap gets decided in Ideas. Bugs go to issues; reports usually get a response within a day, and every release credits its reporters and contributors by name. Want to contribute? CONTRIBUTING.md has the map: skins, translations, and docs make great first PRs, and bigger features start life as a Discussion. And if you're proud of your rig, post it in Show and tell.


Codebase Quality

The codebase went through a comprehensive 7-phase refactoring that eliminated god objects, centralized configuration, and established modular architecture:

PhaseWhat changedImpact
PerformanceCached endpoints, SSE adaptive batching, buffer chunkingSub-16ms terminal latency
Route extractionserver.ts split into 15 domain route modules + auth middleware + port interfaces−67% server.ts LOC (6,736 → 2,254)
Domain splittingtypes.ts → 16 domain files, ralph-tracker → 7 files, respawn-controller → 5 files, session → 6 filesNo more god files
Frontend modulesapp.js → 18 extracted modules across infra, domain & feature layersapp.js core down to ~3.4K LOC
Config consolidation~70 scattered magic numbers → 10 domain-focused config filesZero cross-file duplicates
Test infrastructureShared mock library, 12 route test files, consolidated MockSessionTestable route handlers via app.inject()

Full details: docs/archive/code-structure-findings.md


Published Packages

xterm-zerolag-input

npm

Instant keystroke feedback overlay for xterm.js. Eliminates perceived input latency over high-RTT connections by rendering typed characters immediately as a pixel-perfect DOM overlay. Zero dependencies, 6.1 kB gzipped, configurable prompt detection, CJK/emoji wide-character support, full state machine with 175 tests.

npm install xterm-zerolag-input

Full documentation


Versioning

Codeman follows SemVer. What the version number actually commits to — and what counts as internal (the HTTP/SSE API, on-disk state, experimental features) — is spelled out in docs/versioning-policy.md. If you script against the HTTP API, pin to an exact version.

License

MIT — see LICENSE


Track sessions. Visualize agents. Control respawn. Let it run while you sleep.

If Codeman saves you time, a star helps other people find it.
Bug reports and feature ideas are welcome in Issues.

其他

高风险

  • 来源需自行核对维护者身份。
  • 包含脚本或命令调用,安装前请复核。
  • 可能需要外部 token、网络权限或第三方服务。
  • 存在潜在风险命令,请谨慎安装。
  • 扫描发现:5 条。

Codex — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/Ark0N/Codeman.git
  3. 将 "skills/codeman" 文件夹复制到 Codex 的 skills 目录中。
  4. 重启 Codex 让新的 skill 生效。

Codex — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Codex 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Codex 让新的 skill 生效。

Claude Code — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/Ark0N/Codeman.git
  3. 将 "skills/codeman" 文件夹复制到 Claude Code 的 skills 目录中。
  4. 重启 Claude Code 让新的 skill 生效。

Claude Code — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Claude Code 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Claude Code 让新的 skill 生效。

Cursor — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/Ark0N/Codeman.git
  3. 将 "skills/codeman" 文件夹复制到 Cursor 的 skills 目录中。
  4. 重启 Cursor 让新的 skill 生效。

Cursor — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Cursor 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Cursor 让新的 skill 生效。

GitHub Copilot — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/Ark0N/Codeman.git
  3. 将 "skills/codeman" 文件夹复制到 GitHub Copilot 的 skills 目录中。
  4. 重启 GitHub Copilot 让新的 skill 生效。

GitHub Copilot — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 GitHub Copilot 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 GitHub Copilot 让新的 skill 生效。

Windsurf — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/Ark0N/Codeman.git
  3. 将 "skills/codeman" 文件夹复制到 Windsurf 的 skills 目录中。
  4. 重启 Windsurf 让新的 skill 生效。

Windsurf — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Windsurf 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Windsurf 让新的 skill 生效。
查看 SKILL.md 原文
name: codeman
description: >-
  Drive Codeman, the session manager this agent is running inside, over its HTTP API:
  list sessions, start worker sessions, send them prompts, block until they finish
  (wait / wait-output / send-and-wait), read their output, and clean up; where
  available, message claude workers directly (Claude Code cross-session messaging).
  Use when asked to orchestrate or parallelize work across Codeman sessions, watch
  another session, or start and manage workers. Only usable inside a Codeman-managed
  session (CODEMAN_MUX=1); refuse to act otherwise.

Driving Codeman from inside a session

You are an agent running inside a Codeman-managed terminal session. Codeman is the server that spawned you; its HTTP API can start, prompt, watch, and delete other sessions.

Read as far as your job needs and no further. §0 is the bootstrap, run once. §1 is the whole fast path: spawn N workers, task them, collect answers. If §1 covers your job, run it and stop there. The sections after it are for jobs it does not cover, and reading them to be thorough is the main reason a ten-second run takes minutes. §2 is the verb table when your job is a different one. §3 and §4 are the rules; §6 is setup and credentials, which you only need when something 401s.

Everything else loads on demand, and is meant to be opened at one section, not read through: the verbs in detail (the old §5) in reference/verbs.md, worked multi-worker flows in reference/recipes.md, endpoint tables and a symptom gallery in reference/endpoints.md, and direct messaging to claude workers in reference/messaging.md.

0. Guard and bootstrap

If CODEMAN_MUX is not 1, stop and say so. Do not guess an API URL; a server you are not part of is not yours to drive.

⚠️ Your shell state does not survive between tool calls. Each Bash call starts a fresh shell, so $API, $SELF, the CURL array and delete_session are all gone by the next call, and $$ is a different pid. The filesystem does survive, so write the preamble to a file once and source it afterwards, rather than re-pasting a hundred-odd lines at the top of every call (a half-re-pasted preamble used to be the single most likely way to break a run).

Codeman seeds the preamble file for you when it spawns a claude session (server 1.18.3+), so the bootstrap is usually nothing at all: these are the two lines every later call opens with, and your first REAL call performs them anyway:

. "${XDG_CACHE_HOME:-$HOME/.cache}/codeman-agent-$CODEMAN_SESSION_ID.sh" 2>/dev/null
[ "${CODEMAN_PREAMBLE:-}" = 1.22.0 ] || { echo "preamble missing or stale; run the full §0 block"; exit 1; }

⚠️ Never spend a Bash call on this check alone. §1's block opens with this same loader, so when §1 is the job, start there: the check rides the spawn call for free, and a standalone "preamble OK" call buys nothing while costing a full model turn (measured live: a lone check plus the deliberation around it added ~6 s to a 28 s two-worker run). §0 is done the moment any job call passes its opening check. Only when a call reports missing or stale, run the full block below once — and run it verbatim: paste it as-is, never re-type it, trim it, or "extract the parts you need". A hand-assembled preamble is the documented failure mode of this skill: one live run rebuilt it "minimally" and lost the X-Codeman-Parent-Session header (every worker spawned with no lineage arc in the web UI) and the fast-path functions (the spawn fell back to a serial quick-start loop plus pid polls), turning a ten-second job into a fifty-second one. If your harness directs temporary files into a scratchpad directory, that directive covers task scratch, not this file: it is a per-session cache that every later call re-sources by this exact path, so keep the path below. If you must relocate it anyway, copy the block's content byte-for-byte unchanged and source your path in every later call instead.

test "${CODEMAN_MUX:-}" = 1 || { echo "Not inside a Codeman-managed session; refusing to act."; exit 1; }
: "${CODEMAN_SESSION_ID:?CODEMAN_SESSION_ID not set}" "${HOME:?HOME not set}"
PRE="${XDG_CACHE_HOME:-$HOME/.cache}/codeman-agent-$CODEMAN_SESSION_ID.sh"
mkdir -p "$(dirname "$PRE")"
# Rewrite unless the file already ends with THIS version's stamp, so a stale or a
# half-written file self-heals here instead of costing you a round trip to rm it.
grep -qs '^CODEMAN_PREAMBLE=1.22.0$' "$PRE" || (umask 077; cat > "$PRE" <<'PREAMBLE'
# ---- Codeman agent preamble 1.22.0 (seeded by Codeman at session spawn; the SKILL.md §0 bootstrap rewrites it when missing or stale) ----
API="${CODEMAN_API_URL:?CODEMAN_API_URL not set; refusing to guess}"
SELF="${CODEMAN_SESSION_ID:?CODEMAN_SESSION_ID not set}"
# Credentials, cheapest first. Your session has usually INHERITED the server's
# CODEMAN_PASSWORD already (§6 explains why, and what to do when it has not);
# the data dir's .env is the documented fallback, the same one `codeman attach`
# reads. The data dir is wherever the hook-secret file lives. Values may be
# quoted or `export`-prefixed.
ENV_FILE="${CODEMAN_HOOK_SECRET_FILE:+${CODEMAN_HOOK_SECRET_FILE%hook-secret}.env}"
envval() { sed -n "s/^\(export \)\{0,1\}$1=//p" "$ENV_FILE" | tail -1 | sed 's/^"\(.*\)"$/\1/; s/^'\''\(.*\)'\''$/\1/'; }
if [ -z "${CODEMAN_PASSWORD:-}" ] && [ -n "$ENV_FILE" ] && [ -f "$ENV_FILE" ]; then
  CODEMAN_USERNAME=$(envval CODEMAN_USERNAME)
  CODEMAN_PASSWORD=$(envval CODEMAN_PASSWORD)
fi
AUTH=(); [ -n "${CODEMAN_PASSWORD:-}" ] && AUTH=(-u "${CODEMAN_USERNAME:-admin}:$CODEMAN_PASSWORD")
# -k: harmless on http, required on https (self-signed cert).
# X-Codeman-Parent-Session: tags workers YOU spawn as your children, so the web UI can
# draw the lineage. Set once here and every present and future create call carries it;
# it is ignored on every other endpoint. Purely cosmetic (see §5.1) and it can never
# fail a spawn, so there is no case where you would want to leave it off.
# X-Codeman-Agent-Origin: marks a case directory a spawn CREATES as agent scratch, so the
# user can find and delete it long after your workers are gone (§5.14). Same deal: set
# once, cosmetic, never fails a spawn, and it labels only directories Codeman creates.
CURL=(curl -sk "${AUTH[@]}" -H "X-Codeman-Parent-Session: $SELF" -H "X-Codeman-Agent-Origin: codeman-skill")
CID=codeman-agent-1            # FIXED literal, never "agent-$$": see below

# Fail-CLOSED session delete. The DELETE lives INSIDE the guard on purpose: the older
# `is_self "$SID" || curl -X DELETE ...` shape failed OPEN, because an undefined
# is_self exits 127 and the `||` branch then ran the delete completely unguarded.
# Undefined delete_session is "command not found", which deletes nothing.
delete_session() {
  local id="${1:-}"
  [ -n "$id" ] || { echo "refusing: empty session id"; return 1; }
  [ "${#SELF}" -ge 8 ] || { echo "refusing: \$SELF unset or too short to prove this is not me"; return 1; }
  # ids appear in full AND 8-char form (Docker exports a truncated $SELF; mux names and
  # UI surfaces carry 8-char ids), so compare by prefix in BOTH directions. Equality or
  # a one-directional check each miss a real combination, and the miss deletes you.
  case "$id" in "$SELF"*) echo "refusing: $id is me"; return 1 ;; esac
  case "$SELF" in "$id"*) echo "refusing: $id is me"; return 1 ;; esac
  "${CURL[@]}" -X DELETE "$API/api/v1/sessions/$id"
}

# ---- fast path: the four verbs, already written. §1 composes them. ----
_composer_up() {   # <sid> <timeoutMs> -> "true"/"false". `shift+tab` is the one token
  "${CURL[@]}" -G "$API/api/v1/sessions/$1/wait-output" \
    --data-urlencode 'match=shift+tab' --data-urlencode 'from=buffer' \
    --data-urlencode "timeout=$2" | jq -r '.data.wait.matched // false'
}
_dsh_up() {        # <sid> <timeoutMs> -> "true"/"false". The DeepSeek Harness TUI's
  # composer glyph. Override with DSH_READY_MARK for a profile that draws another one.
  "${CURL[@]}" -G "$API/api/v1/sessions/$1/wait-output" \
    --data-urlencode "match=${DSH_READY_MARK:-❯}" --data-urlencode 'from=buffer' \
    --data-urlencode "timeout=$2" | jq -r '.data.wait.matched // false'
}
# ---- the workspace-trust dialog: READ the screen, never press Enter blind ----
# Claude Code 2.1.252 dropped the option numbers, REVERSED them, and highlights
# "No, exit" by default:
#     Security guide
#   ❯ No, exit
#     Yes, I trust this folder
#   Enter to confirm . Esc to cancel
# so the bare \r that answered the old layout now answers *exit* and the pane is
# dead (`status 1`) seconds after the spawn -- measured on a live 2.1.252 case.
# These two read the rendered pane and steer onto the trust option instead.
_trust_key() {     # <sid> -> "confirm" | "move" | "" (nothing safe to press)
  # full=1 returns the RENDERED pane; a claude pane keeps no tmux history, so that
  # is the current frame rather than every repaint since launch. tail -1 anyway,
  # because the freshest marked row is the only one still true.
  "${CURL[@]}" -G "$API/api/v1/sessions/$1/terminal" --data-urlencode 'full=1' \
    | jq -r '.data.terminalBuffer // empty' \
    | sed -e "s/$(printf '\033')\[[0-9;?]*[a-zA-Z]//g" -e "s/$(printf '\033')[()][AB0]//g" \
    | tr -d ' \t' | grep -i '❯[0-9.]*\(yes,itrustthisfolder\|no,exit\)' | tail -1 \
    | sed -e 's/.*[Yy]es,.*/confirm/' -e 's/.*[Nn]o,.*/move/'
}
_accept_trust() {  # <sid> -> 0 once it has answered the dialog, 1 if it could not
  local sid="$1" k i=1
  while [ "$i" -le 6 ]; do
    k=$(_trust_key "$sid")
    [ -n "$k" ] || return 1   # no dialog on screen, or a layout this cannot read
    # A SEPARATE clientId for these keys. seq is monotonic per clientId, so
    # spending prompt numbers here would make the next sendwait -- whose default
    # seq is the epoch second -- look like a stale duplicate and vanish silently.
    "${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" -H 'Content-Type: application/json' \
      -d "$(jq -nc --arg k "$([ "$k" = confirm ] && printf '\r' || printf '\033[B')" \
              --arg c "$CID-trust-$sid" --argjson s "$i" \
              '{input:$k,useMux:true,clientId:$c,seq:$s}')" >/dev/null
    [ "$k" = confirm ] && return 0
    sleep 1; i=$((i+1))   # re-read: the arrow is CONFIRMED before Enter goes out
  done
  return 1
}
# spawn_worker <caseName> [mode] -> session id on stdout, diagnostics on stderr.
# quick-start AND readiness in one call, with a strict contract: NON-EMPTY stdout means
# a READY worker whose end-of-turn signal can be trusted -- a claude worker in a
# hook-carrying case, or a `deepseek` worker whose harness TUI drew its composer.
# Anything less is rc 1 with EMPTY stdout, and the half-spawned session is deleted here
# rather than handed back, because a worker that never drew its composer would eat the
# task prompt with its trust dialog. There is deliberately no pid poll: wait-output
# already blocks until the composer draws, and pid!=null proved startup, never readiness.
spawn_worker() {
  local name="${1:?spawn_worker needs a case name}" mode="${2:-claude}" q sid cp r
  # parentSessionId doubles the CURL header, so a spawn_worker copied off the shared
  # curl (or a body someone rebuilt from this recipe) still carries its lineage.
  # deepseek: ask for the same permission posture the Run button sends, because the
  # harness's own default (`workspace-write`) still ASKS, and a worker that stops on
  # an approval row is a worker no fan-out can finish. It is not an escalation --
  # claude workers already spawn with permissions skipped, and in multi-user mode the
  # server clamps this back to `workspace-write` for an owner without the grant.
  # Spawn by hand (§5.1) when you want a worker that asks.
  q=$("${CURL[@]}" -X POST "$API/api/v1/quick-start" -H 'Content-Type: application/json' \
      -d "$(jq -nc --arg n "$name" --arg m "$mode" --arg p "$SELF" \
        '{caseName:$n,mode:$m,parentSessionId:$p}
         + (if $m == "deepseek" then {deepSeekConfig:{permissionMode:"danger-full-access"}} else {} end)')")
  sid=$(jq -r 'if .success then .data.sessionId else empty end' <<<"$q")
  # NOT retryable in a loop: every quick-start failure code is terminal (§5.1).
  [ -n "$sid" ] || { jq -c '{error,errorCode}' <<<"$q" >&2; return 1; }
  if [ "$mode" = deepseek ]; then
    # The one non-claude mode with REAL end-of-turn signals: its TUI reports
    # idle/working/blocked to Codeman, so sendwait, until=stop and the Approvals
    # Inbox all work here exactly as they do for claude. No hook file to vet
    # (the bridge is env-injected, not a workspace file) and no trust dialog.
    # ⚠️ Readiness is still not optional, and NOT interchangeable with the stop
    # signal: the harness's boot report lands ~300ms BEFORE the composer paints
    # (measured 2.26s vs 2.56s after spawn), so a sendwait fired straight after
    # quick-start returns on that BOOT signal, reports a turn that never ran, and
    # strands the prompt in a pane that was not yet taking input.
    r=$(_dsh_up "$sid" 45000)
    [ "$r" = true ] || { echo "dsh worker $sid never drew a composer: no pane-capable profile, a profile whose composer is not '${DSH_READY_MARK:-❯}' (set DSH_READY_MARK), or a harness that failed to boot -- check GET /api/v1/deepseek/status. Deleted it" >&2
      delete_session "$sid" >/dev/null; return 1; }
    printf '%s\n' "$sid"; return 0
  fi
  [ "$mode" = claude ] || { printf '%s\n' "$sid"; return 0; }   # no other mode draws a composer to wait on
  # The server installs hooks into every claude workspace now, so this grep normally
  # passes; it stays because the install is gated on a setting the operator can turn
  # off, remote sessions never get hooks, and a session created by an older server
  # still has none. No marker means sendwait would false-resolve on flapping idle,
  # possibly inside the user's REAL repo: refuse rather than run the job there.
  cp=$(jq -r '.data.casePath // empty' <<<"$q")
  grep -qs '/api/hook-event' "$cp/.claude/settings.local.json" || {
    echo "case '$name' resolved to '$cp', which has no Codeman hooks (workspaceHooksEnabled off, remote, or an older server?): turn the setting on, or work §5.1+§5.5 by hand with markers" >&2
    delete_session "$sid" >/dev/null; return 1; }
  # Short composer wait FIRST, then the trust dialog: a case still showing the
  # dialog can never pass the composer wait, so acting early keeps a cold case from
  # paying the whole long wait before the fallback even runs (§5.2). A warm case
  # matches in under a second and never reaches it, and _accept_trust returns in a
  # blink when there is no dialog, so this costs nothing in the ordinary slow case.
  r=$(_composer_up "$sid" 5000)
  if [ "$r" != true ]; then
    # Codeman answers this dialog itself and normally wins the race; this is the
    # bounded fallback for when its 90 s window / 6-keystroke cap has run out.
    _accept_trust "$sid"
    r=$(_composer_up "$sid" 45000)
  fi
  [ "$r" = true ] || { echo "worker $sid never drew a composer; deleted it. Retry by hand via the §5.2 ladder (its billed stage-4 probe included)" >&2
    delete_session "$sid" >/dev/null; return 1; }
  printf '%s\n' "$sid"
}
# spawn_workers <caseName[:mode]>... -> one "<caseName> <sessionId>" line per worker, in
# order; the sessionId column is EMPTY for a spawn that failed (stderr has why).
# CONCURRENT: N workers cost about what one costs. Spawning them one Bash call at a time
# is the single biggest avoidable delay in this skill. A bare name is a claude worker;
# `beta:deepseek` makes that one a DeepSeek Harness worker, and a mixed fleet is one
# call. Case names must be UNIQUE: two workers in one case directory co-edit the same
# tree (§4), so a repeat is an error here, not a race (the mode never disambiguates two
# workers, since they would still share the directory).
spawn_workers() {
  local d spec n m i=0
  [ "$#" -gt 0 ] || { echo "spawn_workers: no case names given" >&2; return 1; }
  [ -z "$(printf '%s\n' "$@" | sed 's/:.*//' | sort | uniq -d)" ] || { echo "spawn_workers: duplicate case names" >&2; return 1; }
  d=$(mktemp -d "${TMPDIR:-/tmp}/codeman-spawn.XXXXXX") || return 1
  for spec in "$@"; do
    n=${spec%%:*}; m=${spec#*:}; [ "$m" = "$spec" ] && m=claude
    ( spawn_worker "$n" "$m" > "$d/$i" ) & i=$((i+1))
  done
  wait
  i=0; for spec in "$@"; do printf '%s %s\n' "${spec%%:*}" "$(cat "$d/$i" 2>/dev/null)"; i=$((i+1)); done
  rm -rf "$d"
}
# sendwait <sid> <prompt> [seq] -> blocks until that worker's turn ENDS (~10 min ceiling
# across its two waits). One billed turn. The \r and the per-worker clientId are applied
# here, which is why you never hand-build this body. seq defaults to the CURRENT EPOCH
# SECOND so that every new prompt is a new frame: the server drops any (clientId,seq)
# pair it has already applied, so a fixed default would make every later prompt to that
# worker a silent no-op that still "succeeds" and reports the previous turn's state.
# Pass seq explicitly for exactly one reason: resending a possibly-delivered frame as a
# deliberate duplicate, at the SAME number (§5.3).
# Delivery is SELF-HEALING: an Ink repaint occasionally eats the Enter, leaving the
# typed prompt stranded on the composer while a long wait runs its whole timeout
# (observed live). So the first wait is short; on its timeout a bare \r goes out (the
# missing Enter when the prompt is stranded, a no-op when the turn is genuinely
# running), then the ORIGINAL frame is resent unchanged, which the server takes as a
# tagged duplicate: it re-waits without retyping (§5.3). Trustworthy for a worker
# spawn_worker handed back -- claude (hooks vetted) or deepseek (status bridge) --
# and for those only. Hook-less workspaces and the other modes resolve on flapping
# idle: markers instead (§5.5). ⚠️ A dsh worker running a profile that does not
# implement the status contract is the one case that LOOKS like claude but is not:
# it accepts the send and then burns both waits. One timeout on a dsh worker whose
# pane clearly finished means that profile, so switch that worker to markers.
sendwait() {
  local sid="${1:?}" p="${2:?}" seq="${3:-$(date +%s)}" body r
  # `wait:"stop,exit"`, never the `wait:true` default set: that set also carries
  # `idle`, which is INFERRED from output stabilization and flaps mid-turn. On a
  # dsh worker whose TUI repaints rarely the session reads `idle` while the model
  # is still answering, and the re-wait below then resolved in 0 ms with
  # `signal:"idle"` on a turn that had another three minutes to run (measured).
  # A wait named after the end of a turn should only end with the turn, or with
  # the worker. ⚠️ This is also what makes a wrong mode LOUD: the modes that
  # cannot deliver `stop` answer 400 (before writing anything) instead of
  # resolving on a flap, which is the answer that sends you to markers (§5.5).
  body=$(jq -nc --arg p "$p" --arg c "$CID-$sid" --argjson s "$seq" \
    '{input:($p+"\r"),useMux:true,clientId:$c,seq:$s,wait:"stop,exit",waitTimeout:20000}')
  r=$("${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" \
        -H 'Content-Type: application/json' --data-binary "$body")
  if jq -e '.data.delivered and .data.wait.timedOut' <<<"$r" >/dev/null 2>&1; then
    "${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" -H 'Content-Type: application/json' \
      -d "$(jq -nc --arg c "$CID-$sid" --argjson s "$(date +%s)" \
        '{input:"\r",useMux:true,clientId:$c,seq:$s}')" >/dev/null
    # The resend is a tagged DUPLICATE, so the server skips the write and reports
    # `delivered:false` for it -- truthfully, but about the wrong send. The first
    # one delivered, so carry that forward, or §1's cleanup reads a completed turn
    # as an undelivered one and keeps a finished worker forever.
    r=$("${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" \
          -H 'Content-Type: application/json' --data-binary "$(jq -c '.waitTimeout=580000' <<<"$body")" \
        | jq -c 'if .success and (.data.wait.ended | not) then .data.delivered = true else . end')
  fi
  printf '%s\n' "$r"
}
# last_text <sid> [prev] -> that worker's last assistant message (claude, codex and
# deepseek write a real transcript; the other modes have none, so read the terminal
# instead -- §5.4). Polled, because the transcript write LAGS the stop signal, and
# "some text exists" is not "THIS turn's text exists": right after a SECOND turn on the same worker the endpoint still serves
# the previous answer for a beat (observed live). When reading consecutive turns, pass
# the previous answer as [prev]: the poll then holds out for text that differs from it,
# falling back to whatever it last saw if the budget runs dry, so an honestly repeated
# answer still comes back. Non-zero exit means the worker really never wrote one.
last_text() {
  local t="" prev="${2:-}"
  for _ in $(seq 1 15); do
    t=$("${CURL[@]}" "$API/api/v1/sessions/$1/last-response" | jq -r '.data.text // empty')
    [ -n "$t" ] && [ "$t" != "$prev" ] && { printf '%s\n' "$t"; return 0; }
    sleep 1
  done
  [ -n "$t" ] && { printf '%s\n' "$t"; return 0; }
  return 1
}

# The stamp is the LAST line on purpose (a truncated write leaves it unset) and is kept
# bare on purpose: the write condition above anchors on it with $, so an inline comment
# here would fail that match and rewrite this file on every single bootstrap.
CODEMAN_PREAMBLE=1.22.0
PREAMBLE
)
. "$PRE"; [ "${CODEMAN_PREAMBLE:-}" = 1.22.0 ] || { echo "preamble at $PRE is stale or truncated: rm it and re-run this block"; exit 1; }

Every later Bash call that touches the API starts with the same two loader lines from the top of this section.

Why it is built this way, all of it load-bearing:

  • It still fails closed. A missing or truncated file means delete_session is undefined, and an undefined function is "command not found", which deletes nothing. ⚠️ This argument covers accidents, NOT a hostile file: a complete attacker-written preamble can define delete_session and set the stamp, and sourcing executes it. What defends against that is the path choice in the next bullet, not this one. Never hand-roll a DELETE of your own, which is the one thing that would route around this.
  • The version stamp is the LAST line, and the write condition greps for it. That one choice covers staleness and truncation together: an old skill version's file and a half-written one both fail the grep and are rewritten in place, so neither costs you a round trip to diagnose and rm. The older [ -s "$PRE" ] condition could not tell a complete file from a half-written one and left both to the post-source guard, which can only refuse, not repair. That guard stays as the fail-closed backstop: if the rewrite itself is cut short, CODEMAN_PREAMBLE is unset and the call stops.
  • Not /tmp. On a shared machine /tmp is world-writable, so another local user can pre-create the exact path you are about to . and have their code run as you. $HOME-derived paths are not world-writable, and the file is written 0600 anyway. The file holds the credential-recovery code, not a recovered password.
  • Never put $$ in a clientId. It changes per call, so the "resend the identical request" loop in §5.3 would stop being a duplicate and would retype the prompt, submitting the turn twice. Use the fixed literal $CID.
  • Only real environment variables (CODEMAN_*, HOME) survive, which is why the preamble rebuilds $API and $SELF from them on every source rather than baking them in.

If a call comes back as unparseable text instead of JSON, that is almost always a plain-text 401: see §6 and the symptom gallery.

1. The fast path: N workers, one Bash call

If the job is "spawn N claude workers, give them tasks, collect the answers", this block is the whole thing. Run it, report, and stop reading. §2 onward is for jobs this does not cover; you are not being careless by not reading them.

Fill in the case names and the prompts, then run it as your FIRST Bash call: no standalone preamble check before it (line one below IS that check), and no reconnaissance. ls ~/codeman-cases answers nothing this block needs: invented fresh names need no lookup, and spawn_worker refuses a name that already exists rather than silently reusing it. Everything below is spawn_workers / sendwait / last_text / delete_session from the §0 preamble, so there is nothing to assemble and no per-call body to hand-build.

. "${XDG_CACHE_HOME:-$HOME/.cache}/codeman-agent-$CODEMAN_SESSION_ID.sh" 2>/dev/null   # §0 loader
[ "${CODEMAN_PREAMBLE:-}" = 1.22.0 ] || { echo "preamble missing or stale; run the full §0 block"; exit 1; }
N=(alpha beta)                    # INVENT one fresh case name per worker; never list cases first
                                  # (a name may carry a mode: `beta:deepseek`, see below)
T=('reply with one line: the absolute path of your working directory'
   'reply with one line: your model name')            # tasks, same order as N

S=(); while read -r _ s; do S+=("$s"); done < <(spawn_workers "${N[@]}")   # concurrent
for i in "${!N[@]}"; do [ -n "${S[$i]:-}" ] || FAIL=1; done
[ -z "${FAIL:-}" ] || { echo "a spawn failed (stderr says why; §5.1): deleting the siblings"
  for s in "${S[@]}"; do [ -n "$s" ] && delete_session "$s" >/dev/null; done; exit 1; }

D=$(mktemp -d) || { for s in "${S[@]}"; do delete_session "$s" >/dev/null; done; exit 1; }
for i in "${!N[@]}"; do sendwait "${S[$i]}" "${T[$i]}" > "$D/$i" & done; wait
for i in "${!N[@]}"; do
  jq -ce --arg n "${N[$i]}" \
    '{worker:$n,delivered:.data.delivered,timedOut:.data.wait.timedOut,signal:.data.wait.signal}' \
    "$D/$i" || echo "{\"worker\":\"${N[$i]}\",\"error\":\"send produced no result\"}"
  echo "== ${N[$i]}"; last_text "${S[$i]}" || echo "(no response written)"
done
for i in "${!N[@]}"; do   # delete ONLY what finished; a timeout means STILL WORKING (§3 rule 5)
  if jq -e '.success and .data.delivered and (.data.wait.timedOut|not)' "$D/$i" >/dev/null 2>&1
  then delete_session "${S[$i]}" >/dev/null
  else echo "kept ${N[$i]} (${S[$i]}): its line above says why; re-wait or repair (§5.3), then delete_session it"
  fi
done; rm -rf "$D"

Measured against a live 1.18.0 server: two cold workers spawned and ready in 6.3 s, both turns dispatched and both answers read in 4.0 s more. If your run takes minutes, the time went into deliberation, not the API. The four things that actually cost time:

  • Spawning serially. One worker per Bash call is one model turn per worker. & plus wait, as above, makes N workers cost about what one costs.
  • Reconnaissance turns before the spawn. A standalone preamble check, an ls ~/codeman-cases, a list_sessions "to see what is there": each is a whole model turn spent learning something this block already handles (line one performs the preamble check, invented names need no listing, and spawn_worker refuses collisions). A live two-worker run spent ~12 s of its 28 s total on exactly two such turns; the API work in between was under 10 s.
  • Re-deriving the happy path from §5.1 + §5.2 + §5.3 + §5.10. That is what the preamble functions exist to end. Compose them; do not rebuild them. The tells that you are rebuilding anyway: a for loop around quick-start, a poll on .data.pid, a bespoke ready() or spawn() of your own. Each is a worse copy of a function already sitting in your preamble; the live run that wrote them spawned serially, polled pid for nothing, and shipped its workers without lineage.
  • Verifying what is already checked for you. Two verifications specifically are not worth a call here, because spawn_worker carries them: the hooks check (it refuses a name that resolved to a hook-less directory with one local grep, so a worker it hands back always has a working stop and sendwait is trustworthy), and the pid poll, which is dead weight because wait-output already blocks on the composer.

Four things this block leans on, each one link away, no detour needed to run it:

  • Those case names must be fresh scratch names: they create ~/codeman-cases/<name>, not your repo. A name that already means something (a linked case, a pre-existing directory) is refused by spawn_worker rather than silently reused. Spawning where the work actually is (a linked case, a git worktree) is a different call, and picking the wrong one is the costliest mistake in this skill: §5.1. Those workspaces do get hooks now, unless the operator disabled it.
  • sendwait supplies the \r, picks a fresh seq, and self-heals a stranded Enter. A prompt without the \r is never submitted (§3), a reused seq is silently swallowed as an already-applied duplicate, and an Enter eaten by an Ink repaint strands the prompt on the composer until a bare \r follows: all three are reasons to let sendwait build the call rather than hand-rolling it.
  • Each sendwait costs that worker one billed turn, as does every prompt you send it.
  • Deleting the sessions does not remove the case directories. They are marked as agent-created, so GET /api/v1/cases/agent-created lists them for cleanup: §5.14.

DeepSeek Harness workers

The block above spawns claude workers. Any entry in N may instead name a mode (beta:deepseek), and a deepseek worker is driven by the same four verbs, with no change to the rest of the block: spawn_workers waits for its composer, sendwait blocks on its real end-of-turn signal, last_text reads its answer, delete_session removes it.

That is true of no other non-claude mode, and it is worth knowing why: the DeepSeek Harness TUI reports idle/working/blocked to Codeman over the supervisor contract it implements, so dsh is the one external CLI with definitive stop/blocked signals instead of guessed-from-silence ones — and it writes a structured transcript, which is what last-response reads for it. shell, opencode, codex, gemini, antigravity, pi, grok and omp have neither and still need markers (§5.5).

Three things to know before you spawn one:

  • It needs a pane-capable profile. dsh ships only web/headless, so the terminal agent is always an installed profile. GET /api/v1/deepseek/status answers both questions separately (available = the binary, runnable = a profile that can drive a pane); a spawn without one fails with OPERATION_FAILED rather than falling back.
  • Do not task it on the strength of a stop alone. The harness reports idle at boot ~300 ms before its composer paints (measured 2.26 s vs 2.56 s), so a sendwait fired straight after quick-start resolves on that boot signal, reports a turn that never ran, and leaves the prompt in a pane that was not yet taking input. Letting spawn_worker gate on readiness is what steps past that edge; it is not optional.
  • A profile that does not implement the contract looks like a hang. Codeman cannot know at spawn time whether one does. The tell is a sendwait that times out on a worker whose pane clearly finished: that profile is one of them, so drive it with markers instead.

2. What do you want to do?

One row per job. Acting on this table alone is correct; the §5 links are the detail.

I want toCallDetail
start a worker where the work isPOST /api/v1/quick-start {"caseName":…}, which creates ~/codeman-cases/<name> unless the name is already a case. Any other path (a git worktree): POST /api/v1/sessions {"workingDir":…} then POST /api/v1/sessions/:id/interactive. Both install hooks by default, so expect full signals in either, and verify rather than assume. N workers means N worktrees§5.1
know a new worker can accept a promptGET .../wait-output?match=shift+tab&from=buffer (urlencode the +); a deepseek worker draws ❯ instead, and its boot stop fires ~300 ms BEFORE that, so never read the signal as readiness§5.2
deliver a task and know when it finishedPOST .../input with "input":"…\r", clientId, seq, "wait":true. Resolves on stop, so it is trustworthy where the signal is real: claude mode with hooks (installed by default, but the operator can disable it and remote sessions never get them) and deepseek mode through its status bridge. Costs the worker one billed turn§5.3
know a hook-less worker finishedit has no stop, and wait:true there resolves on flapping idle without erroring: make it print a split, unique marker and wait-output on that instead§5.5
read the answerGET .../last-response, polled (claude, codex and deepseek write a transcript; empty for the other modes)§5.4
know if it is aliveGET .../wait?until=exit&timeout=1000: an immediate signal:"exit" means dead. status and pid both lie§5.6
know if it is stuckGET .../active-tools and GET .../run-summary are structured and free; two terminal?tail= samples are the crude fallback§5.6
make a runaway worker stopPOST .../input {"input":"\u001b"} (ESC, no \r). Deleting the session would destroy the conversation instead§5.7
resume a worker halted on a usage limitPOST .../auto-resume {"enabled":true}. Respawn and Ralph are not the remedy: respawn runs /clear§5.8
give a worker big inputwrite a file into its workspace with your own tools and send one short line pointing at it. The composer takes 65536 characters, single-line, newlines stripped§5.9
watch N workers at onceone in-flight wait per worker (per-session waiter cap 16); fan-out shapes differ for claude and shell§5.10
find yourself, list what existsGET /api/v1/sessions, match your $SELF by prefix§5.11
read or record what the user wantsGET/PUT .../intent, and POST .../readmymind to predict§5.12
talk to a claude worker directlyListAgents / SendMessage, when the feature is on at both ends§5.13
clean updelete_session "$SID" per id you created. Case directories and git worktrees are not removed with it; GET /api/v1/cases/agent-created lists the scratch case dirs your spawns left behind, for you to report§5.14

3. Rules digest

Ten one-liners. Each breaks something concrete; the reason is one link away.

  1. End every input with \r or Enter is never sent and the text sits unsubmitted (§5.3).
  2. Never branch on .data.status. It reads idle mid-turn and idle on a dead worker (§5.6).
  3. Split your markers. Your typed command echoes into the output stream, so an unsplit marker matches before the command runs (§5.5).
  4. Match single space-free tokens against TUI output. A TUI positions words with cursor moves, so multi-word matches are unreliable there (§5.2).
  5. A wait timeout is a 200, not an error. Loop over short waits; the clamp and the applied wait.timeoutMs are in endpoints.md.
  6. Signals are edge-triggered with no history. Register the waiter before the event can happen; a stop that fires with no waiter is unobservable afterwards (§5.10).
  7. Never delete without delete_session. The server lets a session delete itself (§4).
  8. One in-flight wait per worker. The per-session waiter cap is 16 and abandoned waits count against it (§5.10).
  9. Every message you send a worker costs it a billed turn, including a readiness ping and an interrupted turn (§5.7).
  10. Never answer another session's dialog. Approving a permission prompt you did not raise authorizes an action the user never saw (§4).

4. Safety rules

You are yourself a session on this server, and the API has no undo.

  • Never act on your own session, and know that delete_session is the ONLY guard. The server has no self-protection: a session that DELETEs its own id succeeds and dies silently (verified live). Always delete through delete_session "$SID" from §0; never write a bare curl -X DELETE and never reintroduce the is_self … || curl -X DELETE … shape. That older form failed open: with the function undefined (a missing or truncated preamble file, see §0) bash returns 127, the || branch fires, and the delete runs with no self-check at all. Wrapping the request inside the guard is what makes a lost preamble delete nothing instead of deleting you. Apply the same prefix-both-directions reasoning before any kill, respawn, or input call you write by hand.
  • Mutating calls you may make unprompted (this is an allowlist): POST /api/v1/quick-start; POST /api/v1/sessions + POST /api/v1/sessions/:id/interactive (or /shell) for a directory the user's own task named; POST /api/v1/sessions/:id/input; and DELETE /api/v1/sessions/:id only for a session you created in this conversation, by exact id. Keep a list of the ids you create. Everything else mutating needs the user to have asked for it.
  • Never call these unless the user explicitly asked, naming the target:
    • DELETE /api/cases/:name recursively deletes a real directory of the user's code from disk. One wrong case name destroys work that was never yours.
    • DELETE /api/sessions (no id) is a bulk kill of every session, the user's real work included. DELETE /api/subagents/:agentId kills one background agent; DELETE /api/subagents (no id) does not kill anything, it clears the watcher's map and timers, which blinds every subagent surface in the UI until they are rediscovered. Neither is yours to call.
    • respawn / ralph / orchestrator / cron mutations: respawn runs /clear (wipes a conversation), orchestrator state is a single global slot, cron jobs outlive you.
    • PUT /api/settings, POST /api/system/update: global UI settings; server restart.
    • POST /api/approvals/:id/answer. It types a digit, an Esc or free text into whichever session raised the prompt. Approving another session's permission dialog authorizes a tool call the user never saw, from a session that is not yours. Answer only a prompt raised by a worker you created, and only when the user asked you to.
  • Never spawn a worker into the directory you are editing, and give N workers N git worktrees rather than one shared checkout. Two agents in one working tree interleave writes and each reads the other's half-finished files; a git checkout in one yanks the tree out from under the other. Creating worktrees changes the user's repository state, so say that you did; removing one discards any uncommitted work inside it, so ask first (§5.1).
  • Never tmux kill-session, pkill tmux, pkill claude. The API is the only interface.
  • Sessions count against a global cap of 50 (and, in multi-user mode, a per-user cap of 25 that fires the same 409). Case creation is uncapped and writes real directories. Clean up every session you start, and never retry quick-start in a loop.

5. Recipes → reference/verbs.md

The per-verb detail lives in reference/verbs.md, loaded on demand so it is not paid for on every skill load. Section numbers and anchors are unchanged, so a §5.4 reference still resolves. §1 already covers the common job without any of these; open the one row you actually hit.

OpenWhen
5.1 Where to spawnthe work is not a fresh scratch case: a linked case, a git worktree, any path that already existed. Hooks are absent there, which silently breaks send-and-wait. The costliest mistake in this skill
5.2 Readinessa worker never drew its composer, or you need the trust-dialog ladder by hand
5.3 Send a task and waitthe sendwait body, its signals, and the duplicate-resend loop
5.4 Read the answerlast_text came back empty, or the mode is not claude/codex/deepseek
5.5 Markers for hook-less workersthe worker has no stop hook: synchronize on a split, unique printed marker
5.6 Alive and stuckis it dead or just slow? status and pid both lie
5.7 Interrupt without destroyinga runaway worker you want to stop but keep
5.8 Usage limitsa worker halted on a subscription limit
5.9 Big input via the workspacethe prompt is larger than one composer line
5.10 Fan outmany workers at once: waiter caps, and why signals are edge-triggered
5.11 List and find yourselfenumerate sessions, or match $SELF by prefix
5.12 Read My Mindread or record what the user wants for a case
5.13 Messaging claude workersListAgents / SendMessage instead of the HTTP path
5.14 Clean upwhat deleting a session does not remove, and how to list the case dirs you left

6. Setup and auth

You need this section only when the API answers something jq cannot parse, or when you are on a server old enough to lack the wait endpoints. Endpoint-level detail lives in endpoints.md.

Credentials

Auth is active only when the server has CODEMAN_PASSWORD (or is in multi-user mode). Your session has usually inherited that password already, which is why the §0 preamble tries $CODEMAN_PASSWORD first: Codeman does not strip it. buildClaudeEnv() (src/session-cli-builder.ts) spreads the server's entire process.env into the session and deletes only COLORTERM and CLAUDECODE, and the tmux spawn path applies no denylist either. On a stock password-protected install (install.sh writes the password into the systemd unit or launchd plist, so the server process carries it) the value is simply in your environment.

It is not guaranteed, though, which is what the fallbacks are for. A tmux pane inherits the tmux server's environment, and that server can predate the password; and the data dir's .env is only ever read by the codeman CLI itself, never loaded into the web server's environment.

Fallback 1, in the §0 preamble already: the data dir's .env, the same file codeman attach reads. It is hand-authored; nothing ever writes it.

Fallback 2, for a stock install where the supervisor definition is the only copy on disk. Append this to the preamble file (before its version-stamp line) and re-source:

if [ -z "${CODEMAN_PASSWORD:-}" ]; then    # install.sh puts it in the service definition
  UNIT="$HOME/.config/systemd/user/codeman-web.service"
  PLIST="$HOME/Library/LaunchAgents/com.codeman.web.plist"
  if [ -f "$UNIT" ]; then
    # install.sh backslash-escapes " and \ in the unit value; undo it or a password
    # containing either recovers wrong and auth fails.
    CODEMAN_PASSWORD=$(sed -n 's/^Environment="CODEMAN_PASSWORD=\(.*\)"$/\1/p' "$UNIT" | head -1 | sed 's/\\\(["\\]\)/\1/g')
  elif [ -f "$PLIST" ]; then
    # install.sh XML-escapes the plist value; undo it (&amp; LAST, mirroring escape order).
    CODEMAN_PASSWORD=$(awk '/<key>CODEMAN_PASSWORD<\/key>/{getline; print}' "$PLIST" | sed -n 's/.*<string>\(.*\)<\/string>.*/\1/p' \
      | sed -e 's/&lt;/</g' -e 's/&gt;/>/g' -e 's/&amp;/\&/g')
  fi
fi

⚠️ A 401 is plain text, not the JSON envelope, so on a password-protected server every jq in these recipes dies with jq: parse error instead of showing UNAUTHORIZED. If that happens, check the status with -w '%{http_code}'; if it is 401 and no fallback found a credential, stop and tell the user you need credentials. The same is true of the guards that run before any handler: the Host allowlist (403 Forbidden: host not allowed), the Origin/CSRF guard, and the auth rate limiter's 429 all answer in plain text. The hook-secret bypass covers only /api/hook-event and /api/status-telemetry, never session control.

In multi-user mode accounts live in users.json and the credential is a real user's name and password. A recovered CODEMAN_PASSWORD still often works: bootstrapInitialAdmin() (user-store.ts:417-427) creates the FIRST admin from CODEMAN_USERNAME/CODEMAN_PASSWORD on first boot when no users exist, so on a stock multi-user install that pair usually IS a valid admin login until someone changes it. Try it once; if it fails, ask the user rather than retrying (ten failures rate-limit the address).

Server version

The wait endpoints first ship in Codeman 1.13.0, but do not gate on the version number: a dev build can serve them while reporting an older version. Probe instead. GET .../wait on a real session id answering 404 with an .error starting Route means the server predates them (fall back to polling GET .../terminal?tail= and say so). Session ... not found means your session id is wrong, not the server.

Where the API is unreachable

  • Remote-SSH cases do not export CODEMAN_MUX/CODEMAN_API_URL into the session, so the §0 guard fails closed and you refuse to act. That is correct behavior, not a bug to work around.
  • Inside a Docker case, a loopback-bound server is unreachable from the container, and CODEMAN_DOCKER_BRIDGE_HOOKS=1 does not fix it: that opens a hooks-only listener, so hook events flow but /api/v1/* stays refused. Report it rather than retrying; making it reachable is an operator decision.

Everything else (endpoint tables, per-mode signal table, error codes, capacity limits, Docker/remote caveats): reference/endpoints.md. Fan-out orchestration and blocked-worker handling: reference/recipes.md.

发现问题?提交给管理员复核

评分:

评论 (0)

暂无评论,成为第一个评论者吧!