复制安装命令
用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。
复制前请先查看来源、License 和安全提示。
Mission control for AI coding agents
用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。
复制前请先查看来源、License 和安全提示。
来源文件:README.md
Claude Code • OpenCode • Codex • Antigravity • Gemini • Pi • Grok • OMP • Terminal - One Dashboard • Any Device
English • 简体中文
Codeman is a self-hosted mission control for AI coding agents. It spawns Claude Code, OpenCode, Codex, Antigravity, Gemini, Pi, Grok, or OMP inside persistent tmux sessions, streams the real terminal to any browser, and keeps agents productive after you walk away: it re-prompts on idle, resumes when a usage limit resets, runs scheduled jobs, and shows every background agent working in real time.
Get started in one line (macOS & Linux, Windows via WSL):
curl -fsSL https://getcodeman.com/install | bash
codeman web
# Open http://localhost:3000 and start your first session
The installer asks before every system change, and re-running the same line updates in place. Full details: Quick Start - Installation.
curl -fsSL https://getcodeman.com/install | bash
This installs Node.js, tmux and a build toolchain if missing (node-pty ships no Linux prebuilds, so it compiles from source), clones Codeman to ~/.codeman/app, and builds it. A few things worth knowing:
tailscale serve, so you get https://<machine>.<tailnet>.ts.net with a real certificate and your tailnet as the login, no password needed), any device on your network (0.0.0.0, with a strongly recommended password prompt), or this machine only (127.0.0.1, safest). Skipping the password on a network bind requires an explicit confirmation and ends with a loud warning. The highlighted default reflects what is already on the machine (Tailscale when it is already in use, your existing binding on a re-run), and a bare Enter never pulls in new software. A bare codeman web started by hand still defaults to loopback.~/.codeman/app are stashed (never discarded), and a running service is restarted and verified. If a first install was interrupted, re-running resumes the full setup instead. install.sh update and install.sh uninstall also exist.CODEMAN_NONINTERACTIVE=1 to approve them for automation.You'll need at least one AI coding CLI installed — Claude Code, OpenCode, Codex, Antigravity, Gemini CLI, Pi, Grok Build, DeepSeek Harness, or OMP (any combination works; Gemini CLI is enterprise-only since Google's consumer cutover, and Antigravity is its successor). The installer detects whichever of the nine is present; if none is found, it offers to install Claude Code or OpenCode, or you can skip and install one yourself later. After install:
codeman web
# Open http://localhost:3000 and start your first session
Sharing with a small team? Start it in multi-user mode instead: each person gets their own login and workspace.
codeman users add alice --admin # create the first admin account
codeman web --multiuser # named logins + per-user case spaces
Prefer Docker Compose? A local-image Compose deployment ships in docker/: copy docker/.env.example to docker/.env, set CODEMAN_PASSWORD, then run bash docker/Start-Codeman.sh on Linux. Codeman runs in a container and spawns Docker cases as sibling containers through the host socket. See the Docker deployment guide for direct Compose commands, storage and networking options.
Details in Multi-User Mode below.
To outlive the shell you started it in, without setting anything up:
codeman web -d # detach; logs to ~/.codeman/web.log
codeman web --status # is it up, and on which pid
codeman web --stop # graceful SIGTERM; agents keep running in tmux
-d waits until the server actually answers before reporting success, and refuses to start a second one on the same data dir (two servers sharing a tmux socket attach to each other's sessions).
To have it come back after a reboot, install it as a service instead. The installer's final menu does this for you (option 2); codeman service is the equivalent for an npm i -g aicodeman install:
codeman service install # systemd user unit (Linux) or LaunchAgent (macOS)
codeman service status
codeman service uninstall
service install writes the unit with your current PATH baked in, which matters more than it sounds: launchd hands a job /usr/bin:/bin:/usr/sbin:/sbin, so a Homebrew or nvm node, tmux or claude is invisible to a hand-written plist. It never copies CODEMAN_PASSWORD into the unit file; add that yourself if the service needs auth.
To write the unit by hand instead:
Linux (systemd):
mkdir -p ~/.config/systemd/user
cat > ~/.config/systemd/user/codeman-web.service << EOF
[Unit]
Description=Codeman Web Server
After=network.target
[Service]
Type=simple
ExecStart=$(which node) $HOME/.codeman/app/dist/index.js web
Restart=always
RestartSec=10
[Install]
WantedBy=default.target
EOF
systemctl --user daemon-reload
systemctl --user enable --now codeman-web
loginctl enable-linger $USER
macOS (launchd):
mkdir -p ~/Library/LaunchAgents
cat > ~/Library/LaunchAgents/com.codeman.web.plist << EOF
<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN"
"http://www.apple.com/DTDs/PropertyList-1.0.dtd">
<plist version="1.0">
<dict>
<key>Label</key>
<string>com.codeman.web</string>
<key>ProgramArguments</key>
<array>
<string>$(which node)</string>
<string>$HOME/.codeman/app/dist/index.js</string>
<string>web</string>
</array>
<key>RunAtLoad</key><true/>
<key>KeepAlive</key><true/>
<key>StandardOutPath</key>
<string>/tmp/codeman.log</string>
<key>StandardErrorPath</key>
<string>/tmp/codeman.log</string>
</dict>
</plist>
EOF
launchctl bootstrap gui/$(id -u) ~/Library/LaunchAgents/com.codeman.web.plist
wsl bash -c "curl -fsSL https://getcodeman.com/install | bash"
Codeman requires tmux, so Windows users need WSL. If you don't have WSL yet: run wsl --install in an admin PowerShell, reboot, open Ubuntu, then install your preferred AI coding CLI inside WSL (Claude Code, OpenCode, Codex, Antigravity, Gemini CLI, Pi, Grok Build, DeepSeek Harness, or OMP). After installing, http://localhost:3000 is accessible from your Windows browser.
The most responsive AI coding agent experience on any phone. Full xterm.js terminal with local echo, swipe navigation, and a touch-optimized interface designed for real remote work — not a desktop UI crammed onto a small screen.
![]() | ![]() |
| Answering prompts by touch | Accessory bar + dedicated Enter button |
| Terminal Apps | Codeman Mobile |
|---|---|
| 200-300ms input lag over remote | Local echo — instant feedback |
| Tiny text, no context | Full xterm.js terminal |
| No session management | Swipe between sessions |
| No notifications | Push alerts for approvals and idle |
| Manual reconnect | tmux persistence |
| No agent visibility | Background agents in real-time |
| Copy-paste slash commands | One-tap /init, /clear, /compact |
| Password typing on phone | QR code scan — instant auth |
/init, /clear, /compact quick-action buttons above the virtual keyboard; destructive commands require a double-press to confirm, so you never fire one by accidentvisualViewport API)codeman web --https
# Open on your phone: https://<your-ip>:3000
localhostworks over plain HTTP. Use--httpswhen accessing from another device, or use Tailscale (recommended): the installer can set it up for you (choose Tailscale at the network-access prompt, or runbash ~/.codeman/app/install.sh tailscaleon an existing install). That gives youhttps://<your-machine>.<tailnet>.ts.netwith a real certificate: private to your tailnet, no password required, and PWA install + push notifications work on your phone.
Typing passwords on a phone keyboard is miserable. Codeman replaces it with cryptographically secure single-use QR tokens — scan the code displayed on your desktop and your phone is authenticated instantly.
Each QR encodes a URL containing a 6-character short code that maps to a 256-bit secret (crypto.randomBytes(32)) on the server. Tokens auto-rotate every 60 seconds, are atomically consumed on first scan (replays always fail), and use hash-based Map.get() lookup that leaks nothing through response timing. The short code is an opaque pointer — the real secret never appears in browser history, Referer headers, or Cloudflare edge logs.
The security design addresses all 6 critical QR auth flaws identified in "Demystifying the (In)Security of QR Code-based Login" (USENIX Security 2025, which found 47 of the top-100 websites vulnerable): single-use enforcement, short TTL, cryptographic randomness, server-side generation, real-time desktop notification on scan (QRLjacking detection), and IP + User-Agent session binding with manual revocation. Dual-layer rate limiting (per-IP + global) makes brute force infeasible across 62^6 = 56.8 billion possible codes. Full security analysis: docs/qr-auth-plan.md
A start-to-finish walkthrough for driving Codeman from the browser. If you just installed, this is where to begin.
codeman web # localhost:3000 (loopback only — safe default)
codeman web --port 8080 # custom port (or set CODEMAN_PORT)
codeman web --https # self-signed TLS (only needed for remote access)
codeman web -H 0.0.0.0 # bind LAN — REQUIRES CODEMAN_PASSWORD (see Security)
codeman web -d # detach: survives closing the shell (--status, --stop)
codeman service install # systemd/launchd service: comes back after reboots
Open the printed URL. The page is a single dashboard; everything below happens there.
Click + New Session (or Quick Start). A session is one AI CLI running in its own tmux-backed terminal. You choose:
| Field | What it does |
|---|---|
| Working directory / case | The folder the agent operates in. A "case" is just a named working dir Codeman remembers. Add Case creates one from scratch, links an existing folder, or clones a GitHub repo straight into one (Clone Repo). |
| CLI / run mode | Claude (default), OpenCode, Codex, Antigravity, Gemini, Pi, Grok, OMP, or Terminal (plain shell). |
| Model | Per-session model (App Settings → Models → New Claude sessions). A soft default — /model still works in-session. |
| Effort / Ultracode | Reasoning effort (low–max) or ultracode for dynamic multi-agent workflows. Switchable anytime with /effort. |
Hit start — Codeman spawns the CLI via a real PTY and streams it to your browser over SSE.
Alt+1-9 to jump, Ctrl+Tab for next, drag to reorder (tab order syncs across your devices).xterm.js terminal; full TUIs render correctly. Type directly and press Enter to send. Shift+Enter inserts a newline.Ctrl+Shift+V (Deepgram Nova-3, with auto-silence stop).| Mode | Use it for | Where |
|---|---|---|
| Respawn | Long unattended runs — auto-restarts the CLI on idle/limit, with adaptive timing. Presets: solo-work, overnight-autonomous, … | Respawn tab |
| Orchestrator | Turn one goal into a phased plan and drive it to completion across agents. | Orchestrator panel |
| Cron | Saved, named jobs on a schedule (once/interval/daily/weekly) that spawn a session and send a prompt when due. | ⏰ Cron button (opt-in: App Settings → Header & Panels → Scheduling) |
| Auto-resume | Automatically continue after a subscription rate-limit resets. | Respawn tab (top) |
./scripts/tunnel.sh start opens a Cloudflare tunnel (set CODEMAN_PASSWORD first).codeman tui is a full-screen dashboard in the terminal (codeman tui --list to list, codeman tui 2 to attach straight to one).codeman web -d detaches from your shell (--status, --stop); codeman service install makes it a systemd user unit / macOS LaunchAgent that survives reboots. Both verify the server actually answers before reporting success, and both refuse to start a second server on one data dir. See Keep it running in the background.⚠️ Safety: if you're working inside a Codeman-managed session (
echo $CODEMAN_MUX→1), never runtmux kill-session/pkill claudedirectly — use the web UI or./scripts/tmux-manager.sh.
When accessing your coding agent remotely (VPN, Tailscale, SSH tunnel), every keystroke normally takes 200-300ms to round-trip. Codeman implements a Mosh-inspired local echo system that makes typing feel instant regardless of latency.
A pixel-perfect DOM overlay inside xterm.js renders keystrokes at 0ms. Background forwarding silently sends every character to the PTY in 50ms debounced batches, so Tab completion, Ctrl+R history search, and all shell features work normally. When the server echo arrives 200-300ms later, the overlay seamlessly disappears and the real terminal text takes over — the transition is invisible.
<span> at z-index 7 inside .xterm-screen, completely immune to Ink's constant screen redraws (two previous attempts using terminal.write() failed because Ink corrupts injected buffer content)fontFamily, fontSize, fontWeight, and letterSpacing from xterm.js computed styles so overlay text is visually indistinguishable from real terminal outputExtracted as a standalone library:
xterm-zerolag-input— see Published Packages.
Watch background agents work in real-time. Codeman monitors agent activity and displays each agent in a draggable floating window with animated Matrix-style connection lines back to the parent session.
Multi-agent Workflow runs ("ultracode") get the same treatment: a floating run window tracks the whole workflow live, with phases, per-agent token counts, and the current tool of every agent:
Agent Teams — first-class support for Claude Code's native multi-agent teams (CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1). TeamWatcher polls ~/.claude/teams/, matches teammates to their lead session, and surfaces them as live subagent windows with team-aware idle detection — so the Respawn Controller won't fire while teammates are still working. See docs/agent-teams/.
The core of autonomous work. When the agent goes idle, the Respawn Controller detects it, sends a continue prompt, cycles context management commands for fresh context, and resumes — running 24+ hours completely unattended.
WATCHING → IDLE DETECTED → SEND UPDATE → /clear → /init → CONTINUE → WATCHING
continue — so an overnight run survives the 5-hour window instead of stalling until morning. Recognizes every Claude Code limit-message format, retries if still limited, survives Codeman restarts, and holds respawn cycles while paused so /clear can't wipe the waiting conversation. Enable per session at the top of the Respawn tabsolo-work (3s idle, 60min), subagent-workflow (45s, 240min), team-lead (90s, 480min), ralph-todo (8s, 480min), overnight-autonomous (10s, 480min)Beyond single-session respawn, the Orchestrator turns a high-level goal into a phased plan and drives it to completion across multiple agents — a state machine that runs idle → planning → approval → executing → verifying → (replanning) → completed.
orchestrator key in state.json, so it survives restartsPOST /api/orchestrator/start → /approve → /status (10 endpoints)Full design:
docs/orchestrator-loop-architecture.md.
Run 20 parallel sessions with full visibility — real-time xterm.js terminals at 60fps, per-session token and cost tracking, tab-based navigation, and one-click management.
Every session runs inside tmux — sessions survive server restarts, network drops, and machine sleep. Auto-recovery on startup with dual redundancy. Ghost session discovery finds orphaned tmux sessions. Managed sessions are environment-tagged so the agent won't kill its own session.
Ctrl/Cmd/Alt+K opens a fuzzy session palette; Browse all sessions opens the Session Manager: one deduped list of everything Codeman knows about (live sessions, past sessions from state and lifecycle history, and Claude transcripts), each row showing its first and most recent prompt.
Running Codeman on multiple hosts (laptop, dev box, NAS)? The browser tab title is codeman:<hostname> so you can tell which backend each tab points at without clicking in:
codeman web # codeman:<os.hostname()>
codeman web --title-hostname dev-box # codeman:dev-box (manual override for noisy hostnames)
The title is templated into the served HTML on first byte, so it's correct from the very first paint and works without JavaScript. The same hostname prefix is applied to the tab-flash format (⚠️ (N) codeman:<host>) and to OS-level desktop notifications (codeman:<host>: <event>), so cross-host alerts in the system notification center are also unambiguous.
| Threshold | Action | Result |
|---|---|---|
| 110k tokens | Auto /compact | Context summarized, work continues |
| 140k tokens | Auto /clear | Fresh start with /init |
Every tab tells you its state at a glance. A running session keeps its green status dot. When a session stops and waits for input, its tab turns yellow: steady ring, tinted background, yellow dot, with a slow breathing glow on top. When a permission prompt or question is blocking the agent, the tab turns red with a faster pulse. The base tint never blinks off, so even a split-second glance (or a screenshot) reads the true state; the ring stays visible while the tab is selected, and a page reload re-arms pending alerts from the server, so a blocked session can never hide behind a fresh-looking tab.
Real-time desktop alerts when sessions need attention — permission_prompt and elicitation_dialog trigger critical red tab blinks, idle_prompt triggers yellow blinks. Click any notification to jump directly to the affected session. Hooks auto-configured per case directory.
Click the chart icon on any session tab to see a timeline of everything that happened — respawn cycles, token milestones, auto-compact triggers, idle/working transitions, hook events, errors, and more.
Terminal-based AI agents (Claude Code's Ink, OpenCode's Bubble Tea) redraw the screen on every state change. Codeman implements a 6-layer anti-flicker pipeline for smooth 60fps output across all sessions:
PTY Output → 16ms Server Batch → DEC 2026 Wrap → SSE → Client rAF → xterm.js (60fps)
codeman web -d runs the server detached with a pidfile, ~/.codeman/web.log, and verified startup (it polls the server until it answers, so a port clash never reads as success); codeman service install writes a systemd user unit (Linux) or LaunchAgent (macOS) with your shell's PATH baked in, so an nvm or Homebrew node, tmux and claude are actually found. Secrets are never written into unit files~/codeman-cases/<name> and registers it as a normal case, ready to run an agent in. It preflights the URL while you type (tells you whether it can be cloned anonymously and offers the repo's real branches and tags for the optional branch/tag field), fills the case name in from the URL, and lets you pick which CLI the Run button should use. Public repositories over https://; Codeman never collects or stores credentialsCLAUDE_CODE_* vs OPENCODE_* vs CODEX_* vs ANTIGRAVITY_* vs GEMINI_*/GOOGLE_* vs PI_* vs GROK_*/XAI_* vs OMP_*). See docs/opencode-integration.md, docs/pi-integration.md, docs/grok-integration.md and docs/omp-integration.md.tar.gz to move it to another machine. See docs/docker-cases.mddocs/remote-sessions.mdlow–max) or enable ultracode (dynamic multi-agent workflows). Soft defaults only — switchable anytime with /effort in-session. Extended-thinking budget is configurable tooCtrl+Shift+V)CODEMAN_GESTURE=1 + App Settings → Terminal & Inputcodeman:<host> so multi-host setups stay unambiguousRun a case inside its own hardened Docker container instead of directly on your host — for security isolation, reproducible toolchains, and one-click portability.
docker exec into the same container; killing one session never tears the container out from under the others.--cap-drop ALL, no-new-privileges, PID/memory caps, never --privileged or the docker socket; a sealed profile (no host credentials, network off) is one toggle away..tar.gz, docker load it on the other side, and import it into a fresh case.Prerequisite: just Docker (or Podman). The agent base image builds itself automatically on first use, with progress streamed to the UI (or pre-build it with node scripts/build-agent-image.mjs). Full guide: docs/docker-cases.md.
Point a case at another machine and run the agent there, over SSH, with the same dashboard, mobile UI, and autonomy features. Your laptop is just a window onto a session that lives on the remote host.
codeman-* sessions already running on a host (started by that machine's own Codeman, or by another operator) and attach to one. Attached sessions you don't own detach on tab close, never kill.Set it up under New Case → Remote (host, user, identity file, optional jump host). Full design: docs/remote-sessions.md.
Share one Codeman with a small trusted team, each person getting their own login and workspace. Off by default — without the flag, nothing changes.
Enable with codeman web --multiuser (or CODEMAN_MULTIUSER=1). Create the first admin, then manage users from the CLI or the Users tab in App Settings:
codeman users add alice --admin # prompts for a password (or --password-stdin)
codeman users add bob # a regular user
codeman users list
~/codeman-users/<name>/cases; sessions, cases, search, and real-time events are scoped to their owner. Admins see everything.~/.codeman/users.json; disable, reset (one-time password), or delete an account at any time. Admin actions are audited to ~/.codeman/admin-audit.jsonl.--permission-mode auto (Anthropic's classifier-guarded mode); raw shell sessions, cron launchCommand, and skip-permissions require an explicit per-user grant.⚠️ This separates workspaces; it does not sandbox users from each other. Every session runs as the same OS account, so a determined user's agent can still reach another user's files. For real isolation, pair users with Docker cases or run separate instances under separate OS accounts. See
docs/multi-user-plan.mdand the multi-user section ofdocs/security-architecture.md.
Access Codeman from your phone or any device outside your local network using a free Cloudflare quick tunnel — no port forwarding, no DNS, no static IP required.
Browser (phone/tablet) → Cloudflare Edge (HTTPS) → cloudflared → localhost:3000
Prerequisites: Install cloudflared and set CODEMAN_PASSWORD in your environment.
# Quick start
./scripts/tunnel.sh start # Start tunnel, prints public URL
./scripts/tunnel.sh url # Show current URL
./scripts/tunnel.sh stop # Stop tunnel
./scripts/tunnel.sh status # Service status + URL
The script auto-installs a systemd user service on first run. The tunnel URL is a randomly generated *.trycloudflare.com address that changes each time the tunnel restarts.
# Enable as a persistent service
systemctl --user enable codeman-tunnel
loginctl enable-linger $USER
# Or via the Codeman web UI: App Settings → System → Remote access → Cloudflare Tunnel
admin or CODEMAN_USERNAME)codeman_session cookie (24h TTL, auto-extends on activity)Always set CODEMAN_PASSWORD before exposing via tunnel — without it, anyone with the URL has full access to your sessions.
Typing a password on a phone keyboard is terrible. Codeman solves this with ephemeral single-use QR tokens — scan the code on your desktop, and your phone is instantly authenticated. No password prompt, no typing, no clipboard.
Desktop displays QR → Phone scans → GET /q/Xk9mQ3 → Server validates
→ Token atomically consumed (single-use) → Session cookie issued → 302 to /
→ Desktop notified: "Device authenticated via QR" → New QR auto-generated
Someone who only has the bare tunnel URL (without the QR) still hits the standard password prompt. The QR is the fast path; the password is the fallback.
The server maintains a rotating pool of short-lived, single-use tokens. Each token consists of a 256-bit secret (crypto.randomBytes(32)) paired with a 6-character base62 short code used as an opaque lookup key in the URL path. The QR code encodes a URL like https://abc-xyz.trycloudflare.com/q/Xk9mQ3 — the short code is a pointer, not the secret itself, so it never leaks through browser history, Referer headers, or Cloudflare edge logs.
Every 60 seconds, the server automatically rotates to a fresh token. The previous token remains valid for a 90-second grace period to handle the race where you scan right as rotation happens — after that, it's dead. Each token is single-use: the moment a phone successfully scans it, the token is atomically consumed and a new one is immediately generated for the desktop display.
The design is informed by "Demystifying the (In)Security of QR Code-based Login" (USENIX Security 2025), which found 47 of the top-100 websites vulnerable to QR auth attacks due to 6 critical design flaws across 42 CVEs. Codeman addresses all six:
| USENIX Flaw | Mitigation |
|---|---|
| Flaw-1: Missing single-use enforcement | Token atomically consumed on first scan — replays always fail |
| Flaw-2: Long-lived tokens | 60s TTL with 90s grace, auto-rotation via timer |
| Flaw-3: Predictable token generation | crypto.randomBytes(32) — 256-bit entropy. Short codes use rejection sampling to eliminate modulo bias |
| Flaw-4: Client-side token generation | Server-side only — tokens never leave the server until embedded in the QR |
| Flaw-5: Missing status notification | Desktop toast: "Device [IP] authenticated via QR (Safari). Not you? [Revoke]" — real-time QRLjacking detection |
| Flaw-6: Inadequate session binding | IP + User-Agent stored for audit. Manual session revocation via API. HttpOnly + Secure + SameSite=lax cookies |
Short codes are stored in a Map<shortCode, TokenRecord>. Validation uses Map.get() — a hash-based O(1) lookup that reveals nothing about the target string through response timing. There is no character-by-character string comparison anywhere in the hot path, eliminating timing side-channel attacks entirely.
QR auth has its own rate limiting, completely independent from password auth:
The URL is kept deliberately short (/q/ path + 6-char code = ~53-56 total characters) to target QR Version 4 (33x33 modules) instead of Version 5 (37x37). Smaller QR codes scan faster on budget phones — modern devices read Version 4 in 100-300ms. The /q/ prefix saves 7 bytes compared to /qr-auth/, which alone is the difference between QR versions.
The QR display auto-refreshes every 60 seconds via SSE with the SVG embedded directly in the event payload (~2-5KB) — no extra HTTP fetch, sub-50ms refresh. A countdown timer shows time remaining. A "Regenerate" button instantly invalidates all existing tokens and creates a fresh one (useful if you suspect the QR was photographed).
When someone authenticates via QR, the desktop shows a notification toast with the device's IP and browser — if it wasn't you, one click revokes all sessions.
| Threat | Why it doesn't work |
|---|---|
| QR screenshot shared | Single-use: consumed on first scan. 60s TTL: expired before the attacker can act. Desktop notification alerts you immediately. |
| Replay attack | Atomic single-use consumption + 60s TTL. Old URLs always return 401. |
| Cloudflare edge logs | Short code is an opaque 6-char lookup key, not the real 256-bit token. Single-use means replaying from logs always fails. |
| Brute force | 56.8 billion combinations, ~2 valid at any time, dual-layer rate limiting blocks well before statistical feasibility. |
| QRLjacking | 60s rotation forces real-time relay. Desktop toast provides instant detection. Self-hosted single-user context makes phishing implausible. |
| Timing attack | Hash-based Map lookup — no string comparison timing leak. |
| Session cookie theft | HttpOnly + Secure + SameSite=lax + 24h TTL. Manual revocation at POST /api/auth/revoke. |
| Platform | Model | Comparison |
|---|---|---|
| Discord | Long-lived token, no confirmation, repeatedly exploited | Codeman: single-use + TTL + notification |
| WhatsApp Web | Phone confirms "Link device?", ~60s rotation | Comparable rotation; WhatsApp adds explicit confirmation (acceptable tradeoff for single-user) |
| Signal | Ephemeral public key, E2E encrypted channel | Stronger crypto, but exploited by Russian state actors in 2025 via social engineering despite it |
Full design rationale, security analysis, and implementation details:
docs/qr-auth-plan.md
By default Codeman launches sessions with --dangerously-skip-permissions, so the web UI is by design a remote-code-execution surface for whoever can reach it — the whole security model exists to control who that is. (The startup permission mode is configurable; see below.) Recent hardening (v0.9.0 + v0.9.5) closes the browser-driven attack paths that bite self-hosted dev tools. Full model: docs/security-architecture.md. Found a vulnerability? See SECURITY.md for private disclosure and the list of known limitations.
127.0.0.1, reachable only from the same machine, so the no-password default is safe out of the box (the guided installer asks about network access and configures the binding + password for you). Binding a non-loopback host without CODEMAN_PASSWORD starts but prints a loud warning with three concrete fixes (set a password, loopback + an authenticated tunnel, or explicitly acknowledge with --allow-unauthenticated-network)CODEMAN_USERNAME (default admin) / CODEMAN_PASSWORD. Success issues an opaque 256-bit codeman_session cookie (randomBytes(32)) — validated server-side, not client-signed, so it can't be forged offline (24h TTL, auto-extend, device-context audit log)429 with Retry-After (15-min decay). A valid cookie or correct password recovers immediately even while an attacker hammers the same IP — important because all tunnel traffic shares one loopback IP. QR auth has its own separate limiter--dangerously-skip-permissions is only the default. App Settings → Agents & CLIs → Claude → Startup Mode can switch new sessions to Anthropic's classifier-guarded auto mode (low-prompt, needs Claude Code 2.1.207+), normal prompting, or an explicit allowed-tools list. In multi-user mode, non-granted users are forced to auto, and shell sessions / skip-permissions require an explicit per-user grantThese run for every request — before auth, even on the default no-password loopback install:
127.0.0.1 is rejected with 403 host not allowed before any handler runs. Allowed: localhost, any IP literal, the bind host, .ts.net / .trycloudflare.com / .cfargotunnel.com, the active managed tunnel, and CODEMAN_ALLOWED_HOSTS (add custom reverse-proxy domains here — comma-separated; exact host or leading-dot .suffix for subdomains)POST/PUT/PATCH/DELETE) the Origin must pass the same allowlist, else 403 cross-site request blocked. A missing Origin is allowed (so curl, the CLI, and Claude Code hooks keep working); only a present-but-foreign or opaque null origin is rejectedtext/plain bodies. The global parser no longer JSON-parses text/plain, closing the CORS "simple request" CSRF vector where a cross-site fetch could smuggle JSON into a write route with no preflight4003 on failure (anti-CSWSH)CLAUDE_CODE_* / OPENCODE_* / CODEX_* / ANTIGRAVITY_* / GEMINI_* / GOOGLE_* / PI_* env-prefix allowlist gates which settings each CLI can receiverealpath before boundary checks (no TOCTOU); .., absolute paths, and symlinks resolving outside the working dir are rejected. Caps: 10 MB text preview / 50 MB raw & download; /api/download blocklists sensitive paths (.env, *credentials*, ~/.ssh/, .aws/credentials). SVG/HTML is served octet-stream + nosniff + attachment so it downloads rather than executesContent-Security-Policy (default-src 'self', every exception enumerated), X-Content-Type-Options: nosniff, X-Frame-Options: SAMEORIGIN, HSTS over HTTPS, and CORS reflected only for localhost / 127.0.0.1 / ::1overrides; lockfile integrity is checked on every commit/PR (all entries resolve to registry.npmjs.org with sha512 hashes). Public assets are NUL-byte-scanned and node --check-validated in CICODEMAN_INSTANCE scopes both the tmux socket (-L codeman-<name>) and data dir (~/.codeman-<name>) so two instances never attach each other's live sessionsMobile login uses single-use, 60-second QR tokens — see QR Code Authentication above for the full design (it addresses all 6 flaws from USENIX Security 2025's QR-login study).
codeman tui)A full-screen dashboard for your sessions, in the terminal. Same states as the web UI, because it is a client of the same server:
codeman tui # the dashboard
codeman tui --list # numbered session list, then exit (scriptable)
codeman tui 2 # attach straight to session 2 of that list
Sessions are grouped NEEDS YOU → WORKING → IDLE → RECENT, longest-waiting first. ↑↓/j/k select, 1-9 and [/] switch between sessions, Enter attaches into the tmux pane (F1 to come back). Inside a pane the bar across the top keeps the session strip visible and Alt+1-Alt+9 switch without leaving. y/n/digit answer a pending permission dialog right from the list, p sends a one-line prompt, n starts a session and opens straight into it, x kills one (y confirms), / searches, g shows the away digest, ? is help, q quits. Below 72 columns it drops the preview pane and becomes a single-column list, so it stays usable in Termius on a phone. With no server running it still starts in attach-only degraded mode.
The web UI remains the primary surface; see docs/tui.md for the full guide.
Ctrl bindings also accept Cmd on macOS.
| Shortcut | Action |
|---|---|
Ctrl/Cmd+W | Kill active session |
Ctrl/Cmd/Option+K | Find open session or start a new one |
Ctrl/Cmd+Tab | Next session |
Alt/Option+[ / Alt/Option+] | Previous / next session |
Alt/Option+1-Alt/Option+9 | Switch to tab N (physical keys, so macOS Option layouts work) |
Alt/Option+B | Collapse / expand the session sidebar (sidebar layout only) |
Ctrl+Shift+{ / Ctrl+Shift+} | Move active tab left / right |
Ctrl/Cmd+C | Copy selection, or interrupt when nothing is selected |
Ctrl+Shift+C | Copy selection (never interrupts) |
Ctrl/Cmd+V | Paste, or upload a clipboard image and paste its path |
Ctrl/Cmd+L | Clear terminal |
Ctrl+Shift+R | Restore terminal size |
Ctrl+Shift+V | Toggle voice input |
Ctrl/Cmd + / - | Font size |
Ctrl/Cmd+? | Keyboard help |
Shift+Enter | Insert newline (sent to terminal) |
Escape | Close panels & modals |
For AI agents and automation that control Codeman without a browser: an agent that spins up worker sessions, a CI bot, or Claude Code running inside a Codeman session orchestrating other sessions. Everything the UI does is HTTP + a CLI, so an agent can do it too.
Everything in this section also ships as a Claude Code skill in skills/codeman. Install it once and you never paste API docs into a prompt again. You ask for what you want in plain English, and the agent already sitting inside a Codeman session loads the recipes and drives the API itself.
| How | Command | Scope |
|---|---|---|
| Skills CLI | npx skills add Ark0N/Codeman --skill codeman -g | Global, works for any skills-aware agent |
| Claude Code plugin | /plugin marketplace add Ark0N/Codeman then /plugin install codeman@codeman | Global, through Claude Code's plugin manager; /plugin update codeman follows releases. Pick this OR a codeman skill install, not both: a Claude Code with both lists the skill twice (codeman and codeman:codeman) |
| Bundled CLI | codeman skill install | Global (~/.claude/skills/codeman), for npm installs that never cloned the repo |
| Bundled CLI | codeman skill install --case <name> | One case only |
| Web UI | App Settings → Agents & CLIs → Claude → Agent Skill | Auto-injects into each case on Claude session create (agentSkillEnabled, SYNCED, default off) |
codeman skill uninstall [--case <name>] reverses the CLI installs, and never touches a skills/codeman you wrote yourself.
That is the entire interface. No curl, no endpoint names, no session ids. These prompts work as written:
| You say | The skill does |
|---|---|
| "What sessions are running right now?" | Lists them with name, mode and status. Read-only, safe to ask anytime. |
"Start a shell worker on the myapp case, run the test suite, tell me if it passes." | Spawns, waits on a split completion marker, reads back the exit code, cleans up. |
| "Spin up 3 workers for lint, typecheck and tests. Run them in parallel, report failures." | The fan-out flow: one session per task, all started first, then gathered as each finishes. |
"Have a claude worker on refactor-auth summarize src/session.ts, then close it." | Spawns, runs the readiness ladder (first-run trust dialog included), send-and-wait, reads the clean transcript answer, deletes. |
| "Watch session w4 and tell me if it gets stuck on a permission prompt." | Blocks on the blocked signal and surfaces the question to you. It never answers another session's prompt itself. |
The agent deletes every session it started. Watch the tabs appear and disappear in the dashboard while it works.
You: spin up 3 shell workers, run lint / typecheck / the frontend syntax check in parallel, and tell me which failed.
lint -> 9f2d8e5f dispatched
typecheck -> aff9c691 dispatched 3 tabs appear in the dashboard
syntax -> be9f1f15 dispatched
lint DONE_lint_17909 rc=0
typecheck DONE_typecheck_3409 rc=0 gathered as each one finishes
syntax DONE_syntax_18501 rc=0
deleted 9f2d8e5f, aff9c691, be9f1f15 tabs disappear
Those DONE_<task>_<random> strings are the skill's split marker trick, and they are why the fan-out is reliable on hook-less shell sessions: the typed line contains ${M}_17909, so only the command's real output ever contains DONE_17909. An unsplit marker would match the echo of your own keystrokes before the command had even run.
| File | Contents |
|---|---|
SKILL.md | Safety rules, the ready-made fast path (spawn N workers, task them, collect), and the verb index. Always loaded. |
reference/verbs.md | The 14 verbs in detail: readiness, send-and-wait, markers, interrupts, cleanup. On demand. |
reference/recipes.md | 6 worked multi-worker flows (fan-out, blocked-worker watch, messaging fan-out). On demand. |
reference/endpoints.md | Full endpoint tables, error codes, per-mode signal table, capacity limits. On demand. |
reference/messaging.md | Talking to claude workers directly via Claude Code cross-session messaging. On demand. |
Every recipe in there was verified against a live server, and the comments record the failure modes that were measured rather than guessed.
CODEMAN_MUX unset) the skill refuses to act and does not guess an API URL, so a global install costs an unrelated Claude Code session nothing.⚠️ Turning agentSkillEnabled back off does not remove already-injected copies (a create-time sweep would yank the skill out from under other live sessions sharing that .claude/ dir). Remove them per case with codeman skill uninstall --case <name>.
The rest of this section is the manual path: the same operations as raw HTTP, for a CI bot, a shell script, or any agent without skill support.
When a CLI runs in a Codeman-managed session, these environment variables are set — read them instead of hardcoding anything:
| Variable | Meaning |
|---|---|
CODEMAN_MUX=1 | You're in a managed tmux session. Never tmux kill-session / pkill claude / pkill tmux — you'll kill yourself or a sibling. |
CODEMAN_API_URL | Base URL of the API (e.g. https://127.0.0.1:3000). Use it for every call below. |
CODEMAN_SESSION_ID | Your own session id. Use it to avoid acting on yourself. |
CODEMAN_HOOK_SECRET_FILE | Path to the hook secret (required on /api/hook-event while a managed tunnel is up). |
\r. Programmatic input is sent as literal text, and Enter fires only when the input contains a carriage return: {"input":"run tests\r"}. Without the \r the text sits on the session's prompt unsubmitted (and a combined wait runs its full timeout on a turn that never started). Embedded newlines are stripped rather than rejected, so "echo A\necho B\r" runs the joined command echo Aecho B: send one line per call.clientId and a monotonic per-session seq on POST …/input. The server de-duplicates, so a retry after a dropped connection can't double-deliver a prompt.CODEMAN_PASSWORD is set, send HTTP Basic auth (user admin or CODEMAN_USERNAME) or a codeman_session cookie. The default loopback install is passwordless. A missing Origin header is allowed, so plain curl works; cross-site browser origins are rejected (CSRF guard). ⚠️ A 401 replies with the bare string Unauthorized, not the JSON envelope, so piping it into jq throws a parse error instead of showing the failure: check the status before parsing.{ "success": true, "data": … } (errors: { "success": false, "error", "errorCode" }). A few legacy GETs return bare bodies — handle both (body.data ?? body)./api/v1/* is a stable alias of /api/*.200 and wait.timedOut: true when nothing happened in time, so loop over short waits (60s is the default) rather than issuing one long call, because tunnels cut idle connections. wait.timeoutMs tells you the timeout the server actually applied after clamping (600s ceiling).claude sessions emit stop and blocked. Those two come from Claude Code hooks; shell and the external CLIs (opencode/codex/gemini/antigravity/pi) accept only idle, working and exit. Asking for stop explicitly on those is a 400; omitting until is always safe. ⚠️ On a shell session idle fires once, at startup, and never again, so send-and-wait there can only time out; synchronize hook-less sessions with a wait-output marker.claude sessions emit stop and blocked. Those two come from Claude Code hooks; shell and the external CLIs (opencode/codex/gemini/antigravity/omp) accept only idle, working and exit. Asking for stop explicitly on those is a 400; omitting until is always safe. ⚠️ On a shell session idle fires once, at startup, and never again, so send-and-wait there can only time out; synchronize hook-less sessions with a wait-output marker.8. Nothing reports "ready", so wait for it explicitly. A new session answers {"signal":"exit","immediate":true} (that means not started, not crashed) until its PID exists, and a claude worker in a fresh case then sits on the CLI's trust dialog. Prompt it there and the wait resolves on idle in ~2s looking exactly like a finished turn, while the text sits stuck in the dialog. Recipe 2b below is the sequence that avoids it.# CODEMAN_API_URL is auto-set inside every Codeman session, correct scheme included.
# The fallback below fits a stock install; on a --https install set the https:// URL
# yourself and add -k to each curl (self-signed cert).
API="${CODEMAN_API_URL:-http://127.0.0.1:3000}"
# (add -u admin:"$CODEMAN_PASSWORD" to each call if a password is set)
# 1. See what's running
curl -s "$API/api/sessions" | jq '.data // .'
# 2. Spin up a worker session (a "case" = named working dir)
curl -s -X POST "$API/api/quick-start" \
-H 'Content-Type: application/json' \
-d '{"caseName":"refactor-auth","mode":"claude","effort":"high"}' | jq
# 2b. Wait until that worker is actually READY (see rule 8): composer marker first,
# first-run trust dialog only as the fallback. (Probing trust first and sending
# a blind Enter misfires on re-runs: the dialog text stays in the buffer forever,
# so the probe matches stale text and the Enter lands in a ready composer.)
# Match single tokens: TUI text can reach the matcher without its spaces.
until [ "$(curl -s "$API/api/sessions/$SID" | jq '.data.pid')" != null ]; do sleep 1; done
R=$(curl -sG "$API/api/sessions/$SID/wait-output" --data-urlencode 'match=bypass' \
--data-urlencode 'from=buffer' --data-urlencode 'timeout=5000')
if ! jq -e '.data.wait.matched' <<<"$R" >/dev/null; then
T=$(curl -sG "$API/api/sessions/$SID/wait-output" --data-urlencode 'match=trust' \
--data-urlencode 'from=buffer' --data-urlencode 'timeout=2000')
jq -e '.data.wait.matched' <<<"$T" >/dev/null && \
curl -s -X POST "$API/api/sessions/$SID/input" -H 'Content-Type: application/json' \
-d '{"input":"\r","useMux":true}' # accept the first-run trust dialog
curl -sG "$API/api/sessions/$SID/wait-output" --data-urlencode 'match=bypass' \
--data-urlencode 'from=buffer' --data-urlencode 'timeout=45000' >/dev/null
fi
# 3. Send a prompt into a session (exactly-once: clientId + seq)
curl -s -X POST "$API/api/sessions/$SID/input" \
-H 'Content-Type: application/json' \
-d '{"input":"Run the test suite and summarize failures\r","useMux":true,"clientId":"agent-1","seq":1}'
# 4. Send a prompt and BLOCK until that turn is done (registers the wait before
# writing, so it can't answer with the previous turn's idle state)
curl -s -X POST "$API/api/sessions/$SID/input" \
-H 'Content-Type: application/json' \
-d '{"input":"Run the test suite and summarize failures\r","useMux":true,
"clientId":"agent-1","seq":2,"wait":"stop,exit","waitTimeout":60000}' \
| jq '.data.wait' # -> {"signal":"stop","timedOut":false,"waitedMs":41230,...}
# (`stop` is the definitive end-of-turn hook. Adding `idle` makes it resolve on a
# spinner pause too, and on anything that redraws a ❯ prompt — like a dialog.)
# 4b. Timed out? That's a 200, not a failure. Loop over short waits.
curl -s "$API/api/sessions/$SID/wait?until=stop,exit&timeout=60000" | jq '.data.wait'
# 4c. Or wait for a marker in the output (works for shell sessions too).
# ⚠️ Unique per call (tmux repaints replay old screen text), and SPLIT so the
# typed line never contains it: your own keystrokes echo into the output
# stream, so an unsplit marker matches before the command has run. from=buffer
# catches a marker that printed before the wait landed.
N=$RANDOM
curl -s -X POST "$API/api/sessions/$SID/input" -H 'Content-Type: application/json' \
-d "{\"input\":\"M=DONE; npm test; echo \${M}_$N rc=\$?\r\",\"useMux\":true}"
curl -sG "$API/api/sessions/$SID/wait-output" \
--data-urlencode "match=DONE_$N" --data-urlencode 'from=buffer' \
--data-urlencode 'timeout=60000' | jq '.data.wait'
# 5. Read the terminal back. ⚠️ Use terminal?tail=, NOT /output: the latter's
# textOutput is empty for every tmux-backed (i.e. every interactive) session.
# tail counts BYTES, and what comes back is terminal data, ANSI included.
curl -s "$API/api/sessions/$SID/terminal?tail=8000" | jq -r '.data.terminalBuffer'
# 6. Stream live events (session output, agent activity, status)
curl -sN "$API/api/events" # Server-Sent Events
# 7. Schedule recurring work (cron-style job)
curl -s -X POST "$API/api/cron/jobs" \
-H 'Content-Type: application/json' \
-d '{"name":"nightly-deps","agentType":"claude","workingDir":"/home/me/proj",
"promptMode":"inline_text","promptText":"Update dependencies and open a PR",
"inputMode":"typed","scheduleType":"daily","dailyTime":"03:00",
"enabled":true,"concurrencyPolicy":"warn_only"}' | jq
# 8. Inspect background sub-agents and their transcripts
curl -s "$API/api/subagents" | jq '.data // .'
curl -s "$API/api/subagents/$AID/transcript" | jq -r '.data // .'
# 9. Whole-system snapshot (sessions, settings, respawn, stats)
curl -s "$API/api/status" | jq
The same operations are available as commands (codeman <cmd>, aliases in parentheses) — handy from a shell tool inside a session:
codeman session start -d /path/to/repo # (s) start a session
codeman session list # list sessions
codeman session logs <id> # tail output
codeman task add "fix the failing test" # (t) queue a task
codeman attach <path> # show an attachment card for a local file
codeman tui --list # numbered session list (plain text when piped)
codeman tui 3 # attach to session 3 of that list
Codeman registers Claude Code hooks that POST /api/hook-event (permission_prompt, idle_prompt, stop, task_completed, …) so the dashboard reacts in real time. This endpoint is auth-exempt on loopback but, under a managed tunnel, requires the X-Codeman-Hook-Secret header (read it from $CODEMAN_HOOK_SECRET_FILE). You normally don't call this by hand — Codeman wires it up — but it's how the autonomy layers "see" what the agent is doing.
Full endpoint list and request/response shapes follow.
REST over Fastify — ~200 handlers across 21 route modules, plus an SSE stream and a WebSocket terminal channel. All responses use the ApiResponse<T> envelope ({success, data} / {success, error, errorCode}); /api/v1/* is a stable alias. A representative subset:
| Method | Endpoint | Description |
|---|---|---|
GET | /api/sessions | List all |
POST | /api/quick-start | Create case + start session ({caseName?, mode?, effort?, envOverrides?}) |
POST | /api/sessions/:id/input | Send input ({input, useMux?, clientId?, seq?, wait?, waitTimeout?}: clientId+seq = exactly-once; wait blocks until the turn ends) |
GET | /api/sessions/:id/terminal | Read terminal output (?tail=<bytes>, ?full=1); the read path for interactive sessions |
GET | /api/sessions/:id/output | Parsed one-shot output (textOutput is empty for tmux-backed sessions) |
GET | /api/sessions/:id/wait | Block until a signal fires (?until=stop,idle,exit&timeout=&fresh=); a timeout is a 200 |
GET | /api/sessions/:id/wait-output | Block until a literal string appears (?match=&nocase=&from=now|buffer&timeout=) |
GET | /api/sessions/unified | Unified live + history list (Session Manager) — ?q=&limit= |
POST | /api/sessions/:id/pin | Pin/unpin in the Session Manager ({pinned}) |
PUT | /api/session-order | Sync tab order across devices ({order: [ids]}) |
DELETE | /api/sessions/:id | Delete session |
| Method | Endpoint | Description |
|---|---|---|
POST | /api/sessions/:id/respawn/enable | Enable with config + timer |
POST | /api/sessions/:id/respawn/stop | Stop controller |
PUT | /api/sessions/:id/respawn/config | Update config |
| Method | Endpoint | Description |
|---|---|---|
POST | /api/orchestrator/start | Start orchestration from a goal |
POST | /api/orchestrator/approve | Approve the generated plan |
GET | /api/orchestrator/status | Current phase + progress |
POST | /api/orchestrator/stop | Stop and clean up |
| Method | Endpoint | Description |
|---|---|---|
GET / POST | /api/cron/jobs | List / create cron jobs |
PUT / DELETE | /api/cron/jobs/:id | Update / delete a job |
PUT | /api/cron/jobs/:id/enabled | Enable / disable |
POST | /api/cron/jobs/:id/run | Run now |
GET | /api/cron/jobs/:id/runs | Run history |
| Method | Endpoint | Description |
|---|---|---|
GET | /api/subagents | List all background agents |
GET | /api/subagents/:id | Agent info and status |
GET | /api/subagents/:id/transcript | Full activity transcript |
DELETE | /api/subagents/:id | Kill agent process |
| Method | Endpoint | Description |
|---|---|---|
GET | /api/events | SSE stream |
GET | /api/status | Full app state |
POST | /api/hook-event | Hook callbacks |
GET | /api/system/update/check | Check for a new release |
POST | /api/system/update | Self-update (git-clone installs) |
POST | /api/clipboard | Push text to all connected browsers ({text}) |
GET | /api/sessions/:id/run-summary | Timeline + stats |
Building something on top of Codeman?
docs/extending-codeman.mdis the integration guide: render your own UI as a tab, subscribe to the SSE event stream to react when an agent needs you, drive Codeman from a script, and the traps worth knowing before you start. Codeman has no plugin runtime on purpose, so an integration is just your own process talking HTTP.
flowchart TB
subgraph Codeman["CODEMAN"]
subgraph Frontend["Frontend Layer"]
UI["Web UI<br/><small>xterm.js + Agent Windows</small>"]
API["REST API<br/><small>Fastify</small>"]
SSE["SSE Events<br/><small>/api/events</small>"]
end
subgraph Core["Core Layer"]
SM["Session Manager"]
S1["Session (PTY)"]
S2["Session (PTY)"]
RC["Respawn Controller"]
ORC["Orchestrator Loop"]
end
subgraph Detection["Detection Layer"]
SW["Subagent Watcher<br/><small>~/.claude/projects/*/subagents</small>"]
TW["Team Watcher<br/><small>~/.claude/teams/*</small>"]
end
subgraph Persistence["Persistence Layer"]
SCR["Mux Manager<br/><small>(tmux)</small>"]
SS["State Store<br/><small>state.json</small>"]
end
subgraph External["External"]
CLI["AI CLI<br/><small>Claude Code / OpenCode / Codex / Antigravity / Gemini / Pi</small>"]
CLI["AI CLI<br/><small>Claude Code / OpenCode / Codex / Antigravity / Gemini / OMP</small>"] BG["Background Agents<br/><small>(Task tool)</small>"]
end
end
UI <--> API
API <--> SSE
API --> SM
SM --> S1
SM --> S2
SM --> RC
SM --> ORC
SM --> SS
S1 --> SCR
S2 --> SCR
RC --> SCR
ORC --> SCR
SCR --> CLI
SW --> BG
SW --> SSE
TW --> SSE
npm install
npx tsx src/index.ts web # Dev mode
npm run build # Production build
npm test # Run tests (same suite CI runs; browser/mobile/perf suites have their own commands)
See CLAUDE.md for full documentation.
Questions, setup help, and ideas live in GitHub Discussions: the Q&A section answers the most common ones (phone access, overnight runs, updating), and the roadmap gets decided in Ideas. Bugs go to issues; reports usually get a response within a day, and every release credits its reporters and contributors by name. Want to contribute? CONTRIBUTING.md has the map: skins, translations, and docs make great first PRs, and bigger features start life as a Discussion. And if you're proud of your rig, post it in Show and tell.
The codebase went through a comprehensive 7-phase refactoring that eliminated god objects, centralized configuration, and established modular architecture:
| Phase | What changed | Impact |
|---|---|---|
| Performance | Cached endpoints, SSE adaptive batching, buffer chunking | Sub-16ms terminal latency |
| Route extraction | server.ts split into 15 domain route modules + auth middleware + port interfaces | −67% server.ts LOC (6,736 → 2,254) |
| Domain splitting | types.ts → 16 domain files, ralph-tracker → 7 files, respawn-controller → 5 files, session → 6 files | No more god files |
| Frontend modules | app.js → 18 extracted modules across infra, domain & feature layers | app.js core down to ~3.4K LOC |
| Config consolidation | ~70 scattered magic numbers → 10 domain-focused config files | Zero cross-file duplicates |
| Test infrastructure | Shared mock library, 12 route test files, consolidated MockSession | Testable route handlers via app.inject() |
Full details: docs/archive/code-structure-findings.md
xterm-zerolag-inputInstant keystroke feedback overlay for xterm.js. Eliminates perceived input latency over high-RTT connections by rendering typed characters immediately as a pixel-perfect DOM overlay. Zero dependencies, 6.1 kB gzipped, configurable prompt detection, CJK/emoji wide-character support, full state machine with 175 tests.
npm install xterm-zerolag-input
Codeman follows SemVer. What the version number actually
commits to — and what counts as internal (the HTTP/SSE API, on-disk state,
experimental features) — is spelled out in
docs/versioning-policy.md. If you script against
the HTTP API, pin to an exact version.
MIT — see LICENSE
Track sessions. Visualize agents. Control respawn. Let it run while you sleep.
If Codeman saves you time, a star helps other people find it.
Bug reports and feature ideas are welcome in Issues.
name: codeman
description: >-
Drive Codeman, the session manager this agent is running inside, over its HTTP API:
list sessions, start worker sessions, send them prompts, block until they finish
(wait / wait-output / send-and-wait), read their output, and clean up; where
available, message claude workers directly (Claude Code cross-session messaging).
Use when asked to orchestrate or parallelize work across Codeman sessions, watch
another session, or start and manage workers. Only usable inside a Codeman-managed
session (CODEMAN_MUX=1); refuse to act otherwise.You are an agent running inside a Codeman-managed terminal session. Codeman is the server that spawned you; its HTTP API can start, prompt, watch, and delete other sessions.
Read as far as your job needs and no further. §0 is the bootstrap, run once. §1 is the whole fast path: spawn N workers, task them, collect answers. If §1 covers your job, run it and stop there. The sections after it are for jobs it does not cover, and reading them to be thorough is the main reason a ten-second run takes minutes. §2 is the verb table when your job is a different one. §3 and §4 are the rules; §6 is setup and credentials, which you only need when something 401s.
Everything else loads on demand, and is meant to be opened at one section, not read through: the verbs in detail (the old §5) in reference/verbs.md, worked multi-worker flows in reference/recipes.md, endpoint tables and a symptom gallery in reference/endpoints.md, and direct messaging to claude workers in reference/messaging.md.
If CODEMAN_MUX is not 1, stop and say so. Do not guess an API URL; a server
you are not part of is not yours to drive.
⚠️ Your shell state does not survive between tool calls. Each Bash call starts a
fresh shell, so $API, $SELF, the CURL array and delete_session are all gone by
the next call, and $$ is a different pid. The filesystem does survive, so write
the preamble to a file once and source it afterwards, rather than re-pasting a
hundred-odd lines at the top of every call (a half-re-pasted preamble used to be the
single most likely way to break a run).
Codeman seeds the preamble file for you when it spawns a claude session (server 1.18.3+), so the bootstrap is usually nothing at all: these are the two lines every later call opens with, and your first REAL call performs them anyway:
. "${XDG_CACHE_HOME:-$HOME/.cache}/codeman-agent-$CODEMAN_SESSION_ID.sh" 2>/dev/null
[ "${CODEMAN_PREAMBLE:-}" = 1.22.0 ] || { echo "preamble missing or stale; run the full §0 block"; exit 1; }
⚠️ Never spend a Bash call on this check alone. §1's block opens with this same
loader, so when §1 is the job, start there: the check rides the spawn call for free,
and a standalone "preamble OK" call buys nothing while costing a full model turn
(measured live: a lone check plus the deliberation around it added ~6 s to a 28 s
two-worker run). §0 is done the moment any job call passes its opening check. Only
when a call reports missing or stale, run the full block below once — and run it
verbatim: paste it as-is, never re-type it, trim it, or "extract the parts you
need". A hand-assembled
preamble is the documented failure mode of this skill: one live run rebuilt it
"minimally" and lost the X-Codeman-Parent-Session header (every worker spawned with
no lineage arc in the web UI) and the fast-path functions (the spawn fell back to a
serial quick-start loop plus pid polls), turning a ten-second job into a fifty-second
one. If your harness directs temporary files into a scratchpad directory, that
directive covers task scratch, not this file: it is a per-session cache that every
later call re-sources by this exact path, so keep the path below. If you must relocate
it anyway, copy the block's content byte-for-byte unchanged and source your path in
every later call instead.
test "${CODEMAN_MUX:-}" = 1 || { echo "Not inside a Codeman-managed session; refusing to act."; exit 1; }
: "${CODEMAN_SESSION_ID:?CODEMAN_SESSION_ID not set}" "${HOME:?HOME not set}"
PRE="${XDG_CACHE_HOME:-$HOME/.cache}/codeman-agent-$CODEMAN_SESSION_ID.sh"
mkdir -p "$(dirname "$PRE")"
# Rewrite unless the file already ends with THIS version's stamp, so a stale or a
# half-written file self-heals here instead of costing you a round trip to rm it.
grep -qs '^CODEMAN_PREAMBLE=1.22.0$' "$PRE" || (umask 077; cat > "$PRE" <<'PREAMBLE'
# ---- Codeman agent preamble 1.22.0 (seeded by Codeman at session spawn; the SKILL.md §0 bootstrap rewrites it when missing or stale) ----
API="${CODEMAN_API_URL:?CODEMAN_API_URL not set; refusing to guess}"
SELF="${CODEMAN_SESSION_ID:?CODEMAN_SESSION_ID not set}"
# Credentials, cheapest first. Your session has usually INHERITED the server's
# CODEMAN_PASSWORD already (§6 explains why, and what to do when it has not);
# the data dir's .env is the documented fallback, the same one `codeman attach`
# reads. The data dir is wherever the hook-secret file lives. Values may be
# quoted or `export`-prefixed.
ENV_FILE="${CODEMAN_HOOK_SECRET_FILE:+${CODEMAN_HOOK_SECRET_FILE%hook-secret}.env}"
envval() { sed -n "s/^\(export \)\{0,1\}$1=//p" "$ENV_FILE" | tail -1 | sed 's/^"\(.*\)"$/\1/; s/^'\''\(.*\)'\''$/\1/'; }
if [ -z "${CODEMAN_PASSWORD:-}" ] && [ -n "$ENV_FILE" ] && [ -f "$ENV_FILE" ]; then
CODEMAN_USERNAME=$(envval CODEMAN_USERNAME)
CODEMAN_PASSWORD=$(envval CODEMAN_PASSWORD)
fi
AUTH=(); [ -n "${CODEMAN_PASSWORD:-}" ] && AUTH=(-u "${CODEMAN_USERNAME:-admin}:$CODEMAN_PASSWORD")
# -k: harmless on http, required on https (self-signed cert).
# X-Codeman-Parent-Session: tags workers YOU spawn as your children, so the web UI can
# draw the lineage. Set once here and every present and future create call carries it;
# it is ignored on every other endpoint. Purely cosmetic (see §5.1) and it can never
# fail a spawn, so there is no case where you would want to leave it off.
# X-Codeman-Agent-Origin: marks a case directory a spawn CREATES as agent scratch, so the
# user can find and delete it long after your workers are gone (§5.14). Same deal: set
# once, cosmetic, never fails a spawn, and it labels only directories Codeman creates.
CURL=(curl -sk "${AUTH[@]}" -H "X-Codeman-Parent-Session: $SELF" -H "X-Codeman-Agent-Origin: codeman-skill")
CID=codeman-agent-1 # FIXED literal, never "agent-$$": see below
# Fail-CLOSED session delete. The DELETE lives INSIDE the guard on purpose: the older
# `is_self "$SID" || curl -X DELETE ...` shape failed OPEN, because an undefined
# is_self exits 127 and the `||` branch then ran the delete completely unguarded.
# Undefined delete_session is "command not found", which deletes nothing.
delete_session() {
local id="${1:-}"
[ -n "$id" ] || { echo "refusing: empty session id"; return 1; }
[ "${#SELF}" -ge 8 ] || { echo "refusing: \$SELF unset or too short to prove this is not me"; return 1; }
# ids appear in full AND 8-char form (Docker exports a truncated $SELF; mux names and
# UI surfaces carry 8-char ids), so compare by prefix in BOTH directions. Equality or
# a one-directional check each miss a real combination, and the miss deletes you.
case "$id" in "$SELF"*) echo "refusing: $id is me"; return 1 ;; esac
case "$SELF" in "$id"*) echo "refusing: $id is me"; return 1 ;; esac
"${CURL[@]}" -X DELETE "$API/api/v1/sessions/$id"
}
# ---- fast path: the four verbs, already written. §1 composes them. ----
_composer_up() { # <sid> <timeoutMs> -> "true"/"false". `shift+tab` is the one token
"${CURL[@]}" -G "$API/api/v1/sessions/$1/wait-output" \
--data-urlencode 'match=shift+tab' --data-urlencode 'from=buffer' \
--data-urlencode "timeout=$2" | jq -r '.data.wait.matched // false'
}
_dsh_up() { # <sid> <timeoutMs> -> "true"/"false". The DeepSeek Harness TUI's
# composer glyph. Override with DSH_READY_MARK for a profile that draws another one.
"${CURL[@]}" -G "$API/api/v1/sessions/$1/wait-output" \
--data-urlencode "match=${DSH_READY_MARK:-❯}" --data-urlencode 'from=buffer' \
--data-urlencode "timeout=$2" | jq -r '.data.wait.matched // false'
}
# ---- the workspace-trust dialog: READ the screen, never press Enter blind ----
# Claude Code 2.1.252 dropped the option numbers, REVERSED them, and highlights
# "No, exit" by default:
# Security guide
# ❯ No, exit
# Yes, I trust this folder
# Enter to confirm . Esc to cancel
# so the bare \r that answered the old layout now answers *exit* and the pane is
# dead (`status 1`) seconds after the spawn -- measured on a live 2.1.252 case.
# These two read the rendered pane and steer onto the trust option instead.
_trust_key() { # <sid> -> "confirm" | "move" | "" (nothing safe to press)
# full=1 returns the RENDERED pane; a claude pane keeps no tmux history, so that
# is the current frame rather than every repaint since launch. tail -1 anyway,
# because the freshest marked row is the only one still true.
"${CURL[@]}" -G "$API/api/v1/sessions/$1/terminal" --data-urlencode 'full=1' \
| jq -r '.data.terminalBuffer // empty' \
| sed -e "s/$(printf '\033')\[[0-9;?]*[a-zA-Z]//g" -e "s/$(printf '\033')[()][AB0]//g" \
| tr -d ' \t' | grep -i '❯[0-9.]*\(yes,itrustthisfolder\|no,exit\)' | tail -1 \
| sed -e 's/.*[Yy]es,.*/confirm/' -e 's/.*[Nn]o,.*/move/'
}
_accept_trust() { # <sid> -> 0 once it has answered the dialog, 1 if it could not
local sid="$1" k i=1
while [ "$i" -le 6 ]; do
k=$(_trust_key "$sid")
[ -n "$k" ] || return 1 # no dialog on screen, or a layout this cannot read
# A SEPARATE clientId for these keys. seq is monotonic per clientId, so
# spending prompt numbers here would make the next sendwait -- whose default
# seq is the epoch second -- look like a stale duplicate and vanish silently.
"${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" -H 'Content-Type: application/json' \
-d "$(jq -nc --arg k "$([ "$k" = confirm ] && printf '\r' || printf '\033[B')" \
--arg c "$CID-trust-$sid" --argjson s "$i" \
'{input:$k,useMux:true,clientId:$c,seq:$s}')" >/dev/null
[ "$k" = confirm ] && return 0
sleep 1; i=$((i+1)) # re-read: the arrow is CONFIRMED before Enter goes out
done
return 1
}
# spawn_worker <caseName> [mode] -> session id on stdout, diagnostics on stderr.
# quick-start AND readiness in one call, with a strict contract: NON-EMPTY stdout means
# a READY worker whose end-of-turn signal can be trusted -- a claude worker in a
# hook-carrying case, or a `deepseek` worker whose harness TUI drew its composer.
# Anything less is rc 1 with EMPTY stdout, and the half-spawned session is deleted here
# rather than handed back, because a worker that never drew its composer would eat the
# task prompt with its trust dialog. There is deliberately no pid poll: wait-output
# already blocks until the composer draws, and pid!=null proved startup, never readiness.
spawn_worker() {
local name="${1:?spawn_worker needs a case name}" mode="${2:-claude}" q sid cp r
# parentSessionId doubles the CURL header, so a spawn_worker copied off the shared
# curl (or a body someone rebuilt from this recipe) still carries its lineage.
# deepseek: ask for the same permission posture the Run button sends, because the
# harness's own default (`workspace-write`) still ASKS, and a worker that stops on
# an approval row is a worker no fan-out can finish. It is not an escalation --
# claude workers already spawn with permissions skipped, and in multi-user mode the
# server clamps this back to `workspace-write` for an owner without the grant.
# Spawn by hand (§5.1) when you want a worker that asks.
q=$("${CURL[@]}" -X POST "$API/api/v1/quick-start" -H 'Content-Type: application/json' \
-d "$(jq -nc --arg n "$name" --arg m "$mode" --arg p "$SELF" \
'{caseName:$n,mode:$m,parentSessionId:$p}
+ (if $m == "deepseek" then {deepSeekConfig:{permissionMode:"danger-full-access"}} else {} end)')")
sid=$(jq -r 'if .success then .data.sessionId else empty end' <<<"$q")
# NOT retryable in a loop: every quick-start failure code is terminal (§5.1).
[ -n "$sid" ] || { jq -c '{error,errorCode}' <<<"$q" >&2; return 1; }
if [ "$mode" = deepseek ]; then
# The one non-claude mode with REAL end-of-turn signals: its TUI reports
# idle/working/blocked to Codeman, so sendwait, until=stop and the Approvals
# Inbox all work here exactly as they do for claude. No hook file to vet
# (the bridge is env-injected, not a workspace file) and no trust dialog.
# ⚠️ Readiness is still not optional, and NOT interchangeable with the stop
# signal: the harness's boot report lands ~300ms BEFORE the composer paints
# (measured 2.26s vs 2.56s after spawn), so a sendwait fired straight after
# quick-start returns on that BOOT signal, reports a turn that never ran, and
# strands the prompt in a pane that was not yet taking input.
r=$(_dsh_up "$sid" 45000)
[ "$r" = true ] || { echo "dsh worker $sid never drew a composer: no pane-capable profile, a profile whose composer is not '${DSH_READY_MARK:-❯}' (set DSH_READY_MARK), or a harness that failed to boot -- check GET /api/v1/deepseek/status. Deleted it" >&2
delete_session "$sid" >/dev/null; return 1; }
printf '%s\n' "$sid"; return 0
fi
[ "$mode" = claude ] || { printf '%s\n' "$sid"; return 0; } # no other mode draws a composer to wait on
# The server installs hooks into every claude workspace now, so this grep normally
# passes; it stays because the install is gated on a setting the operator can turn
# off, remote sessions never get hooks, and a session created by an older server
# still has none. No marker means sendwait would false-resolve on flapping idle,
# possibly inside the user's REAL repo: refuse rather than run the job there.
cp=$(jq -r '.data.casePath // empty' <<<"$q")
grep -qs '/api/hook-event' "$cp/.claude/settings.local.json" || {
echo "case '$name' resolved to '$cp', which has no Codeman hooks (workspaceHooksEnabled off, remote, or an older server?): turn the setting on, or work §5.1+§5.5 by hand with markers" >&2
delete_session "$sid" >/dev/null; return 1; }
# Short composer wait FIRST, then the trust dialog: a case still showing the
# dialog can never pass the composer wait, so acting early keeps a cold case from
# paying the whole long wait before the fallback even runs (§5.2). A warm case
# matches in under a second and never reaches it, and _accept_trust returns in a
# blink when there is no dialog, so this costs nothing in the ordinary slow case.
r=$(_composer_up "$sid" 5000)
if [ "$r" != true ]; then
# Codeman answers this dialog itself and normally wins the race; this is the
# bounded fallback for when its 90 s window / 6-keystroke cap has run out.
_accept_trust "$sid"
r=$(_composer_up "$sid" 45000)
fi
[ "$r" = true ] || { echo "worker $sid never drew a composer; deleted it. Retry by hand via the §5.2 ladder (its billed stage-4 probe included)" >&2
delete_session "$sid" >/dev/null; return 1; }
printf '%s\n' "$sid"
}
# spawn_workers <caseName[:mode]>... -> one "<caseName> <sessionId>" line per worker, in
# order; the sessionId column is EMPTY for a spawn that failed (stderr has why).
# CONCURRENT: N workers cost about what one costs. Spawning them one Bash call at a time
# is the single biggest avoidable delay in this skill. A bare name is a claude worker;
# `beta:deepseek` makes that one a DeepSeek Harness worker, and a mixed fleet is one
# call. Case names must be UNIQUE: two workers in one case directory co-edit the same
# tree (§4), so a repeat is an error here, not a race (the mode never disambiguates two
# workers, since they would still share the directory).
spawn_workers() {
local d spec n m i=0
[ "$#" -gt 0 ] || { echo "spawn_workers: no case names given" >&2; return 1; }
[ -z "$(printf '%s\n' "$@" | sed 's/:.*//' | sort | uniq -d)" ] || { echo "spawn_workers: duplicate case names" >&2; return 1; }
d=$(mktemp -d "${TMPDIR:-/tmp}/codeman-spawn.XXXXXX") || return 1
for spec in "$@"; do
n=${spec%%:*}; m=${spec#*:}; [ "$m" = "$spec" ] && m=claude
( spawn_worker "$n" "$m" > "$d/$i" ) & i=$((i+1))
done
wait
i=0; for spec in "$@"; do printf '%s %s\n' "${spec%%:*}" "$(cat "$d/$i" 2>/dev/null)"; i=$((i+1)); done
rm -rf "$d"
}
# sendwait <sid> <prompt> [seq] -> blocks until that worker's turn ENDS (~10 min ceiling
# across its two waits). One billed turn. The \r and the per-worker clientId are applied
# here, which is why you never hand-build this body. seq defaults to the CURRENT EPOCH
# SECOND so that every new prompt is a new frame: the server drops any (clientId,seq)
# pair it has already applied, so a fixed default would make every later prompt to that
# worker a silent no-op that still "succeeds" and reports the previous turn's state.
# Pass seq explicitly for exactly one reason: resending a possibly-delivered frame as a
# deliberate duplicate, at the SAME number (§5.3).
# Delivery is SELF-HEALING: an Ink repaint occasionally eats the Enter, leaving the
# typed prompt stranded on the composer while a long wait runs its whole timeout
# (observed live). So the first wait is short; on its timeout a bare \r goes out (the
# missing Enter when the prompt is stranded, a no-op when the turn is genuinely
# running), then the ORIGINAL frame is resent unchanged, which the server takes as a
# tagged duplicate: it re-waits without retyping (§5.3). Trustworthy for a worker
# spawn_worker handed back -- claude (hooks vetted) or deepseek (status bridge) --
# and for those only. Hook-less workspaces and the other modes resolve on flapping
# idle: markers instead (§5.5). ⚠️ A dsh worker running a profile that does not
# implement the status contract is the one case that LOOKS like claude but is not:
# it accepts the send and then burns both waits. One timeout on a dsh worker whose
# pane clearly finished means that profile, so switch that worker to markers.
sendwait() {
local sid="${1:?}" p="${2:?}" seq="${3:-$(date +%s)}" body r
# `wait:"stop,exit"`, never the `wait:true` default set: that set also carries
# `idle`, which is INFERRED from output stabilization and flaps mid-turn. On a
# dsh worker whose TUI repaints rarely the session reads `idle` while the model
# is still answering, and the re-wait below then resolved in 0 ms with
# `signal:"idle"` on a turn that had another three minutes to run (measured).
# A wait named after the end of a turn should only end with the turn, or with
# the worker. ⚠️ This is also what makes a wrong mode LOUD: the modes that
# cannot deliver `stop` answer 400 (before writing anything) instead of
# resolving on a flap, which is the answer that sends you to markers (§5.5).
body=$(jq -nc --arg p "$p" --arg c "$CID-$sid" --argjson s "$seq" \
'{input:($p+"\r"),useMux:true,clientId:$c,seq:$s,wait:"stop,exit",waitTimeout:20000}')
r=$("${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" \
-H 'Content-Type: application/json' --data-binary "$body")
if jq -e '.data.delivered and .data.wait.timedOut' <<<"$r" >/dev/null 2>&1; then
"${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" -H 'Content-Type: application/json' \
-d "$(jq -nc --arg c "$CID-$sid" --argjson s "$(date +%s)" \
'{input:"\r",useMux:true,clientId:$c,seq:$s}')" >/dev/null
# The resend is a tagged DUPLICATE, so the server skips the write and reports
# `delivered:false` for it -- truthfully, but about the wrong send. The first
# one delivered, so carry that forward, or §1's cleanup reads a completed turn
# as an undelivered one and keeps a finished worker forever.
r=$("${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" \
-H 'Content-Type: application/json' --data-binary "$(jq -c '.waitTimeout=580000' <<<"$body")" \
| jq -c 'if .success and (.data.wait.ended | not) then .data.delivered = true else . end')
fi
printf '%s\n' "$r"
}
# last_text <sid> [prev] -> that worker's last assistant message (claude, codex and
# deepseek write a real transcript; the other modes have none, so read the terminal
# instead -- §5.4). Polled, because the transcript write LAGS the stop signal, and
# "some text exists" is not "THIS turn's text exists": right after a SECOND turn on the same worker the endpoint still serves
# the previous answer for a beat (observed live). When reading consecutive turns, pass
# the previous answer as [prev]: the poll then holds out for text that differs from it,
# falling back to whatever it last saw if the budget runs dry, so an honestly repeated
# answer still comes back. Non-zero exit means the worker really never wrote one.
last_text() {
local t="" prev="${2:-}"
for _ in $(seq 1 15); do
t=$("${CURL[@]}" "$API/api/v1/sessions/$1/last-response" | jq -r '.data.text // empty')
[ -n "$t" ] && [ "$t" != "$prev" ] && { printf '%s\n' "$t"; return 0; }
sleep 1
done
[ -n "$t" ] && { printf '%s\n' "$t"; return 0; }
return 1
}
# The stamp is the LAST line on purpose (a truncated write leaves it unset) and is kept
# bare on purpose: the write condition above anchors on it with $, so an inline comment
# here would fail that match and rewrite this file on every single bootstrap.
CODEMAN_PREAMBLE=1.22.0
PREAMBLE
)
. "$PRE"; [ "${CODEMAN_PREAMBLE:-}" = 1.22.0 ] || { echo "preamble at $PRE is stale or truncated: rm it and re-run this block"; exit 1; }
Every later Bash call that touches the API starts with the same two loader lines from the top of this section.
Why it is built this way, all of it load-bearing:
delete_session is
undefined, and an undefined function is "command not found", which deletes nothing.
⚠️ This argument covers accidents, NOT a hostile file: a complete attacker-written
preamble can define delete_session and set the stamp, and sourcing executes it. What
defends against that is the path choice in the next bullet, not this one. Never
hand-roll a DELETE of your own, which is the one thing that would route around this.rm. The older [ -s "$PRE" ] condition could not tell a
complete file from a half-written one and left both to the post-source guard, which can
only refuse, not repair. That guard stays as the fail-closed backstop: if the rewrite
itself is cut short, CODEMAN_PREAMBLE is unset and the call stops./tmp. On a shared machine /tmp is world-writable, so another local user
can pre-create the exact path you are about to . and have their code run as you.
$HOME-derived paths are not world-writable, and the file is written 0600 anyway.
The file holds the credential-recovery code, not a recovered password.$$ in a clientId. It changes per call, so the "resend the identical
request" loop in §5.3 would stop being a duplicate and would retype the prompt,
submitting the turn twice. Use the fixed literal $CID.CODEMAN_*, HOME) survive, which is why the
preamble rebuilds $API and $SELF from them on every source rather than baking
them in.If a call comes back as unparseable text instead of JSON, that is almost always a plain-text 401: see §6 and the symptom gallery.
If the job is "spawn N claude workers, give them tasks, collect the answers", this block is the whole thing. Run it, report, and stop reading. §2 onward is for jobs this does not cover; you are not being careless by not reading them.
Fill in the case names and the prompts, then run it as your FIRST Bash call: no
standalone preamble check before it (line one below IS that check), and no
reconnaissance. ls ~/codeman-cases answers nothing this block needs: invented
fresh names need no lookup, and spawn_worker refuses a name that already exists
rather than silently reusing it. Everything below is spawn_workers / sendwait /
last_text / delete_session from the §0 preamble, so there is nothing to assemble
and no per-call body to hand-build.
. "${XDG_CACHE_HOME:-$HOME/.cache}/codeman-agent-$CODEMAN_SESSION_ID.sh" 2>/dev/null # §0 loader
[ "${CODEMAN_PREAMBLE:-}" = 1.22.0 ] || { echo "preamble missing or stale; run the full §0 block"; exit 1; }
N=(alpha beta) # INVENT one fresh case name per worker; never list cases first
# (a name may carry a mode: `beta:deepseek`, see below)
T=('reply with one line: the absolute path of your working directory'
'reply with one line: your model name') # tasks, same order as N
S=(); while read -r _ s; do S+=("$s"); done < <(spawn_workers "${N[@]}") # concurrent
for i in "${!N[@]}"; do [ -n "${S[$i]:-}" ] || FAIL=1; done
[ -z "${FAIL:-}" ] || { echo "a spawn failed (stderr says why; §5.1): deleting the siblings"
for s in "${S[@]}"; do [ -n "$s" ] && delete_session "$s" >/dev/null; done; exit 1; }
D=$(mktemp -d) || { for s in "${S[@]}"; do delete_session "$s" >/dev/null; done; exit 1; }
for i in "${!N[@]}"; do sendwait "${S[$i]}" "${T[$i]}" > "$D/$i" & done; wait
for i in "${!N[@]}"; do
jq -ce --arg n "${N[$i]}" \
'{worker:$n,delivered:.data.delivered,timedOut:.data.wait.timedOut,signal:.data.wait.signal}' \
"$D/$i" || echo "{\"worker\":\"${N[$i]}\",\"error\":\"send produced no result\"}"
echo "== ${N[$i]}"; last_text "${S[$i]}" || echo "(no response written)"
done
for i in "${!N[@]}"; do # delete ONLY what finished; a timeout means STILL WORKING (§3 rule 5)
if jq -e '.success and .data.delivered and (.data.wait.timedOut|not)' "$D/$i" >/dev/null 2>&1
then delete_session "${S[$i]}" >/dev/null
else echo "kept ${N[$i]} (${S[$i]}): its line above says why; re-wait or repair (§5.3), then delete_session it"
fi
done; rm -rf "$D"
Measured against a live 1.18.0 server: two cold workers spawned and ready in 6.3 s, both turns dispatched and both answers read in 4.0 s more. If your run takes minutes, the time went into deliberation, not the API. The four things that actually cost time:
& plus
wait, as above, makes N workers cost about what one costs.ls ~/codeman-cases, a list_sessions "to see what is there": each is a whole
model turn spent learning something this block already handles (line one performs
the preamble check, invented names need no listing, and spawn_worker refuses
collisions). A live two-worker run spent ~12 s of its 28 s total on exactly two
such turns; the API work in between was under 10 s.for loop around quick-start, a poll on .data.pid,
a bespoke ready() or spawn() of your own. Each is a worse copy of a function
already sitting in your preamble; the live run that wrote them spawned serially,
polled pid for nothing, and shipped its workers without lineage.spawn_worker carries them: the hooks check (it refuses a
name that resolved to a hook-less directory with one local grep, so a worker it hands
back always has a working stop and sendwait is trustworthy), and the pid poll,
which is dead weight because wait-output already blocks on the composer.Four things this block leans on, each one link away, no detour needed to run it:
~/codeman-cases/<name>, not your repo. A name that already means something (a
linked case, a pre-existing directory) is refused by spawn_worker rather than
silently reused. Spawning where the work actually is (a linked case, a git worktree)
is a different call, and picking the wrong one is the costliest mistake in this
skill: §5.1. Those workspaces do get hooks now, unless the operator disabled it.sendwait supplies the \r, picks a fresh seq, and self-heals a stranded Enter.
A prompt without the \r is never submitted (§3), a reused seq is silently
swallowed as an already-applied duplicate, and an Enter eaten by an Ink repaint
strands the prompt on the composer until a bare \r follows: all three are reasons
to let sendwait build the call rather than hand-rolling it.sendwait costs that worker one billed turn, as does every prompt you send it.GET /api/v1/cases/agent-created lists them for cleanup: §5.14.The block above spawns claude workers. Any entry in N may instead name a mode
(beta:deepseek), and a deepseek worker is driven by the same four verbs, with no
change to the rest of the block: spawn_workers waits for its composer, sendwait
blocks on its real end-of-turn signal, last_text reads its answer, delete_session
removes it.
That is true of no other non-claude mode, and it is worth knowing why: the DeepSeek
Harness TUI reports idle/working/blocked to Codeman over the supervisor contract it
implements, so dsh is the one external CLI with definitive stop/blocked signals
instead of guessed-from-silence ones — and it writes a structured transcript, which is
what last-response reads for it. shell, opencode, codex, gemini, antigravity,
pi, grok and omp have neither and still need markers (§5.5).
Three things to know before you spawn one:
dsh ships only web/headless, so the terminal
agent is always an installed profile. GET /api/v1/deepseek/status answers both
questions separately (available = the binary, runnable = a profile that can drive a
pane); a spawn without one fails with OPERATION_FAILED rather than falling back.stop alone. The harness reports idle at
boot ~300 ms before its composer paints (measured 2.26 s vs 2.56 s), so a sendwait
fired straight after quick-start resolves on that boot signal, reports a turn that
never ran, and leaves the prompt in a pane that was not yet taking input. Letting
spawn_worker gate on readiness is what steps past that edge; it is not optional.sendwait that times out on a
worker whose pane clearly finished: that profile is one of them, so drive it with
markers instead.One row per job. Acting on this table alone is correct; the §5 links are the detail.
| I want to | Call | Detail |
|---|---|---|
| start a worker where the work is | POST /api/v1/quick-start {"caseName":…}, which creates ~/codeman-cases/<name> unless the name is already a case. Any other path (a git worktree): POST /api/v1/sessions {"workingDir":…} then POST /api/v1/sessions/:id/interactive. Both install hooks by default, so expect full signals in either, and verify rather than assume. N workers means N worktrees | §5.1 |
| know a new worker can accept a prompt | GET .../wait-output?match=shift+tab&from=buffer (urlencode the +); a deepseek worker draws ❯ instead, and its boot stop fires ~300 ms BEFORE that, so never read the signal as readiness | §5.2 |
| deliver a task and know when it finished | POST .../input with "input":"…\r", clientId, seq, "wait":true. Resolves on stop, so it is trustworthy where the signal is real: claude mode with hooks (installed by default, but the operator can disable it and remote sessions never get them) and deepseek mode through its status bridge. Costs the worker one billed turn | §5.3 |
| know a hook-less worker finished | it has no stop, and wait:true there resolves on flapping idle without erroring: make it print a split, unique marker and wait-output on that instead | §5.5 |
| read the answer | GET .../last-response, polled (claude, codex and deepseek write a transcript; empty for the other modes) | §5.4 |
| know if it is alive | GET .../wait?until=exit&timeout=1000: an immediate signal:"exit" means dead. status and pid both lie | §5.6 |
| know if it is stuck | GET .../active-tools and GET .../run-summary are structured and free; two terminal?tail= samples are the crude fallback | §5.6 |
| make a runaway worker stop | POST .../input {"input":"\u001b"} (ESC, no \r). Deleting the session would destroy the conversation instead | §5.7 |
| resume a worker halted on a usage limit | POST .../auto-resume {"enabled":true}. Respawn and Ralph are not the remedy: respawn runs /clear | §5.8 |
| give a worker big input | write a file into its workspace with your own tools and send one short line pointing at it. The composer takes 65536 characters, single-line, newlines stripped | §5.9 |
| watch N workers at once | one in-flight wait per worker (per-session waiter cap 16); fan-out shapes differ for claude and shell | §5.10 |
| find yourself, list what exists | GET /api/v1/sessions, match your $SELF by prefix | §5.11 |
| read or record what the user wants | GET/PUT .../intent, and POST .../readmymind to predict | §5.12 |
| talk to a claude worker directly | ListAgents / SendMessage, when the feature is on at both ends | §5.13 |
| clean up | delete_session "$SID" per id you created. Case directories and git worktrees are not removed with it; GET /api/v1/cases/agent-created lists the scratch case dirs your spawns left behind, for you to report | §5.14 |
Ten one-liners. Each breaks something concrete; the reason is one link away.
\r or Enter is never sent and the text sits unsubmitted
(§5.3)..data.status. It reads idle mid-turn and idle on a dead
worker (§5.6).wait.timeoutMs are in
endpoints.md.stop that fires with no waiter is unobservable afterwards
(§5.10).delete_session. The server lets a session delete itself
(§4).You are yourself a session on this server, and the API has no undo.
delete_session is the ONLY guard.
The server has no self-protection: a session that DELETEs its own id succeeds and
dies silently (verified live). Always delete through delete_session "$SID" from
§0; never write a bare curl -X DELETE and never reintroduce the
is_self … || curl -X DELETE … shape. That older form failed open: with the
function undefined (a missing or truncated preamble file, see §0) bash returns 127,
the || branch fires, and the delete runs with no self-check at all. Wrapping the
request inside the guard is what makes a lost preamble delete nothing instead of
deleting you. Apply the same prefix-both-directions reasoning before any kill,
respawn, or input call you write by hand.POST /api/v1/quick-start; POST /api/v1/sessions + POST /api/v1/sessions/:id/interactive
(or /shell) for a directory the user's own task named; POST /api/v1/sessions/:id/input;
and DELETE /api/v1/sessions/:id only for a session you created in this
conversation, by exact id. Keep a list of the ids you create. Everything else
mutating needs the user to have asked for it.DELETE /api/cases/:name recursively deletes a real directory of the user's
code from disk. One wrong case name destroys work that was never yours.DELETE /api/sessions (no id) is a bulk kill of every session, the user's
real work included. DELETE /api/subagents/:agentId kills one background agent;
DELETE /api/subagents (no id) does not kill anything, it clears the watcher's
map and timers, which blinds every subagent surface in the UI until they are
rediscovered. Neither is yours to call./clear (wipes a
conversation), orchestrator state is a single global slot, cron jobs outlive you.PUT /api/settings, POST /api/system/update: global UI settings; server restart.POST /api/approvals/:id/answer. It types a digit, an Esc or free text into
whichever session raised the prompt. Approving another session's permission
dialog authorizes a tool call the user never saw, from a session that is not
yours. Answer only a prompt raised by a worker you created, and only when the
user asked you to.git checkout
in one yanks the tree out from under the other. Creating worktrees changes the
user's repository state, so say that you did; removing one discards any
uncommitted work inside it, so ask first (§5.1).tmux kill-session, pkill tmux, pkill claude. The API is the only interface.quick-start in a
loop.The per-verb detail lives in reference/verbs.md, loaded on demand
so it is not paid for on every skill load. Section numbers and anchors are unchanged, so
a §5.4 reference still resolves. §1 already covers the common job without any of
these; open the one row you actually hit.
| Open | When |
|---|---|
| 5.1 Where to spawn | the work is not a fresh scratch case: a linked case, a git worktree, any path that already existed. Hooks are absent there, which silently breaks send-and-wait. The costliest mistake in this skill |
| 5.2 Readiness | a worker never drew its composer, or you need the trust-dialog ladder by hand |
| 5.3 Send a task and wait | the sendwait body, its signals, and the duplicate-resend loop |
| 5.4 Read the answer | last_text came back empty, or the mode is not claude/codex/deepseek |
| 5.5 Markers for hook-less workers | the worker has no stop hook: synchronize on a split, unique printed marker |
| 5.6 Alive and stuck | is it dead or just slow? status and pid both lie |
| 5.7 Interrupt without destroying | a runaway worker you want to stop but keep |
| 5.8 Usage limits | a worker halted on a subscription limit |
| 5.9 Big input via the workspace | the prompt is larger than one composer line |
| 5.10 Fan out | many workers at once: waiter caps, and why signals are edge-triggered |
| 5.11 List and find yourself | enumerate sessions, or match $SELF by prefix |
| 5.12 Read My Mind | read or record what the user wants for a case |
| 5.13 Messaging claude workers | ListAgents / SendMessage instead of the HTTP path |
| 5.14 Clean up | what deleting a session does not remove, and how to list the case dirs you left |
You need this section only when the API answers something jq cannot parse, or when
you are on a server old enough to lack the wait endpoints. Endpoint-level detail lives
in endpoints.md.
Auth is active only when the server has CODEMAN_PASSWORD (or is in multi-user mode).
Your session has usually inherited that password already, which is why the §0
preamble tries $CODEMAN_PASSWORD first: Codeman does not strip it. buildClaudeEnv()
(src/session-cli-builder.ts) spreads the server's entire process.env into the
session and deletes only COLORTERM and CLAUDECODE, and the tmux spawn path applies
no denylist either. On a stock password-protected install (install.sh writes the
password into the systemd unit or launchd plist, so the server process carries it) the
value is simply in your environment.
It is not guaranteed, though, which is what the fallbacks are for. A tmux pane
inherits the tmux server's environment, and that server can predate the password;
and the data dir's .env is only ever read by the codeman CLI itself, never loaded
into the web server's environment.
Fallback 1, in the §0 preamble already: the data dir's .env, the same file
codeman attach reads. It is hand-authored; nothing ever writes it.
Fallback 2, for a stock install where the supervisor definition is the only copy on disk. Append this to the preamble file (before its version-stamp line) and re-source:
if [ -z "${CODEMAN_PASSWORD:-}" ]; then # install.sh puts it in the service definition
UNIT="$HOME/.config/systemd/user/codeman-web.service"
PLIST="$HOME/Library/LaunchAgents/com.codeman.web.plist"
if [ -f "$UNIT" ]; then
# install.sh backslash-escapes " and \ in the unit value; undo it or a password
# containing either recovers wrong and auth fails.
CODEMAN_PASSWORD=$(sed -n 's/^Environment="CODEMAN_PASSWORD=\(.*\)"$/\1/p' "$UNIT" | head -1 | sed 's/\\\(["\\]\)/\1/g')
elif [ -f "$PLIST" ]; then
# install.sh XML-escapes the plist value; undo it (& LAST, mirroring escape order).
CODEMAN_PASSWORD=$(awk '/<key>CODEMAN_PASSWORD<\/key>/{getline; print}' "$PLIST" | sed -n 's/.*<string>\(.*\)<\/string>.*/\1/p' \
| sed -e 's/</</g' -e 's/>/>/g' -e 's/&/\&/g')
fi
fi
⚠️ A 401 is plain text, not the JSON envelope, so on a password-protected server
every jq in these recipes dies with jq: parse error instead of showing
UNAUTHORIZED. If that happens, check the status with -w '%{http_code}'; if it is
401 and no fallback found a credential, stop and tell the user you need
credentials. The same is true of the guards that run before any handler: the Host
allowlist (403 Forbidden: host not allowed), the Origin/CSRF guard, and the auth
rate limiter's 429 all answer in plain text. The hook-secret bypass covers only
/api/hook-event and /api/status-telemetry, never session control.
In multi-user mode accounts live in users.json and the credential is a real user's
name and password. A recovered CODEMAN_PASSWORD still often works: bootstrapInitialAdmin()
(user-store.ts:417-427) creates the FIRST admin from CODEMAN_USERNAME/CODEMAN_PASSWORD
on first boot when no users exist, so on a stock multi-user install that pair usually IS
a valid admin login until someone changes it. Try it once; if it fails, ask the user
rather than retrying (ten failures rate-limit the address).
The wait endpoints first ship in Codeman 1.13.0, but do not gate on the version
number: a dev build can serve them while reporting an older version. Probe instead.
GET .../wait on a real session id answering 404 with an .error starting Route
means the server predates them (fall back to polling GET .../terminal?tail= and say
so). Session ... not found means your session id is wrong, not the server.
CODEMAN_MUX/CODEMAN_API_URL into the session,
so the §0 guard fails closed and you refuse to act. That is correct behavior, not a
bug to work around.CODEMAN_DOCKER_BRIDGE_HOOKS=1 does not fix it: that opens a hooks-only
listener, so hook events flow but /api/v1/* stays refused. Report it rather than
retrying; making it reachable is an operator decision.Everything else (endpoint tables, per-mode signal table, error codes, capacity limits, Docker/remote caveats): reference/endpoints.md. Fan-out orchestration and blocked-worker handling: reference/recipes.md.
评论 (0)
暂无评论,成为第一个评论者吧!