SkillAtlasSkill 详情

codex

109 Cross-Runtime Skills | 7 Claude Code Agents | One Command Install

审核状态:已审核Quality 72Security 52

复制安装命令

用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。

复制前请先查看来源、License 和安全提示。

项目 README

来源文件:README.md

抓取于 2026年10月1日

Spellbook

109 Cross-Runtime Skills | 7 Claude Code Agents | One Command Install

A cross-runtime skill library for Claude Code, Codex, and multi-agent workflows.

Stars License Skills Agents

Quick Start • Runtime Targets • Pick a Workflow • Skills • Agents • Changelog • Release Status • Contributing • 中文


Rename notice: Spellbook was formerly Claude Arsenal. Claude Code remains a first-class target; the new name reflects the broader roadmap for Claude Code, Codex, and cross-runtime agent skills. See the migration note for details.


Quick Start

Start with one job-shaped workflow. The maintained skills CLI lets you choose the supported coding agents during installation and installs only these four skills:

npx skills add majiayu000/spellbook --skill frontend-design --skill app-ui-design --skill ui-design-system --skill figma-to-react

Use npx skills add majiayu000/spellbook --list to inspect the catalog before installing. See Pick a Workflow for four other focused starting points.

Advanced Cross-Runtime Installer

install.sh remains available when you want explicit Claude Code/Codex target paths or need to install the repository's Claude Code agents as well as skills.

# Install all skills and supported agents into both maintained runtimes
curl -fsSL https://raw.githubusercontent.com/majiayu000/spellbook/main/install.sh | bash -s -- --target all

# Or clone the repository and select skills explicitly
git clone https://github.com/majiayu000/spellbook.git
cd spellbook
./install.sh --target all --skills typescript-project,python-project,devops-excellence

Verify Installation

  • Claude Code: type / to see your installed skills.
  • Codex: restart Codex so it reloads ~/.agents/skills.

Runtime Targets

Spellbook keeps the skill source in one place and installs it into the runtime you use.

TargetInstalled ToStatus
Claude Code~/.claude/skills plus ~/.claude/agentsSkills and agents supported
Codex~/.agents/skillsSkills supported; agents skipped
AllBoth Claude Code and Codex pathsRecommended for multi-tool users

Claude Code remains a first-class target and search entry. The project was formerly known as Claude Arsenal; the new Spellbook name reflects the broader goal: reusable skills that can travel across coding agents. Older Spellbook versions installed Codex skills under ~/.codex/skills; reinstall with the current installer to use the documented Codex user-level skill path.


Pick a Workflow

Start with a small bundle that matches the job, then add more skills when the workflow sticks.

WorkflowInstallGood for
Frontend and UInpx skills add majiayu000/spellbook --skill frontend-design --skill app-ui-design --skill ui-design-system --skill figma-to-reactProduct UI, landing pages, design systems, Figma handoff
Code qualitynpx skills add majiayu000/spellbook --skill codebase-audit --skill flowguard --skill systematic-debugging --skill review-gateAudits, guarded delivery, root-cause debugging, pre-landing review
Ops and releasenpx skills add majiayu000/spellbook --skill release-engineering --skill server-security --skill clash-doctor --skill system-doctorRelease planning, server hardening, and local or network diagnosis
Product and docsnpx skills add majiayu000/spellbook --skill product-discovery --skill prd-master --skill technical-spec --skill product-analyticsDiscovery, PRDs, technical specs, metrics plans
Agent workflowsnpx skills add majiayu000/spellbook --skill codex-agent --skill multi-ai-research --skill flowguard --skill vibeguardCross-review, multi-AI research, context handoff, anti-hallucination checks

High-signal individual skills to try first: github-trending, harmonyos-app, app-ui-design, product-discovery, xiaohongshu, codebase-audit, and server-security.

See Showcase for copy-paste prompts and expected outputs. Use the Spellbook Skill Browser for curated first-party skills, or the Claude Skills Registry for broader community discovery. Release history lives in Changelog.


Why Spellbook

  • Cross-runtime install: one source tree can install into Claude Code and Codex.
  • Validated registry: every installable skill is checked by python3 scripts/validate_skills.py --check.
  • Progressive disclosure: larger skills use references/, templates/, scripts/, and eval files instead of one giant prompt.
  • Practical coverage: engineering, operations, product, UI, content, and agent workflows live in one catalog.

Skills

The generated full skill inventory lives in Skill Registry. Skill layout rules live in Skill Format Policy. Skill authoring quality rules live in Skill Quality Playbook.

Search the Registry

# Free-text query (AND semantics across name, description, category, tags)
python3 scripts/validate_skills.py search rust testing

# Filter by tag
python3 scripts/validate_skills.py search --tag agent

# Restrict to a description language
python3 scripts/validate_skills.py search --language zh deploy

# Machine-readable output
python3 scripts/validate_skills.py search --tag react --json

The tag index lives in registry/tags.json for tooling and dashboards. Curated overrides for skills the keyword heuristic cannot infer live in registry/tag_overrides.yml.

Audit non-blocking skill quality signals:

python3 scripts/audit_skill_quality.py
python3 scripts/audit_skill_quality.py skill-creator

AI & Agent Workflow

Skills for orchestrating, guarding, and maintaining AI agent workflows — the core of Spellbook's cross-runtime mission.

SkillDescription
multi-model-orchestratorCoordinate multi-agent tasks via a centralized handoff document
flowguardGuard long, ambiguous, or stateful agent tasks from drift
skill-lifeguardAdd reliable-skill contracts, checkpoints, smoke hooks, and drift signals
review-gateProduce review packs and require human approval before landing agent changes
skill-auditAudit, design, categorize, and measure agent skills
skill-ecosystem-doctorGovern canonical sources, projections, retirement, quarantine, and cross-runtime verification
threadsCodex-native subagents and parallel GitHub queue lanes
codex-fluentCodex session hygiene, archive strategy, and handoff discipline
codex-retrospectiveCodex self-review of recent history to improve behavior
brainstormingSocratic dialogue for design refinement and architecture exploration

See docs/agent-reliability-trio.md for the Reliable Skill + Context Engineering + Review Gate workflow.

Development Architecture

Build production-ready projects with language-specific best practices.

SkillLanguageKey Features
typescript-projectTypeScriptESM, Zod, Biome, Clean Architecture
python-projectPythonuv, Pydantic, Ruff, FastAPI
rust-projectRustCargo workspace, error handling, async
golang-webGoChi/Echo, sqlc, structured logging
zig-projectZigBuild system, memory management
architecture-foundationCross-languageRuntime, state ownership, adapters, and convergence specs
elegant-architectureCross-languageClean architecture with strict 200-line file limits

Product Lifecycle

End-to-end product development from discovery to deployment.

SkillPhaseWhat You Get
product-discoveryDiscoveryJTBD, user interviews, market research
prd-masterDefinitionPRD writing, user stories, RICE prioritization
technical-specDesignDesign docs, ADR, C4 diagrams
product-analyticsGrowthEvent tracking, A/B testing, AARRR
devops-excellenceDeploymentCI/CD, Docker, Kubernetes, GitOps
observability-sreOperationsMonitoring, logging, tracing, SLO/SLI
product-manager-toolkitDefinitionRICE, customer interviews, PRD templates, discovery frameworks

API & Backend

SkillDescription
api-designREST/GraphQL/gRPC patterns, OpenAPI 3.2
auth-securityOAuth 2.1, JWT, security best practices
database-patternsPostgreSQL, Redis, migrations, optimization
codebase-auditDeep adaptive repository audit with severity-ranked findings and repair roadmap
structured-logging-liteCentralized logging, field standards, and distributed tracing

Development Practices

SkillDescriptionOrigin
contributorEnd-to-end open source contribution workflow from issue discovery to PR submissionCustom
repo-agent-context-auditAudit and scaffold repo agent context across AGENTS, skills, and specsCustom
skill-creatorCreate, improve, and benchmark reusable skillsCustom
humanizerRemove obvious AI writing patterns from user-facing textExternal guide + custom adaptation

Delivery Workflow

Disciplined end-to-end delivery: testing, commits, health checks, and contribution flow.

SkillDescription
app-user-story-qaEnd-to-end app feature inventory, canonical tracker, user-story testing, fixes, and retest loop
test-driven-developmentEnforce RED-GREEN-REFACTOR TDD discipline
comprehensive-testingTest pyramid, unit/integration/E2E/property testing, framework best practices
git-commit-smartGenerate meaningful conventional commit messages from diff
push-allStage, commit, and push all changes after safety checks
project-health-auditorCodebase health, tech debt, dependency, and project risk analysis
contribution-architectMove from bug fixes to architectural improvements and debt discovery

Cross-Tool Interop

Skills for using multiple coding agents and CLI tools together.

SkillDescription
codexInvoke Codex CLI sessions from another agent workflow
codex-agentOptional second-opinion review, cross-verification, and alternatives through Codex CLI
sol-luna-routerKeep GPT-5.6 Sol as commander/reviewer while GPT-5.6 Luna performs bounded implementation
ask-opencliAsk Grok or Gemini through opencli and an existing browser session
multi-ai-researchParallel research across multiple AI tools and internal agents

UI/UX & Design

SkillDescription
app-ui-designiOS/Android UI design, Material Design 3, HIG
product-ux-expertUX evaluation, heuristics, accessibility
frontend-designWeb frontend design patterns
illustrated-galleryIllustrated website template with full-frame ASCII transitions
ui-designerExtract design systems from UI screenshots and references
ui-design-systemDesign system toolkit and design-dev handoff support
web-artifacts-builderClaude.ai HTML artifacts
react-best-practicesReact and Next.js performance patterns distilled from Vercel guidance
react-hooks-best-practicesReact hooks, effects, refs, and component design patterns
blender-editable-3dEditable Blender models from references or concepts, with part-level edits and verified exports
blender-reference-to-3dReference analysis, Blender character modeling, color, multiview review, and editable delivery
slidesSpeech-friendly slide deck and background slide generation
ui-ux-pro-maxCompact UI/UX tables for product patterns, landing pages, charts, and 9 stacks
figma-to-codeFigma designs to production React/Next.js with TypeScript and Tailwind
css-debugDiagnose CSS/layout issues, Tailwind conflicts, z-index stacking
playwright-automationBrowser automation and testing with Playwright

Tooling & Automation

SkillDescription
web-asset-generatorFavicons, app icons, OG images
github-trendingGitHub trending analysis
vibeguardTask contracts, finding scoring, and lightweight anti-hallucination reviews
clash-doctorClash proxy & network diagnostics
clash-routesInspect active proxy routes for specific processes via Mihomo API
optimize-networkSafe local network speed, latency, DNS, Wi-Fi, and bufferbloat diagnostics with VPN/proxy guardrails
disk-cleanerScan and reclaim disk space with interactive cleanup guidance
system-doctorDiagnose CPU, memory, and process-level system slowdowns
codex-log-guardDiagnose and mitigate excessive Codex local SQLite diagnostic log writes
server-securityAudit and harden Linux server SSH, firewall, and exposed services
cliproxy-newapi-stackAdd a loopback-first NewAPI metering layer to an independently verified CLIProxyAPI upstream

Operations & Deploy

Deploy models and diagnose local and remote environments.

SkillDescription
gemma4-local-deployDeploy Gemma 4 12B locally on Mac/Apple Silicon via llama.cpp or Ollama
gpu-useInspect remote server GPU usage (per-card VRAM, processes, containers)
rustdesk-doctorDiagnose RustDesk connection issues
vscode-doctorDiagnose slow or freezing VS Code-compatible editors

Content & Social Media

SkillDescription
xiaohongshuXiaohongshu content creation & publishing
trip-plannerTravel itinerary planning
weeklyWeekly report from Git, Claude Code, and Codex sessions
xiaohongshu-netfeel-guardianRemove translation-tone from Claude's Chinese content for native readability

Mobile & Cross-Platform

SkillDescription
harmonyos-appHarmonyOS with ArkTS, ArkUI, Stage Model

Rust Specific

SkillDescription
rust-best-practicesMicrosoft Rust guidelines, error handling

Agents

Specialized agents for complex tasks.

AgentExpertiseUse Case
tech-lead-orchestratorCoordinationMulti-step tasks, delegation
code-archaeologistExplorationLegacy codebase documentation
backend-typescript-architectArchitectureBun/Node.js, API design
senior-code-reviewerReviewSecurity, performance, architecture
kubernetes-specialistInfrastructureK8s, Helm, GitOps
security-auditorSecurityOWASP Top 10, SAST
opensource-contributorContributionOpen source workflow

Plugins

Spellbook is also a Claude Code plugin marketplace. Install the repo as a marketplace, then install plugins from it:

/plugin marketplace add majiayu000/spellbook
/plugin install idea-coach
/plugin install rust-dev
PluginDescription
idea-coachOpinionated product coach (idea -> PRD -> clickable HTML prototype) + multi-role idea group chat; plugin commands are /idea-coach:idea and /idea-coach:idea-team
rust-devRust best practices, code review, performance, and async patterns

Plugin skills are packaged copies of catalog skills; the catalog (installed by install.sh) remains the cross-runtime source of truth.


Skill Design Philosophy

Every skill in Spellbook follows these principles:

  1. Hard Rules - Mandatory constraints with FORBIDDEN / REQUIRED markers
  2. Practical Examples - Real code, not just theory
  3. Verification Checklists - Actionable validation steps
  4. Battle-Tested - Used in production environments

Documentation

DocumentDescription
ChangelogRelease history and current release status
Installation GuideDetailed setup instructions
Runtime TargetsClaude Code and Codex installation targets
ShowcaseCopy-paste workflow demos
Spellbook Operating ContractAgent behavior rules for autonomy, escalation, pushback, feedback loops, and done-when checks
Skill Format PolicyDirectory vs file skill layout rules
Skill Quality PlaybookTrigger descriptions, gotchas, progressive disclosure, and verification
Skill Testing GuideHow to validate skills work
Creating PluginsBuild your own skills
Product Lifecycle (EN)Full lifecycle coverage
Product Lifecycle (中文)产品生命周期覆盖

Release Status

Spellbook is in pre-1.0 release-readiness mode. The first numbered tag is v0.1.0. The install path remains the repository main branch. See Changelog for release history.

Current limitations:

  • Codex installs skills only; Claude Code agents are skipped for Codex targets.
  • Some skills depend on external CLIs, accounts, credentials, or platform access that are not bundled by the installer.
  • The registry validator checks installable skill structure, not every external workflow end to end.

Support paths:


Credits

Built on the shoulders of giants:


Contributing

Contributions welcome! Please read our Contributing Guide first.


The Agent Infra Stack

This project is one layer of an open-source stack for running coding agents (Claude Code, Codex) as serious infrastructure. Every piece works standalone; together they close the loop:

spellbook sits in the Extend layer — the authoring side of the skill story: write once, run on Claude Code and Codex. Discovery and distribution live in claude-skill-registry.

LayerProjectWhat it does
Extendclaude-skill-registryDiscover and search community Claude Code skills
Extendspellbook ◀ you are hereCross-runtime skills for Claude Code, Codex, and multi-agent workflows
TrustargusStatic install-time scanner for supply-chain attacks (npm / PyPI / crates.io)
TrustvibeguardRules, hooks, and guards against hallucinated or unverified agent changes
RememberrememLocal-first persistent memory for Claude Code and Codex sessions
OrchestrateharnessRust agent orchestration platform — rules, skills, GC, observability
Routelitellm-rsHigh-performance Rust AI gateway — 100+ LLM APIs via OpenAI format
KeepkeeplineSession command center — monitor, recover, never lose agent work

License

MIT License - Use freely in your projects.


If this helps you, consider giving it a ⭐

Made for builders using Claude Code, Codex, and multi-agent workflows

开发与工程数据与 AI

高风险

  • 来源需自行核对维护者身份。
  • 包含脚本或命令调用,安装前请复核。
  • 可能需要外部 token、网络权限或第三方服务。
  • 存在潜在风险命令,请谨慎安装。
  • 扫描发现:3 条。

Codex — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/majiayu000/spellbook.git
  3. 将 "skills/codex" 文件夹复制到 Codex 的 skills 目录中。
  4. 重启 Codex 让新的 skill 生效。

Codex — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Codex 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Codex 让新的 skill 生效。

Claude Code — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/majiayu000/spellbook.git
  3. 将 "skills/codex" 文件夹复制到 Claude Code 的 skills 目录中。
  4. 重启 Claude Code 让新的 skill 生效。

Claude Code — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Claude Code 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Claude Code 让新的 skill 生效。

Cursor — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/majiayu000/spellbook.git
  3. 将 "skills/codex" 文件夹复制到 Cursor 的 skills 目录中。
  4. 重启 Cursor 让新的 skill 生效。

Cursor — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Cursor 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Cursor 让新的 skill 生效。

GitHub Copilot — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/majiayu000/spellbook.git
  3. 将 "skills/codex" 文件夹复制到 GitHub Copilot 的 skills 目录中。
  4. 重启 GitHub Copilot 让新的 skill 生效。

GitHub Copilot — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 GitHub Copilot 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 GitHub Copilot 让新的 skill 生效。

Windsurf — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/majiayu000/spellbook.git
  3. 将 "skills/codex" 文件夹复制到 Windsurf 的 skills 目录中。
  4. 重启 Windsurf 让新的 skill 生效。

Windsurf — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Windsurf 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Windsurf 让新的 skill 生效。
查看 SKILL.md 原文
name: codex
description: Use when the user asks to run Codex CLI (codex exec, codex resume) or references OpenAI Codex for code analysis, refactoring, or automated editing
compatibility: {runtimes: [claude_code]}

Codex Skill Guide

Operating Contract

  1. Run codex --version and codex exec --help first. Stop if Codex is unavailable, and only use flags shown by the installed CLI.
  2. If the user did not specify a model or reasoning effort, use the installed default. Do not hardcode a model list; model names and supported reasoning levels change over time.
  3. Select the smallest sandbox needed: --sandbox read-only for inspection and --sandbox workspace-write for requested local edits. Network access, extra writable roots, and broader modes are separate grants.
  4. Assemble the command from supported options such as:
    • -m, --model <MODEL>
    • --config model_reasoning_effort="<LEVEL>"
    • --sandbox <read-only|workspace-write|danger-full-access>
    • -C, --cd <DIR>
    • --add-dir <DIR>
    • --json
    • --ephemeral
    • --skip-git-repo-check
    • --dangerously-bypass-approvals-and-sandbox
  5. Do not use --skip-git-repo-check by default. Use it only when the user explicitly asks to run outside a Git repository or has approved that boundary bypass for this command.
  6. Do not use deprecated compatibility shortcuts such as --full-auto; use the explicit sandbox shown by current help.
  7. Preserve stderr. Keep it out of the parent context by writing it to a bounded artifact and reading only the exit status, short tail, or targeted diagnostics. Never redirect it to /dev/null.
  8. For automation, batch work, cost investigation, or any run that needs measured usage, add --json and save stdout as JSONL. Read the final usage event and report input, cached input, output, and reasoning tokens when present.
  9. Enforce a hard wall-clock timeout with the supervising runtime. Stop on non-zero exit or timeout; do not automatically retry an expensive run.
  10. When continuing a genuinely conversational task, use codex exec resume --last via stdin. Do not resume a session for homogeneous record batches; start a fresh bounded --ephemeral run per tranche so accumulated history is not resent on every model call.

Safe Prompt Passing

Do not build Codex commands with echo "user prompt" | ...; user text can contain quotes, substitutions, or newlines. Prefer a quoted heredoc so the shell never reinterprets prompt contents:

(
for required_command in python3 head wc mkfifo; do
  command -v "$required_command" >/dev/null 2>&1 || {
    printf 'Missing required command: %s\n' "$required_command" >&2
    exit 127
  }
done
codex_skill_dir=${CODEX_SKILL_DIR:-$HOME/.claude/skills/codex}
[ -f "$codex_skill_dir/scripts/run_with_timeout.py" ] || {
  printf 'Missing Codex timeout helper: %s\n' \
    "$codex_skill_dir/scripts/run_with_timeout.py" >&2
  exit 1
}
codex_artifacts=$(mktemp -d) || exit 1
codex_stderr_max_bytes=1048576
codex_stderr_pipe="$codex_artifacts/stderr.pipe"
mkfifo "$codex_stderr_pipe" || exit 1
head -c "$codex_stderr_max_bytes" <"$codex_stderr_pipe" \
  >"$codex_artifacts/stderr.log" &
codex_stderr_limiter_pid=$!
if python3 "$codex_skill_dir/scripts/run_with_timeout.py" 1800 \
  codex exec resume --last \
  2>"$codex_stderr_pipe" <<'EOF'
Your follow-up prompt goes here.
EOF
then
  codex_status=0
else
  codex_status=$?
fi
stderr_limiter_status=0
wait "$codex_stderr_limiter_pid" || stderr_limiter_status=$?
rm -- "$codex_stderr_pipe" || exit 1
artifact_status=0
stderr_bytes=$(wc -c <"$codex_artifacts/stderr.log") || exit 1
if [ "$stderr_bytes" -ge "$codex_stderr_max_bytes" ]; then
  printf 'Codex stderr reached its %s-byte artifact limit\n' \
    "$codex_stderr_max_bytes" >&2
  artifact_status=125
fi
if [ "$stderr_limiter_status" -ne 0 ]; then artifact_status=$stderr_limiter_status; fi
tail -c 4000 -- "$codex_artifacts/stderr.log"
tail_status=$?
if [ "$codex_status" -eq 0 ] && [ "$artifact_status" -eq 0 ] && [ "$tail_status" -eq 0 ]; then
  rm -R -- "$codex_artifacts" || exit $?
else
  printf 'Codex artifacts retained: %s\n' "$codex_artifacts" >&2
fi
if [ "$codex_status" -ne 0 ]; then exit "$codex_status"; fi
if [ "$artifact_status" -ne 0 ]; then exit "$artifact_status"; fi
exit "$tail_status"
)

Quick Reference

Use caseSandbox modeKey flags
Read-only review or analysisread-only--sandbox read-only
Apply local editsworkspace-write--sandbox workspace-write
Apply edits that need network accessworkspace-write plus config--sandbox workspace-write -c 'sandbox_workspace_write.network_access=true' after approval
Machine-readable usageMatch task--json; save JSONL stdout and stderr separately
Independent batch trancheMatch task--ephemeral --json; do not resume the previous tranche
Permit extra write scopePrefer --add-dirAsk before adding extra writable directories
Permit broad file accessdanger-full-access only after approvalAsk before adding --sandbox danger-full-access
Resume a conversational taskInherited from originalcodex exec resume --last via quoted heredoc
Run from another directoryMatch task needs-C <DIR> plus other flags

Batch Cost Gate

Before more than one similar model call:

  1. Define a small calibration ceiling for records, model calls, wall-clock time, and checkpoint cadence; ask for confirmation first when the user has not authorized even that bounded calibration.
  2. Run one representative calibration tranche with --json inside that ceiling.
  3. Measure actual usage from the JSONL event stream; do not estimate from record count alone.
  4. Project the remaining calls and tokens from the measured tranche.
  5. State the full-run maximum calls, maximum records, wall-clock budget, and checkpoint cadence.
  6. Ask for confirmation when the projected full run is materially larger than the calibration or the user did not already authorize that concrete budget.

Stop at every checkpoint if measured usage exceeds the projection. Cached input is still token usage: a high cached-input share usually means the same large prefix or accumulated session context is being sent repeatedly, not that the run is free.

Following Up

  • Resume only when prior conversational context is necessary. For independent records, pass only the tranche instructions and compact artifacts needed for that tranche.
  • Restate the model, reasoning effort, sandbox, measured usage, and remaining budget before proposing another costly tranche.
  • Reaching the requested done_when condition ends the run. A blocker is not permission to install, upgrade, restart services, migrate or reindex data, edit global config/hooks, write to another repository, or perform GitHub writes unless the current request explicitly authorizes that action.

Error Handling

  • Stop and report failures whenever codex --version or a codex exec command exits non-zero; request direction before retrying.
  • Before you use high-impact flags (--sandbox danger-full-access, --dangerously-bypass-approvals-and-sandbox, --dangerously-bypass-hook-trust, --skip-git-repo-check) ask the user for permission using AskUserQuestion unless it was already given.
  • When output includes warnings, partial results, missing usage, or a timeout, preserve the evidence and ask how to adjust. Do not silently degrade to an unmetered or broader run.

Gotchas

  • --skip-git-repo-check bypasses an important cwd/worktree guard. Treat it like a boundary exception, not a default.
  • danger-full-access and the --dangerously-* bypass flags are high-impact modes. Prefer read-only, then workspace-write, then modes explicitly listed by the installed CLI, then specific --add-dir grants before considering full access.
  • If a prompt came from the user or another model, pass it as stdin or as a single already-quoted CLI argument. Never interpolate it into a shell string.
  • Suppressing stderr hides failure and progress evidence; pasting all stderr into the parent wastes context. Save it, then inspect a bounded tail.
  • A small output does not imply a cheap run. Repeated large cached prefixes can dominate usage across many short calls.

发现问题?提交给管理员复核

评分:

评论 (0)

暂无评论,成为第一个评论者吧!