复制安装命令
用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。
复制前请先查看来源、License 和安全提示。
An agentic development harness for Claude Code & Codex: agent-routed workflows from raw requirem...
用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。
复制前请先查看来源、License 和安全提示。
来源文件:README.md
An agentic development harness for Claude Code & Codex: agent-routed workflows from raw requirement to green PR.
Ship helps agents choose and run the right amount of software delivery process: one standalone phase, a grouped quality/build bundle, or the full raw-input-to-green-PR flow.

Ship is a harness, not a copilot. It doesn't help AI write code — it constrains AI to produce reliable results through mechanically enforced quality gates.
The problem Ship solves: AI coding agents are capable but unreliable. They skip tests, hallucinate about code they haven't read, review their own work and call it good, and declare victory without evidence. Ship makes these failure modes structurally impossible.
/ship:use-ship decides whether the task needs one skill, a phase bundle, or the full /ship:auto workflow.docs/ship/<task-id>/ folder for requirements, design, engineering, quality, delivery, and archive notes./ship:dev, /ship:e2e, /ship:review, /ship:qa, /ship:refactor, and /ship:handoff work directly without a full workflow.input/. The orchestrator keeps only minimal run state. Markdown artifacts and repository code are the deliverables./plugin marketplace add heliohq/ship
/plugin install ship@heliohq
/plugins
Search for Ship, then install it. In Codex App, open Plugins in the sidebar and install Ship from there. Codex loads Ship's skills, MCP config, and hooks from .codex-plugin/plugin.json — the same routing hint and quality gates as Claude Code.
Open a fresh session and confirm the /ship:* skills are available — for example, run /ship:use-ship plan out a user authentication system.
/plugin update ship
Run /ship:use-ship when you want the agent to choose the right Ship route. Run /ship:auto when you explicitly want the full staged workflow. Or run individual phases when you only need one; atomic skills do not require an active auto run.
| Skill | Description |
|---|---|
/ship:use-ship | Route the request to a standalone skill, phase bundle, or full flow |
/ship:auto | Staged workflow: input → design/spec+plan → dev → E2E → review → QA → refactor → handoff |
/ship:design | Adversarial spec + plan with peer challenge rounds |
/ship:dev | Host implements, peer cross-validates; parallel waves for file-independent stories |
/ship:e2e | Codify the change's acceptance criteria as persistent E2E tests, detect or scaffold the framework, run them against the real app |
/ship:review | Bug-focused diff review — no style nits |
/ship:qa | Exploratory sweep against the running app, finds what codified tests missed |
/ship:handoff | PR creation + CI fix loop until checks green |
/ship:refactor | Four-lens scan, classify by risk, apply with verification |
/ship:arch-design | System-design thinking — nine falsifiable lenses, self-interview method, red-team pass — hands off to write-docs |
/ship:write-docs | Project documentation with frontmatter, lifecycle, and indexing, incl. design docs and ADRs |
Skills are available through the host plugin catalog and direct /ship:* commands. At startup, Ship injects only a tiny hint to consult /ship:use-ship when Ship may apply; it does not inject docs, memory, or artifact content.
See docs/skills.md for detailed guides.
Ship is built on ideas from:
/ship:refactor's four-lens scanname: e2e
version: 1.0.0
description: >
Add durable end-to-end tests for user/API-visible behavior. Detect or scaffold
the E2E framework, write tests, run the app, and store evidence. Use for E2E,
Playwright/Cypress, regression tests, or quality gates. Not exploratory QA.
allowed-tools:
- Bash
- Read
- Write
- Edit
- Glob
- Grep
- Agent
- AskUserQuestionYou are the first automated verification gate after dev. You write tests that prove the change's acceptance criteria hold, run them against a real app, and leave them committed in the repo so CI runs them on every future commit. Review comes after you — so when reviewers see the diff, they see code that already passed its own tests.
"Trust me, it works" vs durable verification. Dev just finished writing code. The naïve next step is to ask a reviewer to read it. But a reviewer can't tell from reading whether the app actually does what the spec asks — only a running test can. Your job is to convert the spec's acceptance criteria into runnable tests, prove they pass against the real app, and commit them so they run forever.
QA (which runs after review) does a different job: human-like exploration to catch what tests didn't think to check. You are the codified baseline; QA is the creative sweep above it.
CODIFY WHAT THE USER OBSERVES, NOT WHAT THE CODE DOES INTERNALLY.
ONE GOOD TEST PER ACCEPTANCE CRITERION > FIVE NOISY ONES.
MATCH THE REPO'S EXISTING STYLE BEFORE INVENTING A NEW ONE.
Path note: ../shared/*.md references resolve against this skill's base
directory (announced as "Base directory for this skill" when the skill
loaded), not your working directory.
1. Understand Read spec + diff to know what behavior to codify
2. Detect Find the existing E2E framework, or scaffold one
3. Author Write/extend tests that cover the change
4. Run Execute the suite, iterate until green or a real failure
5. Cleanup Kill anything you started (../shared/cleanup.md)
6. Report Summarize tests added, results, and any regressions
Never:
skip / xfail to
make a test pass. If the app is broken, report it as a FAIL — don't hide it..env.example values or env vars.The inputs decide everything. Read two things:
BASE=$(git symbolic-ref refs/remotes/origin/HEAD 2>/dev/null | sed 's|refs/remotes/origin/||')
[ -z "$BASE" ] && BASE=$(git rev-parse --verify origin/main >/dev/null 2>&1 && echo main || echo master)
git diff "$BASE"...HEAD --stat
git diff "$BASE"...HEAD --name-only
<task_dir>/plan/spec.md (acceptance criteria you must codify)That's it. In the staged workflow you run right after dev and before
review/QA, so there is no earlier verification report to read. If you're
in re-run mode after an e2e_fix, the previous <task_dir>/e2e/report.md
may exist — useful for knowing which tests already failed.
Some changes don't need E2E coverage. Decide early:
| Diff shape | Decision |
|---|---|
Docs-only (*.md, LICENSE, comments) | SKIP |
| Internal refactor with no user-observable change, fully covered by existing tests | SKIP (say so explicitly in the report) |
| CI / formatter / tooling config with no runtime effect | SKIP |
| New feature, bug fix, or behavior change that a user/API caller would notice | PROCEED |
| UI change (even minor) | PROCEED — visual regression and interaction flows matter |
If skipping, write a one-paragraph justification to
<task_dir>/e2e/report.md and emit the SKIP report card. Don't scaffold
frameworks or touch the test dir.
Two-step: use what exists, or scaffold the default for this stack.
Read references/frameworks.md for:
Read references/scaffolding.md only when step 2 applies — it has the
install recipes per framework.
Read references/authoring.md for patterns, selectors, data setup, and
assertion guidelines.
Scope = every acceptance criterion (automate any flow QA verified manually)
references/authoring.md.Match the repo's convention. Common patterns:
| Framework | Location |
|---|---|
| Playwright | tests/e2e/, e2e/, playwright/tests/ |
| Cypress | cypress/e2e/ |
| pytest-playwright | tests/e2e/, tests/integration/ |
| Capybara | spec/system/, spec/features/ |
If the repo already has one of these directories, use it. If scaffolding from
scratch, prefer tests/e2e/ (readable, language-agnostic).
Bring the app up via the shared startup reference:
Read ../shared/startup.md. Set EVIDENCE_DIR=".ship/tasks/<task_id>/e2e"
before running its commands so logs and PIDs land under the e2e folder.
Start services → run migrations → verify readiness.
Track PIDs in <task_dir>/e2e/pids.txt (the shared startup reference does
this automatically via $EVIDENCE_DIR). Phase 5 reads the same file.
Then run the suite. The exact command depends on the framework, but the workflow is constant:
references/authoring.md). Real bug → report it as a FAIL, never weaken
the test to make it pass. In auto mode this triggers e2e_fix, which
routes back to /ship:dev to fix the code.Playwright/Cypress produce traces, videos, and screenshots on failure. Copy
them into <task_dir>/e2e/ so debuggers (human or agent) have evidence:
# $EVIDENCE_DIR was set before entering ../shared/startup.md — reuse it here
mkdir -p "$EVIDENCE_DIR/artifacts"
# Framework-specific examples — adapt to whatever the runner actually produces
[ -d playwright-report ] && cp -r playwright-report "$EVIDENCE_DIR/artifacts/" 2>/dev/null
[ -d test-results ] && cp -r test-results "$EVIDENCE_DIR/artifacts/" 2>/dev/null
[ -d cypress/screenshots ] && cp -r cypress/screenshots "$EVIDENCE_DIR/artifacts/" 2>/dev/null
[ -d cypress/videos ] && cp -r cypress/videos "$EVIDENCE_DIR/artifacts/" 2>/dev/null
Mandatory — never skip, even on failure or timeout. Follow
../shared/cleanup.md with the same EVIDENCE_DIR you set in Phase 4.
It kills tracked PIDs (graceful then forceful), stops any docker compose
stack, and verifies ports are free. Do not inline your own cleanup logic —
the shared contract is the single source of truth.
Write <task_dir>/e2e/report.md with:
Keep the report tight — the tests themselves are the durable artifact; the report is for the pipeline to route decisions.
When invoked with --recheck (after e2e_fix made code changes):
When invoked outside /ship:auto (user types /ship:e2e directly):
<task_dir>. Pick one: .ship/e2e-<date>/ works as a
fallback evidence directory, or write directly next to the repo's test
directory if no evidence is needed.git diff alone (no spec, no QA
report). Use AskUserQuestion if the diff's intent is unclear — what
flow does the user want locked in?<task_dir>/
e2e/
report.md — run summary & test inventory
pids.txt — tracked PIDs for cleanup
artifacts/ — framework traces, videos, screenshots on failure
<repo>/tests/e2e/ — actual test files (committed to repo)
or framework-idiomatic path depending on detection
../shared/startup.md — bring the app up (shared with /ship:qa)../shared/cleanup.md — mandatory cleanup contract (shared with /ship:qa)references/frameworks.md — detection checks + framework selection matrixreferences/scaffolding.md — install recipes for each default frameworkreferences/authoring.md — writing good E2E tests (selectors, data,
assertions, parallelization, stability)Output the report card (read ../shared/report-card.md for the standard
format):
## [E2E] Report Card
| Field | Value |
|-------|-------|
| Status | <DONE / FAIL / BLOCKED / SKIP> |
| Summary | <N> tests added, <M>/<total> passing |
### Metrics
| Metric | Value |
|--------|-------|
| Framework | <name> (<pre-existing | scaffolded>) |
| Tests added | <N> |
| Tests modified | <N> |
| Suite pass rate | <N>/<total> |
| Regressions | <N> |
| Failures (real bugs) | <N> |
### Artifacts
| File | Purpose |
|------|---------|
| <task_dir>/e2e/report.md | Run summary |
| <task_dir>/e2e/artifacts/ | Traces, videos, screenshots (on failure) |
| <repo>/tests/e2e/*.spec.ts | New/modified test files (committed) |
### Next Steps
1. **Fix failures** — /ship:dev to address real bugs found by new tests
2. **Review next (if green)** — /ship:review to check correctness of the code
3. **Iterate tests** — /ship:e2e --recheck after fixes
评论 (0)
暂无评论,成为第一个评论者吧!