SkillAtlasSkill 详情

arch-design

An agentic development harness for Claude Code & Codex: agent-routed workflows from raw requirem...

审核状态:已审核Quality 72Security 90

复制安装命令

用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。

复制前请先查看来源、License 和安全提示。

项目 README

来源文件:README.md

抓取于 2026年9月13日

Ship: AI-Powered Software Development Harness

An agentic development harness for Claude Code & Codex: agent-routed workflows from raw requirement to green PR.

Ship helps agents choose and run the right amount of software delivery process: one standalone phase, a grouped quality/build bundle, or the full raw-input-to-green-PR flow.

Ship workflow: gated stages, disk artifacts, fresh subagents

How It Works

Ship is a harness, not a copilot. It doesn't help AI write code — it constrains AI to produce reliable results through mechanically enforced quality gates.

The problem Ship solves: AI coding agents are capable but unreliable. They skip tests, hallucinate about code they haven't read, review their own work and call it good, and declare victory without evidence. Ship makes these failure modes structurally impossible.

  • Use Ship chooses the right route. /ship:use-ship decides whether the task needs one skill, a phase bundle, or the full /ship:auto workflow.
  • Production artifacts stay organized. When a task needs durable docs, agents use the repo's existing convention or create a focused docs/ship/<task-id>/ folder for requirements, design, engineering, quality, delivery, and archive notes.
  • Atomic skills stay standalone. Focused skills like /ship:dev, /ship:e2e, /ship:review, /ship:qa, /ship:refactor, and /ship:handoff work directly without a full workflow.
  • Input, state, and outputs are separate. Raw requirements live under input/. The orchestrator keeps only minimal run state. Markdown artifacts and repository code are the deliverables.
  • Every phase is isolated. The reviewer has never seen the implementation context. The QA evaluator can only see the spec, the diff, and the running application. Fresh context per phase means no accumulated bias.
  • Plans are adversarially tested. An independent peer challenger produces code-grounded objections with file paths and snippets. The planner must respond with evidence, not hand-waving. Two rounds before you see anything.
  • Evidence is hierarchical. L1 (screenshot, curl response, console log) is the only acceptable proof. L2 (HTTP 200, "tests passed") is insufficient. L3 ("should work based on the code") is an automatic FAIL.
  • State lives on disk, not in memory. The current phase is tracked in local state, and dev keeps a per-story ledger. On resume — or after context compaction — the orchestrator reads disk and picks up where it left off instead of redoing finished work. A stop-gate hook blocks session exit while the workflow is active.
  • Context moves as files, judgment stays expensive. Story briefs, implementer reports, and review diffs are handed to subagents as file paths, not pasted text — nothing bulky parks in the host's context. Every subagent dispatch names its model tier: mechanical transcription can go a tier down, reviewers have a mid-tier floor, and judgment calls never leave the host (adopted from superpowers v6's measured results).
  • The host can't game its own reviewers. Reviewer dispatches carry the spec's constraints verbatim, never "don't flag X" or pre-rated severity. Reviews are read-only, implementer rationales don't downgrade findings, and a defect the plan itself mandates still gets reported — the user decides.
  • The finish line is checks green, not PR created. After opening the PR, Ship enters a goal-directed fix loop — read CI failures, fix the smallest real cause, address review comments, resolve merge conflicts — and keeps going while each round makes progress. It escalates on evidence, not a counter: the same failure surviving a fix aimed at it, an issue needing human judgment, or an external blocker.
  • Test-driven implementation. Stories follow a RED-GREEN-REFACTOR cycle with per-story code review before merge.
image

Installation

Claude Code

/plugin marketplace add heliohq/ship
/plugin install ship@heliohq

Codex

/plugins

Search for Ship, then install it. In Codex App, open Plugins in the sidebar and install Ship from there. Codex loads Ship's skills, MCP config, and hooks from .codex-plugin/plugin.json — the same routing hint and quality gates as Claude Code.

Verify Installation

Open a fresh session and confirm the /ship:* skills are available — for example, run /ship:use-ship plan out a user authentication system.

Updating

/plugin update ship

Skills

Run /ship:use-ship when you want the agent to choose the right Ship route. Run /ship:auto when you explicitly want the full staged workflow. Or run individual phases when you only need one; atomic skills do not require an active auto run.

SkillDescription
/ship:use-shipRoute the request to a standalone skill, phase bundle, or full flow
/ship:autoStaged workflow: input → design/spec+plan → dev → E2E → review → QA → refactor → handoff
/ship:designAdversarial spec + plan with peer challenge rounds
/ship:devHost implements, peer cross-validates; parallel waves for file-independent stories
/ship:e2eCodify the change's acceptance criteria as persistent E2E tests, detect or scaffold the framework, run them against the real app
/ship:reviewBug-focused diff review — no style nits
/ship:qaExploratory sweep against the running app, finds what codified tests missed
/ship:handoffPR creation + CI fix loop until checks green
/ship:refactorFour-lens scan, classify by risk, apply with verification
/ship:arch-designSystem-design thinking — nine falsifiable lenses, self-interview method, red-team pass — hands off to write-docs
/ship:write-docsProject documentation with frontmatter, lifecycle, and indexing, incl. design docs and ADRs

Skills are available through the host plugin catalog and direct /ship:* commands. At startup, Ship injects only a tiny hint to consult /ship:use-ship when Ship may apply; it does not inject docs, memory, or artifact content.

See docs/skills.md for detailed guides.

License

MIT

Acknowledgments

Ship is built on ideas from:

  • agent-browser — Browser automation CLI for AI agents
  • Superpowers — Jesse Vincent's agentic skills framework for Claude Code
  • gstack — Garry Tan's opinionated Claude Code setup
  • Claude Code — Agent workflows and the cleanup pattern that inspired /ship:refactor's four-lens scan
其他

低风险

  • 来源需自行核对维护者身份。
  • 未检测到明显脚本安装指令。
  • 可能需要外部 token、网络权限或第三方服务。
  • 未检测到高风险命令。
  • 扫描发现:0 条。

Codex — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/heliohq/ship.git
  3. 将 "skills/arch-design" 文件夹复制到 Codex 的 skills 目录中。
  4. 重启 Codex 让新的 skill 生效。

Codex — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Codex 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Codex 让新的 skill 生效。

Claude Code — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/heliohq/ship.git
  3. 将 "skills/arch-design" 文件夹复制到 Claude Code 的 skills 目录中。
  4. 重启 Claude Code 让新的 skill 生效。

Claude Code — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Claude Code 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Claude Code 让新的 skill 生效。

Cursor — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/heliohq/ship.git
  3. 将 "skills/arch-design" 文件夹复制到 Cursor 的 skills 目录中。
  4. 重启 Cursor 让新的 skill 生效。

Cursor — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Cursor 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Cursor 让新的 skill 生效。

GitHub Copilot — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/heliohq/ship.git
  3. 将 "skills/arch-design" 文件夹复制到 GitHub Copilot 的 skills 目录中。
  4. 重启 GitHub Copilot 让新的 skill 生效。

GitHub Copilot — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 GitHub Copilot 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 GitHub Copilot 让新的 skill 生效。

Windsurf — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/heliohq/ship.git
  3. 将 "skills/arch-design" 文件夹复制到 Windsurf 的 skills 目录中。
  4. 重启 Windsurf 让新的 skill 生效。

Windsurf — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Windsurf 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Windsurf 让新的 skill 生效。
查看 SKILL.md 原文
name: arch-design
version: 1.0.0
description: >
  System-design thinking before any doc or code: goals/non-goals,
  back-of-envelope numbers, components and contracts, failure modes,
  operability, security, trade-offs. Use for "design this system",
  "architecture for X", "trade-offs for X", "how should we architect",
  "API design", "data model for", "service boundaries", or before an ADR.
  Hands off to /ship:write-docs to record the decision. Not implementation
  planning (/ship:design — that turns a decided design into stories).

Ship: Architecture Design

Think the design through before anything is written or built. The lenses below apply to any system: a service, a CLI, a frontend, a data pipeline, an agent runtime. Scale the depth to the decision — a single component with clear constraints needs the frame, a contract sketch, and trade-offs; a new system with unknowns needs every lens. Skipping a lens is a judgment call you record with its reason ("no trust boundary crossed — internal tool, single user"), never a silent omission.

Path note: ../shared/*.md references resolve against this skill's base directory (announced as "Base directory for this skill" when the skill loaded), not your working directory.

The quality bar for every lens: statements someone could prove wrong — named numbers, named failure behaviors, named rejected alternatives. Virtue words ("scalable", "robust", "flexible") with no test attached are filler.

Red Flag

Never:

  • Jump to a solution before the frame (goals, non-goals, numbers) exists
  • Present a design with zero numbers and zero rejected alternatives — that's a description, not a design
  • Resolve a question from memory when the codebase can answer it
  • Interrogate the user question-by-question — self-interview instead; user-owned calls become recorded assumptions
  • Open many design branches at once — resolve dependencies in order
  • Skip a lens silently — skip with a recorded reason, or don't skip

Method — Walk the Lenses as a Self-Interview

Interview yourself relentlessly about every aspect of the design until no unresolved question remains — the lenses below are the branches of the tree. Take questions one at a time, in dependency order: resolve a decision before opening the ones that build on it (storage before schema, contract before internals) — an answer stacked on an unresolved dependency is a guess, and opening many branches at once produces shallow parallel guesses. For each question, state your recommended answer, then resolve it with the strongest means available:

  • Codebase evidence — if exploring the repo can answer it, explore; most questions die here. Never ask anyone what the code already says.
  • Arithmetic — if a number decides it, compute the number (lens 2).
  • Judgment — adopt your recommended answer and record it under Assumptions with what breaks if it's wrong.

Do not interview the user. Questions only they could answer (product intent, priorities, external constraints) do not become blocking prompts — adopt the recommended answer, mark the assumption, and surface that short list when you present the analysis. The Q&A itself is scaffolding, not deliverable: the analysis records the decisions it produced.

The Lenses

  1. Frame: goals, non-goals, requirements. Functional capabilities; the non-functional numbers that bind the design (latency, throughput, availability, consistency, data volume); real constraints (existing stack, team, timeline, compliance, backward compatibility). Write explicit non-goals — a later reader must be able to tell deliberate exclusion from oversight. Never conflate "what we want" with "what exists": state the gap.
  2. Do the arithmetic. Back-of-envelope estimates for the numbers that drive the shape: request rates, data growth, fan-out, budget per hop. Show the math — a computed number is falsifiable; "should scale fine" is not. Where a number is unknown, state the assumption and what breaks if it is 10× off.
  3. The boring option first. Before designing anything new: does an existing utility, library, managed service, or simpler shape already cover this? Classify the decision — two-way door (reversible: pick the simple option fast and note the revisit trigger) or one-way door (hard to undo: full rigor). Most decisions are two-way doors; treating them all as one-way is how designs bloat.
  4. Components and contracts. Responsibilities, data flow, data model, storage choices driven by access patterns. Define the contracts — interfaces, protocols, schemas — before the internals; they are what other people and later phases build against. Verify assumptions against the actual codebase, not memory.
  5. Failure modes. For each component and dependency: what happens when it is down, slow, or returning garbage? What do concurrent access, retries, and duplicate delivery do? Name the blast radius and the degraded behavior a user sees ("if the queue dies, writes buffer locally for 10 minutes, then reject with a clear error" — never "handles failures gracefully"). Idempotency and partial-failure recovery are design decisions, not implementation details.
  6. Operability and rollout. How does the system get from the current state to this design — the migration path, and whether it must be zero-downtime? What is the rollback story? Which metric or log line says it is working, and which signal says it is not? What does it cost to run? A design you cannot observe or roll back is not finished.
  7. Security and trust boundaries. Where does data cross a trust line? AuthN/authZ at each boundary, secrets and credential handling, which data is sensitive and who may see it. One sentence when nothing crosses; a real section when anything does.
  8. Trade-off analysis. For each major decision: at least two alternatives considered, concrete pros/cons, the deciding factor, and what you are giving up. A design with no rejected alternatives has not been designed.
  9. Revisit triggers. Flag decisions that will not age well, with their trigger: load-dependent ("rethink at 10k rps"), time-bound ("chose X because Y isn't ready"), assumption-sensitive ("multi-region breaks this"). These are honest engineering, not weaknesses.

Red-Team Before Presenting

Re-read the design as a skeptical staff engineer: what is the first thing that breaks in production? Which number is least defensible? Which alternative was dismissed too fast? If an attack lands, fix the design, not the wording. For high-stakes or contested decisions, dispatch a fresh peer challenge with only the draft (see ../shared/runtime-resolution.md) — the same adversarial pattern /ship:design applies to specs.

Execution Handoff

The analysis is the deliverable: the decision first, then the load-bearing numbers, the failure modes that shaped it, the rejected alternatives, and the assumptions the user should confirm (the user-owned questions from the self-interview).

To make it durable, hand off to /ship:write-docs — it records the decision as a design doc (design category: Boundaries required, recommended body shape in that skill's conventions).

Output the report card (read ../shared/report-card.md for the standard format):

## [Arch Design] Report Card

| Field | Value |
|-------|-------|
| Status | <DONE / BLOCKED> |
| Summary | <the decision, one line> |

### Metrics
| Metric | Value |
|--------|-------|
| Lenses applied | <N>/9 (<skipped lenses + recorded reasons>) |
| Alternatives rejected | <N> |
| Assumptions recorded | <N> |
| Revisit triggers | <N> |

### Next Steps
1. **Record it** — /ship:write-docs to write the design doc / ADR
2. **Plan implementation** — /ship:design to turn it into executable stories
3. **Full workflow** — /ship:auto for end-to-end delivery

发现问题?提交给管理员复核

评分:

评论 (0)

暂无评论,成为第一个评论者吧!