SkillAtlasSkill 详情

api

A self-contained Claude Code plugin that carries a feature from a one-line idea to

审核状态:已审核Quality 72Security 70

复制安装命令

用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。

复制前请先查看来源、License 和安全提示。

项目 README

来源文件:README.md

抓取于 2026年8月13日

SDD — Spec-Driven Development for Claude Code

A self-contained Claude Code plugin that carries a feature from a one-line idea to reviewed, verified, shipped code through 19 atomic, stack-agnostic skills and a TDD implementation engine — with a living roadmap above the per-feature flow.

Every skill is Socratic (it walks decisions with you, it doesn't dump a wall of output), gated (a stage hard-refuses when its prerequisite artifact is missing), and stack-agnostic (no language, tracker, or test tool is hard-coded — the skills detect what your repo uses). The Q&A skills (specify / clarify / design) are also depth-tunable — an easy / medium / hard dial decides how much the skill decides for you vs. interrogates you with trade-offs.

Requirements

SDD runs on any Claude plan / model tier — no skill requires a specific model:

  • Skills declare model: inherit and run on your session model — run your SDD session on the strongest tier your account has.
  • The judgment agents (reviewer / critic / devils-advocate / strategist / analyst) default to opus. No Opus access? Set judgment_model: sonnet in .claude/sdd.local.md — the supported path. Fable access? The floor rule lifts judgment to your session model automatically, or pin judgment_model: fable explicitly.
  • A model tier that turns out unavailable degrades, never blocks: the dispatch retries once on the session model and says so (skills/_shared/agent-roster.md).

Install

Claude Code — native plugin:

/plugin marketplace add genkovich/sdd
/plugin install sdd@sdd

After updating to a new release: re-run /plugin install sdd@sdd, then /reload-plugins.

Codex CLI — cd into your project first: the script installs into the current directory (.agents/skills/ + .codex/agents/). Add --global after codex to install under ~ instead, or --prefix DIR to install under an arbitrary directory (useful for trying it out in a sandbox):

cd your-project
curl -fsSL https://raw.githubusercontent.com/genkovich/sdd/main/install.sh | bash -s -- codex

Then restart codex (skills are discovered at session start) and type $sdd-specify.

Alternative — the plugin marketplace. Note that add only registers the marketplace, it installs nothing by itself:

codex plugin marketplace add genkovich/sdd

then inside codex run /plugins, switch to the sdd marketplace tab and pick Install plugin. One naming nuance: the marketplace install registers the original skill names ($specify), while the installer script prefixes them — $sdd-specify — because bare names like review / design / api collide with generic skills. Pick one of the two paths, not both — they register different names for the same skills, so running both shows every skill twice. To undo the script install: re-run install.sh codex --uninstall from the same directory (or with the same --global / --prefix). To undo the marketplace install: /plugins → the sdd tab → uninstall (or remove the [plugins."sdd@…"] entry from ~/.codex/config.toml). The script warns when it detects a marketplace install already registered.

Windows note. The installer is a bash script — run it from Git Bash or WSL. The directories it writes (.agents/, .codex/, .cursor/) start with a dot, which Explorer hides by default — enable «Hidden items» (or dir /a) to see them.

Cursor (2.4+) — the same script; cd into your project first (installs into .cursor/skills/ + .cursor/agents/ of the current directory; --global for ~, --prefix DIR for an arbitrary directory):

cd your-project
curl -fsSL https://raw.githubusercontent.com/genkovich/sdd/main/install.sh | bash -s -- cursor

Then restart Cursor (or run Developer: Reload Window) and invoke a stage by typing / in the chat and picking sdd-specify. (Cursor also reads .agents/skills/, so a Codex install is already visible to Cursor.) Once the plugin is listed on the Cursor marketplace, installing from the in-app marketplace panel works too — project- or user-scoped.

How every Claude-specific mechanism — AskUserQuestion, subagents, /clear, the implement engine modes — maps to Codex / Cursor is one table: skills/_shared/tool-adapters.md.

Start here

The flow is a straight line: each stage writes a file the next one reads. Run them in order (the diagram + table are just below).

/sdd:survey                         ← once per repo: map an existing codebase, OR bootstrap an empty one
/sdd:specify checkout-discounts     ← interviews you, writes the spec (you don't bring one)
/sdd:design … → /sdd:implement … → /sdd:review … → /sdd:ship

Two things to know up front: survey runs once per repo — on an existing codebase it maps the current architecture to docs/architecture-map.md (every later stage reads it); on an empty repo it runs a short foundation session and scaffolds the skeleton (detail below). And specify creates the spec from a short interview — you bring the idea, not the document.

From there you walk the backbone in order. Each step reads the previous step's file and refuses if it's missing, so you can't skip ahead by accident.

Every stage ends with a copy-ready handoff block (skills/_shared/handoff.md): What I did + Review before continuing (links to the files it wrote, so you can eyeball them at the gate) + Run next — /clear, then the next /sdd:… command in a fenced block you copy in one click. The /clear matters because each stage is gated and re-reads its inputs from disk, so it needs no carryover — clearing keeps the context small and stops one stage's chatter from drifting into the next. (Loop-backs are the exception — when review bounces back to implement, you stay in context to iterate; utilities make /clear optional.) It looks like this:

## ✅ specify — checkout-discounts

**What I did**
- wrote docs/features/checkout-discounts/spec.md — size M (from .size); proposed commit `spec: checkout-discounts`

**Review before continuing**
- docs/features/checkout-discounts/spec.md — goals, user stories, the §5 acceptance criteria

**Run next**
1. /clear — mandatory (fresh context; the next stage re-reads its inputs from disk)
2. then run:  /sdd:clarify checkout-discounts

The flow

There are three kinds of skill. Most of your time is the backbone — a straight line you walk in order. A few are utilities you call whenever you need them. Two close the loop after the code is written.

flowchart LR
    IV[interview<br/>optional] -.-> S
    SV[survey<br/>once per repo] --> S
    subgraph backbone["BACKBONE — run in order"]
        S[specify] --> CL[clarify] --> D[design] --> SQ[sequences] --> DM[data-model] --> API[api] --> T[tasks] --> PT[plan-tests] --> IM[implement]
    end
    IM --> RV[review] --> SH[ship]
    subgraph util["UTILITIES — call anytime"]
        CS[classify-size]
        GL[glossary]
        ADR[decide-adr]
        FX[fix]
    end
    CL -.-> GL -.-> D
    SH --> done([shipped: PR + changelog])

Step 0 — survey (once per repo, before the backbone)

#SkillWhat it doesReads → Produces
0surveyExisting repo → scans once, persists the current architecture. Empty repo → level-adaptive foundation session → fixes the foundation + emits a scaffold tasks.json for implement.the repo → docs/architecture-map.md (+ scaffold tasks.json on greenfield)

Backbone — the straight line (run in order)

#SkillWhat it doesReads → Produces
1specifyInterviews you to capture the idea, writes the product spec + acceptance criteria (reads the architecture map for constraints)your idea, architecture-map.md → spec.md
2clarifySweeps the spec for ambiguities (a devil's-advocate pass), closes or defers eachspec.md → tightened spec.md
3designMatches the feature to your existing architecture (see below) + declares the target surfaces, writes the Arc42 SAD + C4 + ADRsspec.md (+ CONTEXT.md if present) → sad.md, adr/*
4sequencesDraws the runtime flows as Mermaid sequence diagramssad.md → sad.md §6
5data-modelDesigns the schema and writes the actual forward+rollback migrations — staged under the feature folder, not the live tree (implement promotes them)spec.md, sad.md, sequences → data-model.md, staged migrations/*.up/down.sql
6apiDerives the OpenAPI contract from the data model (or the existing schema on the fast lane) + sequences + specdata-model.md, sequences, spec.md → contracts/openapi.yaml
7tasksBreaks the work into atomic ≤1-day tasks + a tasks.json dependency DAGall of the above → tasks/*, tasks.json
8plan-testsMaps every acceptance criterion to ≥1 test (inline in the spec for XS/S)spec.md, data-model.md → test-plan.md (M+) or an inline ## Test plan in spec.md (XS/S)
9implementThe TDD engine: writes a failing test, makes it pass, gates, commits — per task; promotes each staged migration into the live migrations/ as it buildstasks.json + all artifacts → code + tests + promoted migrations, committed

Close the loop (after the code is written)

#SkillWhat it doesReads → Produces
10reviewAn independent, clean-context code review of the whole change against spec/AC + qualitythe diff + spec.md → review record, PASS / CHANGES REQUESTED
11shipVerifies the feature actually runs (not just green tests), writes the changelog, opens the PRthe reviewed change → changelog + PR (never auto-merges)

review can bounce back to implement if it finds an unmet acceptance criterion. ship is the end: a reviewed, verified change with a changelog and an open PR — merging to main stays your call.

"We test and review, right?" Yes — in two places. implement runs a per-task gate (unit + integration + lint + vet) on every task as it goes, so each task is green before it's committed. Then review does the independent, whole-change code review a human reviewer would do on the PR, and ship runs the feature for real against its acceptance criteria. Tests-pass happens continuously inside implement; the cross-cutting review + real-world verification are the explicit review and ship steps.

Utilities — call whenever you need them (not part of the line)

  • interview (before specify) — stress-test a raw idea before you commit to a spec: a Socratic pass that surfaces hidden assumptions, names tradeoffs, and proposes sharper angles, ending with the weakest spot + the next step (usually /sdd:specify). Any idea, not just features; optional — reach for it when the idea itself isn't settled.
  • classify-size — size the feature XS/S/M/L/XL (writes .size); later skills read it to decide MVP vs full depth. Run it at the start, or any time scope changes.
  • glossary — capture a domain term in CONTEXT.md with a definition. Run it whenever a new term shows up; design and the spec read the glossary.
  • decide-adr — write a standalone ADR after the fact, when tasks (or a review) flags a decision that needs recording but wasn't captured during design.
  • fix — the bugfix entry point: reproduce, trace the symptom to the spec's acceptance criteria (regression / ambiguous AC / uncovered gap), pin it with a failing test, apply the minimal fix through the same gate implement runs, then patch the spec and write a fix record under _fixes/. Works on a repo with no specs at all (fixes code-first, recommends survey).

Interview depth (easy / medium / hard)

The Q&A skills open by setting a depth dial — one AskUserQuestion per run that tunes how much the skill decides on its own vs. interrogates you. It changes how many questions you get, never what gets covered:

  • easy — the skill makes the reversible, low-stakes calls itself with sensible defaults, asks only the irreversible / high-blast-radius ones, and lists every assumption it made so you can veto. Minimal analyses; diagrams written + summarized (no per-item question).
  • medium (default) — the balanced Socratic walk: one question per real decision.
  • hard — walk every decision with the trade-off foregrounded, run the full ideation analysis suite (competitive research, three strategic approaches, multi-perspective review, devil's-advocate), and probe edge cases harder.

The default is interview_depth in .claude/sdd.local.md (else medium); override it per run, or pass --depth=easy|medium|hard. Full semantics: skills/_shared/interview-depth.md.

Two things the dial never weakens — they hold at every level:

  • Readable diagrams. design and sequences confirm each diagram in prose (a plain-language walk of the flow + branches) and write the source to the file (where Obsidian renders it) — they never dump raw Mermaid into the terminal as the thing to approve. If mmdc is installed, an image is rendered too. (skills/_shared/diagram-presentation.md)
  • Full use-case + acceptance-criteria coverage. Every spec §4 user story and §5 AC is covered end-to-end: specify enforces a use-case floor (every user story carries ≥1 AC) and clarify re-catches a story that lost it; sequences maps each user story to a flow and each AC to a flow, a branch, or an explicit non-runtime N/A (no flow cap); and review traces the whole set through spec → sequences → data-model → api → tasks → implement, flagging anything that dropped out. Even easy/XS covers every use-case + AC — it just asks fewer questions about how.

Target surfaces (what's being built)

design opens §4 by declaring the feature's target surface(s) — what's being built — grounded in C4 container types: backend-service, web-frontend (SSR or SPA), mobile-app, desktop-app, cli, worker, library-sdk. The choice is derived from the spec's "for whom" (the spec stays product-level — it never names a surface), gated by the blast-radius gate (multi-surface usually spawns an ADR), drawn as one C4 container per surface in SAD §5, and written to the SAD frontmatter target_surfaces: [...]. Downstream stages read that declaration and gate their output by it — they never re-derive it:

  • api picks the contract form from the surface (HTTP/OpenAPI · gRPC · events · cli.md · public-api.md); a UI surface consumes the backend contract rather than authoring one.
  • sequences draws UI-driven flows (<user> → <ui> → <service>) for a UI surface.
  • tasks adds a ui task layer for a UI surface (backend-only stays domain/infra/app/ports).
  • plan-tests adds the component / visual-regression / e2e-through-UI tiers (the frontend "testing trophy") for a UI surface; implement detects the actual tools (Playwright / Storybook / …).
  • review traces every acceptance criterion through its surface — a UI AC to a component / e2e-through-UI test, not only a backend one.
  • Reuse, don't reinvent. survey inventories the existing design system / components / tokens / styling into architecture-map.md §Frontend; design / tasks / implement compose and extend it (modelled on the closest existing screen) instead of hand-rolling new UI — the frontend echo of the backend's match-the-repo + copy-the-closest-precedent.

It's Option B — frontend-awareness threaded through the existing stages (a ui layer, UI-architecture ADRs, UI flows, frontend test tiers); there is deliberately no separate component-tree / design-token / screen artifact. Full semantics: skills/_shared/surfaces.md.

Where the spec comes from

It's not an input you have to write — specify produces it. Its interview front asks 3–5 questions about the problem, the users, and what success looks like, then drafts the spec, validates each acceptance criterion with you, and runs a clean-context critic before writing spec.md. The idea is the input; the spec is the output.

Where we study the codebase / hold the current architecture

The existing system is studied once, in survey (Step 0), which persists docs/architecture-map.md — the current architecture: module layout, layering, datastores, conventions, and a C4 of what exists. That map is the single source of "what's already here":

  • specify reads it so the spec's constraints / non-goals reflect the real system (without leaking tech into the acceptance criteria).
  • design reads it and matches the feature to that reality — the SAD describes your system extended, not a greenfield design in a vacuum. It re-scans (via explorer) only if the map is missing or stale.
  • data-model and implement read it for the persistence + wiring conventions the new code must follow, instead of each re-discovering them.

So you don't re-open "what's the current architecture?" at every stage — survey answers it once and the map carries it. Refresh the map (survey again) when the repo has drifted past the reflects_commit it records. In design, decisions expensive to reverse cross a blast-radius gate and become ADRs.

On an empty project there's no current architecture to study — so survey establishes one. Its greenfield mode gauges how you want to engage, then picks the stack / structure / data approach / conventions with you (defaults-heavy), fixes them as the foundation (the same map, marked mode: greenfield-bootstrap, + foundational ADRs for the irreversible choices), and emits a scaffold tasks.json. implement then materializes the skeleton — anchored on a smoke test («builds + boots + the test and migration tooling run») rather than per-folder TDD. After that the repo is real and the per-feature flow builds into it normally.

The roadmap (the portfolio layer)

The backbone builds one feature at a time. roadmap is the layer above it — one living docs/roadmap.md that shows the work across features, kept at outcome altitude (the "why", not a feature-and-date list, which is the biggest source of planning waste):

  • Now — committed, spec'd, in progress. Each item links to its docs/features/<slug>/ (it doesn't restate the spec) + a status.
  • Next — problems/opportunities, deliberately not yet spec'd, ordered by a light RICE score (Reach × Impact × Confidence ÷ Effort). This is the candidate pool.
  • Later — directional outcomes/themes, no detail.
  • Shipped — what landed, with a link.

It stays current because the pipeline updates it: specify promotes a feature to Now, and ship moves it to Shipped — delivery itself keeps the roadmap in sync, so it doesn't rot. It carries a one-line "direction, not a promise" disclaimer and never carries dates.

The implementation engine

implement reads tasks.json, builds a dependency DAG, and runs a TDD cycle per task — SELECT → RED → GREEN → REFACTOR → GATE → COMMIT. It writes a failing test first, proves the failure is for the right reason, writes the minimal code to pass, keeps refactors green, runs the gate, and commits with SDD-Task / SDD-AC trailers.

Three execution modes, chosen automatically from settings + DAG shape (with graceful fallback):

  • Sequential single-agent TDD — the default and the floor everything degrades to.
  • Agent team (team_mode: true) — test-author → implementer → reviewer over the DAG, coordinated through a shared task list, one git worktree per agent.
  • Dynamic workflow (workflow_mode: auto) — a generated Workflow pipeline that fans out independent tasks up to a parallelism cap.

Models, effort & agents

Every skill and every agent declares an execution profile in its frontmatter — which model, how much reasoning effort, and which agents it spawns:

# a skill's frontmatter
model: inherit     # skills run on the session model; agents pin role-fit tier-alias defaults (haiku|sonnet|opus), overridable at dispatch
effort: high       # low | medium | high | xhigh | max
agents: [critic]   # the agents this skill spawns

Model is chosen by the kind of work, not by taste:

Kind of workModelEffortWho
Judgment (spec, design, review, critique, ambiguity, strategy)opus — the agents' default, one switch via judgment_model; the dispatching skills (specify, clarify, design, review) run on the session modelhighreviewer / critic / devils-advocate / strategist / analyst
Execution (write tests, write code)sonnetmedium → high on escalationtest-author, implementer
Research / gathering (+ web)sonnetmediumresearcher (competitive / adjacent-solution research)
Search / scan / derivationhaiku / inheritlow / mediumexplorer; data-model, api, sequences, tasks

The nine agents (agents/): explorer (brownfield scan), test-author (failing tests), implementer (makes them pass), reviewer (independent review), critic (coherence critique), devils-advocate (ambiguity + failure-mode hunt), researcher (competitive / web research), strategist (three strategic approaches), analyst (multi-perspective review) — the read-only ones run in clean isolated context (fresh eyes) and emit only cited findings. The last three are the ideation analyses, dispatched by specify and gated by the depth dial (easy skips them; hard runs the full suite).

Two policy levers sit on top of the table. judgment_model (.claude/sdd.local.md) moves all judgment agents (reviewer / critic / devils-advocate / strategist / analyst) in one switch — its value is open: a tier alias (haiku | sonnet | opus | fable), inherit, or a full model id (default opus; sonnet is the supported path without Opus access) — agents/*.md keep their tier-alias defaults; a per-role model_<role> key still wins. The default is a floor, not a pin: with the key unset, a session on a stronger tier than opus dispatches judgment at the session model (an explicit value is always honored literally), and an unavailable tier degrades — one retry on the session model, never a blocked stage. And on L/XL features the critical verifications — the reviewer in review and the critic in design/specify — run at effort: xhigh (via CLAUDE_CODE_EFFORT_LEVEL); the rest of the judgment work stays high.

The full policy — override precedence (env > invocation > model_<role> > judgment_model > frontmatter > session), the .size scaling, and the env-var fallback for the effort: no-op some builds have — lives in one place: skills/_shared/agent-roster.md. Short version: if a run feels under-reasoned, set CLAUDE_CODE_EFFORT_LEVEL.

Configuration — .claude/sdd.local.md

The pipeline auto-creates this per-project settings file (YAML frontmatter) with documented defaults the first time a skill needs it — normally specify at the start — and adds it to .gitignore (it's per-developer). The file is self-documenting: every key carries its default, its allowed values, and a one-line explanation inline. Edit it to change behaviour. Two keys are plugin-wide — interview_depth is read by the Q&A skills (specify / clarify / design) to pre-select the depth dial, and artifact_language is read by every artifact-writing skill: it sets the language pipeline documents are written in — prose only, while section headings, frontmatter and machine tokens stay English (full rule → skills/_shared/artifact-language.md); the rest configure the implement engine:

interview_depth: medium    # easy | medium | hard — default depth for specify/clarify/design
artifact_language: en      # en | uk — the language pipeline documents are written in (headings + machine tokens stay English)
tdd: true                  # enforce red→green→refactor
team_mode: false           # true → agent team via TeamCreate
workflow_mode: auto        # auto → dynamic Workflow; off → never
max_parallel_agents: 3
isolation: worktree        # worktree | inplace (parallel>1 ⇒ forces worktree)
stop_on_red: true
max_red_retries: 3
gate_lint: true
gate_vet: true
require_integration: auto  # auto | always | never (Docker-probed)
auto_commit: per_task      # per_task | per_phase | off
branch_strategy: feature   # feature | current
cmd_test_unit: ""          # empty = autodetect (escape hatch)
cmd_test_integration: ""
cmd_lint: ""
cmd_vet: ""
model_test_author: sonnet  # per-role model + effort (see Models, effort & agents)
model_implementer: sonnet
model_reviewer: opus
judgment_model: opus       # tier alias (haiku|sonnet|opus|fable), inherit, or a full model id — one switch for all judgment agents; sonnet = supported path without Opus access
effort_test_author: medium # raised to high on escalation / for L-XL features
effort_implementer: medium
effort_reviewer: high

Command detection is a stack-agnostic cascade: settings override → Makefile targets → package.json scripts → language manifests (go.mod, Cargo.toml, pyproject.toml, …) → Docker probe for the integration tier.

Quick start (idea → shipped)

The argument every stage takes is the feature slug — a kebab-case name you make up once at the start (here checkout-discounts). It becomes the folder every artifact lands in — docs/features/checkout-discounts/ — and is how each stage finds the previous stage's files, so use the same slug at every stage.

/sdd:survey                             # once per repo: map the current architecture
/sdd:specify       checkout-discounts   # interview → spec (reads the architecture map)
/sdd:clarify       checkout-discounts
/sdd:design        checkout-discounts
/sdd:sequences     checkout-discounts
/sdd:data-model    checkout-discounts
/sdd:api           checkout-discounts
/sdd:tasks         checkout-discounts
/sdd:plan-tests    checkout-discounts
/sdd:implement     checkout-discounts
/sdd:review        checkout-discounts   # independent review of the whole change
/sdd:ship          checkout-discounts   # verify it runs, changelog, PR

/clear between stages — each stage is gated, re-reads its inputs from disk, and ends by printing the next /sdd:… command to copy (the handoff block). Loop-backs (review → implement) stay in context; utilities make /clear optional.

Three notes on the first run:

  • You don't need classify-size to start — specify classifies the feature and writes .size itself when it's absent. Run /sdd:classify-size <slug> only to size it before specifying, or to re-classify when scope changes.
  • Skip the depth question by passing the dial inline: /sdd:specify checkout-discounts --depth=easy (also on clarify / design; values easy|medium|hard — see Interview depth).
  • Artifacts land in docs/features/<slug>/.

Routes — quick / standard / full

A small feature doesn't need the full backbone — and it shouldn't need a confirmation at every stage either. Alongside .size, classification writes a route to docs/features/<slug>/.route (one word: quick / standard / full; defaults XS/S → quick, M → standard, L/XL → full, confirmed together with the size in the same single question — you can always pick a different route). The route decides how each handoff treats the optional stages (clarify, sequences, data-model, api, plan-tests):

  • quick — the stage checks the skip condition itself: if the stage's work doesn't exist, it's auto-skipped with the reason stated («auto-skipped clarify: zero open questions»), and the ↳ or … line inverts to offer the full path instead. If the work does exist, the stage runs.
  • standard — today's behaviour: the handoff offers the skip as ↳ or … and you pick.
  • full — every optional stage runs; no skip alternatives are printed.

Example — a config-toggle-sized feature (quick route) in one session:

/sdd:specify  rate-limit-bump --depth=easy   # size XS + route quick confirmed in one question →
                                             #   zero open questions → auto-skips clarify (says why)
/sdd:design   rate-limit-bump                # one actor, no multi-step flow, no schema change →
                                             #   auto-skips sequences + data-model → next: api or tasks
/sdd:tasks    rate-limit-bump                # never skipped: implement consumes tasks.json
/sdd:implement rate-limit-bump               # test plan lives inline in spec.md on quick
/sdd:review   rate-limit-bump
/sdd:ship     rate-limit-bump

The skip conditions (clarify — zero open questions; sequences — no multi-step flow; data-model — no schema change; api — no contract change; plan-tests — inline in the spec) are canonical in skills/_shared/size-matrix.md — they're N/A conditions, not size defaults: an XS feature with a migration still runs data-model, on every route. The route steers handoffs only, it never locks a door: re-run /sdd:classify-size <slug> to switch routes mid-flight, or just invoke a skipped stage directly — it always runs.

When a stage refuses

Stages are gated: each one hard-refuses when the artifact it consumes is missing and names the stage to run first. A refusal is not an error — it's the pipeline telling you which step was skipped. The ones you're most likely to meet:

RefusalWhat it meansWhat to do
design: «run specify first»there's no spec.md for this slug yet (or the slug is spelled differently)run /sdd:specify <slug>; check the slug matches the folder under docs/features/
api: «run data-model first»the feature changes the schema but has no data-model.md — the contract can't be invented field-by-field. (No schema change → api doesn't refuse: it derives from the existing schema — the legal fast-lane skip)run /sdd:data-model <slug>
tasks: «no Accepted ADR»design spawned no ADR (rare — usually a sign the SAD walk was cut short)run /sdd:decide-adr <slug> for the key decision, or re-run /sdd:design <slug>

Repository layout

.claude-plugin/   plugin.json + marketplace.json (self-marketplace)
.codex-plugin/    Codex CLI plugin manifest (+ .agents/plugins/marketplace.json — its self-marketplace)
.cursor-plugin/   Cursor plugin manifest (skills/ + agents/ auto-discovered from the root)
install.sh        Codex CLI / Cursor installer — copies the subtree, prefixes skill names, generates functional agents
agents/           explorer, test-author, implementer, reviewer, critic, devils-advocate, researcher, strategist, analyst
scripts/          validate_plugin.py (CI gate: manifests + skill/agent frontmatter + the consistency invariants — links resolve, /sdd: form, handoff block, single-source taxonomy, no _shared orphans)
skills/_shared/   canonical socratic-loop / critic / size-matrix / ask-style / interview-depth / diagram-presentation / surfaces / handoff / tool-adapters (referenced, not duplicated)
skills/<name>/    SKILL.md spine + references/ (heavy detail) + templates/ (output scaffolds)
.mcp.json         declares the sdd-dashboard MCP server (auto-starts at session open; opt-in via dashboard_enabled)
server/           the dashboard MCP server (Bun + TypeScript): server.ts (MCP stdio + Bun.serve HTTP/WS), http.ts (routing + gating, testable), state.ts (disk→pipeline derivation), channel.ts (dashboard_* tools + command allowlist), paths.ts (docs/ scoping), frontmatter.ts (shared parser) + tests/ (bun test)
dashboard/        the browser UI (vanilla JS, terminal-green, read-only): index.html + app.js + style.css + vendor/ (marked, mermaid — vendored, offline; mermaid lazy-loads)

Roadmap

Directions under consideration — not promises, no dates:

  • sync — spec↔code drift detection: re-derive what the code actually does and diff it against the spec/SAD, so long-lived features don't quietly outgrow their documents.
  • Traceability matrix + adherence score — review/ship emit a single AC × (flow / contract / task / test / commit) matrix with a coverage score, instead of prose-only tracing.
  • Tracker integration — tasks.json ⇄ Jira / Linear / GitHub Issues two-way sync (today the export is one-shot and copy-paste).
  • Constitution file — a repo-level set of inviolable rules (security, compliance, style) every stage reads and the validator enforces, complementing the per-feature artifacts.

Shipped: MCP exposure → see The visual dashboard below.

The visual dashboard (opt-in)

The roadmap's "MCP exposure — pipeline state served over MCP so external tools and dashboards can read where every feature stands" has shipped — and gained a control surface. The plugin carries an sdd-dashboard MCP server (server/, Bun + TypeScript) that auto-starts with every Claude Code session (declared in .mcp.json) and, when enabled, serves a local browser dashboard (dashboard/) on 127.0.0.1. It reads every feature off disk (docs/features/<slug>/), shows its pipeline as a per-step checklist — done / skipped / pending / blocked — and renders each artifact (markdown + mermaid diagrams from vendored libs, fully offline; OpenAPI as plain YAML). Artifacts render in whatever language they're written — the state derivation reads only the English structural tokens, which never translate (see artifact_language above). Pure-markdown users who never opt in are unaffected — nothing binds, nothing opens.

Launch it — three steps

  1. Install Bun (the server runtime — the same dependency the official Telegram plugin uses): curl -fsSL https://bun.sh/install | bash or brew install bun.
  2. Set dashboard_enabled: true in your project's .claude/sdd.local.md (see Configuration).
  3. Run /sdd:start in your Claude Code session. The server is already running — it auto-started with the session; this step just hands it your project directory, binds the port if needed, and prints the URL: http://127.0.0.1:<port>/?session=<id>&token=<capability-token>. Open that exact URL in a browser — the token in it authorises the session.

A new session (or a server restart) mints a new token, so an old tab goes stale: re-run /sdd:start and open the fresh URL.

How the panel updates

Three mechanisms, layered:

  1. Live, from disk. The server watches docs/ (fs.watch) and pushes a refresh over the WebSocket whenever an artifact changes — no matter who changed it: a dashboard-driven run, a skill you ran in the terminal, or you editing spec.md in vim. Changes appear within ~1 second.
  2. Enriched, from Claude. When Claude runs a stage it also calls dashboard_update / dashboard_log / dashboard_done — that is what feeds the live activity feed, stage transitions, review verdicts and the final handoff. A terminal-only run still refreshes the artifacts (mechanism 1); it just doesn't narrate.
  3. Self-healing connection. The server pings the WebSocket to keep it alive; if it drops anyway, the browser reconnects with backoff and re-syncs everything from disk — nothing stays stale.

How you control it

The ▶ Run next stage / per-stage run / ⚒ Fix (appears on a CHANGES REQUESTED review) / + new buttons drive your live session — with honest asynchronous semantics:

  • A click sends the request to the server, which builds a validated /sdd:<skill> <slug> command from a strict server-side allowlist and queues it into your Claude session — over the same channel mechanism the official Telegram plugin uses (notifications/claude/channel).
  • The session consumes a queued command only while idle at the prompt. If Claude is mid-task, the command waits; every queued command gets its own queued → running → done status line and the UI never fakes synchronous execution.
  • The depth selector (topbar) sets --depth for dashboard-driven runs: easy (default — skills self-decide reversible calls and rarely block on questions), medium, or hard.
  • If a dashboard-driven run genuinely needs a human decision, Claude posts the question into the panel (dashboard_ask): a card with 2–4 option buttons appears in the activity pane and the run pauses; your click sends the answer back through the same queue and the run resumes. The browser only ever sends an option index — the option text was authored by Claude itself. You can always answer in the terminal instead.
  • Free browser text can never become a command — only the validated skill name + slug + depth pass the allowlist.

What the panel does NOT do

  • It never writes to disk — artifacts are edited only by the pipeline in your terminal.
  • It has no chat input, and a blocking AskUserQuestion in the terminal stays terminal-only — the panel's option cards exist precisely so dashboard-driven runs don't block there, but free text never travels from the browser into the session.
  • It doesn't survive a server restart — re-run /sdd:start for a fresh URL/token.

Setup, config & troubleshooting: server/README.md.

Security: binds loopback only; the API is read-only and every read is realpath-contained to docs/ with an extension allowlist; all routes require a per-session capability token; inbound commands are built only from a server-side skill + slug allowlist (browser text never becomes an arbitrary /sdd: command).

License

MIT © Kyrylo Genkov. See LICENSE.

其他

中风险

  • 来源需自行核对维护者身份。
  • 包含脚本或命令调用,安装前请复核。
  • 可能需要外部 token、网络权限或第三方服务。
  • 未检测到高风险命令。
  • 扫描发现:1 条。

Codex — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/genkovich/sdd.git
  3. 将 "skills/api" 文件夹复制到 Codex 的 skills 目录中。
  4. 重启 Codex 让新的 skill 生效。

Codex — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Codex 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Codex 让新的 skill 生效。

Claude Code — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/genkovich/sdd.git
  3. 将 "skills/api" 文件夹复制到 Claude Code 的 skills 目录中。
  4. 重启 Claude Code 让新的 skill 生效。

Claude Code — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Claude Code 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Claude Code 让新的 skill 生效。

Cursor — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/genkovich/sdd.git
  3. 将 "skills/api" 文件夹复制到 Cursor 的 skills 目录中。
  4. 重启 Cursor 让新的 skill 生效。

Cursor — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Cursor 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Cursor 让新的 skill 生效。

GitHub Copilot — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/genkovich/sdd.git
  3. 将 "skills/api" 文件夹复制到 GitHub Copilot 的 skills 目录中。
  4. 重启 GitHub Copilot 让新的 skill 生效。

GitHub Copilot — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 GitHub Copilot 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 GitHub Copilot 让新的 skill 生效。

Windsurf — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/genkovich/sdd.git
  3. 将 "skills/api" 文件夹复制到 Windsurf 的 skills 目录中。
  4. 重启 Windsurf 让新的 skill 生效。

Windsurf — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Windsurf 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Windsurf 让新的 skill 生效。
查看 SKILL.md 原文
name: api
model: inherit
effort: medium
agents: []
description: >
  Use to derive the API contract for a feature — an OpenAPI 3.1 document at
  docs/features/{slug}/contracts/openapi.yaml plus a drift/sync report (and an events doc when
  the feature has async flows). Triggers on "api for {slug}", "openapi for {slug}",
  "API contract for {slug}", "lock the interface for {slug}", "events for {slug}",
  "/sdd:api {slug}", "контракт API для {slug}", "OpenAPI для {slug}", "опиши ендпоінти".
  The contract is never hand-written: it is a derived function of data-model.md (typed fields +
  constraints), the sad.md §6 sequence diagrams (error branches, async actors), and spec.md
  acceptance criteria. Runs an inline drift check (does the contract match the model and the
  sequences?) and a reconcile mode. Hard-refuse if data-model.md is missing AND the feature
  changes the schema → run `data-model {slug}` first; on a legal no-schema-change skip it
  derives from the existing schema instead.

Skill: api

Projects the upstream artifacts into one interface contract. By default that's an HTTP/OpenAPI contract; this skill is interface-kind aware — and the kind comes from the surface(s) design declared in sad.md frontmatter target_surfaces, read here, not re-derived (→ ../_shared/surfaces.md). For a non-HTTP project it produces the matching contract form (or steps aside):

  • HTTP / REST (default) → contracts/openapi.yaml (OpenAPI 3.1) + api-sync-report.md.
  • gRPC / RPC → a .proto (or the repo's IDL) with the same derive-and-drift discipline.
  • CLI → contracts/cli.md — the command/flag/exit-code surface derived from the AC.
  • Library / SDK → contracts/public-api.md — the public signatures/types the feature exposes.
  • Event-only / worker → just contracts/events.md (no request/response surface).
  • No external interface (pure internal logic) → skip with a one-line note in the report; go straight to tasks.

Whatever the form, the contract is derived from data-model.md — or, on a legal no-schema-change skip, the existing schema — plus the sad.md §6 sequences + the spec's AC, never typed by hand — generation that diverges from the model or the sequences is the bug this skill exists to catch. The rest of this file details the HTTP path (the common case); the same derive → drift-check → reconcile loop applies to the other forms with the form-appropriate artifact.

This skill keeps only its own machinery. Question phrasing is shared → ../_shared/ask-style.md. Depth (events doc only when async; one resource vs full surface) follows the size matrix → ../_shared/size-matrix.md. The drift-resolution dialog reuses the shared 4-state actions — keep it short, point the machinery to _shared.

Contract summary/description prose follows artifact_language — paths, operationId, status codes and schema names never translate → ../_shared/artifact-language.md.

Owner

Backend Lead (drives the interface). The PM confirms each endpoint maps to a real user story; a frontend / consumer engineer is the first reader — the contract is locked before they start integration.

Inputs

  • <slug> — same feature slug used by every earlier stage.
  • Gate (conditional — hard-refuse only when a schema change exists): docs/features/<slug>/data-model.md. When present, it is the source of typed fields and constraints. When absent, evaluate data-model's N/A condition (no schema change — size-matrix fast lane) yourself: sad.md §5 declares no new building blocks/entities, no staged docs/features/<slug>/migrations/, and the spec introduces no new entity → proceed, deriving types/constraints from the existing schema (the live migrations/ DDL + architecture-map.md §Migrations/§Conventions) and saying so loudly in the handoff. Absent and a schema change exists → STOP and point: «run data-model <slug> first — the contract is derived from its entities».
  • (Expected) sad.md frontmatter target_surfaces — picks the contract form (step 1). Absent or empty → warn («surfaces undeclared — re-run design, or proceeding as backend-service») and treat as [backend-service], falling back to the architecture-map derivation (→ ../_shared/surfaces.md).
  • (Expected) docs/features/<slug>/sad.md §6 — the Mermaid sequenceDiagram blocks. Their alt/else branches become the error responses; an async participant (<message-bus> / <external-system>) on a mutating flow marks its endpoint Idempotency-Key-required and seeds events.md. Absent → note the gap (error branches derived from spec.md §5 only — likely misses authorization branches) and still generate.
  • (Expected) docs/features/<slug>/spec.md — §4 user stories give the endpoint list; §5 acceptance criteria give the shape of each happy + error outcome. The spec deliberately holds no HTTP/status/error-code/SQL detail — that mapping is this skill's job.
  • (Optional) docs/features/<slug>/.size — depth hint. Absent → default to M (full surface) and say so loudly in the handoff — «size M (default — no .size; run /sdd:classify-size <slug>)». docs/features/<slug>/adr/*.md — override defaults (versioning, error format, auth scheme) when an ADR mandates it; CONTEXT.md (read both repo-root and docs/features/<slug>/ — per-feature wins → ../glossary/SKILL.md) — glossary terms become schema names verbatim. Existing contracts/openapi.yaml → diff and update in place, never overwrite whole-cloth.

Protocol

  1. Gate + interface kind + read. test -f docs/features/<slug>/data-model.md — a three-way gate, not a binary one:
    • present → derive from it (the default path below, unchanged);
    • absent + no schema change (evaluate data-model's N/A condition yourself: sad.md §5 names no new building blocks/entities, no staged docs/features/<slug>/migrations/, spec introduces no new entity) → PROCEED — this is the legal fast-lane skip: derive types/constraints from the existing schema (live migrations/ DDL + architecture-map.md §Migrations/§Conventions), record each field-origin as existing schema — <migration/DDL anchor> with the same confidence scale, and state loudly in the handoff: «data-model.md absent — legal fast-lane skip (no schema change); fields derived from the existing schema»;
    • absent + a schema change exists → refuse with the pointer above. Determine the interface kind — read sad.md frontmatter target_surfaces FIRST (design already declared it; the surface picks the contract form per ../_shared/surfaces.md: backend-service → OpenAPI / gRPC / events per its sub-kind; cli → contracts/cli.md; worker → contracts/events.md; library-sdk → contracts/public-api.md; a UI surface — web-frontend / mobile-app / desktop-app — consumes the backend contract, it does not author one). Fall back to deriving the kind from docs/architecture-map.md + the spec's capabilities only if the SAD or the field is absent (a greenfield run where design was skipped). HTTP/REST → the OpenAPI path below (the default, detailed here); gRPC/CLI/library/event-only → produce the matching contract form (see the intro) with this same derive→drift→reconcile loop; no external interface (pure internal logic) → skip to tasks with a one-line note in the report — this self-skip is api's N/A condition in the size-matrix fast lane. Then read data-model.md (entities, fields, types, constraints) — or, on a legal skip, the existing schema sources named above — plus sad.md §6 (flows + alt-branches + async actors), spec.md §4/§5. Surface a one-line "found / missing" note for sad.md and spec.md — never refuse on their absence, only narrow the derivation and record the gap.
  2. Copy the template. ./templates/openapi.yaml → docs/features/<slug>/contracts/openapi.yaml. If async flows exist, also ./templates/events.md → contracts/events.md. Fill info.description from spec.md §1 (why this API exists).
  3. Derive endpoints + schemas. One endpoint (or more) per §4 user story. Every request/response field traces to a data-model.md entity column — or, on a legal skip, to an existing-schema column (a live migration's DDL) — copy its constraints across (maxLength/pattern/enum from the model's bounded types). Never invent a field with no origin in any input — ask the user where it comes from. $ref every shared schema; no inline duplication. Lists paginate by cursor (?after=&before=&limit=), wrapped in {items, has_next, has_prev, next_cursor}.
  4. Derive error responses from the sequences. Each endpoint covered by a §6 flow: turn every alt … else … end branch into a responses entry. The error body is the unified envelope {code, message, details?}; code follows the neutral convention module.error_name (snake_case, e.g. lesson.not_owned, lesson.invalid_state) — a naming rule, not a language artifact. Map status by class (4xx client / 5xx server). This closes the spec's usual blind spot — §5 lists the happy path + a couple of errors; the sequences enumerate the authorization and concurrent-state branches the spec omits.
  5. Async + idempotency. A mutating endpoint whose §6 flow shows a retry note or an async actor is marked Idempotency-Key-required (state the TTL). For each async message, fill an events.md entry: event name module.action.vN, payload schema, producer, consumers, retry / dead-letter behaviour.
  6. Examples + placeholder data. Every operation carries a request example + a success example + an error example, using placeholder values only (<...>@example.test, +380 00 000 00 00, Test User) — never real PII.
  7. Inline DRIFT CHECK (bidirectional) + write the report. Compare the generated contract against the read artifacts and write docs/features/<slug>/contracts/api-sync-report.md — see ./references/drift-check.md. It has a field-origins table (one row per operation.field: path | origin | confidence) and a checklist. The check runs both directions:
    • forward (contract derived correctly): endpoint↔model, error-code↔repo, validation↔constraint, OpenAPI↔sequence.
    • back-feed (coverage cross-check): every spec.md §5 AC maps to ≥1 operation/response; every operation maps to a §4 user story + ≥1 AC; every sad.md §6 alt-branch has a response, and any error/authorization response the contract needs but no §6 flow shows is a sequence gap. A gap here is not an api bug — it's a hole upstream: surface it and resolve it as Save-as-OQ with the upstream stage as owner — the OQ row names the producing stage as owner (specify for a missing AC, sequences for a missing branch) with due «before the contract is finalized», so the source gets fixed through the standard 4-state machine, not a fifth action. A core finding failing (or ≥3 flags total) pauses the run — resolve each via the shared 4-state actions (../_shared/ask-style.md): Accept / Fix (the contract) / Save-as-OQ / Drop. A fix that belongs upstream (the spec's AC, the sequence) is the Save-as-OQ variant with the upstream stage as owner (see step 7's back-feed) — never a fifth action. Never silently edit the sources — surface the mismatch and let the human pick the right artifact (the contract, the spec's AC, or the sequence).
  8. Lint + write + commit. Suggest spectral lint contracts/openapi.yaml (add it to the project's check target if not yet wired). On a clean check, the files are written; propose commit api: <slug> contract. Then emit the stage-handoff block per ../_shared/handoff.md — What I did + Review (contracts/openapi.yaml, api-sync-report.md, + events.md if async) + Run next (/clear, then /sdd:tasks <slug>).

Reconcile mode

/sdd:api <slug> --reconcile. Re-derives after an upstream artifact changed (typically data-model.md arrived or was tightened after a thinner first pass). It re-reads inputs, tightens loose types where the model now has a constraint, refreshes the field-origins confidence column, and — the load-bearing part — surfaces any field that had an inferred origin but now disagrees with the model. That disagreement is real drift, not stale incompleteness. info.version is never bumped silently; the user does that with a CHANGELOG line.

Definition of Done

  • docs/features/<slug>/contracts/openapi.yaml written: OpenAPI 3.1, BearerAuth global with public endpoints declaring explicit security: [], every error response the {code, message, details?} envelope, every operation with examples, all shared types via $ref.
  • api-sync-report.md written alongside: field-origins table + the 4-point drift checklist, every core finding ✓ or explicitly resolved with the user.
  • Every endpoint maps to a §4 user story; every field traces to a data-model.md column (or, on a legal fast-lane skip, to an existing-schema column named in the field-origins table); every error code exists in the repo's error definitions (checked in the form the repo uses).
  • contracts/events.md present iff the feature has async flows; each event has a payload schema, producer, consumers, retry / DLQ note.
  • The step-7 bidirectional drift check + api-sync-report.md are this skill's structural self-check (../_shared/self-check.md); its result is reported in the handoff.

Anti-patterns

  • Contract written by hand, then the model/sequences bent to fit it. The arrow is one-way: model (or the existing schema on a legal skip) + sequences + spec → contract.
  • Skipping the drift check because "it was just generated, of course it matches". Generation can match the spec-as-read while diverging from the model or the sequences — different files, different authors. A clean 4/4 ✓ is cheap; a silent ✗ in prod is not.
  • Error responses from the spec only. §5 lists happy + a couple of errors; the §6 sequences hold the authorization and concurrent-state branches. Skipping them leaves blind spots.
  • Inventing a field with no origin in any input, or silently dropping one that left data-model.md (keep it with a # stale note and surface it — the human decides).
  • Tripping the hard refuse when the skip was legal — bouncing a no-schema-change feature back to data-model just to produce an empty document. The gate is conditional (step 1): refuse only when a schema change actually exists; otherwise derive from the existing schema and say so.
  • Stack-specific schema or error names. Schemas use the domain language from data-model.md; error codes are the neutral module.error_name convention — not a Go/TS/Python idiom and not tied to any driver's error type.
  • Free-text errors ({"error": "failed"}), ?v=2 query versioning, nullable: true (3.0 style — use type: [string, null]), offset pagination, or real PII in examples.
  • Re-deriving the interface kind when design already declared it. target_surfaces in sad.md is the primary signal — read it; the architecture-map derivation is the fallback only when the SAD/field is absent (greenfield). Silently re-inferring HTTP-vs-events on every run is the double-derivation this skill's surface-awareness removes.

References & template

发现问题?提交给管理员复核

评分:

评论 (0)

暂无评论,成为第一个评论者吧!