复制安装命令
用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。
复制前请先查看来源、License 和安全提示。
A complete, opinionated library of Claude Skills covering the full lifecycle of building, launch...
用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。
复制前请先查看来源、License 和安全提示。
来源文件:README.md
A complete, opinionated library of Claude Skills covering the full lifecycle of building, launching, running, and growing a brand and a website.
103 stack-agnostic skills covering brand, design, content, SEO, dev, ops, growth, and research. Includes an Ahrefs MCP-powered SEO audit suite. Use them on Next.js, WordPress, Shopify, Webflow, plain HTML, or anything else.
Featured in awesome-claude-skills under Business & Marketing.
Add the marketplace, then install the plugin you want:
/plugin marketplace add rampstackco/claude-skills
# full catalog (103 skills)
/plugin install rampstack-skills@rampstack
# focused subsets
/plugin install rampstack-starter@rampstack
/plugin install rampstack-seo@rampstack
/plugin install rampstack-pm@rampstack
Prefer a lighter marketplace that lists only the curated subsets (no full catalog)? Add rampstackco/plugins instead and install the same three plugins from there:
/plugin marketplace add rampstackco/plugins
/plugin install rampstack-starter@rampstack
Skills load on demand: each contributes roughly its name and description until Claude needs it.
Claude Skills are reusable capability packages that teach Claude how to handle a specific kind of task with a consistent framework, vocabulary, and output format. Each skill is a folder containing a SKILL.md (instructions plus YAML metadata) and optional reference files (templates, checklists, worked examples). Claude loads a skill automatically when a user request matches the skill's description.
Skills work across Claude.ai, Claude Code, and the Anthropic API. Once you write a skill, it is portable across all three.
For the official deep dive, see Anthropic's Agent Skills documentation.
This is not a curated list of other people's skills. It is a single, opinionated library where every skill follows the same structure and conventions, so the skills compose cleanly across a real project lifecycle.
What you get:
SKILL.md and at least one reference filecreative-brief points to brand-voice. incident-response points to monitoring-and-alerting. Each skill's "When NOT to use" tells you which sibling fits your adjacent work.Highlight categories: brand strategy and identity, design systems, content production with full Tier 1 and Tier 2 coverage, full SEO suite (foundation plus Ahrefs MCP-powered audit suite), product management with experimentation and gap-closing tracks, growth tooling for interactive web tools, paid media discipline, frontend dev and accessibility, performance and QA, launch and incident ops, UX research, plus a meta-skill that teaches you to write your own.
Six entry-point skills, one per audience track. Run any of these standalone, or compose them with the rest of the catalog.
| Skill | What it does |
|---|---|
creative-direction (Brand and creative) | Four-axis brief (tone, aesthetic, audience, sensory ambition) that gives every downstream skill a coherent direction |
experiment-design (PM, experimentation) | From hypothesis to decision: sample size, duration, segment analysis, and the failure modes that produce wrong shipping calls |
feature-launch-playbook (PM, gap-closing) | The discipline of launching a feature well: positioning, internal alignment, customer comms, enablement, rollout, monitoring |
pillar-content-architecture (Content) | Hub-and-cluster topical authority: pillar selection, cluster planning, internal linking, refresh discipline |
landing-page-copy (Marketing) | Landing pages, sales pages, hero-to-CTA flow with copy that converts |
funnel-flow-architecture (Growth tooling) | Cross-tool conversion flows architected to match the audience and the funnel stage |
The creative-direction skill rendered as a live showcase →
Forty-two fictional brands generated from briefs that all use the same skill. Each is a fully styled brand site, not a mockup. The showcase demonstrates what the four-axis framework produces in practice and lets you filter by axis position to see how each combination renders.
The skill defines four axes: tone, aesthetic, relationship, sensory. The showcase lets you filter by any combination and see which examples match. Pre-filtered URLs deep-link from the SKILL.md and axes-explained reference, so you can read about a position and click straight through to the rendered examples.
The framework is generative. The showcase is illustrative. Most rare-but-powerful combinations are valid creative choices that simply have not been built yet. Set Provocative + Editorial Restrained + Coach + Resonant and the grid is empty.
Same skill, same brief format. Four completely different visual systems. Notice that Pulse and Bloom share identical axis positions yet read as opposite visual languages. The reference brands and aesthetic interpretation do the rest.
![]() | ![]() |
| Pulse · music streaming Sound that moves with you. Playful / Expressive Maximalist / Companion / Resonant See Pulse demo example → | Forge · boutique fitness Show up. Get hammered. Provocative / Expressive Maximalist / Coach / Resonant See Forge demo example → |
![]() | ![]() |
| Bloom · adaptogenic soda Soda that loves you back. Playful / Expressive Maximalist / Companion / Resonant See Bloom demo example → | Observatory Editorial · observability tool An open-source tool that respects engineer time. Conversational / Editorial Restrained / Peer / Considered See Observatory demo example → |
See all the brands in the showcase →
The creative-direction skill lives at skills/creative-direction/. Install it (see below), give Claude a project name and a few inspiration references, and the skill walks you through producing a brief that downstream skills can consume. The brand sites in the showcase were built from briefs of exactly that shape.
The logo-design skill is rendered on rampstack.co as two parallel surfaces. The variant explorer goes deep on one brand at a time: a primary mark, variants across architectures, applied contexts. The taxonomy gallery goes wide across the architecture space: ten fictional marks demonstrating eight mark architectures (wordmark, lockup, monogram, letterform-as-symbol, abstract, pictorial, combination, emblem). Same skill, two different lenses.
Each brand has a primary mark plus variants across architectures and applied contexts. The logo-design skill walks through the discipline of choosing one architecture and rendering it consistently across the system the brand will actually use.
The brands are filterable by architecture, typographic register, and category. The intent is reference work, not consumable templates.
Ten fictional marks across eight mark architectures: wordmark, lockup, monogram, letterform-as-symbol, abstract, pictorial, combination, emblem. The taxonomy makes the architectural distinctions concrete by showing all eight side-by-side, with three wordmarks at three typographic registers so the architectural label does less work than the execution.
Filter by architecture, vertical, or brand voice; click any mark card to read its design rationale.
The full catalog rendered as a 4-phase reference build: blank brief through deployed audited launch site. Threshold is a fictional PLG onboarding analytics product, but the research, brand foundations, build, and audit findings are all real. The reference build is the catalog's single strongest demonstration of how the skills compose end-to-end.
Phase 1: Strategy and research → Real Ahrefs keyword research, competitor analysis, content gap audit, and backlink opportunity mapping applied to a fictional B2B SaaS brief. Live data tables sourced from the Ahrefs API.
Phase 2: Brand and design → Logo system, color and typography tokens, working brand component primitives. The brand system renders live on the walkthrough page in real fonts and tokens, not just described.
Phase 3: Build and ship → The actual launch microsite built with Next.js using Phase 2's brand foundations. Live at rampstack.co/demo/threshold. Persistent demonstration banner; noindex; local-only waitlist form.
Phase 4: Audit and optimize → Real audit on the deployed site using the catalog's audit suite (axe-core, Lighthouse, manual checks). Real findings with severity, real fixes applied, real before/after metrics. The closing chapter where the catalog audits its own output.
The four phases compose into a working microsite at rampstack.co/demo/threshold. Real Next.js code, real brand foundations from Phase 2 referenced via cross-route imports, real working multi-step waitlist form (no data stored), persistent demonstration banner, and the inline data visualizations that came out of the post-audit polish pass.
Real client work cannot be open-sourced; portfolio claims trigger conflict-of-interest concerns in interviews and consulting conversations. A fictional product with a documented brief plus real research, real brand foundations, real working code, and real audit findings produces a teaching artifact that demonstrates methodology without claiming relationships. Threshold is a measurement tool that does not exist; the methodology that built it is the catalog working end-to-end.
Skills install in three different places depending on where you use Claude. Pick the platform that matches your workflow.
If your Claude.ai plan supports custom Skills:
.zip (one zip per skill folder containing SKILL.md and the references/ subfolder).Claude will load the skill automatically when your request matches its description.
For current plan availability and the exact upload UI, see Anthropic's Skills user guide.
Skills are first-class citizens in Claude Code. Drop them into your skills directory and Claude Code picks them up automatically.
User-level skills (available in every project):
# macOS / Linux
mkdir -p ~/.claude/skills
cp -r skills/* ~/.claude/skills/
# Windows (PowerShell)
New-Item -ItemType Directory -Force -Path "$HOME\.claude\skills"
Copy-Item -Recurse skills\* "$HOME\.claude\skills\"
Project-level skills (available only in a specific project):
mkdir -p .claude/skills
cp -r path/to/this-repo/skills/* .claude/skills/
Start (or restart) Claude Code. Skills load automatically.
For exact current paths and config flags, see the Claude Code documentation.
Use Skills programmatically by referencing them in your API calls. Skills must first be uploaded to your workspace (via the Console or API), then referenced by ID when creating messages.
For the current API surface, request format, and limits, see the Agent Skills API documentation.
You do not have to install all 103. Pick the categories that match your work. The library is modular: each skill stands on its own.
Once installed, skills trigger automatically based on your request. You do not have to name the skill or change how you talk to Claude.
You ask:
"Our organic traffic dropped 30% last week. Help me figure out why."
What happens:
Claude recognizes the request matches seo-traffic-diagnosis, loads the skill, and walks through its 5-layer root cause framework: confirm the change is real → localize the change → page-level analysis → technical analysis → external analysis. By the end, you have a hypothesis statement, evidence, and an action plan, structured the same way every time.
Other natural triggers:
creative-briefseo-onpageseo-backlink-auditseo-content-gap-audit plus content-strategyafter-action-reportskill-creation-walkthroughYou can also call a skill explicitly: "Use the seo-audit-orchestration skill to run a full audit on example.com."
The skills compose into a full project flow:
brand-discovery → brand-ideation → brand-identity → brand-style-guide → brand-voice
↓
creative-brief → information-architecture → content-strategy → design-system
↓
seo-keyword → seo-content-audit → content-and-copy → landing-page-copy
↓
seo-onpage → seo-technical → seo-aeo-geo → seo-offpage → seo-competitor
↓
frontend-component-build → accessibility-audit → performance-optimization
↓
code-review-web → qa-testing → security-baseline → launch-runbook
↓
domain-strategy → monitoring-and-alerting → backup-and-disaster-recovery
↓
incident-response → after-action-report
↓
analytics-strategy → cro-optimization → ux-research → usability-testing → journey-mapping
The SEO audit suite (Ahrefs MCP-powered) wraps around the SEO foundation skills:
seo-audit-orchestration
├── seo-site-health-audit
├── seo-backlink-audit
├── seo-keyword-gap-audit
├── seo-content-gap-audit
├── seo-traffic-diagnosis (also runs standalone for incident-style work)
└── seo-rank-tracking (ongoing, feeds the others)
The catalog also includes four audience tracks that compose alongside the foundational lifecycle. Each track has its own internal flow:
Paid media (Marketing track):
paid-media-strategy → ads-creative-development → ads-performance-analytics
Pairs with the paid media platforms in the integrations catalog at rampstack.co (Google Ads, Meta, LinkedIn, TikTok, plus Synter as the multi-platform aggregator).
Growth tooling (interactive web tools):
funnel-flow-architecture (orchestrator)
├── lead-magnet-design (capture)
├── calculator-design (capture / activate)
├── quiz-and-assessment-design (capture / activate)
├── multi-step-form-design (activate)
├── chatbot-flow-design (activate)
├── onboarding-wizard-design (activate)
├── interactive-product-tour (activate / convert)
├── upgrade-flow-design (convert)
├── scheduler-and-booking-design (convert)
├── comparison-tool-design (convert)
└── product-configurator-design (convert)
funnel-flow-architecture is the orchestrator: it sequences which interactive tool fits each audience and funnel stage, distinguishing matched-funnels from kitchen-sink-funnels.
Tier 2 content lifecycle:
content-strategy → pillar-content-architecture → content-brief-authoring
↓
content-and-copy / long-form-content-frameworks / email-sequences
↓
editorial-qa → content-distribution → programmatic-seo
↓
content-refresh-system → content-repurposing → content-migration
ai-content-collaboration is a workflow layer that runs across every phase rather than a single step. documentation-strategy operates continuously alongside the rest.
Tier 2 product management (two parallel tracks):
Experimentation track:
experiment-design → feature-flagging → experimentation-platform-orchestrator
↓
experimentation-analytics → data-warehouse-experimentation
Gap-closing track:
pm-spec-writing → roadmap-planning → feature-launch-playbook
↓
beta-program-management → product-analytics-setup → integration-orchestrator
The experimentation track ships changes with statistical discipline; the gap-closing track ships features with operational discipline. Both compose with the foundational lifecycle above.
Operations, cross-cutting, and team skills (stakeholder-communication, documentation-strategy, vendor-evaluation, team-onboarding-playbook, dependency-management, cost-optimization, etc.) cut across every track.
You can also pull individual skills for one-off work. Need just a backlink audit? Use seo-backlink-audit. Need to write a creative brief? Use creative-brief. Each skill stands on its own.
The skills compose with the tools your team already uses. 103 skills at the center; 35 integrations across 6 integration categories radiating out via MCPs.
This catalog is the open-source methodology layer. Commercial surfaces at rampstack.co extend it:
The skills in this repository remain free, open-source, and stack-agnostic. The surfaces above are how the same methodology is delivered as a product.
claude-skills follows the Agent Skills Specification, the open standard for portable AI agent skills originally developed by Anthropic and adopted across the AI tooling ecosystem (Claude Code, OpenAI Codex, Gemini CLI, GitHub Copilot, Cursor, VS Code, Goose, Spring AI, and 30+ other platforms as of early 2026).
Beyond the format itself, the catalog is designed around three principles aligned with the guidance Anthropic publishes in Building effective agents:
Simplicity. Each skill covers one focused capability rather than trying to be a multi-purpose document. A roadmap-planning skill plans roadmaps. A keyword-research skill researches keywords. Composing them together produces complex workflows; mixing them inside one skill produces unreliable ones.
Transparency. Every skill declares its scope, dependencies, and expected behavior in machine-readable YAML frontmatter. The catalog is inspectable by tooling, not just by humans reading prose.
Quality contracts via tooling. Structural and content quality is enforced through automated checks (run python .github/scripts/lint_skills.py) rather than convention alone. Every skill is validated against a schema. Every catalog change is validated in CI.
Skills in this catalog are designed to compose into the common agentic workflow patterns Anthropic documents: prompt chaining (sequential steps), routing (classify and direct), parallelization (sectioning or voting), orchestrator-workers (dynamic delegation), and evaluator-optimizer (iterative refinement).
Because the catalog conforms to the open Agent Skills standard, skills work across any platform supporting the specification without modification.
claude-skills is the parent catalog. Curated subsets and companion repos focus on specific specialties:
| Repo | Focus | Skills |
|---|---|---|
| claude-skills | Full catalog (you are here) | 103 |
| claude-skills-starter | General-purpose lite | 14 |
| claude-skills-seo | SEO consulting | 12 |
| claude-skills-pm | Product management | 12 |
| claude-skills-widgets | UI patterns + components | 65 + 32 |
| awesome-claude-skills | Curated discovery list | n/a |
Each family repo is MIT-licensed, conforms to the Agent Skills Specification, and is stack-agnostic. Use the full catalog for breadth; use a specialty subset when working in one domain.
The table above covers the skill catalogs. They are one part of a larger set, and the rest of it is below. All of it is public.
Skills. This repo is the canonical home for all skill content. Alongside it sits the workflows tier: fifteen multi-skill runbooks with their connectors, a getting-started guide, and published run records for the ones that have been executed as written.
Subsets. The five curated repos in the table above copy from this catalog with attribution and track it upstream.
Design direction themes. Thirteen sibling repos, each shipping annotated design tokens with their measured contrast ratios, a component layer, two Tailwind adapters, and a demo that opens from a file with nothing installed. They come in three artifact classes: seven surface registers, one layout archetype, and five shells, which ship a structure a site lives inside (a window manager or a board, a taskbar or a dock, an enhancement contract and a focus model) with the register they wear left swappable. VivaOcean, the animated-scene shell, is the showcase flagship. What the shell class settled, and why, is public in its class decision log, which every new shell reads first and continues. All thirteen are linked from the gallery at rampstack.co/themes.
Creative direction. The themes are not thirteen moods. Each one states its coordinates in the creative direction framework, which sets brand direction on four axes, and the showcase renders archetypes at each position on it.
Engines. Krine, Tholo, and Basano run on one runtime: Krine decides, Tholo builds, Basano proves. The engines page covers what the three share.
Research. The SERP event registry is a dated, sourced, confidence-tagged record of AI model releases, search feature changes, and confirmed algorithm updates, rendered on the site from the repository that holds it.
What shipped, and when, is recorded at rampstack.co/updates.
All 103 skills are shipped. Each has a complete SKILL.md plus at least one reference file (template, checklist, or playbook).
| Skill | What it does |
|---|---|
brand-discovery | Audience research, competitive scan, positioning territory exploration |
creative-brief | Project briefs that align stakeholders before work starts |
creative-direction | Four-axis aesthetic brief (tone, aesthetic, audience, sensory ambition) for cross-skill coherence |
information-architecture | Sitemap, navigation, URL structure, content types, taxonomy |
content-strategy | Editorial strategy, content calendar, topical authority planning |
| Skill | What it does |
|---|---|
brand-ideation | Naming, positioning territories, mood directions, narrative angles |
brand-identity | Logo system, color, typography, imagery, iconography, motion |
brand-style-guide | The canonical reference document for the full brand system |
brand-voice | Voice attributes, tone shifts, vocabulary, paired-example library |
brand-archetype-system | 12 archetype defaults across 18 verticals: color, type, voice, imagery starters |
logo-design | Logo variants across architectures (wordmark, lockup, monogram, letterform-as-symbol), with rationale and application specs |
creative-brief-selector | Live-reference-grounded creative briefs with divergence check against prior builds |
| Skill | What it does |
|---|---|
design-system | Component library, design tokens, design system documentation |
design-standards | Production-grade page and component design standards |
art-direction | Photography, illustration, and visual direction for campaigns |
vertical-site-conventions | Vertical page and site composition built to the experience bar |
| Skill | What it does |
|---|---|
pillar-content-architecture | Hub-level content architecture: pillar topic selection, cluster planning, internal linking, URL structure, pillar and cluster page anatomy, topical authority signals, refresh discipline |
content-brief-authoring | Per-piece editorial brief: target keyword, intent, audience, outline, entity coverage, internal linking, success criteria, and the discipline that distinguishes useful briefs from bloat |
content-and-copy | Website copy, blog content, content production frameworks |
landing-page-copy | Landing pages, sales pages, hero-to-CTA flow |
email-sequences | Onboarding flows, lifecycle campaigns, transactional copy |
programmatic-seo | Designing pSEO programs that work: data sources, template design, quality control at scale, internal linking, crawl budget, AEO/GEO patterns, refresh discipline, and when pSEO is and is not the right answer |
editorial-qa | Pre-publish QA framework: brief adherence, voice consistency, fact accuracy, AI-content audit, AEO/SEO compliance, sampling at scale, and the workflow that distinguishes catch-problems QA from process theater |
ai-content-collaboration | How humans and AI compose in content workflows: participation boundaries, hybrid patterns, voice ownership, the AI slop problem, disclosure and transparency, team calibration, and the ethics of honest AI-assisted production |
long-form-content-frameworks | Structural patterns for individual long-form pieces (case studies, whitepapers, research reports, definitive guides, manifestos, ebooks, long-form tutorials) that distinguish publication-quality work from bloggy-long padding or academic bloat |
content-refresh-system | Systematic content refresh: quarterly audits, refresh prioritization, refresh-vs-merge-vs-delete decisions, the lifecycle discipline that distinguishes intentional programs from set-and-forget decay |
content-repurposing | Cross-format content adaptation: one piece becomes many (blog series, email, social, webinar, podcast, video) with per-format adaptation rather than mass-blast that ignores medium constraints |
content-distribution | Content distribution discipline: owned, earned, and paid channels matched to audience and content type. Channel-fit decisions, distribution cadence, the strategic alternative to spam-everywhere or hope-and-pray |
evidence-based-reviews | Evidence tiers, methodology disclosure, honest review claims |
Tool-agnostic SEO skills. These define the conceptual frameworks. The SEO audit suite below adds the Ahrefs MCP-powered execution layer.
| Skill | What it does |
|---|---|
seo-onpage | Single-page audits and optimization across 8 dimensions |
seo-technical | Crawlability, indexability, rendering, schema, page experience |
seo-keyword | Discovery, intent classification, clustering, prioritization |
seo-competitor | SERP overlap, content gaps, backlink gaps, technical comparison |
seo-offpage | Link building, digital PR, citations, linkable assets |
seo-content-audit | Keep/update/merge/redirect/delete decisions across a site |
seo-aeo-geo | AI search optimization, llms.txt, extraction-friendly content |
End-to-end SEO audit workflows that pull data from the Ahrefs MCP and produce concrete deliverables. These skills assume the Ahrefs MCP is connected.
| Skill | What it does |
|---|---|
seo-audit-orchestration | Master orchestrator: sequences the suite, produces a rollup report |
seo-backlink-audit | Profile health, anchor mix, toxic links, reclamation, gap analysis |
seo-keyword-gap-audit | Competitor keyword gaps with opportunity scoring and clustering |
seo-content-gap-audit | Missing topics, thin coverage, outdated content, decay diagnosis |
seo-traffic-diagnosis | Diagnose drops, stalls, or wins via 5-layer root cause analysis |
seo-site-health-audit | Triage Ahrefs Site Audit findings by SEO impact, not severity |
seo-rank-tracking | Setup, baseline, segmentation, alerting, dashboarding |
| Skill | What it does |
|---|---|
pm-spec-writing | PRDs, user stories, acceptance criteria, dev briefs |
roadmap-planning | Quarterly planning, prioritization, dependency mapping |
integration-orchestrator | Sequence creative-direction work across phases, gates, handoffs, and QA verification |
experiment-design | Hypothesis to decision: sample size, duration, segment analysis, interpretation, and the failure modes that produce wrong shipping calls |
feature-flagging | Flags as production infrastructure: types, naming, lifecycle, targeting, rollout, stale flag cleanup, governance |
experimentation-analytics | Read result panels without fooling yourself: confidence intervals, p-values, multiple testing, sequential testing, CUPED, ratio metrics, network effects, dashboard reconciliation |
experimentation-platform-orchestrator | Pick the right experimentation platform, migrate when wrong, coordinate when multi-platform: a decision framework for Statsig, PostHog, GrowthBook, Optimizely, Amplitude, Eppo, Kameleoon |
product-analytics-setup | Instrument product analytics correctly: event taxonomy, properties, naming conventions, schema versioning, funnels, retention cohorts, North Star selection, and the instrumentation debt that compounds without discipline |
data-warehouse-experimentation | Run experiments out of the warehouse: SQL assignment, exposure logs, dbt metric definitions, statistical analysis, variance reduction with CUPED, sequential testing, and the operational tradeoffs vs platforms |
feature-launch-playbook | The operational discipline of launching a feature well: positioning, internal alignment, customer comms, sales enablement, support readiness, rollout strategy, monitoring, and post-launch measurement |
jtbd-framing | Jobs-to-be-Done framework. Job statements, struggling moments, hire/fire criteria, the difference between feature-thinking and job-thinking. Honest about where JTBD earns its keep and where it becomes performative |
okr-design | OKR design discipline. Outcome statements, key results, scoring, mid-quarter recalibration. Distinguishes sandbagged OKRs (always hit, useless) from aspirational fantasy (impossible, demoralizing) from stretch OKRs (genuine ambition with quarterly accountability) |
beta-program-management | Running betas that produce real signal. Participant selection, structured feedback, beta-to-GA decisions. Distinguishes soft-launch (no structure) from kitchen-sink (everyone in) from structured-beta (calibrated cohort with intentional feedback loops) |
| Skill | What it does |
|---|---|
code-review-web | PR review, build error diagnosis, security and quality checks |
frontend-component-build | Component architecture, props design, accessibility from the start |
accessibility-audit | WCAG compliance audit with remediation plan |
performance-optimization | Core Web Vitals, asset optimization, render performance |
| Skill | What it does |
|---|---|
qa-testing | Pre-launch QA, regression testing, cross-browser checks |
| Skill | What it does |
|---|---|
launch-runbook | Go-live runbook, DNS cutover, deploy day procedures |
incident-response | Incident triage, comms, mitigation, escalation |
after-action-report | Post-mortems, retros, learnings documentation |
domain-strategy | DNS architecture, redirects, registrars, multi-domain portfolios |
monitoring-and-alerting | SLO design, uptime checks, alert routing, on-call rotations |
backup-and-disaster-recovery | RPO/RTO targets, backup strategy, restoration drills |
security-baseline | HTTPS, security headers, CSP, secrets management, vulnerability scans |
email-deliverability | DMARC, SPF, DKIM, sender reputation, deliverability monitoring |
media-asset-management | Image pipelines, video hosting, asset libraries, format selection |
| Skill | What it does |
|---|---|
analytics-strategy | Measurement frameworks, dashboard design, event taxonomy |
cro-optimization | Hypothesis-driven testing, conversion optimization |
Interactive web tools that turn visitors into leads. Lead magnets, calculators, quizzes, multi-step forms, chatbots, and the cross-tool funnel architecture that orchestrates them.
| Skill | What it does |
|---|---|
lead-magnet-design | Designing gated content that earns the email. Distinguishes thin-bait (overpromises, underdelivers) from kitchen-sink-resource (everything, helps with nothing) from earned-value-magnet (delivers standalone value while qualifying the lead) |
calculator-design | Designing interactive calculators that deliver decision-support value while qualifying leads. Distinguishes vanity-calculator (no real value) from lead-trap (hides answer behind email) from transparent-decision-tool (gives genuine value, captures leads honestly) |
quiz-and-assessment-design | Designing quizzes and assessments that produce actionable segmentation. Distinguishes clickbait-quiz (engagement only) from vanity-result (entertaining, not useful) from actionable-segmentation (genuine categorization that drives next-step recommendations) |
multi-step-form-design | Designing multi-step forms that respect cognitive load while maintaining completion intent. Distinguishes kitchen-sink-single-page (overwhelms) from progress-theater (steps without genuine staging) from genuinely-staged (each step earns its own page) |
chatbot-flow-design | Designing conversational flows for chatbots and AI agents on websites. Distinguishes scripted-bot (rigid trees, fail edge cases) from hallucinating-bot (LLM without structure, makes things up) from structured-guided-conversation (LLM-powered with intent architecture and fallback discipline) |
funnel-flow-architecture | Architecting cross-tool conversion flows that match audience and stage. Distinguishes silo-funnels (every tool standalone) from kitchen-sink-funnels (every audience squeezed through one path) from matched-funnels (architecture matched to audience-and-stage) |
onboarding-wizard-design | Designing first-run product onboarding wizards. Distinguishes tutorial-overload (dump everything upfront) from skip-friendly-empty (skipped onboarding leads to abandoned product) from earned-progressive-disclosure (right things at the right moments) |
interactive-product-tour | Designing in-product tours and contextual help. Distinguishes tooltip-spam (every button has a tour stop) from one-and-done (tour shows once, never seen again) from contextual-when-needed (surfaces help at the moment friction occurs) |
upgrade-flow-design | Designing free-to-paid conversion flows. Distinguishes paywall-everywhere (gates everything aggressively) from free-forever-trap (no upgrade path surfaces) from value-triggered-upgrade (paywall surfaces at moments of demonstrated value) |
scheduler-and-booking-design | Designing schedulers and booking flows. Distinguishes any-time-friction (no qualification, just a booking link) from interrogation-gate (so much qualification it scares users off) from qualified-fast-path (just enough qualification to set up the call well) |
comparison-tool-design | Designing comparison tools that help users decide. Distinguishes feature-list-dump (every feature in a row, no decision support) from hidden-recommendation (biased comparison pretending to be neutral) from honest-comparison-with-guidance (genuine comparison plus opinionated recommendation) |
product-configurator-design | Designing interactive product configurators. Distinguishes infinite-options (decision paralysis from too many options) from canned-bundles-only (no real customization) from guided-configuration (smart defaults plus meaningful constraints plus escape hatches) |
Paid media discipline: strategy, creative, and performance analytics. Pairs with the paid media platforms in the /integrations catalog at rampstack.co.
| Skill | What it does |
|---|---|
paid-media-strategy | Hypothesis to spend: channel selection, budget allocation, audience targeting, bid strategy, attribution reality, and the failure modes that burn agency-scale budgets |
ads-creative-development | Hook patterns, format selection, video pacing, variation systems, testing methodology, fatigue detection, and the platform-specific creative norms that separate ads from clutter |
ads-performance-analytics | Read paid media dashboards without fooling yourself: attribution models, platform reporting quirks, ROAS vs LTV, multi-platform reconciliation, incrementality testing, and the interpretation failures that compound into wasted budget |
| Skill | What it does |
|---|---|
ux-research | Research planning, user interviews, qualitative synthesis |
usability-testing | Test design, moderation, findings reports |
journey-mapping | Customer journey maps, service blueprints, friction analysis |
discovery-research-synthesis | Synthesizing customer interviews, research notes, and support tickets into actionable PM decisions. Distinguishes data-dump (no synthesis) from insight-theater (overpolished narrative) from actionable synthesis (decision-grade clarity) |
user-feedback-aggregation | Collecting and synthesizing user feedback across channels into continuous decision signal. Triage discipline that distinguishes loudest-voice (whoever complains most) from averaged-noise (every signal weighted equally) from triaged-synthesis (weighted by source quality and decision relevance) |
competitor-experience-audit | Cross-site experience patterns and gaps across a vertical |
| Skill | What it does |
|---|---|
form-strategy | Form design, validation patterns, spam prevention, conversion tuning |
content-migration | Platform migrations with SEO equity preservation |
internationalization | Locale strategy, hreflang, translation workflow, RTL design |
dependency-management | Package updates, security patches, lockfile hygiene |
cost-optimization | Infrastructure spend audits, rightsizing, contract negotiation |
| Skill | What it does |
|---|---|
stakeholder-communication | Status updates, exec readouts, project communications |
documentation-strategy | Documentation systems, what to document, maintenance cadence |
vendor-evaluation | Tool and vendor selection using a structured rubric |
team-onboarding-playbook | 30-60-90 onboarding plans for new hires and contractors |
skill-creation-walkthrough | The meta-skill: how to write your own custom skills |
Skills compose best when Claude has live access to your data and tools. Model Context Protocol (MCP) servers provide that bridge. The skills in this library work without any MCPs, but pair them with the right ones and they go from "frameworks Claude follows" to "workflows Claude executes against your real systems."
Below is the MCP shortlist by skill area. None of these are required (except the Ahrefs MCP for the SEO audit suite). All are categorical recommendations: where multiple options exist for the same job, pick the one that fits your stack.
The SEO audit suite (skills 23-29) is built around Ahrefs as its primary backend; foundation SEO skills (16-22) work with any equivalent. Competitive intelligence MCPs (Ahrefs, Semrush, Similarweb) cover overlapping but distinct data shapes: backlinks and keywords, traffic estimation, audience behavior. Use them in combination for the strongest signal.
A note on MCP costs: many of these MCPs are wrappers around APIs you are already paying for through a subscription, where MCP calls do not add marginal cost. Others (Ahrefs, Semrush, Similarweb, DataForSEO) use paid API credits per call, and long agentic sessions against these platforms can burn meaningful credit volume quickly. The cost model is documented on each integration's landing page at rampstack.co/integrations. Free with rate limits is called out where it applies (Google Search Console, PageSpeed Insights). When in doubt, check the platform's API pricing before running multi-hour agent workflows.
Backlink and keyword data
seo-audit-orchestration and the 6 audit suite skills (backlink, keyword gap, content gap, traffic, site health, rank tracking). Credits-per-call.seo-keyword, seo-competitor, seo-content-gap-audit. Verify the official MCP endpoint at authoring time; Semrush has shipped first-party MCP tooling. Credits-per-call.Traffic estimation and competitive intelligence
seo-competitor, seo-traffic-diagnosis (external-factor layer), brand-discovery (competitive scan), analytics-strategy (industry benchmarks). Where Ahrefs answers "how do they rank" and Semrush answers "what keywords drive what," Similarweb answers "how much traffic, from where, from whom." Credits-per-call.Search Console and Core Web Vitals
seo-traffic-diagnosis and any audit that needs ground-truth click and impression data. Free with rate limits.performance-optimization and seo-site-health-audit for Core Web Vitals field data. Free with rate limits.code-review-web, pm-spec-writing, roadmap-planning, incident-response. Lets Claude read PRs, file issues, search code, and reference real commits.monitoring-and-alerting and incident-response. Real error data turns generic incident frameworks into specific diagnoses.domain-strategy, security-baseline, performance-optimization. DNS records, redirects, page rules, security headers.launch-runbook and incident-response. Deployments, env vars, build logs.code-review-web, pm-spec-writing, backup-and-disaster-recovery. Schema, queries, edge functions.analytics-strategy, cro-optimization, journey-mapping. Event taxonomy review and funnel analysis grounded in real data.monitoring-and-alerting, incident-response. SLO design and alert routing against actual metrics.incident-response, stakeholder-communication, after-action-report. Read channel context, draft updates, post incident comms.pm-spec-writing, roadmap-planning. Spec writing against the actual issue tracker, not a generic template.brand-discovery, seo-keyword, seo-competitor, ux-researchclaude mcp add in Claude Code for direct installationIf a skill in this library would benefit from a tool integration that does not yet exist, the MCP documentation walks through building one. The seo-audit-orchestration skill is a worked example of how to design a skill suite around a specific MCP's capabilities.
Every skill follows the same structure. See SKILL_AUTHORING.md for the full spec.
Highlights:
skills/
skill-name/
SKILL.md
references/
template.md
checklist.md
example.md
SKILL_AUTHORING.md (the authoring guide)
CONTRIBUTING.md (how to contribute)
MAPPING.md (origin notes for skills ported from existing work)
README.md (this file)
LICENSE (MIT)
Skills are instructions and code that run with your agent's permissions, so how
a catalog is maintained matters. Changes reach main only through pull requests
with signed commits and linear history. Each skill is hashed into a checksum
manifest (SKILLS.lock) you can verify against, and reviewed against a
documented safety checklist before it merges.
This process catches known classes of unsafe content and lets you confirm a skill matches the reviewed version. It is not a promise that any skill is risk-free. See SECURITY.md for the full process and how to report an issue.
Contributions are welcome. Whether you want to fix a typo, add a reference file, or propose an entirely new skill, the bar is the same: follow the uniform structure, keep the voice consistent, and prove the skill earns its place.
See CONTRIBUTING.md for the full process.
The fastest path: use the skill-creation-walkthrough skill itself. It teaches the same authoring discipline used across all 103 skills, with worked examples and a blank template.
Thanks to @IgnacioChiaravalle for the community feedback that shaped PR #36: a CONTRIBUTING.md typo fix, a cross-linking pass between SKILL.md files and their reference files, and the new ARIA patterns reference for the accessibility-audit skill.
MIT. Use it. Fork it. Ship things with it.
name: experiment-design
description: "A discipline for designing experiments (A/B tests, multivariate, holdouts) so the results actually answer the question you asked. Hypothesis writing, sample size, duration, segment analysis, running discipline, matching a result to a pre-committed decision rule, and the common failure modes that produce confidently wrong shipping decisions. Use this skill whenever the user is planning a test that has not run yet: framing a hypothesis, sizing the sample, setting duration, choosing guardrails, or deciding whether something is worth testing at all. Triggers on design an experiment, experiment plan, A/B test, split test, multivariate test, holdout, experiment hypothesis, sample size, minimum detectable effect, MDE, test duration, guardrail metric, no peeking, pre-committed decision rule, is this worth testing. Use `experimentation-analytics` instead when the test has already run and the question is how to read the result panel."
category: product
catalog_summary: "Hypothesis to decision: sample size, duration, segment analysis, interpretation, and the failure modes that produce wrong shipping calls"
display_order: 4A senior product manager's playbook for running experiments that produce trustworthy decisions.
The default state of experimentation in most companies is sloppy. PMs run tests against vague hypotheses, look at results too early, ignore guardrails, stratify into noise, and ship features whose lift is mostly measurement error. The cost is real: ship the wrong thing, kill the right thing, learn the wrong lesson, repeat.
This skill is the discipline that prevents most of those mistakes. It assumes you have a working experimentation platform (Statsig, PostHog, GrowthBook, Optimizely, Amplitude, Eppo, Kameleoon; the platform does not matter for the principles). It assumes you have product-design and engineering pipelines that can deliver real treatment changes. The hard part is the thinking, and that is what is here.
When to use this skill: any time you are about to design or interpret an experiment. Read the relevant section before you start, not after the test is running.
The skill spans the full experiment lifecycle. Pre-experiment readiness (is this thing even worth testing). Hypothesis design (cause, effect, magnitude, mechanism). Sample size and minimum detectable effect (do you have enough traffic to learn anything). Duration (how long is long enough, when does the cycle bias the result). Running discipline (no peeking, guardrails, sequential testing). Interpretation (the three buckets and the inconclusive case). Decision-making (matching the result to a pre-committed rule).
The skill does not cover feature flag operational mechanics; those live in the feature-flagging skill, which handles flag taxonomy, environment management, and stale-flag cleanup as a separate discipline. The skill does not cover statistical analysis depth; for delta methods, variance reduction techniques like CUPED, and Bayesian alternatives, see the experimentation-analytics skill. The skill does not cover platform-specific tooling; for MCP commands, auth models, and platform-specific configuration, consult the chosen platform's official documentation. This skill produces the experiment design; the platform implements it.
For the orchestration layer above (which experiments to run, in what order, with what cadence), see the forthcoming experimentation-platform-orchestrator skill. That skill schedules; this skill designs.
A defensible experiment design sits at the intersection of twelve considerations. Each is covered in detail in its own section below.
references/common-failures.md.The sections below cover each consideration in turn. Read the relevant section before running the experiment, not after.
The most important section in the skill. Most experiment failures trace back to a vague hypothesis.
A real hypothesis has four parts: cause, effect, magnitude, mechanism. Cause is the change you are making. Effect is the metric you expect to move. Magnitude is how much you expect it to move and from what baseline. Mechanism is why you expect this change to produce this effect.
Bad hypothesis, common shape: "We think the new pricing page will increase conversions." What is wrong with it: no magnitude (how much), no mechanism (why), and the metric is "conversions" rather than a specific event with a clear definition. The team will run this test, look at the result, and argue about what counts as a win. Pre-commitment is impossible because nothing was committed.
Good hypothesis, same domain: "Replacing the three-tier pricing comparison with a single recommended tier will increase signup-to-paid conversion by 8 percent (currently 12 percent, target 13 percent) by reducing decision friction for users who already know they want to subscribe." Cause is the tier replacement. Effect is signup-to-paid conversion, defined as the user reaches the paywall and completes payment within seven days. Magnitude is 8 percent relative lift, taking the rate from 12 to 13 percent absolute. Mechanism is decision friction reduction. Now the team has something to test, a number to hit, and a story to falsify.
Primary metric vs guardrails. The primary metric is the thing you are trying to move. Guardrails are the things that must not break: revenue, retention, support ticket volume, page load time, error rates. Pick exactly one primary metric. Pick three to five guardrails. Multiple primary metrics destroy the discipline because they let you cherry-pick the favorable one when results come in.
Falsifiability test. Before launching the experiment, write down what would make you NOT ship this. If the answer is "nothing, we are committed to the change regardless," the hypothesis is not real and the experiment is theater. Skip the test, save the engineering time, and just ship the change.
Directional vs magnitude distinction. Knowing the change moves the needle is different from knowing it moves the needle enough to matter. A 0.3 percent absolute lift on signup conversion may be statistically significant with enough traffic and still not justify the engineering cost of maintaining the change. Magnitude matters as much as direction; the hypothesis names the magnitude that would justify shipping.
For templates and worked examples across common metric types, see references/hypothesis-templates.md.
Sample size grows with the inverse square of the effect you want to detect. Detecting a 1 percent lift requires roughly one hundred times the sample needed to detect a 10 percent lift. Most PMs underestimate this.
The basic decision rule: if your minimum detectable effect (MDE) at current traffic and a reasonable test duration is greater than 5 percent absolute lift, you probably need a bigger MDE. Tiny changes that need huge samples to detect are usually not worth shipping anyway. The change is small either because the underlying mechanism is weak or because the implementation is timid. A weak mechanism is not worth a launch. A timid implementation should be made bolder before testing.
The "we do not have enough traffic" trap. Real for very small products. Lazy for everyone else. If you have ten thousand users a week and you are trying to detect a half-percent absolute lift, you are not running an experiment, you are sampling noise. Pick changes whose expected effect is large enough to detect at your traffic level. If the change is genuinely small, ship it without a test (small upside, small downside, low cost) or do not ship it at all.
Power. The test's ability to detect an effect that is actually present. The conventional floor is 80 percent. Below that, you are rolling dice; the test will frequently miss real effects. Higher power costs more sample. Most platforms default to 80; if you change it, document why.
One-sided vs two-sided. Most PM tests are two-sided despite the temptation to claim otherwise. A one-sided test says "I only care about the positive direction; if the change makes things worse, I do not need to detect it." That is rarely true. If the new pricing page tanks conversion, you want to know. Default to two-sided. If you genuinely want one-sided, document the asymmetry before running.
For pre-calculated sample size tables across common conversion rate baselines and MDEs, see references/sample-size-tables.md. The tables are starting points, not substitutes for running the math against your specific traffic and metric.
Minimum duration is the longer of two constraints. Constraint one: the sample size hits the calculated requirement. Constraint two: the test runs at least one full weekly cycle. Testing only Monday through Wednesday misses weekend behavior, which on most consumer products differs meaningfully from weekday behavior.
Novelty effects. New things attract attention. The first few days of a winning test often overstate the lift. Users notice the change, click it, and produce a temporary effect that fades as the novelty wears off. Run long enough to see if the lift survives the novelty period. Two weeks is the conventional minimum for any UI/UX experiment, even if the sample size hits faster.
Primacy effects. The opposite problem. Existing users may resist the change in week one and adapt by week three. Common in UI rearrangement tests. Killing the test in week one because the result looks negative misses the point that primacy is bigger than the underlying effect at that timescale.
Holdout periods. If the experiment changes a permanent feature (notification frequency, default settings, search ranking), keep a holdout group OFF the new behavior for at least a month after launch. The holdout measures long-term effect, not just the day-1 lift. Long-term effects are often different from short-term effects: a notification change that increases day-1 engagement may decrease month-three retention.
Maximum duration. Usually four to six weeks. Beyond that, the world changes around the test. Seasonality shifts. Marketing campaigns launch. The product evolves. The comparison between treatment and control stops being clean because the underlying user population is no longer comparable across the test window. If the test needs to run longer than six weeks to hit power, the MDE is probably wrong; the change is too small to detect cleanly.
This is a section many discussions of experimentation skip. Worth being direct about.
UX bug fixes. If the current behavior is objectively broken (button does not work, copy says the wrong thing, accessibility fails), fix it. A/B testing it is theater. The right answer is not "let's see if our users prefer a working button"; the right answer is to ship the working button.
Legal-required changes. GDPR consent flows, accessibility compliance, regulatory disclaimers. Ship them. The lift is irrelevant; the compliance is the point. A/B testing whether to comply with the law is not a serious question.
Strategic or philosophical brand questions. "Should our voice be playful or serious?" is not an A/B test question; it is a brand strategy question that needs to be made by humans with context, weighed against brand equity, audience expectation, and long-term positioning. Picking the variant with the higher click-through rate does not answer it because click-through rate was not the brand decision. Use the experiment data as one input to the brand decision, not as the decision itself.
Things you have already decided. If leadership has committed to a direction regardless of test result, do not run an experiment. A/B testing as theater (running tests where the outcome does not change the decision) corrodes trust in the experimentation discipline overall. Other PMs see the test, see the result ignored, and conclude that experiment results do not matter at this company. Then they stop pre-committing. Then the discipline collapses.
Things where the test design is impossible. Cross-device experiences, network effects, internal tooling for a ten-person ops team, anything where the sample size is fundamentally too small or the randomization is fundamentally contaminated. Sometimes the right answer is qualitative research, longitudinal cohort analysis, or just shipping and watching. An experiment that cannot be designed cleanly will not produce a clean answer.
Pre-registered segments versus post-hoc segments. Declaring before the test runs that "we will look at the result for new users versus returning users" is fine. Discovering after the test that "users from California who signed up on Tuesdays via mobile" had a huge lift is almost always noise mining.
The multiple comparisons problem. Every additional segment you analyze increases the probability of finding a "significant" result by chance. With twenty independent segments at p equals 0.05, you expect one false positive purely by chance. With fifty segments, two or three. Do not analyze fifty segments and report the one that hit significance.
When stratification helps versus misleads. Stratification is useful when three things are true. The segment was pre-registered. There is a real prior reason to expect different behavior in this segment. There is enough sample within the segment to detect the effect at the chosen power. If any of the three is missing, stratification is noise mining. The default posture should be: report the overall result. Report pre-registered segments as additional context. Do not report unplanned segments.
The "weighted average" reframe. If a treatment is positive for one segment and negative for another, the right question is "what is the weighted average effect across the population we will actually ship to" not "let's just ship to the segment where it works." Shipping to a segment usually requires UI complexity, audience targeting infrastructure, and ongoing maintenance that the segment-specific lift does not justify. The bias against segment-specific shipping is healthy.
The classic problem: you are running five concurrent A/B tests on the checkout flow. Each test individually shows a small lift. Together, the combinations may not multiply cleanly. They may not even sum cleanly. They may interfere in ways the individual tests cannot reveal.
Pre-experiment hygiene. Before launching a new experiment, check what other experiments are running on the same surface. Ask the other PM owners. Coordinate on which tests are mutually exclusive and which can overlap.
Mutex (mutually exclusive) experiment groups. Most platforms support exclusion rules so users in test A are not also in test B. Use them when interactions are likely. The cost is sample size; the benefit is interpretable results. For a small set of high-stakes tests, mutex is the right call. For dozens of lower-stakes tests, full mutex is impractical; coordinate and document overlap instead.
The "we will analyze it later" fallacy. Post-hoc detangling of overlapping experiments is hard, expensive, and usually inconclusive. The factorial design that would isolate interaction effects requires sample sizes most products do not have. Coordinate up front rather than untangle afterwards.
The trap. Conversion rate is a ratio: conversions divided by users. Standard deviation calculations for raw counts do not apply directly to ratios. A naive variance estimate on a ratio metric tends to be too narrow, leading to overstated confidence and false-positive ship decisions.
Why it matters. Many experimentation platforms quietly use the delta method (or bootstrap, or some other ratio-aware estimator) for ratio metrics. If yours does not, your confidence intervals are wrong. Wrong in the direction that matters: you ship things that look significant but are not.
How to check. Ask your platform vendor: "What is your variance estimator for ratio metrics?" If the answer is "standard t-test on proportions," that is wrong for any ratio that is not a simple binary conversion (converted yes or no, with one row per user). If the answer is "delta method" or "bootstrap with re-sampling at the user level" or "linearization with Taylor expansion," that is correct. Other reasonable answers exist; the test is whether the platform team can articulate a ratio-aware estimator at all.
Worked example. Revenue per user is a ratio: total revenue divided by total users. RPU lift estimates that do not use a ratio-aware estimator tend to overstate confidence. A 5 percent reported lift with p equals 0.04 might actually be a 5 percent point estimate with no statistical significance once the variance is computed correctly. Shipping based on the wrong math means shipping changes that do not produce the claimed effect in production.
For deeper coverage of variance reduction (CUPED, stratified sampling, control variates), see the experimentation-analytics skill when it ships.
The interference problem. In a two-sided marketplace (Uber, Airbnb, eBay, DoorDash, any platform with buyers and sellers), a treatment for buyers may affect sellers regardless of which group is in the test. A treatment that increases buyer demand changes seller behavior. The "control" buyers, who are competing for the same sellers, see different supply because of treatment buyers' actions. The test no longer measures the treatment effect cleanly; it measures treatment effect plus interference.
Common pattern. Marketplace experiments where the control group is contaminated. The test reports a small effect because half the effect leaked into the control. The team underestimates the true effect, kills a winning idea, and learns the wrong lesson.
Mitigations, none perfect. Cluster randomization assigns whole markets (cities, regions, market segments) to treatment or control rather than individual users. Eliminates within-cluster interference but reduces effective sample size by orders of magnitude. Switchback experiments alternate the entire population between treatment and control across time windows (week 1 treatment, week 2 control, week 3 treatment, etc.). Eliminates cross-user interference but requires careful temporal modeling. Geographic isolation runs the experiment in one city while keeping the rest of the network on the original behavior. Eliminates interference but is expensive and slow.
When to call qualitative. If interference is severe and randomization cannot isolate it, the test will not tell you what you want to know. Switch to user research, longitudinal cohort analysis, or a phased rollout with careful instrumentation. The decision is "do we believe this works enough to invest in the rollout monitoring" rather than "did the lift hit significance."
The peeking problem. If you check results every day and stop the test as soon as you see significance, your false positive rate is much higher than the nominal 5 percent. Standard statistical tests assume one analysis at the end of the test. Multiple analyses inflate alpha.
The math. With one analysis, false positive rate is 5 percent (at alpha equals 0.05). With three analyses spread across the test, false positive rate climbs toward 14 percent. With daily peeking on a 28-day test, false positive rate can exceed 30 percent. You will see "significant" results that are not real, ship them, and watch them fail to replicate in production.
Sequential testing methods. Most modern platforms (Statsig, Eppo, parts of PostHog, GrowthBook with mSPRT) support sequential testing that adjusts the math to allow daily peeking without inflating alpha. The methods include sequential probability ratio tests, group sequential designs, and always-valid p-values via mixture sequential probability ratio tests (mSPRT). Use them when the platform offers them. They cost some statistical power in exchange for valid mid-test inference.
Pre-registered stopping rules. If the platform does not support sequential testing, declare the planned analysis date before the test runs. Save the pre-commitment. Do not look at results until that date. If you must look (some platforms make it hard not to), do not make decisions based on what you see. The decision happens at the pre-committed analysis date, not when the early peek looks favorable.
The "ship early because results look great" trap. Results almost always look more dramatic on day 3 than they do at day 14. Regression to the mean kicks in. Novelty fades. The metric stabilizes. The tests that survive the full window and still look great are the ones worth shipping. The ones that "looked great early" and were shipped early are disproportionately the ones that disappointed in production.
The p-hacking inventory. Things people do, often unconsciously, when results do not come out the way they hoped:
Each individually feels like a small judgment call. Cumulatively they destroy the discipline. The result is not "we found a real effect"; the result is "we ran enough analytical knobs that something looked significant."
The pre-commitment fix. Before the test runs, write down: the primary metric and how it is computed; the MDE you are powered to detect; the duration in calendar days; the segments you will analyze (if any); the decision rule that maps each possible result to a ship or kill. Save the pre-commitment somewhere with a timestamp. The PR description that ships the experiment configuration. A signed Slack message. A pinned ticket. Anywhere that makes it immutable.
When the results come in, follow the pre-commitment. If the result was clean and you want to ship, file the launch. If the result was bad and you want to kill, kill. If the result was inconclusive, follow the inconclusive resolution path described below; do not invent new analyses.
The "we learned so much" trap. If the test ran and produced an inconclusive result, the answer is "inconclusive." Not "we learned that the underlying mechanism is more nuanced than expected" or "we discovered a fascinating segment dynamic." Inconclusive is inconclusive. The lesson, if there is one, is for the next hypothesis, not retrofitted onto this one.
Three buckets. Clear win: ship. Clear loss: kill. Inconclusive: the hardest case.
Inconclusive resolution paths, ranked from most acceptable to least:
Bayesian thinking layer. If your prior was strong (this should obviously work, given everything we know about user behavior), an inconclusive result should still update your belief somewhat. The prior was wrong, or the effect is smaller than expected, or the implementation was too timid to capture it. If your prior was weak (we genuinely had no idea what would happen), an inconclusive result means inconclusive; the prior was already uncertain and the result did not narrow it.
The hardest version. Positive primary metric, ambiguous guardrail. Revenue went up but support tickets ticked up. Conversion went up but session length went down. Use the pre-committed decision rule. If you did not pre-commit on the guardrail trade-off, default to "do not ship." The guardrails exist because you cared about them before you saw the result. Do not lower the bar after the fact.
For a step-by-step results-reading checklist, see references/results-interpretation-checklist.md.
Rapid-fire reference. Each pattern is described in more detail in references/common-failures.md.
This skill's output depends on data, measurements, or tool results it cannot generate on its own. When a required input, tool, or data source is unavailable or unverifiable, the sanctioned output is the deliverable with the gap stated: what was needed, what was actually obtained or verified, and which parts of the output are affected. Fabricating, estimating, or interpolating a required number to complete the deliverable is never sanctioned. A stated gap is a complete answer.
references/hypothesis-templates.md. Concrete formats for writing hypotheses that pass the cause-effect-magnitude-mechanism test, with worked examples across conversion, engagement, revenue, retention, and funnel-step metric types.references/sample-size-tables.md. Pre-calculated sample size tables for the most common conversion-rate experiments, with how-to-use guidance and the common pitfalls in sample-size planning.references/common-failures.md. Fifteen anti-patterns that produce wrong shipping decisions, each with symptom, root cause, fix, and prevention.references/results-interpretation-checklist.md. Step-by-step checklist for reading results, the three-bucket decision matrix (win, loss, inconclusive), and the post-launch monitoring discipline.references/platform-comparison.md. Profiles of the major experimentation platforms (Statsig, PostHog, GrowthBook, Optimizely, Amplitude, Eppo, Kameleoon) with strengths, gotchas, and a decision matrix for choosing.references/pre-experiment-readiness-checklist.md. Ten-item go/no-go checklist run through before launch.references/post-experiment-decision-framework.md. The moment-of-decision framework: confirm pre-commitment, apply rule mechanically, route to ship/kill/inconclusive paths, write the post-mortem within a week.Default conservative posture. When results are unclear, do not ship. The cost of a false negative (you did not ship a real win) is usually smaller than the cost of a false positive (you shipped something that does not actually work and now you have to maintain it forever). The maintenance cost compounds; the missed-opportunity cost does not, because you can always test again.
The discipline of experiment design is the discipline of saying "I do not know" out loud when you do not know. Saying it is often the most consequential thing a PM does in a given week. The instinct is to pretend the result is more conclusive than it is, ship to look decisive, and absorb the failure quietly when the change does not work in production. The discipline is to say "the test was inconclusive, here is what I would change to get a real answer, here is the new test plan." Saying that is professional. Pretending otherwise is theater.
For platform-specific patterns (which platform handles which experiment type best, what the MCP commands look like, where the gotchas live), consult the chosen platform's documentation. For the operational layer below this skill (managing flags, retiring stale ones, coordinating environments), see the feature-flagging skill. For the analytical layer above this skill (variance reduction, Bayesian alternatives, sequential testing math), see the experimentation-analytics skill.
评论 (0)
暂无评论,成为第一个评论者吧!