SkillAtlasSkill 详情

ai-agent-reliability

Your landlord kept your deposit.

审核状态:已审核Quality 72Security 90

复制安装命令

用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。

复制前请先查看来源、License 和安全提示。

项目 README

来源文件:README.md

抓取于 2026年8月18日

🧠 PM Skills — 1117 Professional Agent Skills for Claude, ChatGPT, Gemini, Cursor, Codex & Hermes

PM Skills — 1117 professional skills your AI assistant can read. Plain markdown, works with Claude, ChatGPT, Gemini, Cursor, and Codex. MIT licensed.

Your landlord kept your deposit. Your mom got a medical bill that makes no sense. You got laid off on a Tuesday. Someone you love died, and no one handed you the checklist.

Generic AI gives you filler for the moments that matter most. PM Skills gives your AI the exact framework a senior professional would use — for 1,117 real tasks, across work and life.

👉 Start with your moment, not the catalogue → Skill Packs

🍼 New parent · 💼 Just laid off · 🌍 New to this country · 👵 Caring for a parent · 🕊️ Losing someone · 💸 Money in crisis · 🔑 Starting over · 🤖 Getting serious about AI

In the official Anthropic plugin directory Stars npm PyPI Skills SkillCheck SkillSpec Security Audit Version License Sponsor Listed in Awesome Claude Skills Skill of the day Free runs served Website & newsletter

What is PM Skills?

PM Skills is an open-source library of 1117 Agent Skills — plain-markdown SKILL.md files that teach an AI assistant to do one professional task to a senior professional's standard, from writing a PRD to decoding a lease or running a blameless postmortem. Each skill bundles the framework, an output template, quality checks, and anti-patterns. It is MIT-licensed and works with Claude, ChatGPT, Gemini, Cursor, and Codex.

Decode a lease before you sign it. Write a PRD your team can execute. Simulate the promotion committee before the real one meets. Check the weather with zero API keys. Generic AI gives you filler; these give you the structure a senior professional actually uses.

Works natively in Claude Code and Hermes Agent, with ready-to-paste exports for ChatGPT, Gemini, Cursor, Codex and 8 more tools. (PM stands for Professional, not just Product Management.)

Claude Code — native ChatGPT exports Gemini exports Cursor, Codex, Windsurf — one command MCP — any client
Telegram bot Slack app Raycast launcher Obsidian plugin n8n connector
Python — pip install pm-skills Hugging Face dataset Docker image on ghcr GitHub Actions

🐣 New here? Pick a door — each takes about 30 seconds

  1. Just looking → open the ▶ Playground and run a skill in your browser. Nothing to install, nothing to sign up for.
  2. You use Claude Code → type /plugin, search pm-skills, install. Done — ask "decode this lease" and watch.
  3. You use anything else → npx pm-claude-skills add and pick your tool from the menu (Cursor, Codex, Windsurf, ChatGPT, Gemini…).
  4. Want the guided tour → browse the searchable catalogue site and subscribe to get an email whenever new skills launch.

Nothing here can scare your setup. A skill is a markdown file your AI reads — no runtime, no telemetry, no accounts. Installing copies text files; uninstalling is deleting them. Skeptical? Good instinct: read one first — it's designed to be read by humans too.

Don't know what to look for? Describe your task in plain words at 🔎 find — "my landlord kept my deposit", "board meeting on Thursday" — and it names the skill.

Subscribe to the PM Skills newsletter

Never miss a new skill. New ones drop regularly — subscribe to the newsletter and get a short email with a real example whenever they launch. No spam, unsubscribe anytime. Prefer no email? Follow via RSS or browse the newsletter archive.


🧠 Not just what to do — how to think

Most skills here answer "do this task." A new family answers "think differently about my life."

LLMs have one big weakness: they're too correct. On open-ended questions they give the safe, average, textbook answer — technically right and completely forgettable. Two new bundles fight that head-on (inspired by parallel-divergent-ideation research):

💭 pm-thinking — think better

Escape the generic answer and stress-test your own decisions:

🎯 pm-focus — get unstuck

ADHD-friendly executive function (useful for everyone):

✨ See it in action

It's not just a folder of files — the whole library is explorable, runnable, and a little bit magic. All of this runs in your browser, free, nothing to install:

The Skill Playground: pick the Executive Update skill, fill in a few notes, hit run, and watch a structured executive briefing stream out — all in the browser
▶ Pick a skill → fill a short form → run it → a senior-grade artifact streams out. No install, your key stays in your browser (or run free with no key).

Galaxy 3D — fly through all 1117 skills as a glowing constellation you orbit and click into
🌌 Galaxy 3D — fly through all 1117 skills as a living constellation. The ones you've run burn brighter.
PM Skills Wrapped — your practice turned into a shareable, Spotify-Wrapped-style story
🎁 Wrapped — your practice, as a shareable story. 100% local — nothing leaves your browser.

▶ Open the Playground to run any of the 1117 skills with your own key — or just browse them all.

💬 What can I ask it to do?

Anything below is a real ask that activates a real skill — say it in your own words, the description does the routing:

🏠 "decode this lease before I sign" → lease-decoder📋 "write the PRD for our referral feature" → prd-template🚨 "blameless postmortem for Friday's outage" → incident-postmortem
💰 "practice my salary negotiation" → salary-negotiation📉 "why is churn up this quarter?" → churn-analysis⚖️ "rank the backlog with RICE" → rice-prioritisation
🛂 "prep me for the visa interview" → the-visa-interview🔨 "is this contractor quote fair?" → home-contractor-quote-decoder🏡 "should we rent or buy?" → rent-vs-buy
📝 "draft my self-review honestly" → performance-review🚀 "are we ready to launch?" → product-launch-checklist📬 "my inbox is 4,000 deep" → email-triage-system

…all 1117 asks live in the catalog.

⚡ Quick start

You want to…Do this
Browse the skillsSKILLS.md — the full catalog · or the searchable web catalog
Install in Claude Code/plugin → search pm-skills (it's in the official Anthropic directory) — or npx pm-claude-skills add --agent claude
Install in Cursor / Codex / Windsurf / Cline…npx pm-claude-skills add --agent cursor (or codex, windsurf, aider, cline, zed…)
Use one skill in ChatGPT / GeminiCopy it from exports/chatgpt/ or exports/gemini/ and paste as instructions
Skills over MCP, in any sessionclaude mcp add pm-skills -- npx -y pm-claude-skills-mcp

No npm install needed — npx pm-claude-skills … always runs the latest. npx pm-claude-skills list shows everything in your terminal. Full per-tool instructions: docs/installation.md.

📚 The skills

Every skill follows the same discipline: what it produces, the inputs it needs, a real framework (severity scales, decision rules — not vibes), a concrete output template, quality checks, and anti-patterns. All 1117 pass the SkillSpec L3 gate and a security audit in CI.

Decoders bundle crestSimulators bundle crestCalculators bundle crestLive data bundle crestCowork bundle crestTokens bundle crestSeatbelt bundle crestEssentials bundle crest
DecodersSimulatorsCalculatorsLive dataCoworkTokensSeatbeltEssentials

Browse all 1,099 → · try one in your browser →

Every category, with examples

For everyone — life's paperwork and decisions

FamilyWhat it doesExamples (of many)
🔍 Decoders (25+)Read the document before you sign it — plain language, 🔴🟡🟢 severity, the money mathlease · medical bill · job offer · severance · insurance policy · contractor quote · timeshare
🎭 SimulatorsFace the adversary early — the real meeting, then an out-of-character debriefsalary negotiation · promotion committee · thesis defense · visa interview · due-diligence call
🧮 CalculatorsDeterministic Python scripts + honest models — assumptions labeled, no false precisionrent vs buy · FIRE number · debt payoff · raise vs jump · daycare vs stay-home
📡 Live data (17)Real-time answers with zero API keys — weather, rates, flights, scores, all over plain curlweather · currency · crypto · flights · earthquakes · is-it-down
🏠 Life adminThe unglamorous logistics, done in orderrelocation · new parent · caregiving · doctor visits · records requests
💼 Career momentsThe weeks that decide yearslayoff kit · resignation kit · PIP response · first 90 days as manager · interview gauntlet
🏛 Dead mentors (5) 🆕History's sharpest operators, resurrected — the real methods from public-domain classics, applied to modern workMachiavelli on office politics · Sun Tzu on picking your fights · Franklin's decision algebra · Marcus Aurelius on bad days · Bennett's 1908 time audit
🏛 Life systems (20) 🆕Navigating the bureaucracies and emergencies people face alone — civic, disability, immigration, disastervoting-navigator · disability-benefit-appeal · arrival-setup · credential-recognition · go-bag-builder · after-the-disaster
🧠 Human edges (20) 🆕The parts of life nobody built tools for — neurodivergence, invisible illness, grief, identity, the hard conversationsmasking-budget · spoon-planner · diagnosis-limbo-kit · coming-out-rehearsal · grief-admin · rabbit-hole-rescue
⚡ New-gen (10) 🆕How the next generation lives and earns — creator deals, clips, D&D, ranked, resale, the attention warcreator-deal-decoder · clip-factory · ttrpg-session-forge · the-vibe-check · ranked-climb-coach · attention-reset
🔮 2027 (10) 🆕Problems you don't have yet, but will — the agent era's operational skillsagent-severance · deepfake-drill · agent-hiring-panel · context-bankruptcy · clone-brief · api-for-yourself · the-org-simulator
🎲 Tabletop (5) 🆕Game night, upgraded — teach, judge, plan, design, and practice the tradesteach-the-game · rules-lawyer · game-night-planner · board-game-designer · tabletop-negotiator
🧾 Freelance & renters & parentsSmall bundles for specific livespricing your services · late invoices · deposit recovery · IEP meetings · students
🎲 Hobbies (12) 🆕Life outside work — the genuinely fun stuffwine pairing · houseplant care · board-game night · D&D campaign · stargazing · chess openings
💪 Wellbeing (12) 🆕Body and mind, sustainably — not another app streakhome workout · sleep reset · habit builder · posture reset · screen-time detox
🔐 Digital self-defense (12) 🆕When your digital life is under attackidentity-theft recovery · phishing triage · account recovery · data-broker removal · doxxing response
👪 Family & relationships (12) 🆕The people who matternew-baby logistics · wedding vows · co-parenting messages · condolences · in-law boundaries
💭 Thinking modes (24) 🆕Change how your AI reasons — escape the generic answer, stress-test decisionsthe-third-answer · five-minds · decision-panel · red-team-my-plan · devils-advocate · poke-holes-in-this
🎯 Focus & executive function (26) 🆕Get unstuck and run your own brain — ADHD-friendly, for everyonewhere-do-i-start · task-to-first-step · overwhelm-triage · the-one-thing · build-my-memory-file · weekly-unstuck
📖 Learning & mastery (10) 🆕Learn anything faster and make it sticklearn-anything-roadmap · feynman-explainer · spaced-repetition-setup · skill-plateau-breaker · deliberate-practice-plan
💰 Wealth-building (10) 🆕Build wealth on purpose — educational, not financial adviceinvesting-for-beginners · index-fund-starter · ask-for-a-raise · first-100k-plan · financial-independence-roadmap
🤝 Social & relationships (10) 🆕The hard conversations and the human onesmake-friends-as-an-adult · networking-for-introverts · boundary-setting-scripts · give-hard-feedback-kindly · repair-after-a-fight
🩺 Caregiving & aging (10) 🆕Care for aging parents and navigate the system — not medical/legal advicemedical-appointment-advocate · care-team-coordinator · caregiver-burnout-check · long-term-care-options · end-of-life-wishes-conversation
🤖 AI-native life (10) 🆕Use AI itself well — the meta-skills that make every tool betterprompt-library-builder · delegate-to-ai · ai-context-primer · spot-ai-mistakes · get-more-from-ai
🤝 Cowork (100)The office knowledge work an AI coworker actually does — the frameworks — the whole bundleemail triage · spreadsheet audit · meeting cost meter · deck outline first · saying no kindly · delegation brief
⚡ Cowork · Live (12)The same jobs, done — Claude Cowork acts on your real data via connectors + sandbox and returns an artifact — the whole bundleinbox triage (live) · meeting prep (live) · spreadsheet audit (live) · deck from doc · thread → decision · PR description (live)

For professionals — 35 fields

Product Management Engineering Marketing & GTM
Customer Success Data & Analytics Leadership & People
Design & UX Legal Finance
Founders Security Government

…plus HR, sales, operations, research, healthcare, educators, writers, social media, and more — the full profession index, or by bundle in plugins/ (121 bundles). Install any bundle: /plugin install pm-decoders@pm-skills.

Meta

Before installing anyone's skills (including these): skill-vetting — a security read for SKILL.md files. The library's own standard lives in SKILLSPEC.md; every skill's level is enforced in CI.

🔍 What does a skill look like?

A skill is a single markdown file with a name, a description that tells the assistant when to activate it, and a body containing the working framework: required inputs, decision rules or severity scales, a concrete output template, quality checks, and anti-patterns. The assistant reads it and gains the judgment; humans can read, audit, and edit the same file. No runtime, no lock-in.

---
name: lease-decoder
description: "Decode a residential lease into plain English and rank the
  clauses that can hurt you. Use when someone asks 'what am I signing'…"
---
## Framework: Severity Scale
- 🔴 Can cost you real money — auto-renewal into a full new term, break
  penalties beyond re-rental costs, deposit conditions written to fail…

That's the whole trick: it's markdown. Your agent reads it and gains the judgment; you can read it too, audit it, edit it, or write your own. No lock-in, no runtime, no telemetry.

💸 What it costs you, and how to prove it

Cut your token bill

The pm-tokens bundle optimizes every stage of your agent's token journey — no API keys, stdlib Python, nothing leaves your machine. Five habits, typically 30–60% off a session's token flow:

# 1. Map the repo instead of reading it (~3% of the cost of reading everything)
python3 skills/repo-map/scripts/repo_map.py .

# 2. Crush bulk before it enters context (98% smaller on uniform JSON; errors always survive)
python3 skills/context-crusher/scripts/context_crush.py --mode json --file response.json

# 3. Measure what anything costs — at YOUR prices, times YOUR call volume
python3 skills/token-cost/scripts/token_cost.py --file CLAUDE.md --price-in 3 --calls 200

Plus the judgment skills: token-diet (output costs 3–5× input — diet it where safe), context-budget (cache-aware layout: stable first, volatile last), and session-handoff (resume at ~5% of transcript size). See your own breakdown in the 🪙 Token Dashboard — paste what rides in your context, get computed per-piece savings, all in-browser. The full how-to: docs/SAVE-TOKENS.md.

🤝 Make the most of the cowork skills

The pm-cowork bundle is 100 skills for the office work an AI coworker actually does. Install it (/plugin install pm-cowork@pm-skills), then — the whole trick — describe your mess, don't name the skill: say "my inbox is 4,000 deep", "nobody reads my status updates", "this spreadsheet came from someone who left" — the right skill activates on the ask.

Start where it hurts:

Your painSay thisThe skill that answers
Drowning in email"triage my inbox and cut the volume at the source"email-triage-system → inbox-unsubscribe-purge
Calendar is all meetings"audit my recurring meetings and price them"standing-meeting-audit + meeting-cost-meter
Inherited a scary spreadsheet"audit this sheet before we trust it"spreadsheet-audit → formula-detangler
Docs get rewritten in review"outline first, get sign-off, then draft"outline-before-prose
Weeks just happen to you"set up my weekly review"weekly-review-ritual — the hub the others plug into

Three habits that compound: (1) The weekly review is the keystone — it feeds task-triage-matrix, deep-work-blocking, and personal-wip-limits automatically. (2) The skills chain on purpose — email-to-tasks feeds the task triage; the meeting audit feeds async-instead; delegation-brief hands off what the triage says to shed — follow the links inside each skill. (3) Teams adopt one norm at a time — start with agenda-or-cancel or working-agreements, let it stick, then add the next; the ten-norms-on-Monday rollout is how none of them survive.

Prove a skill works, and stop paying MCP rent

Two CLI tools for the trust-and-cost problems the ecosystem keeps hand-waving — both keyless-to-inspect, both one command:

# Does your skill actually work? Prove it. Paired A/B — skill on vs off, same tasks,
# REAL token counts from the API's usage fields, optional blind judge, sha-pinned receipt.
npx pm-claude-skills prove --skill ./my-skill --tasks tasks.txt --runs 2 --judge
npx pm-claude-skills prove --skill ./my-skill --tasks tasks.txt --dry-run   # plan + call count, spends nothing

# Your MCP servers are charging you rent. Measure it: per-server token cost,
# unused-in-N-days flags, "disconnect these three, save X tokens per message".
npx pm-claude-skills mcp-audit --connect

prove exists because the ecosystem is full of "65% better!" claims and almost none are measured — it's the honest-broker harness (the JetBrains "advertised 65%, measured 8.5%" story is exactly why). mcp-audit reads your Claude configs, speaks real MCP to each server to count its schema tokens, and scans your session logs for what you actually use. See also the 📊 AI Spend page — every agent's cost (Claude Code, Codex, Copilot) in one meter, all in-browser.

Agent safety: the pm-seatbelt bundle is the pre-flight checklist before an agent touches email, the browser, or files — least-privilege reviews, prompt-injection spotting, and the blast-radius drill for going autonomous. And RFC 0002 — HANDOFF.md is a dead-simple session-handoff convention (your agent, but it remembers Monday) — a file, not a server, with reference hooks.

Quality, not just quantity

  • Every skill passes the SkillSpec L3 gate — structure, framework, quality checks, anti-patterns — enforced in CI on every commit
  • Eval-scored — 208 scored outputs, avg 4.8/5, judged blind
  • Security-audited — a dedicated CI workflow sweeps every skill and script; calculators are stdlib-only and deterministic with byte-exact output tests
  • Honest by design — decoders end with a not-legal-advice line, calculators name what they don't model, simulators debrief out of character, and skills that shouldn't ghostwrite (student statements) coach instead

🎁 Beyond the skills (the bonus material)

The library grew an ecosystem — all optional, all linked from the full showcase:

📄 The one-page cheatsheet — the whole library on one printable poster · ▶ Skill Playground — try any skill in your browser, no install · 📸 the Gallery — the creative side, in screenshots · Anti-Pattern Museum — 2,900+ shareable rules · The Handbook (also a real printed book) · Workflow recipes · Subagents & slash commands · MCP server + REST API · n8n / Slack / Obsidian integrations · The Boardroom · SkillBench · Org Edition · 🇪🇸 🇫🇷 🇨🇳 🇯🇵 translations

Lint your own skills in CI

The validator that keeps these 1,099 honest, as a GitHub Action:

- uses: mohitagw15856/pm-claude-skills@v76
  with:
    path: .claude/skills   # optional — it finds them otherwise

It checks frontmatter, the Use when … trigger clause a model actually matches on, leftover template text, and structure — and annotates each finding inline on the pull request diff, because a finding on the line beats a finding in a log nobody opens. Also available as npx pm-claude-skills skillcheck.

Zero dependencies, no Docker image, no model call.

Companion tools — for the bits a skill shouldn't guess

A skill can tell a model to check the contrast. Only arithmetic can actually check it. Where a question has a right answer rather than a good one, the skill calls out to a tool instead of estimating — both are MIT, zero-dependency, and neither makes a model call, so they cost nothing to run and return the same answer every time.

notugly — design systems that are provably not ugly. #777777 on white is 4.478 and fails AA; #767676 is 4.542 and passes, and no amount of looking at a screenshot separates those. accessibility-audit, design-system-audit, design-handoff-brief, brand-guidelines and the Figma reviews now fill their contrast rows from npx notugly; the MCP server exposes check_contrast directly; and design-system-generate wraps it for the case where there is no design system and something ships on Thursday.

rulebook — 37 games, 203 rulings, and how commonly each house rule is actually played. board-game-night-planner uses it for teach times and for settling the argument, because a rules disagreement is usually two groups who learned it differently and are both partly right.

🆕 Latest

v76.2.1 — SkillCheck as a GitHub Action, and the design skills now compute their contrast numbers instead of estimating them.

Everything else is in the changelog and the releases — a README should say what this is, not what it was.

❓ First-timer questions, straight answers

Is it actually free? Yes — MIT, all 1117 skills, forever. The skills are markdown; there is nothing to gate. Sponsors fund the playground's free model runs, not access.
Do I need an API key? Not to browse, read, install, or use skills inside a tool you already have (Claude Code, ChatGPT, Cursor…). The playground even serves a few sponsor-funded free runs a day. A key only enters the picture for optional extras like running skills from CI.
I'm not a product manager. Is this for me? PM stands for Professional here. Most of the library is decoders for leases and medical bills, salary-negotiation practice, career-moment kits, life admin, and 35 professions from teaching to veterinary. The product-management corner is just where it started.
Will this mess with my existing setup? No. Skills are inert text files in a folder; your assistant reads them when relevant. Remove the folder and it's like they were never there. The CLI never touches anything outside the skills directory it tells you about.
How do I know these are any good? Every skill passes a structural gate (SkillSpec L3) and a security scan in CI; 208 outputs are eval-scored in the open (avg 4.8/5), and the benchmark report publishes the negative findings too. When something's machine-translated or unscored, it's labelled.

🤝 Contributing

The library grows a skill at a time — plant one of your own. One markdown file, one PR.

Add a skill via PR (the standard, CONTRIBUTING), request one via issue, or publish your own repo to the community index and earn the badge. Translations follow the pattern in skills-i18n/.

❤️ Support

If a skill saved you real money or a real mistake, star the repo — it's how others find it. Sponsors fund the playground's free runs and get naming rights, not influence: become a sponsor.

📄 License

MIT — use them, fork them, ship them at work. Skills are judgment, and judgment wants to be free.


Built by Mohit with Claude. 1117 skills · 121 bundles · 35 professions · every commit gated. The long version of this README — every feature, wave, and frontier bet — lives in the Showcase.

Agent / MCP / Skill 创作测试与质量DevOps 与部署

低风险

  • 来源需自行核对维护者身份。
  • 未检测到明显脚本安装指令。
  • 可能需要外部 token、网络权限或第三方服务。
  • 未检测到高风险命令。
  • 扫描发现:0 条。

Codex — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/mohitagw15856/pm-claude-skills.git
  3. 将 "skills/ai-agent-reliability" 文件夹复制到 Codex 的 skills 目录中。
  4. 重启 Codex 让新的 skill 生效。

Codex — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Codex 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Codex 让新的 skill 生效。

Claude Code — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/mohitagw15856/pm-claude-skills.git
  3. 将 "skills/ai-agent-reliability" 文件夹复制到 Claude Code 的 skills 目录中。
  4. 重启 Claude Code 让新的 skill 生效。

Claude Code — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Claude Code 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Claude Code 让新的 skill 生效。

Cursor — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/mohitagw15856/pm-claude-skills.git
  3. 将 "skills/ai-agent-reliability" 文件夹复制到 Cursor 的 skills 目录中。
  4. 重启 Cursor 让新的 skill 生效。

Cursor — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Cursor 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Cursor 让新的 skill 生效。

GitHub Copilot — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/mohitagw15856/pm-claude-skills.git
  3. 将 "skills/ai-agent-reliability" 文件夹复制到 GitHub Copilot 的 skills 目录中。
  4. 重启 GitHub Copilot 让新的 skill 生效。

GitHub Copilot — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 GitHub Copilot 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 GitHub Copilot 让新的 skill 生效。

Windsurf — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/mohitagw15856/pm-claude-skills.git
  3. 将 "skills/ai-agent-reliability" 文件夹复制到 Windsurf 的 skills 目录中。
  4. 重启 Windsurf 让新的 skill 生效。

Windsurf — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Windsurf 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Windsurf 让新的 skill 生效。
查看 SKILL.md 原文
name: ai-agent-reliability
description: "Make an AI agent or automation reliable enough to trust — the tests, checks, and guardrails that catch its failures before they reach anything real. Use when asked how do I test my AI agent, make my automation reliable, my agent works sometimes, or how do I trust an AI workflow in production. Produces a map of where the agent can fail (bad input, hallucination, wrong tool call, edge cases, silent errors), the checks that catch each (validation, evals on real cases, human-in-the-loop gates, monitoring), a right-sized reliability plan scaled to the stakes, and a rollout that earns trust incrementally — so an agent that works in a demo becomes one that works in reality. For builders putting AI agents into real workflows."

AI-Agent Reliability

An AI agent that works in a demo and one you can trust in production are different things — the gap is everything that happens when input is messy, the model hallucinates, a tool call goes wrong, or an error fails silently. This maps where your agent can fail and the specific checks that catch each, scaled to the stakes, plus a rollout that earns trust incrementally — so "works sometimes" becomes "works reliably."

What This Skill Produces

  • A failure map — where this agent can go wrong: bad/unexpected input, hallucinated output, wrong or malformed tool calls, unhandled edge cases, silent failures, and runaway loops
  • The catching checks per failure — input validation, output verification, evals on real cases, schema/format checks on tool calls, human-in-the-loop gates, and monitoring/alerts
  • An eval approach — testing on a real set of cases (including the hard ones) so quality is measured, not assumed, and regressions are caught
  • Human-in-the-loop placement — where a human must approve, scaled to consequence (irreversible/external actions gated, low-stakes automated)
  • A right-sized plan — reliability effort matched to the stakes, not gold-plating a low-risk toy or under-testing a high-risk system
  • A trust-building rollout — shadow mode → low-stakes → expand, with monitoring, rather than shipping it everywhere and hoping

Required Inputs

Ask for these if not provided:

  • The agent — what it does, what tools/actions it takes, what it touches
  • The stakes — what a failure costs (drives how hard to test and gate)
  • Where it fails now — the flakiness you've seen (points at the weak spots)
  • Your setup — the framework/tools, and whether you can add evals/monitoring

Framework: Map Failures, Catch Each, Earn Trust

  1. Enumerate the failure modes. Walk the agent's path — input, reasoning, tool calls, output, actions — and name where each step can break. You can't guard what you haven't named.
  2. Attach a check to each. Validation for input, verification for output, schema checks for tool calls, evals for quality, gates for consequential actions — a specific catch per failure.
  3. Build real evals. A set of representative and hard cases, scored — so you know it works and catch regressions before users do.
  4. Gate by consequence. Irreversible or external actions get a human check; low-stakes steps run free. Match the gate to the cost.
  5. Right-size it. Don't over-engineer a low-risk helper or under-test a system that moves money or data — effort follows stakes.
  6. Roll out to earn trust. Shadow mode, then low-stakes live, then expand — with monitoring and alerts — so reliability is proven, not assumed.

Output Format

Agent reliability: [what it does] · stakes [level]

Failure map: [bad input · hallucination · wrong tool call · edge cases · silent errors · runaway loops]. Catch each: [failure → the check: validation / verification / schema / eval / human gate / monitor]. Evals: [the real + hard cases to test on, scored]. Human gates: [the consequential actions that need approval]. Right-sized: [effort matched to stakes — where to invest, where not]. Rollout: [shadow → low-stakes → expand, with monitoring].

Quality Checks

  • Enumerates failure modes across the agent's whole path
  • Attaches a specific check to each failure
  • Includes evals on real and hard cases, scored
  • Gates consequential actions with a human; automates low-stakes
  • Scales effort to stakes; rolls out to build trust incrementally

Anti-Patterns

  • Shipping a demo as if it's production-ready.
  • No evals — quality assumed, regressions invisible.
  • The same trust level for a summary and a money transfer.
  • Gold-plating a toy or under-testing a high-stakes system.
  • Big-bang launch with no shadow mode or monitoring.

Example Trigger Phrases

  • "How do I test my AI agent so I can actually trust it?"
  • "My automation works sometimes — how do I make it reliable?"
  • "How do I put an AI workflow into production safely?"
  • "What checks does my agent need before I let it run on real data?"
  • "How do I know my agent won't do something dumb and irreversible?"

发现问题?提交给管理员复核

评分:

评论 (0)

暂无评论,成为第一个评论者吧!