SkillAtlasSkill 详情

workflows:work

Security audit: baseline 52/52 CLEAN

审核状态:已审核Quality 80Security 80

复制安装命令

用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。

复制前请先查看来源、License 和安全提示。

项目 README

来源文件:README.md

抓取于 2026年8月5日

Awesome GitHub stars License: CC BY-SA 4.0 PRs Welcome Validate catalog OpenSSF Scorecard Security audit: baseline 52/52 CLEAN Rigor coverage Powered by StatsPAI

Auto-Empirical Research Skills (AERS)

📌 文档结构(2026-07-22 起): 本文件是中文默认入口 —— banner + badges + 信任面 + 9 阶段流水线速览 + 76 行合集总表。 每个合集的完整描述、按用途分组、精确数字、验证方法在 docs/CONTENT_ZH.md(扩展正文,总表行内的 → 直接跳转到对应锚点)。

English version: README-en.md · 中文扩展正文:docs/CONTENT_ZH.md · README-zh-CN.md 已弃用(重定向占位)

🌐 语言: English | 简体中文(默认) | 繁體中文 | 日本語 | 한국어


CoPaper.AI Stanford REAP - Center on China's Economy & Institutions

Stanford REAP × CoPaper.AI · 实证研究 AI 工具的学术工业级产品
由斯坦福实证研究方法论团队打造,覆盖从数据清洗到顶刊投稿的完整工作流



实证研究智能体技能大全封面图

🚀 New here? Open the Skill Search → to filter all 1,096 skills by method, stage, language, and license. The 5-minute tour (make quickstart) prints the same picture in your terminal.

🇨🇳 中文用户从本文件开始(流水线速览 + 76 行总表),每个合集的完整描述见 docs/CONTENT_ZH.md。📖 English readers: see README-en.md.


信任面 · Trust surface (rigor stats)

Rigor laneCountWhere
Numeric benchmark tasks — gold values recomputed from real data each run17benchmark/
Behavioral eval scenarios / rubric items37 / 183eval-harness/

Full trust overview: docs/TRUST.md · docs/RIGOR_COVERAGE.md


中文文档结构

中文内容分两级维护,各司其职:

  • 本文件(README.md,GitHub 默认入口):banner、badges、信任面、9 阶段流水线速览、76 行合集总表。
  • docs/CONTENT_ZH.md(扩展正文):每个合集的完整描述(#skill-NN 锚点)、按用途分组、精确数字、2 分钟验证、三层信任、旗舰流水线详解、贡献与引用。总表行内的 → 直接跳到对应锚点。
  • 其他语言:README-en.md · README-zh-TW.md · README-ja.md · README-ko.md

[!NOTE] 维护规则: 改合集总表 → 本文件与 CONTENT_ZH.md 的锚点表两处同步;改合集详情 / 分组 / 数字 → 只改 docs/CONTENT_ZH.md。统计数字(合集数 / skill 数)以 catalog/skills.json 为准,由 make validate 的 readme-stats 检查器守护。

贡献者(Contributors): 提交前请在本地跑通完整门禁 make check(catalog 校验 + 链接 + 单元测试 + eval-harness + benchmark)。详见 CONTRIBUTING.md。

旧版归档: README-zh-CN.md 已弃用,仅作向后兼容的重定向占位。


🚀 从一个 idea 到一篇论文:社科实证研究 · 端到端流水线(全自动、可介入)

AERS 不只是 76 个散装 skill —— 它能陪你走完一篇论文。 从模糊 idea → 选题精炼 → 文献综述 → 数据获取 → 识别策略 → 估计建模 → 稳健性审计 → 出版级表格 / 图形 → 写作与同行评审 → 降 AIGC → 投稿。端到端、全自动、每一步都可被人介入(中间任何一步你都可以接过去手工改方法、补变量、加稳健性,再让流水线自动接上跑)。

9 阶段流水线 · 每一步都覆盖到具体 skill

#阶段关键 skills(点合集名进目录,→ 进完整说明)
1️⃣选题精炼 — Agent 把模糊想法收紧成"可证伪 + 可执行"的研究问题· 25 Diverga · 33 claude-scholar · 05 research-superpower · 11 compound-science
2️⃣文献综述 — 检索 · 筛选 · PRISMA 流程 · 批判性阅读 · 主题分析· 36 literature-review-skill · 24 academic-research-skills · 59 openalex-skill · 68 research-productivity-skills · 53 thematic-analysis
3️⃣数据获取 — 公开数据库 · API · 网页抓取 · 数据清洗· 33 claude-scholar · 68 research-productivity-skills · 32 stata-skill · 57 edgartools
4️⃣识别策略 — DiD / RD / IV / SCM / DML / matching 全覆盖· ⭐ 00 StatsPAI 🔥 · 10 causal-inference-mixtape · 13 MixtapeTools · 51 CausalPy · 63 scientific-agent-skills
5️⃣估计建模 — Python / Stata / R 三栈,900+ 估计器· ⭐ 00.1 Full Empirical · Python · ⭐ 00.2 Full Empirical · Stata · ⭐ 00.3 Full Empirical · R · 40 pyfixest · 39 marginaleffects · 09 awesome-econ-ai
6️⃣稳健性审计 — 复现包检查 · Honest-DiD · R&R 模拟· 41 sewage-econometrics-check · ⭐ 50 AER-skills · 21 AI-research-feedback
7️⃣表格 & 图形 — 期刊出版级排版 · LaTeX 嵌入· ⭐ 00 StatsPAI · 07 AI-Research-SKILLs · 33 claude-scholar · 08 latex-document-skill
8️⃣写作 & 同行评审 — LaTeX / Quarto · 仿审稿人 · 校对· 06 stats-paper-writing · 04 scientific-writer · 22 christopherkenny-skills · 38 academic-proofreader · 56 econ-writing-skill · 16 clo-author
9️⃣降 AIGC & 投稿 — 知网 / 万方 / Turnitin / 23 类 AI 痕迹模式· ⭐ 48 de-AIGC-skills 🇨🇳🇬🇧 · 44 humanizer_academic · 45 deslop · 46 stop-slop · 47 avoid-ai-writing · 49 humanize-chinese

🎼 元编排:⭐ 69 Paper-WorkFlow —— 一键串起来

Paper-WorkFlow 是 AERS 的"指挥棒",它把上面 9 个阶段的 skill 串成 一条按键即运行的端到端流水线。 你在 IDE 入口给它一句自然语言:

"开一个新论文项目:空气污染与中国劳动力市场,CS 设计 + 省级面板"

它会自动按顺序调:

  1. ⭐ 00 StatsPAI → sp.csdid(...) 给出 CS-DID 估计草案 + 写出估计方程与识别假设
  2. 33 claude-scholar → 抓变量定义 / 数据源候选 / 相关文献
  3. ⭐ 00 StatsPAI → 真跑 sp.feols(...) + sp.honest_did(...)
  4. 41 sewage-econometrics-check → 10 项复现包审计 + 稳健性体检
  5. ⭐ 00 StatsPAI + 07 AI-Research-SKILLs → 出 Table 1–5 + 期刊级图
  6. 38 academic-proofreader → 通读 + §comment 标"审稿人会挑刺的位置"
  7. 56 econ-writing-skill 起草初稿 + ⭐ 48 de-AIGC-skills 🇨🇳🇬🇧 + 45 deslop 过知网 / Turnitin

任何阶段你都可以手动介入 —— 上一阶段的产物全部落盘(产物-幂等 pipeline),你接过去改方法、补控制、加稳健性,再让流水线自动接下去跑。这就是"全自动 + 可介入"。

🏆 7 个 Stanford REAP × CoPaper.AI 自研 skill —— 是整个流水线的主干

⭐ Skill在流水线里的角色
00 StatsPAI 🔥因果引擎:900+ 函数,sp.causal(...) 一行跑闭环(DID / RD / IV / SCM / DML / matching)
00.1 Full Empirical · Python 📘显式 Python 栈(pandas / statsmodels / linearmodels / pyfixest)
00.2 Full Empirical · Stata 📊显式 Stata 栈(reghdfe / ivreg2 / csdid / sdid / rdrobust)
00.3 Full Empirical · R 📗显式 R 栈(tidyverse / fixest / did / HonestDiD)+ Quarto 渲染
48 de-AIGC-skills 🇨🇳🇬🇧中英双语学术降 AIGC(Turnitin AI / GPTZero / 知网 / 万方)
50 AER-skills 📕Top-5 经济学投稿套件:识别 → 稳健性 → R&R
69 Paper-WorkFlow 🧭元编排器,把上面 9 个阶段串成一键流水线

为什么挑这 7 个?因为它们的行为都被基准钉死了 —— 不是营销口径,是对着已知答案反复跑过验证过的(17 项数值 benchmark + 37 项行为评测 ↗)。

看到这里 —— 完整 76 行合集目录

↴ 直跳到下方 76 行总表(每个合集带 #skill-NN 锚点)。如果你更关心"这些 skill 怎么用"而不是"有哪些 skill",看 📘 中文唯一权威正文 里的「按用途分组」与「旗舰流水线」两节。


🧰 76 个核心 Skills 合集一览(00 → 72,编号连续无空缺)

打开仓库 → 看见整座库。 全部 76 个合集 · 1,096 个 skill,每一个都已 vendor 进本仓库,由 catalog/skills.json 跟踪。⭐ = Stanford REAP × CoPaper.AI 团队自研的 skill;其余为精选、经安全审计的社区作品。

主题图例 — 🚀 全流程与编排器 · 🎯 因果推断与计量经济学 · 📚 文献与研究设计 · ✍️ 写作 / 编辑 / 去 AIGC · 📑 引用 / 复现 / 同行评审 · 🛠️ 数据 / 工具 / 基础设施

点击【→】 跳转到 docs/CONTENT_ZH.md 中该合集的完整描述;点击合集名 直接打开其目录。

#合集一句话详情
⭐ 00StatsPAI 🔥因果引擎 · Agent-native Python DSL:sp.causal(...) 一行跑闭环(DID/RD/IV/SCM/DML,900+ 函数)→
⭐ 00.1Full Empirical · Python 📘显式栈:pandas · statsmodels · linearmodels · pyfixest→
⭐ 00.2Full Empirical · Stata 📊reghdfe · ivreg2 · csdid · sdid · rdrobust 复现包→
⭐ 00.3Full Empirical · R 📗tidyverse · fixest · did · HonestDiD + Quarto 渲染→
01academic-paper-skills大纲 → 手稿写作 + 7 维审稿人模拟→
02research-skills医学影像综述、提案、论文转幻灯片→
03scientific-skills假设生成 + 28 个科学数据库→
04scientific-writer引用管理 + 科学写作→
05research-superpower系统化检索、筛选与引文溯源→
06stats-paper-writing端到端 LaTeX 统计论文写作→
07AI-Research-SKILLs发表级 ML 图表、LaTeX、引文核验→
08latex-document-skill创建 / 编译任意 LaTeX 文档为 PDF→
09awesome-econ-aiPython 面板数据分析(linearmodels)→
10causal-inference-mixtapeDID / IV / RDD / SCM 模板(Cunningham)→
11compound-science面向定量社会科学的贝叶斯估计→
12claude-code-my-workflow提交 → PR → 合并的研究工作流(Emory)→
13MixtapeToolsCunningham 的因果推断工具集与讲义→
14research-starterR 中的 IV / DiD / RDD,含完整诊断→
15social-science-researchR 或 Python 端到端数据分析→
16clo-author多代理数据分析(R / Stata / Python)→
17DAAF安全意识代理框架(32 条 deny rule)→
18stata-accounting来自 126 篇 JAR 论文的实测 Stata 范式→
19vera-economic-intelligence经济情报 / 政策研究情报工作流→
20python-econ-skillDSGE / HANK 与定量经济计算→
21AI-research-feedback用 AI 同行评审生成结构化反馈→
22christopherkenny-skills面向 Quarto(.qmd)的 APSA 风格检查器→
23baygent带护栏的 PyMC / Arviz 贝叶斯工作流→
24academic-research-skills5 审稿人多视角论文评审→
25Diverga研究问题精炼器(抗模式坍缩)→
26scholar统计算法设计与文档→
27my_claude_skills经济学摘要写作指南→
28paper-replicate-agent论文复现代理演示→
29project20XXy可复现手稿 + notebook 项目→
30zirui-song-claude-skillsZirui Song 的研究辅助 Claude 技能集→
31claude-code-skillsPython 面板数据分析→
32stata-skill高性能 Stata C/C++ 插件→
33claude-scholar研究全生命周期:选题 → 综述 → 实验 → 审稿回复→
34research-companion头脑风暴、评估并决策研究方向→
35academic-writing-skills面向投稿场所的工业 AI 文献研究→
36literature-review-skill完整文献综述工作流(中文)→
37IlanStrauss-ai-skillsIlan Strauss 经济学研究 AI 工作流→
38academic-proofreader学术校对→
39marginaleffects预测、斜率与比较(R / Python)→
40pyfixestPython 中的快速固定效应估计→
41sewage-econometrics-check10 项复现包审计→
42ARIS自主「research-in-sleep」代理,端到端→
43research-plugins478 个研究插件:数据可视化、领域、基础设施→
44humanizer_academic为医学/学术手稿去 AI 味(23 类模式)→
45deslop去除 AI 写作痕迹(5 维评分)→
46stop-slop三层 AI 痕迹检测与改写→
47avoid-ai-writing审计 → 改写 → 二次审计 AI 味(留痕)→
⭐ 48de-AIGC-skills 🇨🇳🇬🇧中英双语学术降 AIGC(Turnitin AI / GPTZero / 知网 / 万方)→
49humanize-chinese检测并人性化 AI 生成的中文文本→
⭐ 50AER-skills 📕Top-5 经济学投稿套件:识别 → 稳健性 → R&R→
51CausalPy贝叶斯准实验(PyMC Labs)→
52slr-prisma系统文献综述,PRISMA 2020→
53thematic-analysisBraun & Clarke 六阶段定性主题分析→
54open-science-skills引用一致性、DOI 与论据支撑审计→
55r-skillsR 中用 brms 做贝叶斯推断→
56econ-writing-skill综合 50+ 顶级指南的经济学写作→
57edgartools查询与分析 SEC 文件→
58econstack政策简报(UK GES / AU Treasury)→
59openalex-skill通过 OpenAlex 查询 2.4 亿+ 学术作品→
60superpapers综合性实证研究支持套件→
61research-methods与预注册匹配的验证性检验→
62citation-checker对照 CrossRef / S2 / OpenAlex 核验引用→
63scientific-agent-skillsDoWhy 识别–估计–反驳框架→
64mcp-stata20 个 Stata 因果推断与复现 skill→
65game-theory-paper-writer生成并压力测试博弈论论文→
66empirical-research-skills面向大型面板的 R 性能优化→
67econfin-workflow-toolkit中国公司金融实证工作流,从提案到论文→
68research-productivity-skills论文检索、SSRN、DOI 查询、下载→
⭐ 69Paper-WorkFlow 🧭元编排器,串起整个社会科学论文流水线→
70ssci-polish ✍️SSCI / SCI 英文论文语言润色(语法、可读性、学术语气)→
⭐ 71lit-review-agent-tools 🔍文献综述工具选型 + 一键安装运行(MinerU / PaperQA2 / ASReview / STORM / MCP 服务器)→
⭐ 72Kaggle Research 🧪通过官方 CLI 安全检索 Kaggle 资源、限界下载公开数据并保留审计证据→

想看更详细的描述(主题分类、字段、统计)? 见 docs/CONTENT_ZH.md 中标注 #skill-NN 锚点的同一张表 —— 它是每个合集的完整描述所在的扩展正文。

📈 项目历程

自 2026-04 首次发布以来的主干里程碑(完整提交记录见 Commits 与 CHANGELOG.md):

---
config:
  gitGraph:
    rotateCommitLabel: false
---
gitGraph TB:
   commit id: "2026-04 首次发布"
   branch community
   commit id: "2026-05 首个社区 PR"
   checkout main
   merge community
   commit id: "2026-05 更名 AERS"
   commit id: "2026-06 插件市场"
   commit id: "2026-06 全库路由器"
   commit id: "2026-07 首个 tag" tag: "v2026.07"
   branch kaggle
   commit id: "2026-07 Kaggle 集成"
   checkout main
   merge kaggle
   commit id: "2026-08 de-AIGC 双语"
Star History Chart

Star 增长曲线(非提交数)· 由 scripts/build-star-history.py 从 GitHub API 生成并提交入库

如果 AERS 对你的工作有帮助,请引用它(CITATION.cff)并点个 Star,让更多研究者看到。


AI 是放大器,不是替代品。它替你做最耗时的"搬砖",你保留最核心的"判断"。


CoPaper.AI Stanford REAP

Stanford REAP × CoPaper.AI · 实证研究 AI 工具的学术工业级产品


扫码访问 copaper.ai
扫码访问 copaper.ai
CoPaper.AI 公众号
关注公众号「CoPaper.AI」

内置 20 个方法论 skill · 20 分钟完成实证论文 · 自研 StatsPAI(900+ 函数 / MIT 开源)

研究与检索

中风险

  • 来源需自行核对维护者身份。
  • 包含脚本或命令调用,安装前请复核。
  • 未检测到明显外部权限要求。
  • 未检测到高风险命令。
  • 扫描发现:2 条。

Codex — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/brycewang-stanford/Auto-Empirical-Research-Skills.git
  3. 将 "skills/11-James-Traina-compound-science/skills/workflows-work" 文件夹复制到 Codex 的 skills 目录中。
  4. 重启 Codex 让新的 skill 生效。

Codex — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Codex 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Codex 让新的 skill 生效。

Claude Code — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/brycewang-stanford/Auto-Empirical-Research-Skills.git
  3. 将 "skills/11-James-Traina-compound-science/skills/workflows-work" 文件夹复制到 Claude Code 的 skills 目录中。
  4. 重启 Claude Code 让新的 skill 生效。

Claude Code — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Claude Code 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Claude Code 让新的 skill 生效。

Cursor — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/brycewang-stanford/Auto-Empirical-Research-Skills.git
  3. 将 "skills/11-James-Traina-compound-science/skills/workflows-work" 文件夹复制到 Cursor 的 skills 目录中。
  4. 重启 Cursor 让新的 skill 生效。

Cursor — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Cursor 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Cursor 让新的 skill 生效。

GitHub Copilot — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/brycewang-stanford/Auto-Empirical-Research-Skills.git
  3. 将 "skills/11-James-Traina-compound-science/skills/workflows-work" 文件夹复制到 GitHub Copilot 的 skills 目录中。
  4. 重启 GitHub Copilot 让新的 skill 生效。

GitHub Copilot — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 GitHub Copilot 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 GitHub Copilot 让新的 skill 生效。

Windsurf — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/brycewang-stanford/Auto-Empirical-Research-Skills.git
  3. 将 "skills/11-James-Traina-compound-science/skills/workflows-work" 文件夹复制到 Windsurf 的 skills 目录中。
  4. 重启 Windsurf 让新的 skill 生效。

Windsurf — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Windsurf 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Windsurf 让新的 skill 生效。
查看 SKILL.md 原文
name: workflows:work
description: Execute research implementation plans efficiently while maintaining estimation quality and finishing features
argument-hint: "<plan file, estimation specification, or task description>"
allowed-tools: Read, Glob, Edit, Write, Bash

Work Plan Execution Command

Pipeline mode: This command operates fully autonomously. All decisions are made automatically.

Execute a research implementation plan systematically. The focus is on shipping complete, reproducible research code by understanding requirements quickly, following existing patterns, and maintaining estimation quality throughout.

Input Document

<input_document> #$ARGUMENTS </input_document>

If no input document is provided: Look for the most recent plan in docs/plans/ and use it. If no plans exist, state "No plan found. Run /workflows:plan first." and stop.

Execution Workflow

Phase 1: Quick Start

  1. Read Plan

    • Read the work document completely
    • Review any references, brainstorm origins, or linked code paths
    • Identify the estimation method, identification strategy, and key deliverables
    • Note any open questions from planning — resolve by picking the conservative default and documenting the choice
    • Proceed immediately — do not wait for approval
  2. Setup Environment

    First, detect the project environment:

    # Detect estimation language
    if [ -f "requirements.txt" ] || [ -f "setup.py" ] || [ -f "pyproject.toml" ]; then
      echo "LANG=python"
    elif [ -f "DESCRIPTION" ] || [ -f "renv.lock" ] || [ -f ".Rprofile" ]; then
      echo "LANG=R"
    elif [ -f "Project.toml" ]; then
      echo "LANG=julia"
    elif ls *.do >/dev/null 2>&1; then
      echo "LANG=stata"
    fi
    
    # Detect pipeline tools
    ls Makefile Snakefile dvc.yaml 2>/dev/null
    

    Then check the current branch:

    current_branch=$(git branch --show-current)
    default_branch=$(git symbolic-ref refs/remotes/origin/HEAD 2>/dev/null | sed 's@^refs/remotes/origin/@@')
    if [ -z "$default_branch" ]; then
      default_branch=$(git rev-parse --verify origin/main >/dev/null 2>&1 && echo "main" || echo "master")
    fi
    

    If already on a feature branch (not the default branch):

    • Continue working on it. Proceed to step 3.

    If on the default branch:

    Option A: Create a new branch (default)

    git pull origin $default_branch
    git checkout -b <branch-name-from-plan>
    

    Use a meaningful name derived from the plan (e.g., feat/callaway-santanna-did, fix/blp-convergence).

    Option B: Use a worktree (for parallel estimation runs) See references/worktree-patterns.md if the plan involves parallel workstreams or the user has multiple active branches.

    Automatically choose Option A unless the plan explicitly mentions parallel workstreams.

  3. Activate Research Environment

    Read compound-science.local.md for environment configuration. Then activate:

    Python:

    # Activate virtual environment
    if [ -d ".venv" ]; then source .venv/bin/activate
    elif [ -d "venv" ]; then source venv/bin/activate
    elif command -v conda &>/dev/null; then conda activate $(basename $PWD)
    fi
    # Verify key packages
    python -c "import numpy, scipy, pandas; print('Core packages OK')"
    

    R:

    # Check renv status
    Rscript -e "if (file.exists('renv.lock')) renv::status()"
    

    Verify data paths:

    # Check that referenced data files exist
    ls data/ 2>/dev/null | head -5
    
  4. Create Task List

    • Use TodoWrite to break plan into actionable tasks
    • Include dependencies between tasks
    • Prioritize based on the plan's phase structure
    • Include estimation-specific quality check tasks:
      • Convergence verification after each estimation step
      • Standard error computation and diagnostic tests
      • Robustness checks specified in the plan
    • Keep tasks specific and completable

Phase 2: Execute

  1. Task Execution Loop

    For each task in priority order:

    while (tasks remain):
      - Mark task as in_progress in TodoWrite
      - Read any referenced files from the plan
      - Look for similar patterns in codebase
      - Implement following existing conventions
      - Write tests for new functionality
      - Run Estimation Quality Check (see below)
      - Run tests after changes
      - Mark task as completed in TodoWrite
      - Mark off the corresponding checkbox in the plan file ([ ] → [x])
      - Evaluate for incremental commit (see below)
    

    Estimation Quality Check — Before marking an estimation task done:

    CheckWhat to verify
    ConvergenceDid the optimizer converge? Check exit flag, gradient norm, iteration count. Multiple starting values yield consistent results?
    Sensible estimatesAre parameter signs correct? Magnitudes economically reasonable? No values at boundary constraints?
    Standard errorsComputed with appropriate method (robust, clustered, bootstrap)? Positive definite Hessian? No suspiciously small or large SEs?
    DiagnosticsFirst-stage F > 10 (if IV)? Overidentification test (if overidentified)? Hausman or specification tests where relevant?
    Numerical stabilityLog-likelihood (not likelihood) used? Condition number of key matrices acceptable? No NaN/Inf in outputs?
    ReproducibilityRandom seeds set? Results identical across runs? Dependencies pinned?

    When to skip: Pure data cleaning, documentation updates, or pipeline configuration changes that don't involve estimation. If the task is purely additive (new utility function, data loading), the check takes 10 seconds and the answer is "no estimation, skip."

    When this matters most: Any change that touches estimation routines, moment conditions, likelihood functions, or simulation code.

    IMPORTANT: Always update the original plan document by checking off completed items. Use the Edit tool to change - [ ] to - [x] for each task you finish.

  2. Incremental Commits

    After completing each task, evaluate whether to create an incremental commit:

    Commit when...Don't commit when...
    Estimation step complete with verified convergencePartial estimation code that won't run
    Data pipeline stage verifiedIncomplete data transformation
    Tests pass + meaningful progressTests failing
    About to switch contexts (data work → estimation)Purely scaffolding with no behavior
    Robustness check completeWould need a "WIP" commit message

    Heuristic: "Can I write a commit message that describes a complete, verifiable change? If yes, commit."

    Commit workflow:

    # 1. Verify tests pass (use project's test command)
    # Examples: pytest, Rscript tests/run_tests.R, etc.
    
    # 2. Stage only files related to this logical unit
    git add <files related to this logical unit>
    
    # 3. Commit with conventional message
    git commit -m "feat(estimation): description of this unit"
    

    Note: Incremental commits use clean conventional messages. The final Phase 4 commit/PR includes full attribution.

  3. Follow Existing Patterns

    • The plan should reference similar code — read those files first
    • Match naming conventions exactly (variable names, function signatures, file organization)
    • Reuse existing estimation utilities where possible
    • Follow project coding standards (see CLAUDE.md)
    • When in doubt, grep for similar implementations
  4. Test Continuously

    • Run relevant tests after each significant change
    • Don't wait until the end to test
    • Fix failures immediately
    • Add new tests for new functionality
    • For estimation code: verify convergence AND test with known-parameter DGP if feasible
  5. Track Progress

    • Keep TodoWrite updated as you complete tasks
    • Note any convergence issues or unexpected results
    • Create new tasks if scope expands (e.g., new robustness check needed)
    • Log estimation results at milestones (point estimates, standard errors, diagnostics)

Phase 3: Quality Check

  1. Run Core Quality Checks

    Always run before submitting:

    # Run full test suite
    # Python: pytest
    # R: Rscript tests/run_tests.R or testthat::test_dir("tests")
    
    # Run linting (per CLAUDE.md)
    # Python: ruff check . or flake8
    # R: lintr::lint_dir()
    
  2. Estimation-Specific Validation

    For any work involving estimation:

    • Estimation converges with sensible parameters (check all specifications)
    • Standard errors computed correctly (appropriate clustering/robustness)
    • Diagnostic tests run and results documented
    • Multiple starting values checked (if nonlinear estimation)
    • Results reproducible with fixed random seed
    • No numerical warnings (NaN, overflow, singular matrices)
  3. Consider Reviewer Agents (Optional)

    Use for complex or risky changes. Read agents from compound-science.local.md frontmatter (review_agents). If no settings file, create one following the template in workflows-review/references/project-config.md.

    Run configured agents in parallel with Task tool. Address critical issues before proceeding.

    Default agents for estimation work:

    • econometric-reviewer — identification and inference review
    • numerical-auditor — numerical stability and convergence
    • identification-critic — identification argument completeness
  4. Final Validation

    • All TodoWrite tasks marked completed
    • All tests pass
    • Linting passes
    • Estimation converges with sensible results
    • Standard errors and diagnostics computed
    • Code follows existing patterns
    • Random seeds set and documented
    • No console errors or warnings

Phase 4: Ship It

  1. Create Commit

    git add <relevant files>
    git status  # Review what's being committed
    git diff --staged  # Check the changes
    
    git commit -m "$(cat <<'EOF'
    feat(estimation): description of what and why
    
    Brief explanation if needed.
    
    Co-Authored-By: Claude <noreply@anthropic.com>
    EOF
    )"
    
  2. Create Pull Request

    git push -u origin <branch-name>
    
    gh pr create --title "feat(estimation): [Description]" --body "$(cat <<'EOF'
    ## Summary
    - What was implemented
    - Methodological approach and key decisions
    - Estimation results summary (if applicable)
    
    ## Estimation Quality
    - Convergence: [status]
    - Diagnostics: [first-stage F, overid test, specification tests]
    - Robustness: [alternative specifications checked]
    
    ## Testing
    - Tests added/modified
    - Estimation verified with [approach]
    
    ## Reproducibility
    - Random seeds: [set/documented]
    - Pipeline: [runs end-to-end / specific steps]
    - Dependencies: [pinned in requirements.txt/renv.lock]
    
    ## Research Impact
    - Identification: [any changes to assumptions]
    - Estimation: [computational cost, convergence]
    - Robustness: [new checks added/updated]
    - Replication: [package changes]
    EOF
    )"
    
  3. Update Plan Status

    If the input document has YAML frontmatter with a status field, update it:

    status: active  →  status: completed
    
  4. Summary

    • Display what was completed
    • Link to PR
    • Summarize estimation results if applicable
    • Note any follow-up work needed (additional robustness checks, referee suggestions)

Phase 5: Handoff

Pipeline mode (when invoked from /lfg or /slfg):

  • Skip the interactive menu
  • Auto-invoke /workflows:review on the files that were changed

Standalone mode (when invoked directly by the user):

  • After the Phase 4 summary, present options:
    1. Proceed to review (Recommended) — Immediately run /workflows:review in this session
    2. Continue working — Return to Phase 2 task loop for additional implementation
    3. End session — Stop here; changes are committed

Swarm Mode (Optional)

For complex plans with multiple independent workstreams, enable swarm mode for parallel execution.

When to Use Swarm Mode

Use Swarm Mode when...Use Standard Mode when...
Plan has independent estimation specificationsSingle estimation pipeline
Multiple robustness checks can run in parallelSequential estimation steps
Monte Carlo with independent DGP variantsSimple parameter change
Large replication package with separable componentsSmall feature or bug fix

Enabling Swarm Mode

To trigger swarm execution, say:

"Make a Task list and launch an army of agent swarm subagents to build the plan"

See references/orchestration-patterns.md in the slfg skill for detailed swarm patterns and best practices.


Key Principles

Start Fast, Execute Methodically

  • Read the plan, set up environment, then execute
  • Don't wait for perfect understanding — resolve ambiguity by picking conservative defaults
  • The goal is to finish the implementation with verified estimation quality

The Plan is Your Guide

  • Plans reference existing code, methods papers, and brainstorm decisions — load those references
  • Follow the plan's phase structure and acceptance criteria
  • Don't reinvent — match existing patterns in the codebase

Test Estimation Quality Continuously

  • Verify convergence after each estimation step, not at the end
  • Check diagnostics as you go — fix issues immediately
  • Continuous quality checking prevents late-stage surprises

Quality is Built In

  • Follow existing patterns
  • Write tests for new code
  • Verify estimation convergence and diagnostics
  • Run linting before pushing
  • Use reviewer agents for complex or risky estimation changes only

Ship Complete, Reproducible Research

  • Mark all tasks completed before moving on
  • Don't leave estimation code 80% done — partial results are worse than no results
  • A finished, reproducible implementation ships; a perfect but incomplete one doesn't
  • Seeds set, dependencies pinned, pipeline runs end-to-end

Quality Checklist

Before creating PR, verify:

  • All TodoWrite tasks marked completed
  • Tests pass
  • Linting passes
  • Estimation converges with sensible parameters
  • Standard errors computed correctly
  • Diagnostic tests run and documented
  • Random seeds set and results reproducible
  • Code follows existing patterns
  • Commit messages follow conventional format
  • PR description includes estimation quality section
  • PR description includes reproducibility section
  • PR description includes research impact section
  • Plan file updated with completed checkboxes

Common Pitfalls to Avoid

  • Skipping convergence checks — verify estimation converged, don't assume it did
  • Wrong standard errors — check clustering level, robustness to heteroskedasticity, bootstrap if needed
  • Missing seeds — set random seeds BEFORE any stochastic computation
  • Ignoring plan references — the plan has code paths and method citations for a reason
  • Testing at the end — test continuously or discover convergence failures too late
  • 80% done syndrome — finish the estimation, run diagnostics, compute standard errors
  • Hardcoded paths — use relative paths, check data directory structure

Routes To

  • /workflows:review — review the implementation
  • /workflows:compound — document solutions discovered during implementation

发现问题?提交给管理员复核

评分:

评论 (0)

暂无评论,成为第一个评论者吧!