复制安装命令
用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。
复制前请先查看来源、License 和安全提示。
Security audit: baseline 52/52 CLEAN
用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。
复制前请先查看来源、License 和安全提示。
来源文件:README.md
📌 文档结构(2026-07-22 起): 本文件是中文默认入口 —— banner + badges + 信任面 + 9 阶段流水线速览 + 76 行合集总表。 每个合集的完整描述、按用途分组、精确数字、验证方法在
docs/CONTENT_ZH.md(扩展正文,总表行内的→直接跳转到对应锚点)。English version:
README-en.md· 中文扩展正文:docs/CONTENT_ZH.md·README-zh-CN.md已弃用(重定向占位)
🌐 语言: English | 简体中文(默认) | 繁體中文 | 日本語 | 한국어
|
|
Stanford REAP × CoPaper.AI · 实证研究 AI 工具的学术工业级产品
由斯坦福实证研究方法论团队打造,覆盖从数据清洗到顶刊投稿的完整工作流
🚀 New here? Open the Skill Search → to filter all 1,096 skills by method, stage, language, and license. The 5-minute tour (
make quickstart) prints the same picture in your terminal.🇨🇳 中文用户从本文件开始(流水线速览 + 76 行总表),每个合集的完整描述见
docs/CONTENT_ZH.md。📖 English readers: seeREADME-en.md.
| Rigor lane | Count | Where |
|---|---|---|
| Numeric benchmark tasks — gold values recomputed from real data each run | 17 | benchmark/ |
| Behavioral eval scenarios / rubric items | 37 / 183 | eval-harness/ |
Full trust overview:
docs/TRUST.md·docs/RIGOR_COVERAGE.md
把项目 URL 地址 https://github.com/brycewang-stanford/Auto-Empirical-Research-Skills 丢给 Claude Code / Codex,并指定是目录 / 项目 / 全局安装 —— 剩下的让它自己做。例如:
帮我安装 https://github.com/brycewang-stanford/Auto-Empirical-Research-Skills
装到「全局」(~/.claude/skills/),我想在所有项目里都能用
把最后一行换成你要的作用域即可:
| 作用域 | 说给 Agent 的话 | 落到哪里 |
|---|---|---|
| 目录(当前会话临时用) | "只在当前目录用,不要全局安装" | 当前工作目录下的 .claude/skills/ |
| 项目(团队共享,可提交进 git) | "装到本项目" | 项目根目录 .claude/skills/ |
| 全局(所有项目可用) | "装到全局" | ~/.claude/skills/(Codex 为 ~/.codex/skills/) |
A. 插件市场(Claude Code v2.1+,推荐,可升级)
claude plugin marketplace add brycewang-stanford/Auto-Empirical-Research-Skills
claude plugin install aer-skills@auto-empirical-research-skills # 顶刊投稿全流程(9 skills)
claude plugin install empirical-analysis-python@auto-empirical-research-skills # Python 计量流水线
claude plugin install empirical-analysis-stata@auto-empirical-research-skills # Stata 计量流水线
claude plugin install empirical-analysis-r@auto-empirical-research-skills # R + Quarto 流水线
B. 只要某一个 skill —— 直接拷文件夹
git clone --recurse-submodules https://github.com/brycewang-stanford/Auto-Empirical-Research-Skills.git
cd Auto-Empirical-Research-Skills
cp -R skills/00.1-Full-empirical-analysis-skill_Python .claude/skills/ # 项目级
cp -R skills/00.1-Full-empirical-analysis-skill_Python ~/.claude/skills/ # 全局
拷进去的文件夹必须自带 SKILL.md(部分合集的 SKILL.md 在下一层,拷那一层)。
新开一个会话,直接用自然语言说要做什么,Agent 会按 description 自动挑 skill;说不动就点名方法或 skill:
用面板数据跑一个 Callaway–Sant'Anna 事件研究,并出 HonestDiD 稳健性和期刊级表格
完整安装说明(Codex / CodeBuddy 整库导入、
--plugin-dir单次加载、常见故障排查)见INSTALL.md。
中文内容分两级维护,各司其职:
docs/CONTENT_ZH.md(扩展正文):每个合集的完整描述(#skill-NN 锚点)、按用途分组、精确数字、2 分钟验证、三层信任、旗舰流水线详解、贡献与引用。总表行内的 → 直接跳到对应锚点。README-en.md · README-zh-TW.md · README-ja.md · README-ko.md[!NOTE] 维护规则: 改合集总表 → 本文件与 CONTENT_ZH.md 的锚点表两处同步;改合集详情 / 分组 / 数字 → 只改
docs/CONTENT_ZH.md。统计数字(合集数 / skill 数)以catalog/skills.json为准,由make validate的 readme-stats 检查器守护。贡献者(Contributors): 提交前请在本地跑通完整门禁
make check(catalog 校验 + 链接 + 单元测试 + eval-harness + benchmark)。详见CONTRIBUTING.md。旧版归档:
README-zh-CN.md已弃用,仅作向后兼容的重定向占位。
AERS 不只是 76 个散装 skill —— 它能陪你走完一篇论文。 从模糊 idea → 选题精炼 → 文献综述 → 数据获取 → 识别策略 → 估计建模 → 稳健性审计 → 出版级表格 / 图形 → 写作与同行评审 → 降 AIGC → 投稿。端到端、全自动、每一步都可被人介入(中间任何一步你都可以接过去手工改方法、补变量、加稳健性,再让流水线自动接上跑)。
Paper-WorkFlow 是 AERS 的"指挥棒",它把上面 9 个阶段的 skill 串成 一条按键即运行的端到端流水线。
你在 IDE 入口给它一句自然语言:
"开一个新论文项目:空气污染与中国劳动力市场,CS 设计 + 省级面板"
它会自动按顺序调:
sp.csdid(...) 给出 CS-DID 估计草案 + 写出估计方程与识别假设sp.feols(...) + sp.honest_did(...)任何阶段你都可以手动介入 —— 上一阶段的产物全部落盘(产物-幂等 pipeline),你接过去改方法、补控制、加稳健性,再让流水线自动接下去跑。这就是"全自动 + 可介入"。
| ⭐ Skill | 在流水线里的角色 |
|---|---|
| 00 StatsPAI 🔥 | 因果引擎:900+ 函数,sp.causal(...) 一行跑闭环(DID / RD / IV / SCM / DML / matching) |
| 00.1 Full Empirical · Python 📘 | 显式 Python 栈(pandas / statsmodels / linearmodels / pyfixest) |
| 00.2 Full Empirical · Stata 📊 | 显式 Stata 栈(reghdfe / ivreg2 / csdid / sdid / rdrobust) |
| 00.3 Full Empirical · R 📗 | 显式 R 栈(tidyverse / fixest / did / HonestDiD)+ Quarto 渲染 |
| 48 de-AIGC-skills 🇨🇳🇬🇧 | 中英双语学术降 AIGC(Turnitin AI / GPTZero / 知网 / 万方) |
| 50 AER-skills 📕 | Top-5 经济学投稿套件:识别 → 稳健性 → R&R |
| 69 Paper-WorkFlow 🧭 | 元编排器,把上面 9 个阶段串成一键流水线 |
为什么挑这 7 个?因为它们的行为都被基准钉死了 —— 不是营销口径,是对着已知答案反复跑过验证过的(17 项数值 benchmark + 37 项行为评测 ↗)。
↴ 直跳到下方 76 行总表(每个合集带 #skill-NN 锚点)。如果你更关心"这些 skill 怎么用"而不是"有哪些 skill",看 📘 中文唯一权威正文 里的「按用途分组」与「旗舰流水线」两节。
00 → 72,编号连续无空缺)打开仓库 → 看见整座库。 全部 76 个合集 · 1,096 个 skill,每一个都已 vendor 进本仓库,由
catalog/skills.json跟踪。⭐ = Stanford REAP × CoPaper.AI 团队自研的 skill;其余为精选、经安全审计的社区作品。主题图例 — 🚀 全流程与编排器 · 🎯 因果推断与计量经济学 · 📚 文献与研究设计 · ✍️ 写作 / 编辑 / 去 AIGC · 📑 引用 / 复现 / 同行评审 · 🛠️ 数据 / 工具 / 基础设施
点击【→】 跳转到
docs/CONTENT_ZH.md中该合集的完整描述;点击合集名 直接打开其目录。🙏 尊重原作者 — 「来源」列直接链回上游原始仓库(
owner/repo)。本仓库里的社区合集都是上游快照:请去原仓库点 star、提 issue、看 LICENSE。完整的许可证与来源置信度审计见docs/LICENSE_AUDIT.md,机器可读版本在catalog/provenance.json。
| # | 合集 | 一句话 | 详情 | 来源 |
|---|---|---|---|---|
| ⭐ 00 | StatsPAI 🔥 | 因果引擎 · Agent-native Python DSL:sp.causal(...) 一行跑闭环(DID/RD/IV/SCM/DML,900+ 函数) | → | brycewang-stanford/StatsPAI |
| ⭐ 00.1 | Full Empirical · Python 📘 | 显式栈:pandas · statsmodels · linearmodels · pyfixest | → | ⭐ 本仓库 |
| ⭐ 00.2 | Full Empirical · Stata 📊 | reghdfe · ivreg2 · csdid · sdid · rdrobust 复现包 | → | ⭐ 本仓库 |
| ⭐ 00.3 | Full Empirical · R 📗 | tidyverse · fixest · did · HonestDiD + Quarto 渲染 | → | ⭐ 本仓库 |
| 01 | academic-paper-skills | 大纲 → 手稿写作 + 7 维审稿人模拟 | → | lishix520/academic-paper-skills |
| 02 | research-skills | 医学影像综述、提案、论文转幻灯片 | → | luwill/research-skills |
| 03 | scientific-skills | 假设生成 + 28 个科学数据库 | → | K-Dense-AI/claude-scientific-skills |
| 04 | scientific-writer | 引用管理 + 科学写作 | → | K-Dense-AI/claude-scientific-writer |
| 05 | research-superpower | 系统化检索、筛选与引文溯源 | → | kthorn/research-superpower |
| 06 | stats-paper-writing | 端到端 LaTeX 统计论文写作 | → | fuhaoda/stats-paper-writing-agent-skills |
| 07 | AI-Research-SKILLs | 发表级 ML 图表、LaTeX、引文核验 | → | Orchestra-Research/AI-Research-SKILLs |
| 08 | latex-document-skill | 创建 / 编译任意 LaTeX 文档为 PDF | → | ndpvt-web/latex-document-skill |
| 09 | awesome-econ-ai | Python 面板数据分析(linearmodels) | → | meleantonio/awesome-econ-ai-stuff |
| 10 | causal-inference-mixtape | DID / IV / RDD / SCM 模板(Cunningham) | → | Jill0099/causal-inference-mixtape |
| 11 | compound-science | 面向定量社会科学的贝叶斯估计 | → | James-Traina/compound-science |
| 12 | claude-code-my-workflow | 提交 → PR → 合并的研究工作流(Emory) | → | pedrohcgs/claude-code-my-workflow |
| 13 | MixtapeTools | Cunningham 的因果推断工具集与讲义 | → | scunning1975/MixtapeTools |
| 14 | research-starter | R 中的 IV / DiD / RDD,含完整诊断 | → | luischanci/claude-code-research-starter |
| 15 | social-science-research | R 或 Python 端到端数据分析 | → | Felpix-Studios/social-science-research |
| 16 | clo-author | 多代理数据分析(R / Stata / Python) | → | hsantanna88/clo-author |
| 17 | DAAF | 安全意识代理框架(32 条 deny rule) | → | DAAF-Contribution-Community/daaf |
| 18 | stata-accounting | 来自 126 篇 JAR 论文的实测 Stata 范式 | → | jusi-aalto/stata-accounting-research |
| 19 | vera-economic-intelligence | 经济情报 / 政策研究情报工作流 | → | CuellarC05/vera-economic-intelligence |
| 20 | python-econ-skill | DSGE / HANK 与定量经济计算 | → | wenddymacro/python-econ-skill |
| 21 | AI-research-feedback | 用 AI 同行评审生成结构化反馈 | → | claesbackman/AI-research-feedback |
| 22 | christopherkenny-skills | 面向 Quarto(.qmd)的 APSA 风格检查器 | → | christopherkenny/skills |
| 23 | baygent | 带护栏的 PyMC / Arviz 贝叶斯工作流 | → | Learning-Bayesian-Statistics/baygent-skills |
| 24 | academic-research-skills | 5 审稿人多视角论文评审 | → | Imbad0202/academic-research-skills |
| 25 | Diverga | 研究问题精炼器(抗模式坍缩) | → | HosungYou/Diverga |
| 26 | scholar | 统计算法设计与文档 | → | Data-Wise/claude-plugins |
| 27 | my_claude_skills | 经济学摘要写作指南 | → | dariia-m/my_claude_skills |
| 28 | paper-replicate-agent | 论文复现代理演示 | → | maxwell2732/paper-replicate-agent-demo |
| 29 | project20XXy | 可复现手稿 + notebook 项目 | → | quarcs-lab/project20XXy |
| 30 | zirui-song-claude-skills | Zirui Song 的研究辅助 Claude 技能集 | → | zirui-song/claude-skills |
| 31 | claude-code-skills | Python 面板数据分析 | → | thalysandratos/claude-code-skills |
| 32 | stata-skill | 高性能 Stata C/C++ 插件 | → | dylantmoore/stata-skill |
| 33 | claude-scholar | 研究全生命周期:选题 → 综述 → 实验 → 审稿回复 | → | Galaxy-Dawn/claude-scholar |
| 34 | research-companion | 头脑风暴、评估并决策研究方向 | → | andrehuang/research-companion |
| 35 | academic-writing-skills | 面向投稿场所的工业 AI 文献研究 | → | bahayonghang/academic-writing-skills |
| 36 | literature-review-skill | 完整文献综述工作流(中文) | → | taoyunudt/literature-review-skill |
| 37 | IlanStrauss-ai-skills | Ilan Strauss 经济学研究 AI 工作流 | → | IlanStrauss/ai-skills |
| 38 | academic-proofreader | 学术校对 | → | peternka/academic_proofreader |
| 39 | marginaleffects | 预测、斜率与比较(R / Python) | → | vincentarelbundock/marginaleffects |
| 40 | pyfixest | Python 中的快速固定效应估计 | → | py-econometrics/pyfixest |
| 41 | sewage-econometrics-check | 10 项复现包审计 | → | sticerd-eee/sewage |
| 42 | ARIS | 自主「research-in-sleep」代理,端到端 | → | wanshuiyin/Auto-claude-code-research-in-sleep |
| 43 | research-plugins | 478 个研究插件:数据可视化、领域、基础设施 | → | wentorai/research-plugins |
| 44 | humanizer_academic | 为医学/学术手稿去 AI 味(23 类模式) | → | matsuikentaro1/humanizer_academic |
| 45 | deslop | 去除 AI 写作痕迹(5 维评分) | → | stephenturner/skill-deslop |
| 46 | stop-slop | 三层 AI 痕迹检测与改写 | → | hardikpandya/stop-slop |
| 47 | avoid-ai-writing | 审计 → 改写 → 二次审计 AI 味(留痕) | → | conorbronsdon/avoid-ai-writing |
| ⭐ 48 | de-AIGC-skills 🇨🇳🇬🇧 | 中英双语学术降 AIGC(Turnitin AI / GPTZero / 知网 / 万方) | → | ⭐ 本仓库 |
| 49 | humanize-chinese | 检测并人性化 AI 生成的中文文本 | → | swaylq/humanize-chinese |
| ⭐ 50 | AER-skills 📕 | Top-5 经济学投稿套件:识别 → 稳健性 → R&R | → | brycewang-stanford/AER-skills |
| 51 | CausalPy | 贝叶斯准实验(PyMC Labs) | → | pymc-labs/CausalPy |
| 52 | slr-prisma | 系统文献综述,PRISMA 2020 | → | keemanxp/slr-prisma |
| 53 | thematic-analysis | Braun & Clarke 六阶段定性主题分析 | → | keemanxp/thematic-analysis-skill |
| 54 | open-science-skills | 引用一致性、DOI 与论据支撑审计 | → | scdenney/open-science-skills |
| 55 | r-skills | R 中用 brms 做贝叶斯推断 | → | ab604/claude-code-r-skills |
| 56 | econ-writing-skill | 综合 50+ 顶级指南的经济学写作 | → | hanlulong/econ-writing-skill |
| 57 | edgartools | 查询与分析 SEC 文件 | → | dgunning/edgartools |
| 58 | econstack | 政策简报(UK GES / AU Treasury) | → | charlescoverdale/econstack |
| 59 | openalex-skill | 通过 OpenAlex 查询 2.4 亿+ 学术作品 | → | shiquda/openalex-skill |
| 60 | superpapers | 综合性实证研究支持套件 | → | regisely/superpapers |
| 61 | research-methods | 与预注册匹配的验证性检验 | → | phdemotions/research-methods |
| 62 | citation-checker | 对照 CrossRef / S2 / OpenAlex 核验引用 | → | PHY041/claude-skill-citation-checker |
| 63 | scientific-agent-skills | DoWhy 识别–估计–反驳框架 | → | tondevrel/scientific-agent-skills |
| 64 | mcp-stata | 20 个 Stata 因果推断与复现 skill | → | tmonk/mcp-stata |
| 65 | game-theory-paper-writer | 生成并压力测试博弈论论文 | → | 本仓库 PR #17 |
| 66 | empirical-research-skills | 面向大型面板的 R 性能优化 | → | SiyaoZheng/ai4ss-skills |
| 67 | econfin-workflow-toolkit | 中国公司金融实证工作流,从提案到论文 | → | 本仓库 PR #22 |
| 68 | research-productivity-skills | 论文检索、SSRN、DOI 查询、下载 | → | 本仓库 PR #21 |
| ⭐ 69 | Paper-WorkFlow 🧭 | 元编排器,串起整个社会科学论文流水线 | → | brycewang-stanford/Paper-WorkFlow |
| 70 | ssci-polish ✍️ | SSCI / SCI 英文论文语言润色(语法、可读性、学术语气) | → | ⭐ 本仓库 |
| ⭐ 71 | lit-review-agent-tools 🔍 | 文献综述工具选型 + 一键安装运行(MinerU / PaperQA2 / ASReview / STORM / MCP 服务器) | → | brycewang-stanford/lit-review-agent-tools |
| ⭐ 72 | Kaggle Research 🧪 | 通过官方 CLI 安全检索 Kaggle 资源、限界下载公开数据并保留审计证据 | → | ⭐ 本仓库 |
想看更详细的描述(主题分类、字段、统计)? 见
docs/CONTENT_ZH.md中标注#skill-NN锚点的同一张表 —— 它是每个合集的完整描述所在的扩展正文。
自 2026-04 首次发布以来的主干里程碑(完整提交记录见 Commits 与 CHANGELOG.md):
---
config:
gitGraph:
rotateCommitLabel: false
---
gitGraph TB:
commit id: "2026-04 首次发布"
branch community
commit id: "2026-05 首个社区 PR"
checkout main
merge community
commit id: "2026-05 更名 AERS"
commit id: "2026-06 插件市场"
commit id: "2026-06 全库路由器"
commit id: "2026-07 首个 tag" tag: "v2026.07"
branch kaggle
commit id: "2026-07 Kaggle 集成"
checkout main
merge kaggle
commit id: "2026-08 de-AIGC 双语"
Star 增长曲线(非提交数)· 由 scripts/build-star-history.py 从 GitHub API 生成并提交入库
如果 AERS 对你的工作有帮助,请引用它(CITATION.cff)并点个 Star,让更多研究者看到。
AI 是放大器,不是替代品。它替你做最耗时的"搬砖",你保留最核心的"判断"。
|
|
Stanford REAP × CoPaper.AI · 实证研究 AI 工具的学术工业级产品
![]() 扫码访问 copaper.ai |
![]() 关注公众号「CoPaper.AI」 |
内置 20 个方法论 skill · 20 分钟完成实证论文 · 自研 StatsPAI(900+ 函数 / MIT 开源)
name: methods-communicator
description: Effective communication strategies for statistical methodsTranslating complex statistical methodology for applied researchers, practitioners, and students
Use this skill when writing: package vignettes, tutorial materials, workshop content, applied journal articles, interpretation guides, FAQ documentation, or any communication targeting non-methodological audiences.
| Audience | Statistical Background | Primary Needs | Communication Style |
|---|---|---|---|
| Methods Researchers | Advanced | Theory, proofs, efficiency | Technical, precise |
| Applied Statisticians | Intermediate-Advanced | Implementation, assumptions | Technical with examples |
| Quantitative Researchers | Intermediate | When to use, interpretation | Practical, guided |
| Graduate Students | Developing | Step-by-step, intuition | Pedagogical, scaffolded |
| Practitioners | Variable | Point-and-click, templates | Simplified, checklist-based |
| Technical Term | Plain Language | Analogy |
|---|---|---|
| Natural Indirect Effect | How much of treatment's effect works through the mediator | "The portion of medicine that helps by reducing inflammation" |
| Natural Direct Effect | Treatment's effect through all other pathways | "All other ways the medicine helps beyond reducing inflammation" |
| Sequential Ignorability | No unmeasured confounding at each step | "Apples-to-apples comparison at each stage" |
| Positivity | All treatment combinations are possible | "Everyone had a real chance of getting either treatment" |
| Identification | Can estimate causal effect from data | "The data can answer our causal question" |
| Technical | Applied Researcher Version |
|---|---|
| "The estimator is consistent" | "With more data, estimates get closer to the truth" |
| "Asymptotically normal" | "For large samples, you can use normal-theory confidence intervals" |
| "Efficiency bound" | "The best precision you can possibly achieve" |
| "Double robust" | "Correct if either model is right (doesn't need both)" |
| "Bootstrapped confidence interval" | "We resampled the data many times to estimate uncertainty" |
## Template: Interpreting Indirect Effects
**For a standardized indirect effect of 0.15:**
"The treatment increases the outcome by 0.15 standard deviations
through its effect on the mediator.
In practical terms: for every 100 people treated, we would expect
approximately [X] additional positive outcomes that can be attributed
specifically to the pathway through the mediator.
This effect size is considered [small/medium/large] by conventional
standards in [field]."
# Package Vignette: [Feature Name]
## Overview
[1-2 sentence description of what this vignette covers]
**You will learn:**
- [Learning objective 1]
- [Learning objective 2]
- [Learning objective 3]
**Prerequisites:**
- [Required knowledge 1]
- [Required package 2]
## Quick Start
[Minimal working example - copy-pasteable code that runs immediately]
## Detailed Tutorial
### Step 1: [First Action]
[Explanation of what we're doing and why]
```r
# Annotated code
result <- function_name(
data = my_data, # Your dataset
mediator = "M", # Name of mediator variable
outcome = "Y" # Name of outcome variable
)
What this does: [Plain language explanation]
Common issues:
[Continue pattern...]
# Example output
print(result)
Key values to look at:
| Output | What it means | What's "good" |
|---|---|---|
estimate | The indirect effect | Depends on your context |
ci.lower, ci.upper | 95% confidence interval | Doesn't include 0 = significant |
p.value | Probability under null | < 0.05 conventionally significant |
[Walk through interpretation in words someone would actually say]
Q: Why is my confidence interval so wide? A: [Clear, actionable explanation]
Q: What if my mediator is binary? A: [Clear, actionable explanation]
vignette("advanced-models")vignette("sensitivity")
---
## Pedagogical Techniques
### The "Build-Up" Approach
Start simple, add complexity gradually:
```markdown
## Understanding Mediation: A Graduated Approach
### Level 1: The Basic Idea (No Math)
Think of a drug that treats depression. It might work in two ways:
1. **Directly** affecting brain chemistry → improved mood
2. **Indirectly** by improving sleep → which then improves mood
Mediation analysis asks: "How much of the drug's benefit comes from
each pathway?"
### Level 2: With Diagrams (Minimal Math)
Treatment (X) ──────→ Outcome (Y) │ ↑ └────→ Mediator (M) ─┘
- **Direct effect**: X → Y arrow
- **Indirect effect**: X → M → Y pathway
### Level 3: With Simple Formulas
Total Effect = Direct Effect + Indirect Effect
- Direct: $c'$ (effect with M held constant)
- Indirect: $a \times b$ (X→M effect × M→Y effect)
### Level 4: Full Formal Notation
[For those who want the technical version]
Use one consistent example throughout:
# Example dataset used throughout tutorials
# Intervention study: Exercise program for depression
# - treatment: exercise (1) vs. waitlist (0)
# - mediator: self_efficacy (continuous, 1-10)
# - outcome: depression_score (continuous, 0-63 BDI)
# - covariates: age, gender, baseline_depression
data("exercise_depression", package = "mediation")
# We'll use this data for all examples in this vignette
## Common Misconceptions
### Misconception 1: "If the indirect effect is significant, mediation is proven"
**Why it's wrong:** Mediation analysis shows *statistical* association
through the mediator path, not *proof* of causal mediation.
**Better framing:** "Our data are consistent with a mediation process,
assuming our causal assumptions hold."
### Misconception 2: "A non-significant indirect effect means no mediation"
**Why it's wrong:** We may lack power to detect the effect, or the
effect may be small but real.
**Better framing:** "We did not find statistically significant evidence
of mediation (indirect effect = X, 95% CI: [L, U])."
### Misconception 3: "The bootstrapped CI is always better"
**Why it's wrong:** Bootstrap is better for *asymmetric* sampling
distributions (like products). For normally-distributed effects,
delta-method works fine.
**When to use which:** [Decision guide]
# Module: [Topic Name]
## Duration: [X] minutes
### Learning Objectives
By the end of this module, participants will be able to:
1. [Measurable objective 1]
2. [Measurable objective 2]
### Pre-Assessment (2 min)
[Quick poll or question to gauge prior knowledge]
### Lecture Content (15 min)
#### Slide 1: Motivating Question
[Real-world question that motivates the topic]
#### Slide 2-5: Core Concept
[Building up the idea with visuals]
#### Slide 6-7: Worked Example
[Step-by-step with actual data]
### Hands-On Exercise (20 min)
**Setup:**
```r
# Load packages and data
library(mediation)
data("exercise_depression")
Task 1: [Specific task with expected output]
Task 2: [Build on Task 1]
Discussion: [Question to discuss with neighbor]
[Mistakes you see people make, and how to avoid them]
---
## Applied Journal Translation
### Adapting Methods for Applied Journals
| Methodological Paper | Applied Paper |
|---------------------|---------------|
| "We employ a semiparametric efficient estimator that achieves the efficiency bound under the nonparametric model" | "We used an efficient estimation approach that provides optimal precision" |
| "Under the assumption of sequential ignorability (Assumptions 1-3)..." | "Assuming no unmeasured confounding at each step of the mediation process..." |
| "The influence function takes the form..." | [Omit; put in supplement] |
| "Monte Carlo simulations with 1000 replications" | "We verified performance through simulation studies (see Supplementary Materials)" |
### Applied Methods Section Template
```markdown
## Statistical Analysis
### Mediation Model
We examined whether [mediator] explained the relationship between
[treatment] and [outcome] using [method name] (Author, Year). This
approach decomposes the total treatment effect into:
- **Direct effect**: The portion of the effect that operates
independently of [mediator]
- **Indirect effect**: The portion operating through [mediator]
### Assumptions
This analysis requires that:
1. [Plain language assumption 1]
2. [Plain language assumption 2]
3. [Plain language assumption 3]
We assessed the sensitivity of our findings to potential violations
using [sensitivity analysis approach].
### Implementation
Analyses were conducted in R (version X.X) using the [package] package
(Author, Year). Confidence intervals were computed using [method] with
[N] bootstrap resamples. Code for all analyses is available at [URL].
## Frequently Asked Questions
### Getting Started
**Q: What type of data do I need for mediation analysis?**
A: You need:
- A treatment/exposure variable (X)
- A potential mediator variable (M)
- An outcome variable (Y)
- Ideally, covariates that might confound these relationships
The mediator should be measured *after* the treatment but *before*
(or contemporaneously with) the outcome.
---
**Q: How large should my sample be?**
A: For detecting medium-sized indirect effects (standardized ~ 0.26):
- N ≈ 150-200 for good power
- N ≈ 75 minimum for very large effects
- N ≈ 500+ for small effects
Use power analysis tools like `pwr.med` to determine your specific needs.
---
### Interpretation Questions
**Q: My indirect effect is significant but my direct effect is not.
What does this mean?**
A: This pattern suggests "full mediation" - the treatment's effect
appears to operate entirely through the mediator. However:
1. "Full" mediation is rare and often reflects low power for the direct effect
2. Focus on effect sizes, not just significance
3. Report both effects with confidence intervals
---
**Q: Can the indirect effect be larger than the total effect?**
A: Yes! This happens when direct and indirect effects have opposite signs.
For example:
- Direct effect: -0.20 (treatment directly *reduces* outcome)
- Indirect effect: +0.35 (treatment increases mediator, which increases outcome)
- Total effect: +0.15
This is called "inconsistent mediation" or "suppression."
---
### Troubleshooting
**Q: I'm getting an error about convergence. What should I do?**
A: Common solutions:
1. Check for missing data: `sum(is.na(your_data))`
2. Scale your variables: `scale(variable)`
3. Remove outliers or influential observations
4. Simplify your model (fewer covariates)
5. Increase bootstrap iterations
If problems persist, check the package's GitHub issues.
#' User-Friendly Error Messages
#'
#' @examples
#' # Instead of:
#' stop("non-conformable arguments")
#'
#' # Use:
#' stop(paste0(
#' "The mediator and outcome variables have different lengths.\n",
#' " - mediator has ", length(mediator), " observations\n",
#' " - outcome has ", length(outcome), " observations\n",
#' "Check for missing data or subsetting issues."
#' ))
# Wrapper for common checks
check_input <- function(data, treatment, mediator, outcome) {
errors <- character()
# Check variables exist
if (!treatment %in% names(data)) {
errors <- c(errors, sprintf(
"Treatment variable '%s' not found in data.\nAvailable columns: %s",
treatment, paste(names(data), collapse = ", ")
))
}
if (!mediator %in% names(data)) {
errors <- c(errors, sprintf(
"Mediator variable '%s' not found in data.\nAvailable columns: %s",
mediator, paste(names(data), collapse = ", ")
))
}
# Check for missing data
n_missing <- sum(is.na(data[[treatment]]) | is.na(data[[mediator]]) | is.na(data[[outcome]]))
if (n_missing > 0) {
errors <- c(errors, sprintf(
"Found %d observations with missing data in key variables.\n",
"Use `na.omit(data[c('%s', '%s', '%s')])` to remove, or consider multiple imputation.",
n_missing, treatment, mediator, outcome
))
}
if (length(errors) > 0) {
stop(paste(errors, collapse = "\n\n"), call. = FALSE)
}
}
#' Print Method for Mediation Results
#'
#' Designed for applied researchers who need clear interpretation
print.mediation_result <- function(x, ...) {
cat("\n")
cat("======================================\n")
cat(" MEDIATION ANALYSIS RESULTS \n")
cat("======================================\n\n")
# Effect estimates
cat("EFFECT DECOMPOSITION:\n")
cat(sprintf(" Total Effect: %6.3f 95%% CI [%6.3f, %6.3f]\n",
x$total, x$total_ci[1], x$total_ci[2]))
cat(sprintf(" Direct Effect: %6.3f 95%% CI [%6.3f, %6.3f]\n",
x$direct, x$direct_ci[1], x$direct_ci[2]))
cat(sprintf(" Indirect Effect: %6.3f 95%% CI [%6.3f, %6.3f] %s\n",
x$indirect, x$indirect_ci[1], x$indirect_ci[2],
ifelse(x$indirect_ci[1] > 0 | x$indirect_ci[2] < 0, "*", "")))
cat("\n")
# Proportion mediated
if (x$total != 0) {
prop_med <- x$indirect / x$total * 100
cat(sprintf(" Proportion Mediated: %.1f%%\n", prop_med))
}
cat("\n")
# Plain language interpretation
cat("INTERPRETATION:\n")
if (x$indirect_ci[1] > 0) {
cat(sprintf(" There is evidence of positive mediation (p < .05).\n"))
cat(sprintf(" The treatment increases the outcome by %.3f through\n", x$indirect))
cat(sprintf(" its effect on the mediator.\n"))
} else if (x$indirect_ci[2] < 0) {
cat(sprintf(" There is evidence of negative mediation (p < .05).\n"))
} else {
cat(sprintf(" The indirect effect is not statistically significant.\n"))
cat(sprintf(" We cannot conclude that mediation is present.\n"))
}
cat("\n")
# Caveats
cat("IMPORTANT CAVEATS:\n")
cat(" • Results assume no unmeasured confounding\n")
cat(" • See sensitivity analysis with sensitivityAnalysis()\n")
cat(" • Report effect sizes, not just p-values\n")
cat("\n")
invisible(x)
}
Version: 1.0.0 Created: 2025-12-08 Domain: Statistical communication for diverse audiences Target Outputs: Vignettes, tutorials, workshops, applied papers
评论 (0)
暂无评论,成为第一个评论者吧!