复制安装命令
用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。
复制前请先查看来源、License 和安全提示。
Security audit: baseline 52/52 CLEAN
用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。
复制前请先查看来源、License 和安全提示。
来源文件:README.md
📌 文档结构(2026-07-22 起): 本文件是中文默认入口 —— banner + badges + 信任面 + 9 阶段流水线速览 + 76 行合集总表。 每个合集的完整描述、按用途分组、精确数字、验证方法在
docs/CONTENT_ZH.md(扩展正文,总表行内的→直接跳转到对应锚点)。English version:
README-en.md· 中文扩展正文:docs/CONTENT_ZH.md·README-zh-CN.md已弃用(重定向占位)
🌐 语言: English | 简体中文(默认) | 繁體中文 | 日本語 | 한국어
|
|
Stanford REAP × CoPaper.AI · 实证研究 AI 工具的学术工业级产品
由斯坦福实证研究方法论团队打造,覆盖从数据清洗到顶刊投稿的完整工作流
🚀 New here? Open the Skill Search → to filter all 1,096 skills by method, stage, language, and license. The 5-minute tour (
make quickstart) prints the same picture in your terminal.🇨🇳 中文用户从本文件开始(流水线速览 + 76 行总表),每个合集的完整描述见
docs/CONTENT_ZH.md。📖 English readers: seeREADME-en.md.
| Rigor lane | Count | Where |
|---|---|---|
| Numeric benchmark tasks — gold values recomputed from real data each run | 17 | benchmark/ |
| Behavioral eval scenarios / rubric items | 37 / 183 | eval-harness/ |
Full trust overview:
docs/TRUST.md·docs/RIGOR_COVERAGE.md
把项目 URL 地址 https://github.com/brycewang-stanford/Auto-Empirical-Research-Skills 丢给 Claude Code / Codex,并指定是目录 / 项目 / 全局安装 —— 剩下的让它自己做。例如:
帮我安装 https://github.com/brycewang-stanford/Auto-Empirical-Research-Skills
装到「全局」(~/.claude/skills/),我想在所有项目里都能用
把最后一行换成你要的作用域即可:
| 作用域 | 说给 Agent 的话 | 落到哪里 |
|---|---|---|
| 目录(当前会话临时用) | "只在当前目录用,不要全局安装" | 当前工作目录下的 .claude/skills/ |
| 项目(团队共享,可提交进 git) | "装到本项目" | 项目根目录 .claude/skills/ |
| 全局(所有项目可用) | "装到全局" | ~/.claude/skills/(Codex 为 ~/.codex/skills/) |
A. 插件市场(Claude Code v2.1+,推荐,可升级)
claude plugin marketplace add brycewang-stanford/Auto-Empirical-Research-Skills
claude plugin install aer-skills@auto-empirical-research-skills # 顶刊投稿全流程(9 skills)
claude plugin install empirical-analysis-python@auto-empirical-research-skills # Python 计量流水线
claude plugin install empirical-analysis-stata@auto-empirical-research-skills # Stata 计量流水线
claude plugin install empirical-analysis-r@auto-empirical-research-skills # R + Quarto 流水线
B. 只要某一个 skill —— 直接拷文件夹
git clone --recurse-submodules https://github.com/brycewang-stanford/Auto-Empirical-Research-Skills.git
cd Auto-Empirical-Research-Skills
cp -R skills/00.1-Full-empirical-analysis-skill_Python .claude/skills/ # 项目级
cp -R skills/00.1-Full-empirical-analysis-skill_Python ~/.claude/skills/ # 全局
拷进去的文件夹必须自带 SKILL.md(部分合集的 SKILL.md 在下一层,拷那一层)。
新开一个会话,直接用自然语言说要做什么,Agent 会按 description 自动挑 skill;说不动就点名方法或 skill:
用面板数据跑一个 Callaway–Sant'Anna 事件研究,并出 HonestDiD 稳健性和期刊级表格
完整安装说明(Codex / CodeBuddy 整库导入、
--plugin-dir单次加载、常见故障排查)见INSTALL.md。
中文内容分两级维护,各司其职:
docs/CONTENT_ZH.md(扩展正文):每个合集的完整描述(#skill-NN 锚点)、按用途分组、精确数字、2 分钟验证、三层信任、旗舰流水线详解、贡献与引用。总表行内的 → 直接跳到对应锚点。README-en.md · README-zh-TW.md · README-ja.md · README-ko.md[!NOTE] 维护规则: 改合集总表 → 本文件与 CONTENT_ZH.md 的锚点表两处同步;改合集详情 / 分组 / 数字 → 只改
docs/CONTENT_ZH.md。统计数字(合集数 / skill 数)以catalog/skills.json为准,由make validate的 readme-stats 检查器守护。贡献者(Contributors): 提交前请在本地跑通完整门禁
make check(catalog 校验 + 链接 + 单元测试 + eval-harness + benchmark)。详见CONTRIBUTING.md。旧版归档:
README-zh-CN.md已弃用,仅作向后兼容的重定向占位。
AERS 不只是 76 个散装 skill —— 它能陪你走完一篇论文。 从模糊 idea → 选题精炼 → 文献综述 → 数据获取 → 识别策略 → 估计建模 → 稳健性审计 → 出版级表格 / 图形 → 写作与同行评审 → 降 AIGC → 投稿。端到端、全自动、每一步都可被人介入(中间任何一步你都可以接过去手工改方法、补变量、加稳健性,再让流水线自动接上跑)。
Paper-WorkFlow 是 AERS 的"指挥棒",它把上面 9 个阶段的 skill 串成 一条按键即运行的端到端流水线。
你在 IDE 入口给它一句自然语言:
"开一个新论文项目:空气污染与中国劳动力市场,CS 设计 + 省级面板"
它会自动按顺序调:
sp.csdid(...) 给出 CS-DID 估计草案 + 写出估计方程与识别假设sp.feols(...) + sp.honest_did(...)任何阶段你都可以手动介入 —— 上一阶段的产物全部落盘(产物-幂等 pipeline),你接过去改方法、补控制、加稳健性,再让流水线自动接下去跑。这就是"全自动 + 可介入"。
| ⭐ Skill | 在流水线里的角色 |
|---|---|
| 00 StatsPAI 🔥 | 因果引擎:900+ 函数,sp.causal(...) 一行跑闭环(DID / RD / IV / SCM / DML / matching) |
| 00.1 Full Empirical · Python 📘 | 显式 Python 栈(pandas / statsmodels / linearmodels / pyfixest) |
| 00.2 Full Empirical · Stata 📊 | 显式 Stata 栈(reghdfe / ivreg2 / csdid / sdid / rdrobust) |
| 00.3 Full Empirical · R 📗 | 显式 R 栈(tidyverse / fixest / did / HonestDiD)+ Quarto 渲染 |
| 48 de-AIGC-skills 🇨🇳🇬🇧 | 中英双语学术降 AIGC(Turnitin AI / GPTZero / 知网 / 万方) |
| 50 AER-skills 📕 | Top-5 经济学投稿套件:识别 → 稳健性 → R&R |
| 69 Paper-WorkFlow 🧭 | 元编排器,把上面 9 个阶段串成一键流水线 |
为什么挑这 7 个?因为它们的行为都被基准钉死了 —— 不是营销口径,是对着已知答案反复跑过验证过的(17 项数值 benchmark + 37 项行为评测 ↗)。
↴ 直跳到下方 76 行总表(每个合集带 #skill-NN 锚点)。如果你更关心"这些 skill 怎么用"而不是"有哪些 skill",看 📘 中文唯一权威正文 里的「按用途分组」与「旗舰流水线」两节。
00 → 72,编号连续无空缺)打开仓库 → 看见整座库。 全部 76 个合集 · 1,096 个 skill,每一个都已 vendor 进本仓库,由
catalog/skills.json跟踪。⭐ = Stanford REAP × CoPaper.AI 团队自研的 skill;其余为精选、经安全审计的社区作品。主题图例 — 🚀 全流程与编排器 · 🎯 因果推断与计量经济学 · 📚 文献与研究设计 · ✍️ 写作 / 编辑 / 去 AIGC · 📑 引用 / 复现 / 同行评审 · 🛠️ 数据 / 工具 / 基础设施
点击【→】 跳转到
docs/CONTENT_ZH.md中该合集的完整描述;点击合集名 直接打开其目录。🙏 尊重原作者 — 「来源」列直接链回上游原始仓库(
owner/repo)。本仓库里的社区合集都是上游快照:请去原仓库点 star、提 issue、看 LICENSE。完整的许可证与来源置信度审计见docs/LICENSE_AUDIT.md,机器可读版本在catalog/provenance.json。
| # | 合集 | 一句话 | 详情 | 来源 |
|---|---|---|---|---|
| ⭐ 00 | StatsPAI 🔥 | 因果引擎 · Agent-native Python DSL:sp.causal(...) 一行跑闭环(DID/RD/IV/SCM/DML,900+ 函数) | → | brycewang-stanford/StatsPAI |
| ⭐ 00.1 | Full Empirical · Python 📘 | 显式栈:pandas · statsmodels · linearmodels · pyfixest | → | ⭐ 本仓库 |
| ⭐ 00.2 | Full Empirical · Stata 📊 | reghdfe · ivreg2 · csdid · sdid · rdrobust 复现包 | → | ⭐ 本仓库 |
| ⭐ 00.3 | Full Empirical · R 📗 | tidyverse · fixest · did · HonestDiD + Quarto 渲染 | → | ⭐ 本仓库 |
| 01 | academic-paper-skills | 大纲 → 手稿写作 + 7 维审稿人模拟 | → | lishix520/academic-paper-skills |
| 02 | research-skills | 医学影像综述、提案、论文转幻灯片 | → | luwill/research-skills |
| 03 | scientific-skills | 假设生成 + 28 个科学数据库 | → | K-Dense-AI/claude-scientific-skills |
| 04 | scientific-writer | 引用管理 + 科学写作 | → | K-Dense-AI/claude-scientific-writer |
| 05 | research-superpower | 系统化检索、筛选与引文溯源 | → | kthorn/research-superpower |
| 06 | stats-paper-writing | 端到端 LaTeX 统计论文写作 | → | fuhaoda/stats-paper-writing-agent-skills |
| 07 | AI-Research-SKILLs | 发表级 ML 图表、LaTeX、引文核验 | → | Orchestra-Research/AI-Research-SKILLs |
| 08 | latex-document-skill | 创建 / 编译任意 LaTeX 文档为 PDF | → | ndpvt-web/latex-document-skill |
| 09 | awesome-econ-ai | Python 面板数据分析(linearmodels) | → | meleantonio/awesome-econ-ai-stuff |
| 10 | causal-inference-mixtape | DID / IV / RDD / SCM 模板(Cunningham) | → | Jill0099/causal-inference-mixtape |
| 11 | compound-science | 面向定量社会科学的贝叶斯估计 | → | James-Traina/compound-science |
| 12 | claude-code-my-workflow | 提交 → PR → 合并的研究工作流(Emory) | → | pedrohcgs/claude-code-my-workflow |
| 13 | MixtapeTools | Cunningham 的因果推断工具集与讲义 | → | scunning1975/MixtapeTools |
| 14 | research-starter | R 中的 IV / DiD / RDD,含完整诊断 | → | luischanci/claude-code-research-starter |
| 15 | social-science-research | R 或 Python 端到端数据分析 | → | Felpix-Studios/social-science-research |
| 16 | clo-author | 多代理数据分析(R / Stata / Python) | → | hsantanna88/clo-author |
| 17 | DAAF | 安全意识代理框架(32 条 deny rule) | → | DAAF-Contribution-Community/daaf |
| 18 | stata-accounting | 来自 126 篇 JAR 论文的实测 Stata 范式 | → | jusi-aalto/stata-accounting-research |
| 19 | vera-economic-intelligence | 经济情报 / 政策研究情报工作流 | → | CuellarC05/vera-economic-intelligence |
| 20 | python-econ-skill | DSGE / HANK 与定量经济计算 | → | wenddymacro/python-econ-skill |
| 21 | AI-research-feedback | 用 AI 同行评审生成结构化反馈 | → | claesbackman/AI-research-feedback |
| 22 | christopherkenny-skills | 面向 Quarto(.qmd)的 APSA 风格检查器 | → | christopherkenny/skills |
| 23 | baygent | 带护栏的 PyMC / Arviz 贝叶斯工作流 | → | Learning-Bayesian-Statistics/baygent-skills |
| 24 | academic-research-skills | 5 审稿人多视角论文评审 | → | Imbad0202/academic-research-skills |
| 25 | Diverga | 研究问题精炼器(抗模式坍缩) | → | HosungYou/Diverga |
| 26 | scholar | 统计算法设计与文档 | → | Data-Wise/claude-plugins |
| 27 | my_claude_skills | 经济学摘要写作指南 | → | dariia-m/my_claude_skills |
| 28 | paper-replicate-agent | 论文复现代理演示 | → | maxwell2732/paper-replicate-agent-demo |
| 29 | project20XXy | 可复现手稿 + notebook 项目 | → | quarcs-lab/project20XXy |
| 30 | zirui-song-claude-skills | Zirui Song 的研究辅助 Claude 技能集 | → | zirui-song/claude-skills |
| 31 | claude-code-skills | Python 面板数据分析 | → | thalysandratos/claude-code-skills |
| 32 | stata-skill | 高性能 Stata C/C++ 插件 | → | dylantmoore/stata-skill |
| 33 | claude-scholar | 研究全生命周期:选题 → 综述 → 实验 → 审稿回复 | → | Galaxy-Dawn/claude-scholar |
| 34 | research-companion | 头脑风暴、评估并决策研究方向 | → | andrehuang/research-companion |
| 35 | academic-writing-skills | 面向投稿场所的工业 AI 文献研究 | → | bahayonghang/academic-writing-skills |
| 36 | literature-review-skill | 完整文献综述工作流(中文) | → | taoyunudt/literature-review-skill |
| 37 | IlanStrauss-ai-skills | Ilan Strauss 经济学研究 AI 工作流 | → | IlanStrauss/ai-skills |
| 38 | academic-proofreader | 学术校对 | → | peternka/academic_proofreader |
| 39 | marginaleffects | 预测、斜率与比较(R / Python) | → | vincentarelbundock/marginaleffects |
| 40 | pyfixest | Python 中的快速固定效应估计 | → | py-econometrics/pyfixest |
| 41 | sewage-econometrics-check | 10 项复现包审计 | → | sticerd-eee/sewage |
| 42 | ARIS | 自主「research-in-sleep」代理,端到端 | → | wanshuiyin/Auto-claude-code-research-in-sleep |
| 43 | research-plugins | 478 个研究插件:数据可视化、领域、基础设施 | → | wentorai/research-plugins |
| 44 | humanizer_academic | 为医学/学术手稿去 AI 味(23 类模式) | → | matsuikentaro1/humanizer_academic |
| 45 | deslop | 去除 AI 写作痕迹(5 维评分) | → | stephenturner/skill-deslop |
| 46 | stop-slop | 三层 AI 痕迹检测与改写 | → | hardikpandya/stop-slop |
| 47 | avoid-ai-writing | 审计 → 改写 → 二次审计 AI 味(留痕) | → | conorbronsdon/avoid-ai-writing |
| ⭐ 48 | de-AIGC-skills 🇨🇳🇬🇧 | 中英双语学术降 AIGC(Turnitin AI / GPTZero / 知网 / 万方) | → | ⭐ 本仓库 |
| 49 | humanize-chinese | 检测并人性化 AI 生成的中文文本 | → | swaylq/humanize-chinese |
| ⭐ 50 | AER-skills 📕 | Top-5 经济学投稿套件:识别 → 稳健性 → R&R | → | brycewang-stanford/AER-skills |
| 51 | CausalPy | 贝叶斯准实验(PyMC Labs) | → | pymc-labs/CausalPy |
| 52 | slr-prisma | 系统文献综述,PRISMA 2020 | → | keemanxp/slr-prisma |
| 53 | thematic-analysis | Braun & Clarke 六阶段定性主题分析 | → | keemanxp/thematic-analysis-skill |
| 54 | open-science-skills | 引用一致性、DOI 与论据支撑审计 | → | scdenney/open-science-skills |
| 55 | r-skills | R 中用 brms 做贝叶斯推断 | → | ab604/claude-code-r-skills |
| 56 | econ-writing-skill | 综合 50+ 顶级指南的经济学写作 | → | hanlulong/econ-writing-skill |
| 57 | edgartools | 查询与分析 SEC 文件 | → | dgunning/edgartools |
| 58 | econstack | 政策简报(UK GES / AU Treasury) | → | charlescoverdale/econstack |
| 59 | openalex-skill | 通过 OpenAlex 查询 2.4 亿+ 学术作品 | → | shiquda/openalex-skill |
| 60 | superpapers | 综合性实证研究支持套件 | → | regisely/superpapers |
| 61 | research-methods | 与预注册匹配的验证性检验 | → | phdemotions/research-methods |
| 62 | citation-checker | 对照 CrossRef / S2 / OpenAlex 核验引用 | → | PHY041/claude-skill-citation-checker |
| 63 | scientific-agent-skills | DoWhy 识别–估计–反驳框架 | → | tondevrel/scientific-agent-skills |
| 64 | mcp-stata | 20 个 Stata 因果推断与复现 skill | → | tmonk/mcp-stata |
| 65 | game-theory-paper-writer | 生成并压力测试博弈论论文 | → | 本仓库 PR #17 |
| 66 | empirical-research-skills | 面向大型面板的 R 性能优化 | → | SiyaoZheng/ai4ss-skills |
| 67 | econfin-workflow-toolkit | 中国公司金融实证工作流,从提案到论文 | → | 本仓库 PR #22 |
| 68 | research-productivity-skills | 论文检索、SSRN、DOI 查询、下载 | → | 本仓库 PR #21 |
| ⭐ 69 | Paper-WorkFlow 🧭 | 元编排器,串起整个社会科学论文流水线 | → | brycewang-stanford/Paper-WorkFlow |
| 70 | ssci-polish ✍️ | SSCI / SCI 英文论文语言润色(语法、可读性、学术语气) | → | ⭐ 本仓库 |
| ⭐ 71 | lit-review-agent-tools 🔍 | 文献综述工具选型 + 一键安装运行(MinerU / PaperQA2 / ASReview / STORM / MCP 服务器) | → | brycewang-stanford/lit-review-agent-tools |
| ⭐ 72 | Kaggle Research 🧪 | 通过官方 CLI 安全检索 Kaggle 资源、限界下载公开数据并保留审计证据 | → | ⭐ 本仓库 |
想看更详细的描述(主题分类、字段、统计)? 见
docs/CONTENT_ZH.md中标注#skill-NN锚点的同一张表 —— 它是每个合集的完整描述所在的扩展正文。
自 2026-04 首次发布以来的主干里程碑(完整提交记录见 Commits 与 CHANGELOG.md):
---
config:
gitGraph:
rotateCommitLabel: false
---
gitGraph TB:
commit id: "2026-04 首次发布"
branch community
commit id: "2026-05 首个社区 PR"
checkout main
merge community
commit id: "2026-05 更名 AERS"
commit id: "2026-06 插件市场"
commit id: "2026-06 全库路由器"
commit id: "2026-07 首个 tag" tag: "v2026.07"
branch kaggle
commit id: "2026-07 Kaggle 集成"
checkout main
merge kaggle
commit id: "2026-08 de-AIGC 双语"
Star 增长曲线(非提交数)· 由 scripts/build-star-history.py 从 GitHub API 生成并提交入库
如果 AERS 对你的工作有帮助,请引用它(CITATION.cff)并点个 Star,让更多研究者看到。
AI 是放大器,不是替代品。它替你做最耗时的"搬砖",你保留最核心的"判断"。
|
|
Stanford REAP × CoPaper.AI · 实证研究 AI 工具的学术工业级产品
![]() 扫码访问 copaper.ai |
![]() 关注公众号「CoPaper.AI」 |
内置 20 个方法论 skill · 20 分钟完成实证论文 · 自研 StatsPAI(900+ 函数 / MIT 开源)
name: stata-c-plugins
description: >-
Develop high-performance C/C++ plugins for Stata using the stplugin.h SDK.
Use when the user asks to create a Stata plugin, write C/C++ code for Stata,
accelerate a Stata command with C, build cross-platform Stata plugins,
or translate/port a Python or R package into Stata. Covers the full
lifecycle: SDK setup, data flow, memory safety, .ado wrappers with
preserve/merge, cross-platform compilation, performance optimization
(pthreads, pre-sorted indices, XorShift RNG), debugging, and distribution
via net install. Also includes a translation workflow for porting Python/R
packages to Stata — wrapping existing C++ backends when available, or
writing C from scratch when not.Build high-performance C/C++ plugins for Stata. This skill covers the full lifecycle from SDK setup through cross-platform distribution, based on real experience building production Stata plugins for statistical imputation, random forests, string matching, and causal inference.
This skill assumes macOS (Apple Silicon or Intel) as the development platform. Build commands, cross-compilation workflows, and Docker instructions are all Mac-oriented. The plugins themselves target all four platforms (macOS ARM64, macOS x86_64, Linux x86_64, Windows x86_64), but the development environment is macOS. If you need to develop on Linux or Windows natively, adapt the compilation and Docker sections accordingly.
Before writing any code, enter plan mode. A good plan covers:
references/translation_workflow.md — full translation workflow, test repurposing, fidelity auditreferences/testing_strategy.md — test layers, reference data generation, Layer 0 (repurpose original tests)references/performance_patterns.md — pthreads, XorShift RNG, quickselect, pre-sorted indicesreferences/packaging_and_help.md — .toc/.pkg/.sthlp templates, build scriptsreferences/cpp_plugins.md — C++ wrapping, extern "C", exception safety, compilationtranslation_workflow.md)Implement sequentially across components, in parallel within each component. Once an interface is defined, dispatch independent sub-tasks as parallel subagents (e.g., C plugin implementation, .ado wrapper, and test suite can run simultaneously). Merge their work, run the full test suite, then proceed to the review loop before moving to the next component.
Run the review loop after every component:
When translating a package, always check for an existing C/C++ backend before writing any algorithm code. Many R packages have C++ in src/. Many Python packages have Cython or vendored C/C++ libraries. Standalone C++ libraries exist for string matching, linear algebra, tree algorithms, and more.
If a C++ implementation exists, wrap it. Do not reimplement the algorithm in C. Wrapping gives you identical output (same code path), production-grade performance, and a fraction of the code. The plugin is just a thin extern "C" glue layer between Stata's SDK and the library's API. Binary size is irrelevant — statically link everything (-static-libstdc++ -static-libgcc) and ship whatever size the binary turns out to be, even 10-15 MB on Windows. Users don't care about plugin file size; they care about correct results.
See references/cpp_plugins.md for the full pattern and references/translation_workflow.md for the workflow. Working examples of this approach (wrapping C++ backends, multi-plugin dispatching, save/load for scoring on new data) can be found in the repos listed in the project CLAUDE.md under "Example Applications."
For translation projects, also: repurpose the original package's test suite and data (see references/testing_strategy.md Layer 0), write additional Stata-specific tests, and end the plan with a multi-agent fidelity audit. See references/translation_workflow.md for the complete workflow.
Download stplugin.h and stplugin.c from: https://www.stata.com/plugins/
These two files define the interface between your C code and Stata:
| Function/Macro | Purpose |
|---|---|
SF_vdata(var, obs, &val) | Read variable value (1-indexed!) |
SF_vstore(var, obs, val) | Write variable value (1-indexed!) |
SF_nobs() | Number of observations in current dataset |
SF_nvar() | Number of variables in the entire dataset (not just plugin call) |
SF_is_missing(val) | Check for Stata missing value (.) |
SV_missval | The missing value constant |
SF_display(msg) | Print informational text in Stata |
SF_error(msg) | Print red error text in Stata |
Indexing is 1-based. Both variable indices and observation indices start at 1, not 0. Off-by-one errors here are silent and catastrophic — you read the wrong variable's data with no warning.
A crash in your plugin kills the entire Stata session. No save prompt, no recovery. The user loses all unsaved work. This is the single most important thing to internalize.
malloc()/calloc() return for NULLargc before accessing argv[]-fsanitize=address during developmentstata_call(), free at the endEvery plugin implements one function. Plugins can also be written in C++ — the entry point just needs extern "C" linkage so Stata can find it; everything else can be full C++. The obvious case for C++ is when existing C++ code is available to wrap (e.g., an R package's src/ directory). C++ also helps when you need complex data structures or threading via std::thread. For practical C++ guidance — the extern "C" pattern, exception safety, compilation commands, wrapping libraries — see references/cpp_plugins.md. The rest of this file focuses on C because it's the simpler default.
#include "stplugin.h"
// For C++ plugins, wrap the entry point with extern "C":
// extern "C" {
// STDLL stata_call(int argc, char *argv[]) { ... }
// }
STDLL stata_call(int argc, char *argv[]) {
// 0. Validate arguments BEFORE accessing argv[]
if (argc < 3) {
SF_error("myplugin requires 3 arguments: n_train n_test seed\n");
return 198; // Stata's "syntax error" code
}
// 1. Parse arguments (all strings — use atoi/atof)
int n_train = atoi(argv[0]);
int n_test = atoi(argv[1]);
int seed = atoi(argv[2]);
// 2. Get dimensions
ST_int nobs = SF_nobs();
// CAUTION: SF_nvar() returns ALL variables in the dataset, not just
// the ones passed to `plugin call`. If the .ado creates tempvars
// (touse, merge_id, etc.) the count will be higher than expected.
// Pass the variable count via argv instead of relying on SF_nvar().
int p = atoi(argv[3]); // safer: pass feature count explicitly
// 3. Allocate memory
double *X = calloc(nobs * p, sizeof(double));
double *y = calloc(nobs, sizeof(double));
double *pred = calloc(nobs, sizeof(double));
if (!X || !y || !pred) {
SF_error("myplugin: out of memory\n");
if (X) free(X); if (y) free(y); if (pred) free(pred);
return 909;
}
// 4. Read data from Stata (1-indexed!)
ST_double val;
for (ST_int obs = 1; obs <= nobs; obs++) {
SF_vdata(1, obs, &val); // var 1 = depvar
y[obs-1] = val;
for (int j = 0; j < p; j++) {
SF_vdata(j + 2, obs, &val); // vars 2..nvars-1 = features
X[(obs-1) * p + j] = val;
}
}
// 5. Run your algorithm
int rc = my_algorithm(X, y, pred, n_train, n_test, p, seed);
if (rc != 0) {
SF_error("myplugin: algorithm failed\n");
free(X); free(y); free(pred);
return 909;
}
// 6. Write results back to Stata
for (ST_int obs = 1; obs <= nobs; obs++) {
SF_vstore(nvars, obs, pred[obs-1]); // last var = output
}
free(X); free(y); free(pred);
return 0; // 0 = success
}
0 — success198 — syntax error (bad arguments)909 — insufficient memory601 — file not foundUsers never call plugin call directly. An .ado file provides the Stata-native interface.
This is the core pattern for plugins that operate on a subset of data:
program define mycommand, rclass
syntax varlist(min=2) [if] [in], GENerate(name) [SEED(integer 12345) REPlace]
gettoken depvar indepvars : varlist
if "`replace'" != "" {
capture drop `generate'
}
confirm new variable `generate'
// Mark sample: novarlist ALLOWS missing depvar (critical for imputation)
marksample touse, novarlist
markout `touse' `indepvars' // but DO exclude missing predictors
// Stable merge key — create BEFORE any sorting or subsetting
tempvar merge_id
quietly gen long `merge_id' = _n
// Count subsets
quietly count if `touse' & !missing(`depvar')
local n_train = r(N)
quietly count if `touse' & missing(`depvar')
local n_test = r(N)
// Create output variable (all missing initially)
quietly gen double `generate' = .
// Preserve, subset, call plugin
preserve
quietly keep if `touse'
// Sort if plugin requires it (donors first, test second)
tempvar sort_order
quietly gen `sort_order' = missing(`depvar')
quietly sort `sort_order'
// Call plugin
plugin call myplugin `depvar' `indepvars' `generate', ///
`n_train' `n_test' `seed'
// Save results and restore
tempfile results
quietly keep `merge_id' `generate'
quietly save `results'
restore
// Merge predictions back (update replaces missing with non-missing)
quietly merge 1:1 `merge_id' using `results', nogenerate update
end
Why update works: The generate variable is all-missing before preserve. After restore, it's still all-missing. The update option replaces missing values with non-missing ones from the merge file. The replace option is handled earlier via capture drop, so by merge time the variable is always freshly created.
CRITICAL: Some plugins expect data sorted a specific way (training rows first, test rows second). Others handle missing data internally. Sorting mismatches are among the most dangerous bugs — the plugin silently reads the wrong data, producing garbage output with no error message. A mismatched sort order can drop prediction quality dramatically (e.g., correlation going from 0.99 to 0.38) because the plugin treats test observations as training data and vice versa.
SF_is_missing() internally: do NOT sort in the .ado wrappern_train contiguous rows then n_test rows: sort by missing(depvar) before callingDocument which pattern your plugin uses.
Use the gtools-style OS detection pattern. This detects the OS via c(os) and constructs a bare filename. The bare filename is resolved via Stata's adopath, which is reliable across all platforms.
/* ---- Load plugin (gtools-style: detect OS, bare filename) ---- */
if ( inlist("`c(os)'", "MacOSX") | strpos("`c(machine_type)'", "Mac") ) local c_os_ macosx
else local c_os_: di lower("`c(os)'")
cap program drop myplugin
program myplugin, plugin using("myplugin_`c_os_'.plugin")
This resolves to myplugin_macosx.plugin, myplugin_windows.plugin, or myplugin_unix.plugin depending on platform.
WARNING — DO NOT use findfile + absolute paths. The following pattern is BROKEN on Windows and must never be used:
* BROKEN — DO NOT USE
capture findfile myplugin.plugin
capture program myplugin, plugin using("`r(fn)'")
findfile returns an absolute path (e.g., C:\ado\plus\m\myplugin.plugin). On Windows, Stata's LoadLibrary call fails when given certain absolute paths via using(). The gtools-style pattern avoids this by passing a bare filename (no path), which Stata resolves via the adopath — exactly how gtools, ftools, and other major packages work.
Similarly, do not use a nested if/else cascade trying each platform-arch suffix. This was the old pattern in several packages and fails for the same reason if findfile is involved, plus it's fragile and verbose.
Plugin file naming: pluginname_os.plugin where os is one of macosx, unix, windows. Examples: qrf_plugin_macosx.plugin, grf_plugin_windows.plugin.
Note: clear all wipes loaded plugin definitions. If a test script starts with clear all, all program ... plugin definitions are gone. Reload them.
Build for three platforms (ARM Macs run x86_64 via Rosetta, so one macOS binary suffices). Install the Windows cross-compiler first: brew install mingw-w64.
| Target OS | Output name suffix | Compiler | -D flag | Link flag | pthreads |
|---|---|---|---|---|---|
| macOS (ARM64) | _macosx | gcc -arch arm64 | -DSYSTEM=APPLEMAC | -bundle | -pthread |
| Linux (x86_64) | _unix | gcc | -DSYSTEM=OPUNIX | -shared | -pthread |
| Windows (x86_64) | _windows | x86_64-w64-mingw32-gcc | -DSYSTEM=STWIN32 | -shared | -lwinpthread |
All platforms: -O3 -fPIC for release, add -g -fsanitize=address for development.
For C++ plugins: use g++ instead of gcc. Add -std=c++ at the version the library requires (check its docs — C++11, C++14, and C++17 are all common). Header-only C++ libraries can be vendored into c_source/ and included with -I.. Always use -static-libstdc++ -static-libgcc on Windows and Linux.
Naming convention: pluginname_os.plugin (e.g., qrf_plugin_macosx.plugin, grf_plugin_windows.plugin). The os suffix must match what the gtools-style loader produces: macosx, unix, or windows.
macOS note: use -bundle, NOT -shared. This is a common mistake.
There is no native Linux cross-compiler on macOS. Use Docker via Colima (brew install colima docker, then colima start). Build with a one-liner:
docker run --rm --platform linux/amd64 -v "$(pwd):/build" -w /build ubuntu:18.04 \
bash -c "apt-get update -qq && apt-get install -y -qq g++ gcc make > /dev/null 2>&1 && make linux"
glibc compatibility: Build on Ubuntu 18.04 for maximum compatibility (requires only GLIBC 2.14, works on any Linux from ~2012+). Building on Ubuntu 22.04+ requires GLIBC 2.34, which excludes RHEL 8, Ubuntu 20.04, and many HPC environments.
See references/performance_patterns.md for detailed code examples of:
SF_vdata, SF_vstore, SF_display) from worker threads — read all data on the main thread first, dispatch computation to workers, write results back on the main thread after joining.runiform()). XorShift128+ is fast, statistically sound, and thread-safe (each thread gets its own state). Seed from argv[] for reproducibility.Debugging is hard because you can't attach a debugger to Stata's plugin host.
Printf via SF_display():
char buf[256];
snprintf(buf, sizeof(buf), "Debug: n=%d, p=%d\n", n, p);
SF_display(buf);
Write diagnostic files:
FILE *f = fopen("plugin_debug.log", "w");
fprintf(f, "value at [%d][%d] = %f\n", i, j, val);
fclose(f);
Test standalone first. Write a main() that reads CSV and calls your algorithm. Debug with normal tools (gdb, valgrind, sanitizers). Then adapt for the plugin interface.
Build with sanitizers during development: -g -fsanitize=address
Check SF_vdata() return values. It returns RC (0=success). Non-zero means invalid obs/var index.
| Symptom | Likely Cause |
|---|---|
| Stata crashes silently | Segfault: buffer overflow, bad argv access, NULL deref |
| Plugin returns all missing | Wrong variable count, wrong obs indexing, plugin not loaded |
| Results are garbage | Sorting mismatch, 0-vs-1 indexing error, unnormalized inputs |
| "plugin not found" | Wrong filename, clear all wiped definition, wrong platform |
| Works on Mac, fails on Linux | Integer size difference, use int32_t/int64_t from <stdint.h> |
Use platform-specific .pkg files so users only download the binary for their OS. Stata's net install has no conditional logic, so the way to avoid shipping all 4 binaries to every user is to offer separate packages per platform. All packages install the same .ado and .sthlp files — only the .plugin binary differs.
mypackage/
├── stata.toc # lists all package variants
├── mypackage.pkg # all platforms (for users who don't care)
├── mypackage_mac.pkg # macOS only
├── mypackage_linux.pkg # Linux only
├── mypackage_win.pkg # Windows only
├── mycommand.sthlp # overview help file (short name!)
├── mycommand.ado # user-facing command
├── myplugin_macosx.plugin
├── myplugin_unix.plugin
├── myplugin_windows.plugin
└── c_source/ # NOT distributed, for building
├── build.py
├── stplugin.c
├── stplugin.h
└── algorithm.c
Users install their platform's package:
* macOS
net install mypackage_mac, from("https://raw.githubusercontent.com/user/repo/main") replace
* Linux
net install mypackage_linux, from("https://raw.githubusercontent.com/user/repo/main") replace
* Windows
net install mypackage_win, from("https://raw.githubusercontent.com/user/repo/main") replace
All platform binaries ship via the all-platform .pkg, or users can install platform-specific packages. Stata loads only the matching plugin at runtime via gtools-style OS detection. Windows C++ binaries can be 10-15MB due to static linking, which is normal.
See references/packaging_and_help.md for .toc, .pkg, .sthlp templates and SMCL formatting.
Sorting destroys merge keys. If you sort inside preserve/restore, the merge_id linkage breaks. Always create merge_id BEFORE preserve.
1-indexed everything. SF_vdata(var, obs, &val) — both var and obs start at 1. Off-by-one errors are silent.
marksample excludes missing by default. For imputation (where missing depvar IS the point), use marksample touse, novarlist.
macOS c(os) returns "MacOSX". Use the gtools pattern: inlist("c(os)'", "MacOSX") | strpos("c(machine_type)'", "Mac") to detect Mac. For other platforms, lower(c(os)) gives "windows" or "unix".
argv[] has no bounds checking. Accessing argv[3] when argc == 2 is a segfault. Always check argc first.
clear all wipes plugins. Reload plugin definitions after clear all in test scripts.
Only the first program define in a .ado file is auto-discovered. Subprograms need their own .ado files or explicit run to load.
Normalize inputs when the algorithm requires it (neural networks, gradient-based methods, distance-based methods like KNN). Scale to mean=0, sd=1 in the .ado wrapper, denormalize predictions after. The plugin should receive clean, normalized data — let the .ado handle the scaling.
pthreads on Windows needs -lwinpthread. Use conditional linker flags.
Memory errors crash Stata with no recovery. Pre-allocate everything, check every allocation, build with sanitizers during development.
glibc version mismatch. Building Linux plugins on a modern distro produces binaries that won't load on older systems. Use Ubuntu 18.04 in Docker for maximum compatibility.
SF_nvar() returns total dataset variables. It counts ALL variables in the dataset, not just the ones in the plugin call varlist. If the .ado creates tempvars (touse, merge_id, sort keys), the count will be higher than expected. Never use SF_nvar() to validate argument counts — pass the expected count via argv instead.
findfile + absolute paths breaks on Windows. findfile returns an absolute path that Stata's LoadLibrary can't resolve on Windows. Use the gtools-style OS detection pattern instead (see Plugin Loading section above) — it constructs a bare filename that Stata resolves via the adopath.
method() not model() for method selection optionsgenerate() (abbreviation gen()) for output variable namingreplace as a flag option, not replace()algorithm_plugin_os.plugin where os is macosx, unix, or windowsGENerate, MAXDepth)version 14.0) for plugin supportmypackage_stata, the overview help file should still be mypackage.sthlp (so help mypackage works). Don't append "stata" to help file or command names — the user is already in Stata.
评论 (0)
暂无评论,成为第一个评论者吧!