复制安装命令
用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。
复制前请先查看来源、License 和安全提示。
Security audit: baseline 52/52 CLEAN
用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。
复制前请先查看来源、License 和安全提示。
来源文件:README.md
📌 文档结构(2026-07-22 起): 本文件是中文默认入口 —— banner + badges + 信任面 + 9 阶段流水线速览 + 76 行合集总表。 每个合集的完整描述、按用途分组、精确数字、验证方法在
docs/CONTENT_ZH.md(扩展正文,总表行内的→直接跳转到对应锚点)。English version:
README-en.md· 中文扩展正文:docs/CONTENT_ZH.md·README-zh-CN.md已弃用(重定向占位)
🌐 语言: English | 简体中文(默认) | 繁體中文 | 日本語 | 한국어
|
|
Stanford REAP × CoPaper.AI · 实证研究 AI 工具的学术工业级产品
由斯坦福实证研究方法论团队打造,覆盖从数据清洗到顶刊投稿的完整工作流
🚀 New here? Open the Skill Search → to filter all 1,096 skills by method, stage, language, and license. The 5-minute tour (
make quickstart) prints the same picture in your terminal.🇨🇳 中文用户从本文件开始(流水线速览 + 76 行总表),每个合集的完整描述见
docs/CONTENT_ZH.md。📖 English readers: seeREADME-en.md.
| Rigor lane | Count | Where |
|---|---|---|
| Numeric benchmark tasks — gold values recomputed from real data each run | 17 | benchmark/ |
| Behavioral eval scenarios / rubric items | 37 / 183 | eval-harness/ |
Full trust overview:
docs/TRUST.md·docs/RIGOR_COVERAGE.md
中文内容分两级维护,各司其职:
docs/CONTENT_ZH.md(扩展正文):每个合集的完整描述(#skill-NN 锚点)、按用途分组、精确数字、2 分钟验证、三层信任、旗舰流水线详解、贡献与引用。总表行内的 → 直接跳到对应锚点。README-en.md · README-zh-TW.md · README-ja.md · README-ko.md[!NOTE] 维护规则: 改合集总表 → 本文件与 CONTENT_ZH.md 的锚点表两处同步;改合集详情 / 分组 / 数字 → 只改
docs/CONTENT_ZH.md。统计数字(合集数 / skill 数)以catalog/skills.json为准,由make validate的 readme-stats 检查器守护。贡献者(Contributors): 提交前请在本地跑通完整门禁
make check(catalog 校验 + 链接 + 单元测试 + eval-harness + benchmark)。详见CONTRIBUTING.md。旧版归档:
README-zh-CN.md已弃用,仅作向后兼容的重定向占位。
AERS 不只是 76 个散装 skill —— 它能陪你走完一篇论文。 从模糊 idea → 选题精炼 → 文献综述 → 数据获取 → 识别策略 → 估计建模 → 稳健性审计 → 出版级表格 / 图形 → 写作与同行评审 → 降 AIGC → 投稿。端到端、全自动、每一步都可被人介入(中间任何一步你都可以接过去手工改方法、补变量、加稳健性,再让流水线自动接上跑)。
Paper-WorkFlow 是 AERS 的"指挥棒",它把上面 9 个阶段的 skill 串成 一条按键即运行的端到端流水线。
你在 IDE 入口给它一句自然语言:
"开一个新论文项目:空气污染与中国劳动力市场,CS 设计 + 省级面板"
它会自动按顺序调:
sp.csdid(...) 给出 CS-DID 估计草案 + 写出估计方程与识别假设sp.feols(...) + sp.honest_did(...)任何阶段你都可以手动介入 —— 上一阶段的产物全部落盘(产物-幂等 pipeline),你接过去改方法、补控制、加稳健性,再让流水线自动接下去跑。这就是"全自动 + 可介入"。
| ⭐ Skill | 在流水线里的角色 |
|---|---|
| 00 StatsPAI 🔥 | 因果引擎:900+ 函数,sp.causal(...) 一行跑闭环(DID / RD / IV / SCM / DML / matching) |
| 00.1 Full Empirical · Python 📘 | 显式 Python 栈(pandas / statsmodels / linearmodels / pyfixest) |
| 00.2 Full Empirical · Stata 📊 | 显式 Stata 栈(reghdfe / ivreg2 / csdid / sdid / rdrobust) |
| 00.3 Full Empirical · R 📗 | 显式 R 栈(tidyverse / fixest / did / HonestDiD)+ Quarto 渲染 |
| 48 de-AIGC-skills 🇨🇳🇬🇧 | 中英双语学术降 AIGC(Turnitin AI / GPTZero / 知网 / 万方) |
| 50 AER-skills 📕 | Top-5 经济学投稿套件:识别 → 稳健性 → R&R |
| 69 Paper-WorkFlow 🧭 | 元编排器,把上面 9 个阶段串成一键流水线 |
为什么挑这 7 个?因为它们的行为都被基准钉死了 —— 不是营销口径,是对着已知答案反复跑过验证过的(17 项数值 benchmark + 37 项行为评测 ↗)。
↴ 直跳到下方 76 行总表(每个合集带 #skill-NN 锚点)。如果你更关心"这些 skill 怎么用"而不是"有哪些 skill",看 📘 中文唯一权威正文 里的「按用途分组」与「旗舰流水线」两节。
00 → 72,编号连续无空缺)打开仓库 → 看见整座库。 全部 76 个合集 · 1,096 个 skill,每一个都已 vendor 进本仓库,由
catalog/skills.json跟踪。⭐ = Stanford REAP × CoPaper.AI 团队自研的 skill;其余为精选、经安全审计的社区作品。主题图例 — 🚀 全流程与编排器 · 🎯 因果推断与计量经济学 · 📚 文献与研究设计 · ✍️ 写作 / 编辑 / 去 AIGC · 📑 引用 / 复现 / 同行评审 · 🛠️ 数据 / 工具 / 基础设施
点击【→】 跳转到
docs/CONTENT_ZH.md中该合集的完整描述;点击合集名 直接打开其目录。
| # | 合集 | 一句话 | 详情 |
|---|---|---|---|
| ⭐ 00 | StatsPAI 🔥 | 因果引擎 · Agent-native Python DSL:sp.causal(...) 一行跑闭环(DID/RD/IV/SCM/DML,900+ 函数) | → |
| ⭐ 00.1 | Full Empirical · Python 📘 | 显式栈:pandas · statsmodels · linearmodels · pyfixest | → |
| ⭐ 00.2 | Full Empirical · Stata 📊 | reghdfe · ivreg2 · csdid · sdid · rdrobust 复现包 | → |
| ⭐ 00.3 | Full Empirical · R 📗 | tidyverse · fixest · did · HonestDiD + Quarto 渲染 | → |
| 01 | academic-paper-skills | 大纲 → 手稿写作 + 7 维审稿人模拟 | → |
| 02 | research-skills | 医学影像综述、提案、论文转幻灯片 | → |
| 03 | scientific-skills | 假设生成 + 28 个科学数据库 | → |
| 04 | scientific-writer | 引用管理 + 科学写作 | → |
| 05 | research-superpower | 系统化检索、筛选与引文溯源 | → |
| 06 | stats-paper-writing | 端到端 LaTeX 统计论文写作 | → |
| 07 | AI-Research-SKILLs | 发表级 ML 图表、LaTeX、引文核验 | → |
| 08 | latex-document-skill | 创建 / 编译任意 LaTeX 文档为 PDF | → |
| 09 | awesome-econ-ai | Python 面板数据分析(linearmodels) | → |
| 10 | causal-inference-mixtape | DID / IV / RDD / SCM 模板(Cunningham) | → |
| 11 | compound-science | 面向定量社会科学的贝叶斯估计 | → |
| 12 | claude-code-my-workflow | 提交 → PR → 合并的研究工作流(Emory) | → |
| 13 | MixtapeTools | Cunningham 的因果推断工具集与讲义 | → |
| 14 | research-starter | R 中的 IV / DiD / RDD,含完整诊断 | → |
| 15 | social-science-research | R 或 Python 端到端数据分析 | → |
| 16 | clo-author | 多代理数据分析(R / Stata / Python) | → |
| 17 | DAAF | 安全意识代理框架(32 条 deny rule) | → |
| 18 | stata-accounting | 来自 126 篇 JAR 论文的实测 Stata 范式 | → |
| 19 | vera-economic-intelligence | 经济情报 / 政策研究情报工作流 | → |
| 20 | python-econ-skill | DSGE / HANK 与定量经济计算 | → |
| 21 | AI-research-feedback | 用 AI 同行评审生成结构化反馈 | → |
| 22 | christopherkenny-skills | 面向 Quarto(.qmd)的 APSA 风格检查器 | → |
| 23 | baygent | 带护栏的 PyMC / Arviz 贝叶斯工作流 | → |
| 24 | academic-research-skills | 5 审稿人多视角论文评审 | → |
| 25 | Diverga | 研究问题精炼器(抗模式坍缩) | → |
| 26 | scholar | 统计算法设计与文档 | → |
| 27 | my_claude_skills | 经济学摘要写作指南 | → |
| 28 | paper-replicate-agent | 论文复现代理演示 | → |
| 29 | project20XXy | 可复现手稿 + notebook 项目 | → |
| 30 | zirui-song-claude-skills | Zirui Song 的研究辅助 Claude 技能集 | → |
| 31 | claude-code-skills | Python 面板数据分析 | → |
| 32 | stata-skill | 高性能 Stata C/C++ 插件 | → |
| 33 | claude-scholar | 研究全生命周期:选题 → 综述 → 实验 → 审稿回复 | → |
| 34 | research-companion | 头脑风暴、评估并决策研究方向 | → |
| 35 | academic-writing-skills | 面向投稿场所的工业 AI 文献研究 | → |
| 36 | literature-review-skill | 完整文献综述工作流(中文) | → |
| 37 | IlanStrauss-ai-skills | Ilan Strauss 经济学研究 AI 工作流 | → |
| 38 | academic-proofreader | 学术校对 | → |
| 39 | marginaleffects | 预测、斜率与比较(R / Python) | → |
| 40 | pyfixest | Python 中的快速固定效应估计 | → |
| 41 | sewage-econometrics-check | 10 项复现包审计 | → |
| 42 | ARIS | 自主「research-in-sleep」代理,端到端 | → |
| 43 | research-plugins | 478 个研究插件:数据可视化、领域、基础设施 | → |
| 44 | humanizer_academic | 为医学/学术手稿去 AI 味(23 类模式) | → |
| 45 | deslop | 去除 AI 写作痕迹(5 维评分) | → |
| 46 | stop-slop | 三层 AI 痕迹检测与改写 | → |
| 47 | avoid-ai-writing | 审计 → 改写 → 二次审计 AI 味(留痕) | → |
| ⭐ 48 | de-AIGC-skills 🇨🇳🇬🇧 | 中英双语学术降 AIGC(Turnitin AI / GPTZero / 知网 / 万方) | → |
| 49 | humanize-chinese | 检测并人性化 AI 生成的中文文本 | → |
| ⭐ 50 | AER-skills 📕 | Top-5 经济学投稿套件:识别 → 稳健性 → R&R | → |
| 51 | CausalPy | 贝叶斯准实验(PyMC Labs) | → |
| 52 | slr-prisma | 系统文献综述,PRISMA 2020 | → |
| 53 | thematic-analysis | Braun & Clarke 六阶段定性主题分析 | → |
| 54 | open-science-skills | 引用一致性、DOI 与论据支撑审计 | → |
| 55 | r-skills | R 中用 brms 做贝叶斯推断 | → |
| 56 | econ-writing-skill | 综合 50+ 顶级指南的经济学写作 | → |
| 57 | edgartools | 查询与分析 SEC 文件 | → |
| 58 | econstack | 政策简报(UK GES / AU Treasury) | → |
| 59 | openalex-skill | 通过 OpenAlex 查询 2.4 亿+ 学术作品 | → |
| 60 | superpapers | 综合性实证研究支持套件 | → |
| 61 | research-methods | 与预注册匹配的验证性检验 | → |
| 62 | citation-checker | 对照 CrossRef / S2 / OpenAlex 核验引用 | → |
| 63 | scientific-agent-skills | DoWhy 识别–估计–反驳框架 | → |
| 64 | mcp-stata | 20 个 Stata 因果推断与复现 skill | → |
| 65 | game-theory-paper-writer | 生成并压力测试博弈论论文 | → |
| 66 | empirical-research-skills | 面向大型面板的 R 性能优化 | → |
| 67 | econfin-workflow-toolkit | 中国公司金融实证工作流,从提案到论文 | → |
| 68 | research-productivity-skills | 论文检索、SSRN、DOI 查询、下载 | → |
| ⭐ 69 | Paper-WorkFlow 🧭 | 元编排器,串起整个社会科学论文流水线 | → |
| 70 | ssci-polish ✍️ | SSCI / SCI 英文论文语言润色(语法、可读性、学术语气) | → |
| ⭐ 71 | lit-review-agent-tools 🔍 | 文献综述工具选型 + 一键安装运行(MinerU / PaperQA2 / ASReview / STORM / MCP 服务器) | → |
| ⭐ 72 | Kaggle Research 🧪 | 通过官方 CLI 安全检索 Kaggle 资源、限界下载公开数据并保留审计证据 | → |
想看更详细的描述(主题分类、字段、统计)? 见
docs/CONTENT_ZH.md中标注#skill-NN锚点的同一张表 —— 它是每个合集的完整描述所在的扩展正文。
自 2026-04 首次发布以来的主干里程碑(完整提交记录见 Commits 与 CHANGELOG.md):
---
config:
gitGraph:
rotateCommitLabel: false
---
gitGraph TB:
commit id: "2026-04 首次发布"
branch community
commit id: "2026-05 首个社区 PR"
checkout main
merge community
commit id: "2026-05 更名 AERS"
commit id: "2026-06 插件市场"
commit id: "2026-06 全库路由器"
commit id: "2026-07 首个 tag" tag: "v2026.07"
branch kaggle
commit id: "2026-07 Kaggle 集成"
checkout main
merge kaggle
commit id: "2026-08 de-AIGC 双语"
Star 增长曲线(非提交数)· 由 scripts/build-star-history.py 从 GitHub API 生成并提交入库
如果 AERS 对你的工作有帮助,请引用它(CITATION.cff)并点个 Star,让更多研究者看到。
AI 是放大器,不是替代品。它替你做最耗时的"搬砖",你保留最核心的"判断"。
|
|
Stanford REAP × CoPaper.AI · 实证研究 AI 工具的学术工业级产品
![]() 扫码访问 copaper.ai |
![]() 关注公众号「CoPaper.AI」 |
内置 20 个方法论 skill · 20 分钟完成实证论文 · 自研 StatsPAI(900+ 函数 / MIT 开源)
name: r-python-translation
description: >-
R-to-Python translation for data analysis. Maps R packages (tidyverse, ggplot2, fixest, survey, sf, plm) to Python equivalents (polars, plotnine, pyfixest, svy, geopandas). Use when user has R background or requests R-equivalent code comments.
metadata:
audience: research-coders
domain: research-methodology
skill-last-updated: "2026-03-28"R-to-Python translation reference for quantitative social science data analysis. Maps R ecosystem packages (tidyverse/dplyr, ggplot2, fixest, survey, sf, plm, lme4, marginaleffects, rdrobust) to DAAF Python equivalents (polars, plotnine, pyfixest, statsmodels, linearmodels, svy, geopandas). Use when user mentions R/RStudio background, requests R-equivalent code comments, needs to understand Python analysis code from an R perspective, or wants to translate R data analysis concepts to Python. Covers paradigm differences, verb-by-verb operation translations, regression modeling, causal inference, visualization, and workflow adaptation.
Cross-language translation reference for researchers moving between the R and Python data analysis ecosystems. This skill maps R packages, idioms, and workflows to their DAAF Python equivalents so that R-background users can audit, understand, and learn from DAAF-produced code, and so that code-producing agents can annotate their output with R equivalents when directed.
This skill is a routing hub — it provides overview tables, decision trees, and directs readers to the detailed reference files listed below. The reference files contain the exhaustive verb-by-verb mappings, code examples, and edge-case documentation.
Use cases:
Each topic in ./references/ contains focused documentation:
| File | Purpose | When to Read |
|---|---|---|
paradigm-differences.md | Core language and paradigm differences | Encountering fundamental R-vs-Python confusion |
polars-dplyr.md | Core dplyr/tidyr to polars verb mapping (select, filter, mutate, joins, reshaping, window functions, lazy eval) | Reading or writing data manipulation code |
polars-strings-dates-factors.md | String, date/time, and factor operations (stringr, lubridate, forcats to polars) | Working with string/date/categorical columns |
regression-modeling.md | fixest/stats/plm to pyfixest/statsmodels/linearmodels | Reading or writing regression code |
visualization.md | ggplot2/plotly R to plotnine/plotly Python | Reading or writing visualization code |
causal-inference.md | R causal inference ecosystem to Python equivalents | Working with DiD, RDD, IV, event studies |
survey-spatial-ml.md | survey/sf/tidymodels to svy/geopandas/scikit-learn | Working with surveys, spatial data, or ML |
workflow-environment.md | RStudio/Quarto workflow to DAAF/marimo workflow | Adapting to DAAF's execution model |
external-resources.md | Curated guides and tutorials with provenance | Seeking additional learning materials |
gotchas.md | Common R-user mistakes in Python | Debugging or reviewing code from R perspective |
paradigm-differences.md then the relevant domain file (e.g., polars-dplyr.md for data wrangling, regression-modeling.md for models) then gotchas.mdparadigm-differences.md then polars-dplyr.md then workflow-environment.md then external-resources.mdWhat kind of R operation?
├─ Data wrangling (filter, mutate, join, pivot, summarise)
│ └─ ./references/polars-dplyr.md
├─ Regression / statistical modeling
│ └─ ./references/regression-modeling.md
├─ Plotting / visualization
│ └─ ./references/visualization.md
├─ Causal inference (DiD, RDD, IV, event studies)
│ └─ ./references/causal-inference.md
├─ Surveys / spatial / machine learning
│ └─ ./references/survey-spatial-ml.md
└─ Fundamental language differences (types, syntax, environment)
└─ ./references/paradigm-differences.md
What looks unfamiliar?
├─ Expression syntax (pl.col().method().alias())
│ └─ ./references/paradigm-differences.md
├─ Missing values (None vs NaN vs null vs NA)
│ └─ ./references/paradigm-differences.md
├─ Formula interface (~) behaves differently
│ └─ ./references/regression-modeling.md
├─ Import patterns and namespacing
│ └─ ./references/gotchas.md
└─ No interactive REPL / console workflow
└─ ./references/workflow-environment.md
What does the R script do?
├─ Loads and wrangles data (read_csv, dplyr verbs)
│ └─ ./references/polars-dplyr.md
├─ Runs regressions (lm, feols, plm)
│ └─ ./references/regression-modeling.md
├─ Creates plots (ggplot, plotly)
│ └─ ./references/visualization.md
├─ Uses survey weights (svydesign, svymean)
│ └─ ./references/survey-spatial-ml.md
├─ Spatial operations (sf, st_join)
│ └─ ./references/survey-spatial-ml.md
├─ Multiple of the above
│ └─ Start with ./references/paradigm-differences.md, then each relevant file
└─ Uses a package not listed above
└─ ./references/external-resources.md for broader ecosystem guidance
What went wrong?
├─ 1-indexed access gave wrong element
│ └─ ./references/gotchas.md
├─ Factor/categorical behaves differently
│ └─ ./references/gotchas.md
├─ NA handling surprised me
│ └─ ./references/paradigm-differences.md
├─ Pipe operator (|> or %>%) not available
│ └─ ./references/paradigm-differences.md
├─ library() vs import confusion
│ └─ ./references/gotchas.md
└─ Model output structure is different
└─ ./references/regression-modeling.md
Which R package?
├─ dplyr / tidyr / readr / tibble → polars
│ └─ ./references/polars-dplyr.md
├─ ggplot2 → plotnine
│ └─ ./references/visualization.md
├─ plotly (R) → plotly (Python)
│ └─ ./references/visualization.md
├─ fixest → pyfixest
│ └─ ./references/regression-modeling.md
├─ stats (lm, glm) → statsmodels
│ └─ ./references/regression-modeling.md
├─ plm / lme4 / estimatr → linearmodels
│ └─ ./references/regression-modeling.md
├─ survey → svy
│ └─ ./references/survey-spatial-ml.md
├─ sf / terra → geopandas
│ └─ ./references/survey-spatial-ml.md
├─ tidymodels / caret → scikit-learn
│ └─ ./references/survey-spatial-ml.md
├─ marginaleffects → marginaleffects (Python)
│ └─ ./references/regression-modeling.md
├─ rdrobust / did / synthdid → rdrobust / pyfixest DiD
│ └─ ./references/causal-inference.md
└─ Quarto / RMarkdown → marimo
└─ ./references/workflow-environment.md
| Python Package | R Equivalent | Fidelity | Key Difference |
|---|---|---|---|
| polars | dplyr + tidyr + data.table | Low | Expression system vs verb grammar; method chaining vs pipe |
| pyfixest | fixest | High | Near-identical formula syntax; minor SE default differences |
| plotnine | ggplot2 | High | Same grammar of graphics; Python string quoting for aes |
| plotly | plotly (R) | High | px.scatter() vs plot_ly(); similar output |
| statsmodels | base R stats + lmtest + sandwich | Medium | Three formula dialects; manual vcov specification |
| linearmodels | plm + lme4 + estimatr | Medium | Requires pandas MultiIndex for panel structure |
| scikit-learn | tidymodels / caret | Medium | Imperative fit/predict vs declarative recipe pipeline |
| geopandas | sf + terra | Medium | shapely geometries vs sfc; different CRS handling |
| svy | survey (Lumley) | Medium | Limited GLM family coverage (gaussian/binomial/Poisson only) |
| marimo | Quarto / RMarkdown | Medium | Reactive cells vs knit-based linear execution |
Fidelity key: High = near-direct translation, same mental model. Medium = same capability, different API patterns. Low = fundamentally different paradigm requiring conceptual remapping.
Translations in this skill reference specific library versions. Python versions are pinned in DAAF's Docker environment (Python 3.12). R versions reference CRAN releases as of March 2026. When syntax or behavior has changed between versions, the reference files note the change.
| Python Package | DAAF Version | R Equivalent | R Version (CRAN) |
|---|---|---|---|
| polars | 1.38.1 | dplyr + tidyr + data.table | dplyr 1.2.0, tidyr 1.3.2, data.table 1.18.2 |
| pyfixest | 0.40.0 | fixest | 0.14.0 |
| plotnine | 0.15.3 | ggplot2 | 4.0.2 |
| plotly | 6.5.2 | plotly (R) | 4.12.0 |
| statsmodels | 0.14.6 | base R stats + lmtest + sandwich | lmtest 0.9-40, sandwich 3.1-1 |
| linearmodels | unpinned | plm + lme4 + estimatr | plm 2.6-7, lme4 2.0-1 |
| scikit-learn | 1.8.0 | tidymodels / caret | tidymodels 1.4.1, caret 7.0-1 |
| geopandas | 1.1.3 | sf + terra | sf 1.1-0, terra 1.9-11 |
| svy | 0.13.0 | survey | survey 4.5 |
| marginaleffects | unpinned | marginaleffects (R) | 0.32.0 |
| rdrobust | unpinned | rdrobust (R) | 3.0.0 |
| marimo | 0.19.11 | Quarto / RMarkdown | Quarto 1.6.x |
Unpinned packages: linearmodels, marginaleffects, and rdrobust install the latest version at Docker build time. Translations for these packages reference their documented API as of March 2026.
R version note: R package versions are from CRAN as of March 2026 (R 4.5.3). Check
packageVersion("pkg") in your R installation to verify your local version matches.
These are the friction points R users encounter most frequently when reading or writing DAAF Python code. Each is covered in depth in the referenced file.
| # | Friction Point | R Way | Python Way | Reference |
|---|---|---|---|---|
| 1 | Expression system | df %>% mutate(x = a + b) | df.with_columns((pl.col("a") + pl.col("b")).alias("x")) | paradigm-differences.md |
| 2 | Formula fragmentation | One universal ~ syntax | Three dialects (pyfixest, statsmodels, linearmodels) | regression-modeling.md |
| 3 | Missing values | Single NA type | None, NaN, and null (context-dependent) | paradigm-differences.md |
| 4 | mutate equivalent | mutate(new = expr) | with_columns(expr.alias("new")) | polars-dplyr.md |
| 5 | No row index | Tibbles have row numbers | Polars has no row index; use with_row_index() | paradigm-differences.md |
| 6 | Polars-to-pandas bridge | Data frames go directly into models | Must call .to_pandas() before statsmodels/pyfixest | paradigm-differences.md |
| 7 | Factor vs Categorical | factor() with ordered levels | pl.Categorical / pd.Categorical (different semantics) | gotchas.md |
| 8 | Package fragmentation | One package per domain (fixest does it all) | Multiple packages per domain (statsmodels + linearmodels + pyfixest) | paradigm-differences.md |
| 9 | 1-indexed vs 0-indexed | x[1] is first element | x[0] is first element | gotchas.md |
| 10 | Namespace model | library() exports all names | import requires explicit namespacing | gotchas.md |
This section defines when and how code-producing agents add inline R-equivalent comments to DAAF Python scripts.
Annotations are added only when the orchestrator explicitly passes an R-background directive to the agent. This is not a default behavior.
Trigger conditions (orchestrator activates this when any apply):
How the orchestrator passes the directive: The orchestrator adds the following to the agent prompt:
"User has R background. Load r-python-translation skill. Add inline R-equivalent comments for non-trivial data operations."
# R: df %>% filter(year == 2020)
filtered = df.filter(pl.col("year") == 2020)
# R: df %>% mutate(pct = count / sum(count))
result = df.with_columns(
(pl.col("count") / pl.col("count").sum()).alias("pct")
)
# R: feols(y ~ x1 + x2 | state + year, data = df, cluster = ~state)
fit = pf.feols("y ~ x1 + x2 | state + year", data=pdf, vcov={"CRV1": "state"})
print()/assert validation lines, file I/O boilerplate (pl.read_parquet, df.write_parquet), config sections, section separator comments# R: comment per logical operation, placed on the line immediately above the Python code# INTENT:, # REASONING:, # ASSUMES:), not a replacement| Skill | Relationship |
|---|---|
polars | Python-side data wrangling — detailed API reference for the dplyr/tidyr equivalent |
pyfixest | Python-side fixed effects regression — detailed API for the fixest equivalent |
plotnine | Python-side static visualization — detailed API for the ggplot2 equivalent |
plotly | Python-side interactive visualization — detailed API for plotly R equivalent |
statsmodels | Python-side general modeling — covers base R stats, lmtest, sandwich equivalents |
linearmodels | Python-side panel/IV models — covers plm, lme4, estimatr equivalents |
scikit-learn | Python-side ML — covers tidymodels/caret equivalents |
geopandas | Python-side spatial data — covers sf/terra equivalents |
svy | Python-side survey analysis — covers survey (Lumley) equivalents |
marimo | Python-side notebooks — covers Quarto/RMarkdown workflow equivalents |
stata-python-translation | Parallel skill for Stata-background users — shares the same Python target stack |
Note: Individual tool skills contain library-specific usage guidance (syntax, gotchas, performance). This skill provides the R-to-Python conceptual bridge — use both together when an R-background user is working with a specific library.
| Topic | Reference File |
|---|---|
Pipe operator (%>% / ` | >`) equivalents |
| Expression system (pl.col, .alias) | ./references/paradigm-differences.md |
| Missing value semantics (NA vs None/NaN/null) | ./references/paradigm-differences.md |
| Type system differences | ./references/paradigm-differences.md |
| Package/namespace model | ./references/paradigm-differences.md |
| 0-indexing vs 1-indexing | ./references/paradigm-differences.md |
| Polars-to-pandas conversion for modeling | ./references/paradigm-differences.md |
| Row index differences | ./references/paradigm-differences.md |
| dplyr verb mapping (filter, select, mutate, arrange) | ./references/polars-dplyr.md |
| summarise / group_by equivalents | ./references/polars-dplyr.md |
| tidyr verbs (pivot_longer, pivot_wider, separate, unite) | ./references/polars-dplyr.md |
| Join operations (left_join, inner_join, anti_join) | ./references/polars-dplyr.md |
| String operations (stringr vs polars .str) | ./references/polars-strings-dates-factors.md |
| Date operations (lubridate vs polars .dt) | ./references/polars-strings-dates-factors.md |
| across() / where() equivalents | ./references/polars-dplyr.md |
| case_when equivalent | ./references/polars-dplyr.md |
| readr I/O equivalents | ./references/polars-dplyr.md |
| fixest formula syntax in pyfixest | ./references/regression-modeling.md |
| lm() / glm() in statsmodels | ./references/regression-modeling.md |
| Formula interface comparison (three Python dialects) | ./references/regression-modeling.md |
| Standard error specification differences | ./references/regression-modeling.md |
| plm panel models in linearmodels | ./references/regression-modeling.md |
| lme4 mixed effects equivalents | ./references/regression-modeling.md |
| marginaleffects (R to Python) | ./references/regression-modeling.md |
| Model summary / tidy output | ./references/regression-modeling.md |
| Sandwich / robust SE equivalents | ./references/regression-modeling.md |
| ggplot2 layer mapping to plotnine | ./references/visualization.md |
| aes() string quoting in plotnine | ./references/visualization.md |
| Theme customization | ./references/visualization.md |
| Scale functions | ./references/visualization.md |
| Faceting (facet_wrap, facet_grid) | ./references/visualization.md |
| plotly R vs plotly Python | ./references/visualization.md |
| ggsave equivalent | ./references/visualization.md |
| Difference-in-differences (did, did2s) | ./references/causal-inference.md |
| Regression discontinuity (rdrobust) | ./references/causal-inference.md |
| Instrumental variables (ivreg vs pyfixest IV) | ./references/causal-inference.md |
| Event study designs | ./references/causal-inference.md |
| Synthetic control | ./references/causal-inference.md |
| Matching / propensity scores | ./references/causal-inference.md |
| survey package to svy | ./references/survey-spatial-ml.md |
| svydesign / svymean / svyglm equivalents | ./references/survey-spatial-ml.md |
| sf spatial operations to geopandas | ./references/survey-spatial-ml.md |
| CRS / projection handling | ./references/survey-spatial-ml.md |
| Spatial joins (st_join vs sjoin) | ./references/survey-spatial-ml.md |
| tidymodels pipeline to scikit-learn | ./references/survey-spatial-ml.md |
| RStudio vs DAAF workflow | ./references/workflow-environment.md |
| Quarto / RMarkdown vs marimo | ./references/workflow-environment.md |
| Interactive console vs file-first execution | ./references/workflow-environment.md |
| Package management (renv vs pip/uv) | ./references/workflow-environment.md |
| Project structure conventions | ./references/workflow-environment.md |
| Curated R-to-Python migration guides | ./references/external-resources.md |
| Package documentation links | ./references/external-resources.md |
| Tutorial recommendations with provenance | ./references/external-resources.md |
| 1-indexed list/vector access | ./references/gotchas.md |
| Factor vs Categorical pitfalls | ./references/gotchas.md |
| library() vs import habits | ./references/gotchas.md |
| T/F vs True/False | ./references/gotchas.md |
| Assignment operator (<- vs =) | ./references/gotchas.md |
| Vectorized operations expectations | ./references/gotchas.md |
| NULL vs None differences | ./references/gotchas.md |
| apply family vs map/list comprehension | ./references/gotchas.md |
| Copying semantics (R copy-on-modify vs Python references) | ./references/gotchas.md |
| Logical operators (& / | vs and / or) |
| String interpolation (glue vs f-strings) | ./references/gotchas.md |
| data.table vs polars | ./references/polars-strings-dates-factors.md |
| Lazy evaluation (polars LazyFrame vs R lazy tibble) | ./references/polars-dplyr.md |
| nest/unnest equivalents | ./references/polars-dplyr.md |
| Window functions (over vs mutate + group_by) | ./references/polars-dplyr.md |
| Coordinate systems (coord_flip, coord_polar) | ./references/visualization.md |
| Stat layers (stat_smooth, stat_summary) | ./references/visualization.md |
| Color palette mapping (viridis, brewer) | ./references/visualization.md |
| Multi-panel layouts (patchwork vs subplot) | ./references/visualization.md |
| Staggered DiD estimators | ./references/causal-inference.md |
| Parallel trends testing | ./references/causal-inference.md |
| BRR / jackknife replication weights | ./references/survey-spatial-ml.md |
| Raster data handling (terra vs rasterio) | ./references/survey-spatial-ml.md |
| Feature engineering (recipes vs sklearn Pipeline) | ./references/survey-spatial-ml.md |
| Cross-validation (rsample vs sklearn) | ./references/survey-spatial-ml.md |
| Environment/workspace differences (.RData vs nothing) | ./references/workflow-environment.md |
| Debugging workflow (browser() vs breakpoint()) | ./references/workflow-environment.md |
| R help system (?func) vs Python help(func) | ./references/workflow-environment.md |
| Cheat sheet and quick-reference links | ./references/external-resources.md |
| Community resources (Stack Overflow tags, forums) | ./references/external-resources.md |
评论 (0)
暂无评论,成为第一个评论者吧!