复制安装命令
用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。
复制前请先查看来源、License 和安全提示。
Security audit: baseline 52/52 CLEAN
用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。
复制前请先查看来源、License 和安全提示。
来源文件:README.md
📌 文档结构(2026-07-22 起): 本文件是中文默认入口 —— banner + badges + 信任面 + 9 阶段流水线速览 + 76 行合集总表。 每个合集的完整描述、按用途分组、精确数字、验证方法在
docs/CONTENT_ZH.md(扩展正文,总表行内的→直接跳转到对应锚点)。English version:
README-en.md· 中文扩展正文:docs/CONTENT_ZH.md·README-zh-CN.md已弃用(重定向占位)
🌐 语言: English | 简体中文(默认) | 繁體中文 | 日本語 | 한국어
|
|
Stanford REAP × CoPaper.AI · 实证研究 AI 工具的学术工业级产品
由斯坦福实证研究方法论团队打造,覆盖从数据清洗到顶刊投稿的完整工作流
🚀 New here? Open the Skill Search → to filter all 1,096 skills by method, stage, language, and license. The 5-minute tour (
make quickstart) prints the same picture in your terminal.🇨🇳 中文用户从本文件开始(流水线速览 + 76 行总表),每个合集的完整描述见
docs/CONTENT_ZH.md。📖 English readers: seeREADME-en.md.
| Rigor lane | Count | Where |
|---|---|---|
| Numeric benchmark tasks — gold values recomputed from real data each run | 17 | benchmark/ |
| Behavioral eval scenarios / rubric items | 37 / 183 | eval-harness/ |
Full trust overview:
docs/TRUST.md·docs/RIGOR_COVERAGE.md
中文内容分两级维护,各司其职:
docs/CONTENT_ZH.md(扩展正文):每个合集的完整描述(#skill-NN 锚点)、按用途分组、精确数字、2 分钟验证、三层信任、旗舰流水线详解、贡献与引用。总表行内的 → 直接跳到对应锚点。README-en.md · README-zh-TW.md · README-ja.md · README-ko.md[!NOTE] 维护规则: 改合集总表 → 本文件与 CONTENT_ZH.md 的锚点表两处同步;改合集详情 / 分组 / 数字 → 只改
docs/CONTENT_ZH.md。统计数字(合集数 / skill 数)以catalog/skills.json为准,由make validate的 readme-stats 检查器守护。贡献者(Contributors): 提交前请在本地跑通完整门禁
make check(catalog 校验 + 链接 + 单元测试 + eval-harness + benchmark)。详见CONTRIBUTING.md。旧版归档:
README-zh-CN.md已弃用,仅作向后兼容的重定向占位。
AERS 不只是 76 个散装 skill —— 它能陪你走完一篇论文。 从模糊 idea → 选题精炼 → 文献综述 → 数据获取 → 识别策略 → 估计建模 → 稳健性审计 → 出版级表格 / 图形 → 写作与同行评审 → 降 AIGC → 投稿。端到端、全自动、每一步都可被人介入(中间任何一步你都可以接过去手工改方法、补变量、加稳健性,再让流水线自动接上跑)。
Paper-WorkFlow 是 AERS 的"指挥棒",它把上面 9 个阶段的 skill 串成 一条按键即运行的端到端流水线。
你在 IDE 入口给它一句自然语言:
"开一个新论文项目:空气污染与中国劳动力市场,CS 设计 + 省级面板"
它会自动按顺序调:
sp.csdid(...) 给出 CS-DID 估计草案 + 写出估计方程与识别假设sp.feols(...) + sp.honest_did(...)任何阶段你都可以手动介入 —— 上一阶段的产物全部落盘(产物-幂等 pipeline),你接过去改方法、补控制、加稳健性,再让流水线自动接下去跑。这就是"全自动 + 可介入"。
| ⭐ Skill | 在流水线里的角色 |
|---|---|
| 00 StatsPAI 🔥 | 因果引擎:900+ 函数,sp.causal(...) 一行跑闭环(DID / RD / IV / SCM / DML / matching) |
| 00.1 Full Empirical · Python 📘 | 显式 Python 栈(pandas / statsmodels / linearmodels / pyfixest) |
| 00.2 Full Empirical · Stata 📊 | 显式 Stata 栈(reghdfe / ivreg2 / csdid / sdid / rdrobust) |
| 00.3 Full Empirical · R 📗 | 显式 R 栈(tidyverse / fixest / did / HonestDiD)+ Quarto 渲染 |
| 48 de-AIGC-skills 🇨🇳🇬🇧 | 中英双语学术降 AIGC(Turnitin AI / GPTZero / 知网 / 万方) |
| 50 AER-skills 📕 | Top-5 经济学投稿套件:识别 → 稳健性 → R&R |
| 69 Paper-WorkFlow 🧭 | 元编排器,把上面 9 个阶段串成一键流水线 |
为什么挑这 7 个?因为它们的行为都被基准钉死了 —— 不是营销口径,是对着已知答案反复跑过验证过的(17 项数值 benchmark + 37 项行为评测 ↗)。
↴ 直跳到下方 76 行总表(每个合集带 #skill-NN 锚点)。如果你更关心"这些 skill 怎么用"而不是"有哪些 skill",看 📘 中文唯一权威正文 里的「按用途分组」与「旗舰流水线」两节。
00 → 72,编号连续无空缺)打开仓库 → 看见整座库。 全部 76 个合集 · 1,096 个 skill,每一个都已 vendor 进本仓库,由
catalog/skills.json跟踪。⭐ = Stanford REAP × CoPaper.AI 团队自研的 skill;其余为精选、经安全审计的社区作品。主题图例 — 🚀 全流程与编排器 · 🎯 因果推断与计量经济学 · 📚 文献与研究设计 · ✍️ 写作 / 编辑 / 去 AIGC · 📑 引用 / 复现 / 同行评审 · 🛠️ 数据 / 工具 / 基础设施
点击【→】 跳转到
docs/CONTENT_ZH.md中该合集的完整描述;点击合集名 直接打开其目录。
| # | 合集 | 一句话 | 详情 |
|---|---|---|---|
| ⭐ 00 | StatsPAI 🔥 | 因果引擎 · Agent-native Python DSL:sp.causal(...) 一行跑闭环(DID/RD/IV/SCM/DML,900+ 函数) | → |
| ⭐ 00.1 | Full Empirical · Python 📘 | 显式栈:pandas · statsmodels · linearmodels · pyfixest | → |
| ⭐ 00.2 | Full Empirical · Stata 📊 | reghdfe · ivreg2 · csdid · sdid · rdrobust 复现包 | → |
| ⭐ 00.3 | Full Empirical · R 📗 | tidyverse · fixest · did · HonestDiD + Quarto 渲染 | → |
| 01 | academic-paper-skills | 大纲 → 手稿写作 + 7 维审稿人模拟 | → |
| 02 | research-skills | 医学影像综述、提案、论文转幻灯片 | → |
| 03 | scientific-skills | 假设生成 + 28 个科学数据库 | → |
| 04 | scientific-writer | 引用管理 + 科学写作 | → |
| 05 | research-superpower | 系统化检索、筛选与引文溯源 | → |
| 06 | stats-paper-writing | 端到端 LaTeX 统计论文写作 | → |
| 07 | AI-Research-SKILLs | 发表级 ML 图表、LaTeX、引文核验 | → |
| 08 | latex-document-skill | 创建 / 编译任意 LaTeX 文档为 PDF | → |
| 09 | awesome-econ-ai | Python 面板数据分析(linearmodels) | → |
| 10 | causal-inference-mixtape | DID / IV / RDD / SCM 模板(Cunningham) | → |
| 11 | compound-science | 面向定量社会科学的贝叶斯估计 | → |
| 12 | claude-code-my-workflow | 提交 → PR → 合并的研究工作流(Emory) | → |
| 13 | MixtapeTools | Cunningham 的因果推断工具集与讲义 | → |
| 14 | research-starter | R 中的 IV / DiD / RDD,含完整诊断 | → |
| 15 | social-science-research | R 或 Python 端到端数据分析 | → |
| 16 | clo-author | 多代理数据分析(R / Stata / Python) | → |
| 17 | DAAF | 安全意识代理框架(32 条 deny rule) | → |
| 18 | stata-accounting | 来自 126 篇 JAR 论文的实测 Stata 范式 | → |
| 19 | vera-economic-intelligence | 经济情报 / 政策研究情报工作流 | → |
| 20 | python-econ-skill | DSGE / HANK 与定量经济计算 | → |
| 21 | AI-research-feedback | 用 AI 同行评审生成结构化反馈 | → |
| 22 | christopherkenny-skills | 面向 Quarto(.qmd)的 APSA 风格检查器 | → |
| 23 | baygent | 带护栏的 PyMC / Arviz 贝叶斯工作流 | → |
| 24 | academic-research-skills | 5 审稿人多视角论文评审 | → |
| 25 | Diverga | 研究问题精炼器(抗模式坍缩) | → |
| 26 | scholar | 统计算法设计与文档 | → |
| 27 | my_claude_skills | 经济学摘要写作指南 | → |
| 28 | paper-replicate-agent | 论文复现代理演示 | → |
| 29 | project20XXy | 可复现手稿 + notebook 项目 | → |
| 30 | zirui-song-claude-skills | Zirui Song 的研究辅助 Claude 技能集 | → |
| 31 | claude-code-skills | Python 面板数据分析 | → |
| 32 | stata-skill | 高性能 Stata C/C++ 插件 | → |
| 33 | claude-scholar | 研究全生命周期:选题 → 综述 → 实验 → 审稿回复 | → |
| 34 | research-companion | 头脑风暴、评估并决策研究方向 | → |
| 35 | academic-writing-skills | 面向投稿场所的工业 AI 文献研究 | → |
| 36 | literature-review-skill | 完整文献综述工作流(中文) | → |
| 37 | IlanStrauss-ai-skills | Ilan Strauss 经济学研究 AI 工作流 | → |
| 38 | academic-proofreader | 学术校对 | → |
| 39 | marginaleffects | 预测、斜率与比较(R / Python) | → |
| 40 | pyfixest | Python 中的快速固定效应估计 | → |
| 41 | sewage-econometrics-check | 10 项复现包审计 | → |
| 42 | ARIS | 自主「research-in-sleep」代理,端到端 | → |
| 43 | research-plugins | 478 个研究插件:数据可视化、领域、基础设施 | → |
| 44 | humanizer_academic | 为医学/学术手稿去 AI 味(23 类模式) | → |
| 45 | deslop | 去除 AI 写作痕迹(5 维评分) | → |
| 46 | stop-slop | 三层 AI 痕迹检测与改写 | → |
| 47 | avoid-ai-writing | 审计 → 改写 → 二次审计 AI 味(留痕) | → |
| ⭐ 48 | de-AIGC-skills 🇨🇳🇬🇧 | 中英双语学术降 AIGC(Turnitin AI / GPTZero / 知网 / 万方) | → |
| 49 | humanize-chinese | 检测并人性化 AI 生成的中文文本 | → |
| ⭐ 50 | AER-skills 📕 | Top-5 经济学投稿套件:识别 → 稳健性 → R&R | → |
| 51 | CausalPy | 贝叶斯准实验(PyMC Labs) | → |
| 52 | slr-prisma | 系统文献综述,PRISMA 2020 | → |
| 53 | thematic-analysis | Braun & Clarke 六阶段定性主题分析 | → |
| 54 | open-science-skills | 引用一致性、DOI 与论据支撑审计 | → |
| 55 | r-skills | R 中用 brms 做贝叶斯推断 | → |
| 56 | econ-writing-skill | 综合 50+ 顶级指南的经济学写作 | → |
| 57 | edgartools | 查询与分析 SEC 文件 | → |
| 58 | econstack | 政策简报(UK GES / AU Treasury) | → |
| 59 | openalex-skill | 通过 OpenAlex 查询 2.4 亿+ 学术作品 | → |
| 60 | superpapers | 综合性实证研究支持套件 | → |
| 61 | research-methods | 与预注册匹配的验证性检验 | → |
| 62 | citation-checker | 对照 CrossRef / S2 / OpenAlex 核验引用 | → |
| 63 | scientific-agent-skills | DoWhy 识别–估计–反驳框架 | → |
| 64 | mcp-stata | 20 个 Stata 因果推断与复现 skill | → |
| 65 | game-theory-paper-writer | 生成并压力测试博弈论论文 | → |
| 66 | empirical-research-skills | 面向大型面板的 R 性能优化 | → |
| 67 | econfin-workflow-toolkit | 中国公司金融实证工作流,从提案到论文 | → |
| 68 | research-productivity-skills | 论文检索、SSRN、DOI 查询、下载 | → |
| ⭐ 69 | Paper-WorkFlow 🧭 | 元编排器,串起整个社会科学论文流水线 | → |
| 70 | ssci-polish ✍️ | SSCI / SCI 英文论文语言润色(语法、可读性、学术语气) | → |
| ⭐ 71 | lit-review-agent-tools 🔍 | 文献综述工具选型 + 一键安装运行(MinerU / PaperQA2 / ASReview / STORM / MCP 服务器) | → |
| ⭐ 72 | Kaggle Research 🧪 | 通过官方 CLI 安全检索 Kaggle 资源、限界下载公开数据并保留审计证据 | → |
想看更详细的描述(主题分类、字段、统计)? 见
docs/CONTENT_ZH.md中标注#skill-NN锚点的同一张表 —— 它是每个合集的完整描述所在的扩展正文。
自 2026-04 首次发布以来的主干里程碑(完整提交记录见 Commits 与 CHANGELOG.md):
---
config:
gitGraph:
rotateCommitLabel: false
---
gitGraph TB:
commit id: "2026-04 首次发布"
branch community
commit id: "2026-05 首个社区 PR"
checkout main
merge community
commit id: "2026-05 更名 AERS"
commit id: "2026-06 插件市场"
commit id: "2026-06 全库路由器"
commit id: "2026-07 首个 tag" tag: "v2026.07"
branch kaggle
commit id: "2026-07 Kaggle 集成"
checkout main
merge kaggle
commit id: "2026-08 de-AIGC 双语"
Star 增长曲线(非提交数)· 由 scripts/build-star-history.py 从 GitHub API 生成并提交入库
如果 AERS 对你的工作有帮助,请引用它(CITATION.cff)并点个 Star,让更多研究者看到。
AI 是放大器,不是替代品。它替你做最耗时的"搬砖",你保留最核心的"判断"。
|
|
Stanford REAP × CoPaper.AI · 实证研究 AI 工具的学术工业级产品
![]() 扫码访问 copaper.ai |
![]() 关注公众号「CoPaper.AI」 |
内置 20 个方法论 skill · 20 分钟完成实证论文 · 自研 StatsPAI(900+ 函数 / MIT 开源)
name: education-data-source-eada
description: >-
EADA — college athletics gender equity (~2,000+ institutions, 2002-2021). Participation, coaching, salaries, expenses, revenues, athletic aid by gender. Not Title IX compliance data. No sector column; join IPEDS on unitid for institution type.
metadata:
audience: any-agent
domain: data-source
skill-authored: "2026-02-09"
skill-last-updated: "2026-02-09"Equity in Athletics Disclosure Act (EADA) data for college athletics gender equity analysis covering ~2,000+ institutions (2002-2021). Use when analyzing athletic participation, coaching staff, salaries, expenses, revenues, or athletic aid by gender at colleges/universities, or understanding Title IX context in athletics. EADA is NOT Title IX compliance data. Note: no sector column; join to IPEDS on unitid to filter by institution type.
The EADA provides the only standardized, publicly available dataset on college athletics participation, coaching, finances, and athletic aid by gender across ~2,000+ postsecondary institutions, enabling gender equity analysis in intercollegiate athletics.
CRITICAL: Value Encoding
EADA data from the Education Data Portal uses integer codes for categorical variables. Original EADA web tools use string labels; the Portal converts these to integers. Always verify codes against the codebook (see Truth Hierarchy below).
Context ath_classification_codeMissing values Portal (integers) 1= NCAA DI FBS-1,-2,-3Original EADA String labels Blank / N/A Note: There is no
sectorcolumn in EADA Portal data. To filter by sector, join with IPEDS directory data onunitid.See
./references/variable-definitions.mdfor complete encoding tables.
unitid (6-digit IPEDS institution ID)| File | Purpose | When to Read |
|---|---|---|
title-ix-context.md | Legal framework, gender equity requirements | Understanding policy context |
data-elements.md | Participation, coaches, salaries, expenses, revenues | Identifying available variables |
sport-level-data.md | Data available by individual sport | Sport-specific analysis |
variable-definitions.md | Key variables, codes, special values | Interpreting specific data elements |
limitations.md | Data quality issues, comparability, self-reporting caveats | Assessing data reliability |
fetch-patterns.md | Mirror URLs and fetch code patterns | Fetching data |
Research question?
├─ Gender equity overview → Start with participation + aid ratios
│ └─ See ./references/data-elements.md
├─ Coaching disparities → Coach counts + salaries by gender
│ └─ See ./references/data-elements.md (Coaching section)
├─ Financial investment → Expenses + revenues by team gender
│ └─ See ./references/data-elements.md (Financial section)
├─ Sport-specific analysis → Individual sport data
│ └─ See ./references/sport-level-data.md
├─ Title IX compliance assessment → CAUTION: EADA ≠ compliance data
│ └─ See ./references/limitations.md (Critical)
└─ Trend analysis → Year-over-year comparisons
└─ See ./references/fetch-patterns.md
Variable categories?
├─ Participation counts
│ ├─ Unduplicated by gender → `undup_athpartic_men`, `undup_athpartic_women`
│ ├─ Duplicated (sport-level sum) → `athpartic_men`, `athpartic_women`
│ ├─ Coed teams → `athpartic_coed_men`, `athpartic_coed_women`
│ └─ By sport → See ./references/sport-level-data.md
├─ Coaching
│ ├─ Head coaches → `men_fthdcoach_*`, `women_fthdcoach_*` variables
│ ├─ Assistant coaches → `men_ftascoach_*`, `women_ftascoach_*` variables
│ └─ Salaries → `hdcoach_salary_*`, `ascoach_salary_*` variables
├─ Financial
│ ├─ Expenses → `ath_exp_*` variables
│ ├─ Revenues → `ath_rev_*` variables
│ └─ Athletic aid → `ath_stuaid_*` variables
└─ Detailed definitions → See ./references/variable-definitions.md
Interpretation question?
├─ What counts as "participation"?
│ └─ See ./references/variable-definitions.md
├─ Why don't participation ratios match enrollment?
│ └─ See ./references/limitations.md
├─ Is this institution Title IX compliant?
│ └─ CANNOT determine from EADA data alone
│ └─ See ./references/limitations.md (Critical)
├─ Why are some values missing or zero?
│ └─ See ./references/limitations.md
└─ How do I compare across institutions?
└─ See ./references/limitations.md (Comparability section)
| Metric | Calculation | Interpretation |
|---|---|---|
| Female participation ratio | undup_athpartic_women / (undup_athpartic_men + undup_athpartic_women) | Compare to female enrollment ratio |
| Participation gap | Female enrollment % - Female participation % | Positive = underrepresentation |
| Opportunities per student | undup_athpartic_total / enrollment_total | Athletic opportunity rate |
| Metric | Calculation | Notes |
|---|---|---|
| Aid ratio | ath_stuaid_women / (ath_stuaid_men + ath_stuaid_women) | Should approximate participation ratio |
| Per-participant expense | ath_opexp_perpart_men, ath_opexp_perpart_women | Pre-calculated per-participant operating expense |
| Recruiting investment | recruitexp_men, recruitexp_women | Indicator of program investment |
| Metric | Focus | Variables |
|---|---|---|
| Female coaches of women's teams | % female | women_fthdcoach_fem, women_pthdcoach_fem |
| Salary equity | Avg salary comparison | hdcoach_salary_men, hdcoach_salary_women |
| ID | Format | Level | Example | Notes |
|---|---|---|---|---|
unitid | 6-digit integer | Institution | 110635 | Same as IPEDS; primary join key |
opeid | String | Institution | "00123400" | OPE ID (may be null for early years) |
year | 4-digit integer | Reporting year | 2021 | Fiscal year ending |
fips | Integer | State | 6 (California) | Federal FIPS code |
inst_name | String | Institution | "University of..." | Institution name |
| Filter | Variable | Example Values |
|---|---|---|
| Institution | unitid | 6-digit IPEDS ID |
| Year | year | 2002–2021 |
| State | fips | Integer FIPS code (e.g., 6 = California) |
| Athletic Division | ath_classification_code | Integer codes 1–20 (see below) |
Note: There is no
sectorcolumn in the EADA Portal data. To filter by institutional sector, join with IPEDS directory data onunitid.
| Code | Division | Code | Division |
|---|---|---|---|
| 1 | NCAA Division I FBS | 12 | NJCAA Division I |
| 2 | NCAA Division I FCS | 13 | NJCAA Division II |
| 3 | NCAA Division I (no football) | 14 | NJCAA Division III |
| 4 | NCAA Division II (with football) | 15 | NCCAA Division I |
| 5 | NCAA Division II (no football) | 16 | NCCAA Division II |
| 6 | NCAA Division III (with football) | 17 | CCCAA |
| 7 | NCAA Division III (no football) | 18 | Independent |
| 8 | Other (check ath_classification_other) | 19 | NWAC |
| 9 | NAIA Division I | 20 | USCAA |
| 10 | NAIA Division II | ||
| 11 | NAIA Division III |
Note: Code 1 was historically labeled "NCAA Division I-A" and code 2 "NCAA Division I-AA" in earlier years. The
ath_classification_namestring column reflects the label used at the time of reporting.
| Code | Meaning | When Used |
|---|---|---|
-1 | Missing/not reported | Data not submitted by institution |
-2 | Not applicable | Item doesn't apply (e.g., no men's team) |
-3 | Suppressed | Data suppressed for privacy |
| Topic | Years Available | Update Frequency |
|---|---|---|
| Institution-level | 2002–2021 | Annual |
| Sport-level | 2002–2021 | Annual |
| Coaching details | 2002–2021 | Annual |
| Financial data | 2002–2021 | Annual |
Note: Some columns (e.g.,
num_sports, aggregated totals with_allsuffix) are null for earlier years (2002) and were added in later reporting cycles. Theopeidcolumn is null for 2002.
| Question | Key Variables | Reference |
|---|---|---|
| Are women underrepresented in athletics? | undup_athpartic_*, enrollment_* | data-elements.md |
| How much do institutions invest in women's sports? | ath_exp_*, ath_rev_* | data-elements.md |
| Are coaches of women's teams paid fairly? | hdcoach_salary_* | variable-definitions.md |
| Which sports have most female participants? | Sport-level data | sport-level-data.md |
| Has participation equity improved over time? | Multi-year trend | fetch-patterns.md |
Datasets for EADA are available via the Education Data Portal mirror system. All data fetching uses fetch_from_mirrors() from fetch-patterns.md, with mirrors defined in mirrors.yaml and canonical paths in datasets-reference.md.
Key datasets:
| Dataset | Path | Type | Codebook |
|---|---|---|---|
| Institutional Characteristics | eada/colleges_eada_inst_characteristics | Single | eada/codebook_colleges_eada_inst-characteristics |
EADA naming note: The data path uses
inst_characteristics(underscores) while the codebook path usesinst-characteristics(hyphens). Always use the exact paths fromdatasets-reference.md.
When interpreting EADA variable definitions and coded values, apply this priority:
| Priority | Source | Rationale |
|---|---|---|
| 1 (highest) | Actual data file (parquet) | What you observe IS the truth |
| 2 | Live codebook (.xls via get_codebook_url()) | Authoritative documentation; may lag |
| 3 (lowest) | This skill's reference docs | Summarized; convenient but may drift |
Use get_codebook_url("eada/codebook_colleges_eada_inst-characteristics") from fetch-patterns.md to construct the codebook download URL.
import polars as pl
# Filter by athletic division (NCAA Division I FBS only)
df_d1_fbs = df.filter(pl.col("ath_classification_code") == 1)
# Exclude coded missing values before calculations
df_clean = df.filter(
(pl.col("undup_athpartic_men") >= 0) &
(pl.col("undup_athpartic_women") >= 0)
)
# Note: No `sector` column in EADA data. To filter by sector,
# join with IPEDS directory data on unitid first.
| Pitfall | Issue | Solution |
|---|---|---|
| Including coded missing values | -1, -2, -3 treated as real numbers skew totals and ratios | Filter >= 0 on all numeric columns before aggregation |
| Assuming Title IX compliance | EADA data cannot determine Title IX compliance — it is a disclosure tool, not an enforcement mechanism | Read ./references/limitations.md; use EADA for descriptive analysis only |
| Comparing across institutions naively | Different reporting practices, program sizes, and classification levels make raw comparisons misleading | Normalize by enrollment, filter to same classification, and note caveats |
| Using wrong variable names | Portal variable names differ from EADA source documentation (e.g., undup_athpartic_men not partic_men) | Always verify column names against actual data or codebook; see ./references/variable-definitions.md |
| Self-reported data accuracy | Institutions self-report without independent verification; errors and inconsistencies exist | Cross-check outliers against institution websites or IPEDS data |
| Ignoring zero values | Zero may mean "no team" or "not reported" depending on context | Distinguish between true zeros and missing data using -1/-2 codes |
Assuming sector column exists | EADA data has no sector column | Join with IPEDS directory on unitid to get sector |
EADA Data Title IX Compliance
──────────────────────────────────────────────────────────
Self-reported OCR investigation
Snapshot (Oct 15) Continuous obligation
Participation counts only Participation + interest + ability
No "laundry list" items 13+ treatment areas
Public disclosure Enforcement mechanism
Always read: ./references/limitations.md before drawing compliance conclusions.
| Source | Relationship | When to Use |
|---|---|---|
education-data-source-ipeds | Complementary institution data | Joining enrollment, demographics, finances via unitid |
education-data-explorer | Parent discovery skill | Finding available endpoints across all sources |
education-data-query | Data fetching | Downloading parquet/CSV files from mirrors |
| Topic | Reference File |
|---|---|
| Title IX law | ./references/title-ix-context.md |
| Gender equity requirements | ./references/title-ix-context.md |
| Three-prong test | ./references/title-ix-context.md |
| Participation variables | ./references/data-elements.md |
| Coaching variables | ./references/data-elements.md |
| Salary variables | ./references/data-elements.md |
| Expense variables | ./references/data-elements.md |
| Revenue variables | ./references/data-elements.md |
| Athletic aid | ./references/data-elements.md |
| Sport-specific data | ./references/sport-level-data.md |
| Variable definitions | ./references/variable-definitions.md |
| Integer encoding tables | ./references/variable-definitions.md |
| Data limitations | ./references/limitations.md |
| Self-reporting issues | ./references/limitations.md |
| EADA vs Title IX | ./references/limitations.md |
| Fetch patterns | ./references/fetch-patterns.md |
| Mirror URLs | ./references/fetch-patterns.md |
评论 (0)
暂无评论,成为第一个评论者吧!