SkillAtlasSkill 详情

Evaluating Paper Relevance

Security audit: baseline 52/52 CLEAN

审核状态:已审核Quality 72Security 70

复制安装命令

用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。

复制前请先查看来源、License 和安全提示。

项目 README

来源文件:README.md

抓取于 2026年8月4日

Awesome GitHub stars License: CC BY-SA 4.0 PRs Welcome Validate catalog OpenSSF Scorecard Security audit: baseline 52/52 CLEAN Rigor coverage Powered by StatsPAI

Auto-Empirical Research Skills (AERS)

📌 文档结构(2026-07-22 起): 本文件是中文默认入口 —— banner + badges + 信任面 + 9 阶段流水线速览 + 76 行合集总表。 每个合集的完整描述、按用途分组、精确数字、验证方法在 docs/CONTENT_ZH.md(扩展正文,总表行内的 → 直接跳转到对应锚点)。

English version: README-en.md · 中文扩展正文:docs/CONTENT_ZH.md · README-zh-CN.md 已弃用(重定向占位)

🌐 语言: English | 简体中文(默认) | 繁體中文 | 日本語 | 한국어


CoPaper.AI Stanford REAP - Center on China's Economy & Institutions

Stanford REAP × CoPaper.AI · 实证研究 AI 工具的学术工业级产品
由斯坦福实证研究方法论团队打造,覆盖从数据清洗到顶刊投稿的完整工作流



实证研究智能体技能大全封面图

🚀 New here? Open the Skill Search → to filter all 1,096 skills by method, stage, language, and license. The 5-minute tour (make quickstart) prints the same picture in your terminal.

🇨🇳 中文用户从本文件开始(流水线速览 + 76 行总表),每个合集的完整描述见 docs/CONTENT_ZH.md。📖 English readers: see README-en.md.


信任面 · Trust surface (rigor stats)

Rigor laneCountWhere
Numeric benchmark tasks — gold values recomputed from real data each run17benchmark/
Behavioral eval scenarios / rubric items37 / 183eval-harness/

Full trust overview: docs/TRUST.md · docs/RIGOR_COVERAGE.md


中文文档结构

中文内容分两级维护,各司其职:

  • 本文件(README.md,GitHub 默认入口):banner、badges、信任面、9 阶段流水线速览、76 行合集总表。
  • docs/CONTENT_ZH.md(扩展正文):每个合集的完整描述(#skill-NN 锚点)、按用途分组、精确数字、2 分钟验证、三层信任、旗舰流水线详解、贡献与引用。总表行内的 → 直接跳到对应锚点。
  • 其他语言:README-en.md · README-zh-TW.md · README-ja.md · README-ko.md

[!NOTE] 维护规则: 改合集总表 → 本文件与 CONTENT_ZH.md 的锚点表两处同步;改合集详情 / 分组 / 数字 → 只改 docs/CONTENT_ZH.md。统计数字(合集数 / skill 数)以 catalog/skills.json 为准,由 make validate 的 readme-stats 检查器守护。

贡献者(Contributors): 提交前请在本地跑通完整门禁 make check(catalog 校验 + 链接 + 单元测试 + eval-harness + benchmark)。详见 CONTRIBUTING.md。

旧版归档: README-zh-CN.md 已弃用,仅作向后兼容的重定向占位。


🚀 从一个 idea 到一篇论文:社科实证研究 · 端到端流水线(全自动、可介入)

AERS 不只是 76 个散装 skill —— 它能陪你走完一篇论文。 从模糊 idea → 选题精炼 → 文献综述 → 数据获取 → 识别策略 → 估计建模 → 稳健性审计 → 出版级表格 / 图形 → 写作与同行评审 → 降 AIGC → 投稿。端到端、全自动、每一步都可被人介入(中间任何一步你都可以接过去手工改方法、补变量、加稳健性,再让流水线自动接上跑)。

9 阶段流水线 · 每一步都覆盖到具体 skill

#阶段关键 skills(点合集名进目录,→ 进完整说明)
1️⃣选题精炼 — Agent 把模糊想法收紧成"可证伪 + 可执行"的研究问题· 25 Diverga · 33 claude-scholar · 05 research-superpower · 11 compound-science
2️⃣文献综述 — 检索 · 筛选 · PRISMA 流程 · 批判性阅读 · 主题分析· 36 literature-review-skill · 24 academic-research-skills · 59 openalex-skill · 68 research-productivity-skills · 53 thematic-analysis
3️⃣数据获取 — 公开数据库 · API · 网页抓取 · 数据清洗· 33 claude-scholar · 68 research-productivity-skills · 32 stata-skill · 57 edgartools
4️⃣识别策略 — DiD / RD / IV / SCM / DML / matching 全覆盖· ⭐ 00 StatsPAI 🔥 · 10 causal-inference-mixtape · 13 MixtapeTools · 51 CausalPy · 63 scientific-agent-skills
5️⃣估计建模 — Python / Stata / R 三栈,900+ 估计器· ⭐ 00.1 Full Empirical · Python · ⭐ 00.2 Full Empirical · Stata · ⭐ 00.3 Full Empirical · R · 40 pyfixest · 39 marginaleffects · 09 awesome-econ-ai
6️⃣稳健性审计 — 复现包检查 · Honest-DiD · R&R 模拟· 41 sewage-econometrics-check · ⭐ 50 AER-skills · 21 AI-research-feedback
7️⃣表格 & 图形 — 期刊出版级排版 · LaTeX 嵌入· ⭐ 00 StatsPAI · 07 AI-Research-SKILLs · 33 claude-scholar · 08 latex-document-skill
8️⃣写作 & 同行评审 — LaTeX / Quarto · 仿审稿人 · 校对· 06 stats-paper-writing · 04 scientific-writer · 22 christopherkenny-skills · 38 academic-proofreader · 56 econ-writing-skill · 16 clo-author
9️⃣降 AIGC & 投稿 — 知网 / 万方 / Turnitin / 23 类 AI 痕迹模式· ⭐ 48 de-AIGC-skills 🇨🇳🇬🇧 · 44 humanizer_academic · 45 deslop · 46 stop-slop · 47 avoid-ai-writing · 49 humanize-chinese

🎼 元编排:⭐ 69 Paper-WorkFlow —— 一键串起来

Paper-WorkFlow 是 AERS 的"指挥棒",它把上面 9 个阶段的 skill 串成 一条按键即运行的端到端流水线。 你在 IDE 入口给它一句自然语言:

"开一个新论文项目:空气污染与中国劳动力市场,CS 设计 + 省级面板"

它会自动按顺序调:

  1. ⭐ 00 StatsPAI → sp.csdid(...) 给出 CS-DID 估计草案 + 写出估计方程与识别假设
  2. 33 claude-scholar → 抓变量定义 / 数据源候选 / 相关文献
  3. ⭐ 00 StatsPAI → 真跑 sp.feols(...) + sp.honest_did(...)
  4. 41 sewage-econometrics-check → 10 项复现包审计 + 稳健性体检
  5. ⭐ 00 StatsPAI + 07 AI-Research-SKILLs → 出 Table 1–5 + 期刊级图
  6. 38 academic-proofreader → 通读 + §comment 标"审稿人会挑刺的位置"
  7. 56 econ-writing-skill 起草初稿 + ⭐ 48 de-AIGC-skills 🇨🇳🇬🇧 + 45 deslop 过知网 / Turnitin

任何阶段你都可以手动介入 —— 上一阶段的产物全部落盘(产物-幂等 pipeline),你接过去改方法、补控制、加稳健性,再让流水线自动接下去跑。这就是"全自动 + 可介入"。

🏆 7 个 Stanford REAP × CoPaper.AI 自研 skill —— 是整个流水线的主干

⭐ Skill在流水线里的角色
00 StatsPAI 🔥因果引擎:900+ 函数,sp.causal(...) 一行跑闭环(DID / RD / IV / SCM / DML / matching)
00.1 Full Empirical · Python 📘显式 Python 栈(pandas / statsmodels / linearmodels / pyfixest)
00.2 Full Empirical · Stata 📊显式 Stata 栈(reghdfe / ivreg2 / csdid / sdid / rdrobust)
00.3 Full Empirical · R 📗显式 R 栈(tidyverse / fixest / did / HonestDiD)+ Quarto 渲染
48 de-AIGC-skills 🇨🇳🇬🇧中英双语学术降 AIGC(Turnitin AI / GPTZero / 知网 / 万方)
50 AER-skills 📕Top-5 经济学投稿套件:识别 → 稳健性 → R&R
69 Paper-WorkFlow 🧭元编排器,把上面 9 个阶段串成一键流水线

为什么挑这 7 个?因为它们的行为都被基准钉死了 —— 不是营销口径,是对着已知答案反复跑过验证过的(17 项数值 benchmark + 37 项行为评测 ↗)。

看到这里 —— 完整 76 行合集目录

↴ 直跳到下方 76 行总表(每个合集带 #skill-NN 锚点)。如果你更关心"这些 skill 怎么用"而不是"有哪些 skill",看 📘 中文唯一权威正文 里的「按用途分组」与「旗舰流水线」两节。


🧰 76 个核心 Skills 合集一览(00 → 72,编号连续无空缺)

打开仓库 → 看见整座库。 全部 76 个合集 · 1,096 个 skill,每一个都已 vendor 进本仓库,由 catalog/skills.json 跟踪。⭐ = Stanford REAP × CoPaper.AI 团队自研的 skill;其余为精选、经安全审计的社区作品。

主题图例 — 🚀 全流程与编排器 · 🎯 因果推断与计量经济学 · 📚 文献与研究设计 · ✍️ 写作 / 编辑 / 去 AIGC · 📑 引用 / 复现 / 同行评审 · 🛠️ 数据 / 工具 / 基础设施

点击【→】 跳转到 docs/CONTENT_ZH.md 中该合集的完整描述;点击合集名 直接打开其目录。

#合集一句话详情
⭐ 00StatsPAI 🔥因果引擎 · Agent-native Python DSL:sp.causal(...) 一行跑闭环(DID/RD/IV/SCM/DML,900+ 函数)→
⭐ 00.1Full Empirical · Python 📘显式栈:pandas · statsmodels · linearmodels · pyfixest→
⭐ 00.2Full Empirical · Stata 📊reghdfe · ivreg2 · csdid · sdid · rdrobust 复现包→
⭐ 00.3Full Empirical · R 📗tidyverse · fixest · did · HonestDiD + Quarto 渲染→
01academic-paper-skills大纲 → 手稿写作 + 7 维审稿人模拟→
02research-skills医学影像综述、提案、论文转幻灯片→
03scientific-skills假设生成 + 28 个科学数据库→
04scientific-writer引用管理 + 科学写作→
05research-superpower系统化检索、筛选与引文溯源→
06stats-paper-writing端到端 LaTeX 统计论文写作→
07AI-Research-SKILLs发表级 ML 图表、LaTeX、引文核验→
08latex-document-skill创建 / 编译任意 LaTeX 文档为 PDF→
09awesome-econ-aiPython 面板数据分析(linearmodels)→
10causal-inference-mixtapeDID / IV / RDD / SCM 模板(Cunningham)→
11compound-science面向定量社会科学的贝叶斯估计→
12claude-code-my-workflow提交 → PR → 合并的研究工作流(Emory)→
13MixtapeToolsCunningham 的因果推断工具集与讲义→
14research-starterR 中的 IV / DiD / RDD,含完整诊断→
15social-science-researchR 或 Python 端到端数据分析→
16clo-author多代理数据分析(R / Stata / Python)→
17DAAF安全意识代理框架(32 条 deny rule)→
18stata-accounting来自 126 篇 JAR 论文的实测 Stata 范式→
19vera-economic-intelligence经济情报 / 政策研究情报工作流→
20python-econ-skillDSGE / HANK 与定量经济计算→
21AI-research-feedback用 AI 同行评审生成结构化反馈→
22christopherkenny-skills面向 Quarto(.qmd)的 APSA 风格检查器→
23baygent带护栏的 PyMC / Arviz 贝叶斯工作流→
24academic-research-skills5 审稿人多视角论文评审→
25Diverga研究问题精炼器(抗模式坍缩)→
26scholar统计算法设计与文档→
27my_claude_skills经济学摘要写作指南→
28paper-replicate-agent论文复现代理演示→
29project20XXy可复现手稿 + notebook 项目→
30zirui-song-claude-skillsZirui Song 的研究辅助 Claude 技能集→
31claude-code-skillsPython 面板数据分析→
32stata-skill高性能 Stata C/C++ 插件→
33claude-scholar研究全生命周期:选题 → 综述 → 实验 → 审稿回复→
34research-companion头脑风暴、评估并决策研究方向→
35academic-writing-skills面向投稿场所的工业 AI 文献研究→
36literature-review-skill完整文献综述工作流(中文)→
37IlanStrauss-ai-skillsIlan Strauss 经济学研究 AI 工作流→
38academic-proofreader学术校对→
39marginaleffects预测、斜率与比较(R / Python)→
40pyfixestPython 中的快速固定效应估计→
41sewage-econometrics-check10 项复现包审计→
42ARIS自主「research-in-sleep」代理,端到端→
43research-plugins478 个研究插件:数据可视化、领域、基础设施→
44humanizer_academic为医学/学术手稿去 AI 味(23 类模式)→
45deslop去除 AI 写作痕迹(5 维评分)→
46stop-slop三层 AI 痕迹检测与改写→
47avoid-ai-writing审计 → 改写 → 二次审计 AI 味(留痕)→
⭐ 48de-AIGC-skills 🇨🇳🇬🇧中英双语学术降 AIGC(Turnitin AI / GPTZero / 知网 / 万方)→
49humanize-chinese检测并人性化 AI 生成的中文文本→
⭐ 50AER-skills 📕Top-5 经济学投稿套件:识别 → 稳健性 → R&R→
51CausalPy贝叶斯准实验(PyMC Labs)→
52slr-prisma系统文献综述,PRISMA 2020→
53thematic-analysisBraun & Clarke 六阶段定性主题分析→
54open-science-skills引用一致性、DOI 与论据支撑审计→
55r-skillsR 中用 brms 做贝叶斯推断→
56econ-writing-skill综合 50+ 顶级指南的经济学写作→
57edgartools查询与分析 SEC 文件→
58econstack政策简报(UK GES / AU Treasury)→
59openalex-skill通过 OpenAlex 查询 2.4 亿+ 学术作品→
60superpapers综合性实证研究支持套件→
61research-methods与预注册匹配的验证性检验→
62citation-checker对照 CrossRef / S2 / OpenAlex 核验引用→
63scientific-agent-skillsDoWhy 识别–估计–反驳框架→
64mcp-stata20 个 Stata 因果推断与复现 skill→
65game-theory-paper-writer生成并压力测试博弈论论文→
66empirical-research-skills面向大型面板的 R 性能优化→
67econfin-workflow-toolkit中国公司金融实证工作流,从提案到论文→
68research-productivity-skills论文检索、SSRN、DOI 查询、下载→
⭐ 69Paper-WorkFlow 🧭元编排器,串起整个社会科学论文流水线→
70ssci-polish ✍️SSCI / SCI 英文论文语言润色(语法、可读性、学术语气)→
⭐ 71lit-review-agent-tools 🔍文献综述工具选型 + 一键安装运行(MinerU / PaperQA2 / ASReview / STORM / MCP 服务器)→
⭐ 72Kaggle Research 🧪通过官方 CLI 安全检索 Kaggle 资源、限界下载公开数据并保留审计证据→

想看更详细的描述(主题分类、字段、统计)? 见 docs/CONTENT_ZH.md 中标注 #skill-NN 锚点的同一张表 —— 它是每个合集的完整描述所在的扩展正文。


AI 是放大器,不是替代品。它替你做最耗时的"搬砖",你保留最核心的"判断"。


CoPaper.AI Stanford REAP

Stanford REAP × CoPaper.AI · 实证研究 AI 工具的学术工业级产品


扫码访问 copaper.ai
扫码访问 copaper.ai
CoPaper.AI 公众号
关注公众号「CoPaper.AI」

内置 20 个方法论 skill · 20 分钟完成实证论文 · 自研 StatsPAI(900+ 函数 / MIT 开源)

数据与 AI

中风险

  • 来源需自行核对维护者身份。
  • 包含脚本或命令调用,安装前请复核。
  • 可能需要外部 token、网络权限或第三方服务。
  • 未检测到高风险命令。
  • 扫描发现:3 条。

Codex — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/brycewang-stanford/Auto-Empirical-Research-Skills.git
  3. 将 "skills/05-kthorn-research-superpower/research/evaluating-paper-relevance" 文件夹复制到 Codex 的 skills 目录中。
  4. 重启 Codex 让新的 skill 生效。

Codex — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Codex 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Codex 让新的 skill 生效。

Claude Code — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/brycewang-stanford/Auto-Empirical-Research-Skills.git
  3. 将 "skills/05-kthorn-research-superpower/research/evaluating-paper-relevance" 文件夹复制到 Claude Code 的 skills 目录中。
  4. 重启 Claude Code 让新的 skill 生效。

Claude Code — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Claude Code 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Claude Code 让新的 skill 生效。

Cursor — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/brycewang-stanford/Auto-Empirical-Research-Skills.git
  3. 将 "skills/05-kthorn-research-superpower/research/evaluating-paper-relevance" 文件夹复制到 Cursor 的 skills 目录中。
  4. 重启 Cursor 让新的 skill 生效。

Cursor — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Cursor 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Cursor 让新的 skill 生效。

GitHub Copilot — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/brycewang-stanford/Auto-Empirical-Research-Skills.git
  3. 将 "skills/05-kthorn-research-superpower/research/evaluating-paper-relevance" 文件夹复制到 GitHub Copilot 的 skills 目录中。
  4. 重启 GitHub Copilot 让新的 skill 生效。

GitHub Copilot — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 GitHub Copilot 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 GitHub Copilot 让新的 skill 生效。

Windsurf — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/brycewang-stanford/Auto-Empirical-Research-Skills.git
  3. 将 "skills/05-kthorn-research-superpower/research/evaluating-paper-relevance" 文件夹复制到 Windsurf 的 skills 目录中。
  4. 重启 Windsurf 让新的 skill 生效。

Windsurf — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Windsurf 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Windsurf 让新的 skill 生效。
查看 SKILL.md 原文
name: Evaluating Paper Relevance
description: Two-stage paper screening - abstract scoring then deep dive for specific data extraction
when_to_use: After literature search returns results. When need to determine if paper contains specific data. When screening papers for relevance. When extracting methods, results, data from papers.
version: 1.0.0

Evaluating Paper Relevance

Overview

Two-stage screening process: quick abstract scoring followed by deep dive into promising papers.

Core principle: Precision over breadth. Find papers that actually contain the specific data/methods user needs, not just topically related papers.

When to Use

Use this skill when:

  • Have list of papers from search
  • Need to determine which papers have relevant data
  • User asks for specific information (measurements, protocols, datasets, etc.)
  • Screening papers one-by-one
  • Any research domain (medicinal chemistry, genomics, ecology, computational methods, etc.)

Choosing Your Approach

Small searches (<50 papers):

  • Manual screening with progress reporting
  • Use papers-reviewed.json + SUMMARY.md only
  • No helper scripts needed
  • Report progress to user for every paper

Large searches (50-150 papers):

  • Consider helper scripts (screen_papers.py + deep_dive_papers.py)
  • Use Progressive Enhancement Pattern (see Helper Scripts section)
  • Create README.md with methodology
  • May want TOP_PRIORITY_PAPERS.md for quick reference
  • Use richer JSON structure (evaluated-papers.json categorized by relevance)
  • Consider using subagent-driven-review skill for parallel screening

Very large searches (>150 papers):

  • Definitely use helper scripts with Progressive Enhancement Pattern
  • Create full auxiliary documentation suite (README.md, TOP_PRIORITY_PAPERS.md)
  • Consider citation network analysis
  • Plan for multi-week timeline
  • Strongly consider subagent-driven-review skill for parallelization
  • May need multiple consolidation checkpoints

Two-Stage Process

Stage 1: Abstract Screening (Fast)

Goal: Quickly identify promising papers

Score 0-10 based on:

  • Keywords match (0-3 points): Does abstract mention key terms relevant to the query?
  • Data type match (0-4 points): Does it mention the specific information user needs?
    • Examples: measurements (IC50, expression levels, population sizes), protocols, datasets, structures, sequences, code
  • Specificity (0-3 points): Is it specific to user's question or just general background/review?

Decision rules:

  • Score < 5: Skip (not relevant)
  • Score 5-6: Note in summary as "possibly relevant" but skip for now
  • Score ≥ 7: Proceed to Stage 2 (deep dive)

IMPORTANT: Report to user for EVERY paper:

📄 [N/Total] Screening: "Paper Title"
   Abstract score: 8 → Fetching full text...

or

📄 [N/Total] Screening: "Paper Title"
   Abstract score: 4 → Skipping (insufficient relevance)

Never screen silently - user needs to see progress happening

Stage 2: Deep Dive (Thorough)

Goal: Extract specific data/methods from promising papers

1. Check ChEMBL (for medicinal chemistry papers)

If paper describes medicinal chemistry / SAR data:

Use skills/research/checking-chembl to check if paper is in ChEMBL database:

curl -s "https://www.ebi.ac.uk/chembl/api/data/document.json?doi=$doi"

If found in ChEMBL:

  • Note ChEMBL ID and activity count in SUMMARY.md
  • Report to user: "✓ ChEMBL: CHEMBL3870308 (45 data points)"
  • Structured SAR data available without PDF parsing

Continue to full text fetch for context, methods, discussion.

2. Fetch Full Text

Try in order:

A. PubMed Central (free full text):

# Check if available in PMC
curl "https://eutils.ncbi.nlm.nih.gov/entrez/eutils/esearch.fcgi?db=pmc&term=PMID[PMID]&retmode=json"

# If found, fetch full text XML via API
curl "https://eutils.ncbi.nlm.nih.gov/entrez/eutils/efetch.fcgi?db=pmc&id=PMCID&rettype=full&retmode=xml"

# Or fetch HTML directly (note: use pmc.ncbi.nlm.nih.gov, not www.ncbi.nlm.nih.gov/pmc)
curl "https://pmc.ncbi.nlm.nih.gov/articles/PMCID/"

B. DOI resolution:

# Try publisher link
curl -L "https://doi.org/10.1234/example.2023"
# May hit paywall - check response

C. Unpaywall (MANDATORY if paywalled): CRITICAL: If step B hits a paywall, you MUST immediately try Unpaywall before giving up.

Use skills/research/finding-open-access-papers to find free OA version:

curl "https://api.unpaywall.org/v2/DOI?email=USER_EMAIL"
# Often finds versions in repositories, preprint servers, author copies
# IMPORTANT: Ask user for their email if not already provided - do NOT use claude@anthropic.com

Report to user:

⚠️  Paper behind paywall, checking Unpaywall...
✓ Found open access version at [repository/preprint server]

or

⚠️  Paper behind paywall, checking Unpaywall...
✗ No open access version available - continuing with abstract only

D. Preprints (direct):

  • Check bioRxiv: https://www.biorxiv.org/content/10.1101/{doi}
  • Check arXiv (for computational papers)

If full text unavailable AFTER trying Unpaywall:

  • Note in SUMMARY.md: "⚠️ Full text behind paywall - no OA version found via Unpaywall"
  • Continue with abstract-only evaluation (limited)

CRITICAL: Do NOT skip Unpaywall check. Many paywalled papers have free versions in repositories.

2. Scan for Relevant Content

Focus on sections:

  • Methods: Experimental procedures, protocols
  • Results: Data tables, figures, measurements
  • Tables/Figures: Often contain the specific data user needs
  • Supplementary Information: Additional data, extended methods

What to look for (adapt to research domain):

  • Specific data user requested
    • Medicinal chemistry: IC50 values, compound structures, SAR data
    • Genomics: Gene expression levels, sequences, variant data
    • Ecology: Population measurements, species counts, environmental parameters
    • Computational: Algorithms, code availability, performance benchmarks
    • Clinical: Patient outcomes, treatment protocols, sample sizes
  • Methods/protocols described in detail
  • Statistical analysis and significance
  • Data availability statements
  • Code/data repositories mentioned

Use grep/text search (adapt search terms):

# Examples for different domains
grep -i "IC50\|Ki\|MIC" paper.xml                    # Medicinal chemistry
grep -i "expression\|FPKM\|RNA-seq" paper.xml        # Genomics
grep -i "abundance\|population\|sampling" paper.xml  # Ecology
grep -i "algorithm\|github\|code" paper.xml          # Computational

3. Extract Findings

Create structured extraction (adapt to research domain):

Example 1: Medicinal chemistry

{
  "doi": "10.1234/medchem.2023",
  "title": "Novel kinase inhibitors...",
  "relevance_score": 9,
  "findings": {
    "data_found": [
      "IC50 values for compounds 1-12 (Table 2)",
      "Selectivity data (Figure 3)",
      "Synthesis route (Scheme 1)"
    ],
    "key_results": [
      "Compound 7: IC50 = 12 nM",
      "10-step synthesis, 34% yield"
    ]
  }
}

Example 2: Genomics

{
  "doi": "10.1234/genomics.2023",
  "title": "Gene expression in disease...",
  "relevance_score": 8,
  "findings": {
    "data_found": [
      "RNA-seq data for 50 samples (GEO: GSE12345)",
      "Differential expression results (Table 1)",
      "Gene set enrichment analysis (Figure 4)"
    ],
    "key_results": [
      "123 genes upregulated (FDR < 0.05)",
      "Pathway enrichment: immune response"
    ]
  }
}

Example 3: Computational methods

{
  "doi": "10.1234/compbio.2023",
  "title": "Novel alignment algorithm...",
  "relevance_score": 9,
  "findings": {
    "data_found": [
      "Algorithm pseudocode (Methods)",
      "Code repository (github.com/user/tool)",
      "Benchmark results (Table 2)"
    ],
    "key_results": [
      "10x faster than BLAST",
      "98% accuracy on test dataset"
    ]
  }
}

4. Download Materials

PDFs:

# If PDF available
curl -L -o "papers/$(echo $doi | tr '/' '_').pdf" "https://doi.org/$doi"

Supplementary data:

# Download SI files if URLs found
curl -o "papers/${doi}_supp.zip" "https://publisher.com/supp/file.zip"

5. Update Tracking Files

CRITICAL: Use ONLY papers-reviewed.json and SUMMARY.md. Do NOT create custom tracking files.

CRITICAL: Add EVERY paper to papers-reviewed.json, regardless of score. This prevents re-reviewing papers and tracks complete search history.

Add to papers-reviewed.json:

For relevant papers (score ≥7):

{
  "10.1234/example.2023": {
    "pmid": "12345678",
    "status": "relevant",
    "score": 9,
    "source": "pubmed_search",
    "timestamp": "2025-10-11T10:30:00Z",
    "found_data": ["IC50 values", "synthesis methods"],
    "has_full_text": true,
    "chembl_id": "CHEMBL1234567"
  }
}

For not-relevant papers (score <7):

{
  "10.1234/another.2023": {
    "pmid": "12345679",
    "status": "not_relevant",
    "score": 4,
    "source": "pubmed_search",
    "timestamp": "2025-10-11T10:31:00Z",
    "reason": "no activity data, review paper"
  }
}

Always add papers even if skipped - this prevents re-processing and documents what was already checked.

Add to SUMMARY.md (examples for different domains):

Medicinal chemistry example:

### [Novel kinase inhibitors with improved selectivity](https://doi.org/10.1234/medchem.2023) (Score: 9)

**DOI:** [10.1234/medchem.2023](https://doi.org/10.1234/medchem.2023)
**PMID:** [12345678](https://pubmed.ncbi.nlm.nih.gov/12345678/)
**ChEMBL:** [CHEMBL1234567](https://www.ebi.ac.uk/chembl/document_report_card/CHEMBL1234567/)

**Key Findings:**
- IC50 values for 12 inhibitors (Table 2)
- Compound 7: IC50 = 12 nM, >80-fold selectivity
- Synthesis route (Scheme 1, page 4)

**Files:** PDF, supplementary data

Genomics example:

### [Transcriptomic analysis of disease progression](https://doi.org/10.1234/genomics.2023) (Score: 8)

**DOI:** [10.1234/genomics.2023](https://doi.org/10.1234/genomics.2023)
**PMID:** [23456789](https://pubmed.ncbi.nlm.nih.gov/23456789/)
**Data:** [GEO: GSE12345](https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE12345)

**Key Findings:**
- RNA-seq data: 50 samples, 3 conditions
- 123 differentially expressed genes (FDR < 0.05)
- Immune pathway enrichment (Figure 3)

**Files:** PDF, supplementary tables with gene lists

Computational methods example:

### [Fast sequence alignment with novel algorithm](https://doi.org/10.1234/compbio.2023) (Score: 9)

**DOI:** [10.1234/compbio.2023](https://doi.org/10.1234/compbio.2023)
**Code:** [github.com/user/tool](https://github.com/user/tool)

**Key Findings:**
- New alignment algorithm (pseudocode in Methods)
- 10x faster than BLAST, 98% accuracy
- Benchmark datasets available

**Files:** PDF, code repository linked

IMPORTANT: Always make DOIs and PMIDs clickable links:

  • DOI format: [10.1234/example.2023](https://doi.org/10.1234/example.2023)
  • PMID format: [12345678](https://pubmed.ncbi.nlm.nih.gov/12345678/)
  • Makes papers easy to access directly from SUMMARY.md

Progress Reporting

CRITICAL: Report to user as you work - never work silently!

For every paper, report:

  1. Start screening: 📄 [N/Total] Screening: "Title..."
  2. Abstract score: Abstract score: X/10
  3. Decision: What you're doing next (fetching full text / skipping / etc)

For relevant papers, report findings immediately (adapt to domain):

Medicinal chemistry example:

📄 [15/127] Screening: "Selective BTK inhibitors..."
   Abstract score: 8 → Fetching full text...
   ✓ Found IC50 data for 8 compounds (Table 2)
   ✓ Selectivity data vs 50 kinases (Figure 3)
   → Added to SUMMARY.md

Genomics example:

📄 [23/89] Screening: "Gene expression in liver disease..."
   Abstract score: 9 → Fetching full text...
   ✓ RNA-seq data available (GEO: GSE12345)
   ✓ 123 DEGs identified (Table 1, FDR < 0.05)
   → Added to SUMMARY.md

Computational methods example:

📄 [7/45] Screening: "Novel phylogenetic algorithm..."
   Abstract score: 8 → Fetching full text...
   ✓ Code available (github.com/user/tool)
   ✓ Benchmark results (10x faster, Table 2)
   → Added to SUMMARY.md

Update user every 5-10 papers with summary:

📊 Progress: Reviewed 30/127 papers
   - Highly relevant: 3
   - Relevant: 5
   - Currently screening paper 31...

Why this matters: User needs to see work happening and provide feedback/corrections early

Integration with Other Skills

For medicinal chemistry papers:

  • Use skills/research/checking-chembl to find curated SAR data
  • Check BEFORE attempting to parse activity tables from PDFs
  • ~30-40% of medicinal chemistry papers have ChEMBL data

During full text fetching:

  • If paywalled: MANDATORY to use skills/research/finding-open-access-papers (Unpaywall)
  • Do NOT skip this step - Unpaywall finds ~50% of paywalled papers for free

After finding relevant paper:

  1. Check ChEMBL (if medicinal chemistry)
  2. Extract findings to SUMMARY.md
  3. Download files to papers/ folder
  4. Call traversing-citations skill to find related papers
  5. Update papers-reviewed.json to avoid re-processing

Scoring Rubric

ScoreMeaningAction
0-4Not relevantSkip, brief note in summary
5-6Possibly relevantNote for later, skip deep dive for now
7-8RelevantDeep dive, extract data, add to summary
9-10Highly relevantDeep dive, extract data, follow citations, highlight in summary

Helper Scripts (Optional)

When screening many papers (>20), consider creating a helper script:

Benefits:

  • Batch processing with rate limiting
  • Consistent scoring logic
  • Save intermediate results
  • Resume after interruption

Create in research session folder:

# research-sessions/YYYY-MM-DD-query/screen_papers.py

Key components:

  1. Fetch abstracts - PubMed efetch with error handling
  2. Score abstracts - Implement scoring rubric (0-10)
  3. Rate limiting - 500ms delay between API calls (or longer if running parallel subagents)
  4. Save results - JSON with scored papers categorized by relevance
  5. Progress reporting - Print status as it runs

Progressive Enhancement Pattern (Recommended for 50+ papers)

For large-scale screening, use two-script pattern:

Script 1: Abstract Screening (screen_papers.py)

  • Batch fetch abstracts
  • Score using rubric (0-10)
  • Categorize by relevance
  • Output: evaluated-papers.json with basic metadata

Script 2: Deep Dive (deep_dive_papers.py)

  • Read Script 1 output
  • Fetch full text for highly relevant papers (score ≥8)
  • Extract domain-specific data (measurements, protocols, datasets, etc.)
  • Update same JSON file with enhanced metadata

Benefits:

  • Can run steps independently - Score abstracts once, re-run deep dive multiple times
  • Resume if interrupted - No need to re-fetch abstracts if deep dive fails
  • Re-run deep dive without re-scoring abstracts - Adjust extraction logic, keep scores
  • Consistent and reproducible - Same scoring logic applied to all papers
  • Save API calls - Abstract screening happens once, deep dive only on relevant papers

Script design:

  • Parameterize keywords and data types for specific query
  • Progressive enhancement - add detail to same JSON file
  • Include rate limiting (500ms between API calls for single script, longer if parallel)
  • Keep scripts with research session for reproducibility

When NOT to create helper script:

  • Few papers (<20)
  • One-off quick searches
  • Manual screening is faster

Common Mistakes

Not tracking all papers: Only adding relevant papers to papers-reviewed.json → Add EVERY paper regardless of score to prevent re-review Skipping Unpaywall: Hitting paywall and giving up → ALWAYS check Unpaywall first, many papers have free versions Creating unnecessary files for small searches: For <50 papers, use ONLY papers-reviewed.json and SUMMARY.md. For large searches (>100 papers), structured evaluated-papers.json and auxiliary files (README.md, TOP_PRIORITY_PAPERS.md) add significant value and should be used. Too strict: Skipping papers that mention data indirectly → Re-read abstract carefully Too lenient: Deep diving into tangentially related papers → Focus on specific data user needs Missing supplementary data: Many papers hide key data in SI → Always check for supplementary files Silent screening: User can't see progress → Report EVERY paper as you screen it No periodic summaries: User loses big picture → Update every 5-10 papers Non-clickable DOIs/PMIDs: Plain text identifiers → Always use markdown links Re-reviewing papers: Wastes time → Always check papers-reviewed.json first Not using helper scripts: Manually screening 100+ papers → Consider batch script

Quick Reference

TaskAction
Check if reviewedLook up DOI in papers-reviewed.json
Score abstractKeywords (0-3) + Data type (0-4) + Specificity (0-3)
Get full textTry PMC → DOI → Unpaywall → Preprints
Find dataGrep for terms, focus on Methods/Results/Tables
Download PDFcurl -L -o papers/FILE.pdf URL
Update trackingAdd to papers-reviewed.json + SUMMARY.md

Next Steps

After evaluating paper:

  • If score ≥ 7: Call skills/research/traversing-citations
  • Continue to next paper in search results
  • Check if reached 50 papers or 5 minutes → ask user to continue or stop

Auxiliary Files (for large searches >100 papers)

README.md Template

Use this structure for research projects with 100+ papers:

  1. Project Overview

    • Query description
    • Target molecules/topics
    • Date completed
  2. Quick Start Guide

    • Where to start reading
    • Priority lists
  3. File Inventory

    • Description of each file
    • What each is used for
  4. Key Findings Summary

    • Statistics
    • Top findings
    • Coverage by category
  5. Methodology

    • Scoring rubric
    • Decision rules
    • Data sources
  6. Next Steps

    • Recommended actions
    • Priority order

TOP_PRIORITY_PAPERS.md Template

For datasets with >50 relevant papers, create curated priority list:

  • Organized by tier (Tier 1: Must-read, Tier 2: High-value, etc.)
  • Include score, DOI, key findings summary
  • Note full text availability
  • Suggest reading order

Example structure:

# Top Priority Papers

## Tier 1: Must-Read (Score 10)

### [Paper Title](https://doi.org/10.xxxx/yyyy) (Score: 10)

**DOI:** [10.xxxx/yyyy](https://doi.org/10.xxxx/yyyy)
**PMID:** [12345678](https://pubmed.ncbi.nlm.nih.gov/12345678/)
**Full text:** ✓ PMC12345678

**Key Findings:**
- Finding 1
- Finding 2

---

## Tier 2: High-Value (Score 8-9)

[Additional papers organized by priority...]

发现问题?提交给管理员复核

评分:

评论 (0)

暂无评论,成为第一个评论者吧!