SkillAtlasSkill 详情

nature-downloader

nature-skills 面向全球 AI 学者收录可复用科研技能,强调真实问题解决、可验证工作流与可直接使用的科研产物。

审核状态:已审核Quality 72Security 52

复制安装命令

用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。

复制前请先查看来源、License 和安全提示。

项目 README

来源文件:README.md

抓取于 2026年7月31日

nature-skills 面向全球 AI 学者收录可复用科研技能,强调真实问题解决、可验证工作流与可直接使用的科研产物。

目录

1. 项目发起人与运营信息

1.1 创始人介绍

大家好,我是 nature-skills 的创立者袁一哲。感谢大家持续关注本项目。我们在抖音更新了许多视频教程,大家可以根据名称检索查看,希望能够真正帮助到科研工作。

1.2 知识星球

知识星球名称:Nature Skills 以及背后的哲学。

Nature Skills 知识星球

1.3 自营 GPT / Claude 代充与成品号

严格筛选渠道商,提供完全正规的充值渠道与服务。欢迎访问 Nature AI 充值卡网(已上线plus一年代充,Pro5x,20x等等):

Nature AI GPT 与 Claude 代充及成品号服务
Nature AI 充值卡网
https://apiciyuan.top/
ea3af1aadda16b0ddc18565450715d7c
  扫码添加微信客服

1.4 商务合作

如有商务合作意向,欢迎发送邮件至 natureskills2026@outlook.com。

2. Skills 主要开发者

开发者项目角色主要方向与贡献主页与联系
袁一哲创始人 / 维护者项目发起、技能体系设计与社区运营—
马昕瑞核心开发者Skills 日常维护Gmail
胡彬主要贡献者Agentic AgentEmail

3. 项目理念与社区

3.1 自己的一些浅薄观点

  • 最近发现,我设计的 Nature Skills 被谷歌 DeepMind 关注并借鉴,他们参考了其中的引用体系、脚本思路以及技能设计哲学,推出了 Science Skills。说实话,这让我挺欣慰的——当国外的顶尖 AI 机构开始从我们的工作中汲取灵感时,说明中国开发者的原创思想正在被世界看见。这不是被复制的失落,而是中国力量在开源土壤里生根后,自然向外生长出的影响力。
  • 我们设计 Skills 的重心,从来不是要求每个人都来啃透这套思想,而是这套思想本身就具备被机器理解并复用的能力。你如果想创立一个全新的 Skill,或者把它适配到自己的专属领域,根本不需要从头学起——直接把 Nature Skills 的 GitHub 地址发给 Codex,它就能自动学习其中的设计精髓,帮你完成新 Skill 的创建和修改。这才是思想的真正解放:它不再依赖口口相传,而是通过 AI 直接流淌进每一个需要它的角落。
  • Nature Skills 真正的价值,或许并不在于那些具体的技能模块,而在于它悄悄推开了一扇新的大门——它让很多人第一次意识到,原来可以借助 Codex 或智能体来操控本地电脑做科研。我有幸见证并陪伴了许多人完成科研范式的转变。当他们惊叹“原来科研还可以这样去做”的那一刻,这种认知上的破壁和解放,远比 Skills 本身更让我觉得有意义。它不是一个工具的成功,而是一种新的思考方式开始在人群中蔓延。
  • 在当下,几乎所有实用的工具都可以被提炼为标准化流程,而标准化流程恰好可以封装成可复用技能。

3.2 视频教程与社区

视频教程请关注抖音
抖音视频教程
Agent 科研交流群
Agent 科研交流群
袁博个人微信
袁博个人微信

4. 快速开始

安装完成后,可以直接把论文、段落、审稿意见或任务描述交给 Agent。下面这些提示词可以直接复制使用:

想做什么直接这样说
读论文 / 中英文对照把这篇 PDF 做成图文对应的中英文对照 Markdown reader。
生成文献汇报 PPT把这篇论文做成中文组会汇报 PPT,保留关键图件和来源标注。
润色或翻译论文段落把这段中文改写成 Nature 风格英文,保持学术含义不变。
写摘要、引言或讨论根据这些结果和图件,帮我起草 Nature 风格的摘要和引言。
预投稿审稿模拟从 Nature 审稿人视角评估这篇稿件,给出三份 reviewer reports。
回复审稿意见根据这封返修邮件,帮我写逐点回复、cover letter,并标出修改稿需要标红的位置。
查文献、他引和引用者画像整理这篇文章的引用数、严格他引数、DOI,并看引用者里有没有院士、Fellow 或领域大牛。
做科研图或论文示意图根据这段方法和结果,帮我生成投稿级科研图或论文示意图草稿。

如果你不确定该用哪个技能,直接描述任务即可;如果已经知道技能名,可以在提示词中明确写“使用 nature-reader”或“使用 nature-response”。

5. 安装

nature-skills 是一组围绕 SKILL.md 组织的可复用技能包。skills/ 下的每个顶层技能目录都是一个可安装单元,例如 nature-*;nature-shared 是供其他技能读取的共享支持包,默认不作为独立触发技能计入技能索引。

5.1 npx skills 安装方式

需要先安装 Node.js 18 或更高版本。无需全局安装 CLI;先查看仓库中可安装的技能名:

npx skills add Yuan1z0825/nature-skills --list

把全部技能全局安装到 Codex。nature-shared 会随全量安装一起加入,因此依赖共享参考资料的技能也能正常工作:

npx skills add Yuan1z0825/nature-skills --global --agent codex --skill '*' --yes --copy

只为当前项目安装一个独立技能时,省略 --global。例如:

npx skills add Yuan1z0825/nature-skills --agent codex --skill nature-figure --yes --copy

单独安装 nature-reader、nature-paper2ppt、nature-polishing 或 nature-writing 时,同时选择共享支持包:

npx skills add Yuan1z0825/nature-skills --global --agent codex \
  --skill nature-reader --skill nature-shared --yes --copy

也可以把全部技能安装到 CLI 支持的所有 agent:

npx skills add Yuan1z0825/nature-skills --all

检查全局安装结果并更新:

npx skills list --global --agent codex --json
npx skills update --global --yes

只更新一个技能,或只更新当前项目中的技能:

npx skills update nature-reader --global --yes
npx skills update --project --yes

技能选择参数使用 --list 显示的 frontmatter 名称;例如目录 nature-proposal-writer 当前显示为 researchwrite。npx skills 管理的是技能文件,Python、R、浏览器或 MCP 等可选运行依赖仍需按下文说明单独配置。

5.2 Claude Code 安装方式

Claude Code 不能直接使用 scripts/update-codex-skills.sh,因为这个脚本只负责同步到 Codex 的 ~/.codex/skills/。用于 Claude Code 时,推荐保留一个稳定的本地 clone,再用 subagent 或 slash command 指向真实的 skills/*/SKILL.md。这样不会破坏技能目录结构,也能继续读取 references/、static/、manifest.yaml、脚本、资产和 skills/nature-shared/。

如果还没有安装 Claude Code:

npm install -g @anthropic-ai/claude-code
claude

先把仓库 clone 到一个稳定路径:

mkdir -p ~/ai-skills
cd ~/ai-skills
git clone https://github.com/Yuan1z0825/nature-skills.git

推荐方式:为常用技能创建 Claude Code subagent wrapper。以 nature-reader 为例:

mkdir -p ~/.claude/agents
cat > ~/.claude/agents/nature-reader.md <<'EOF'
---
name: nature-reader
description: Use for Chinese-English paper reading, figure-aware translation, and source-grounded paper notes.
---

When invoked, first read `~/ai-skills/nature-skills/skills/nature-reader/SKILL.md` and follow it as the governing workflow.
Read supporting files from `~/ai-skills/nature-skills/skills/nature-reader/` and `~/ai-skills/nature-skills/skills/nature-shared/` only when needed.
Do not replace this skill with a generic paper-reading response.
EOF

然后开启新的 Claude Code 会话,直接请求使用这个 subagent:

Use the nature-reader subagent to turn this paper into a Chinese-English Markdown reader.

如果你更喜欢 slash command,也可以创建命令 wrapper:

mkdir -p ~/.claude/commands
cat > ~/.claude/commands/nature-reader.md <<'EOF'
Read `~/ai-skills/nature-skills/skills/nature-reader/SKILL.md` first and follow it strictly.
Read directly needed supporting files from `~/ai-skills/nature-skills/skills/nature-reader/` and `~/ai-skills/nature-skills/skills/nature-shared/`.

$ARGUMENTS
EOF

在 Claude Code 中使用:

/nature-reader 把这篇论文做成中英文对照的完整 Markdown reader。

安装其他技能时,把示例中的 nature-reader 换成对应目录名即可,例如 nature-polishing、nature-writing、nature-reviewer、nature-response 或 nature-figure。后续更新只需要:

cd ~/ai-skills/nature-skills
git pull

只要 wrapper 仍然指向这个稳定 clone 路径,就不需要重复复制技能文件。

自动更新(可选)

如果你希望 Claude Code 每次开启会话时自动拉取上游更新,可以用 scripts/autoupdate-skills.sh 配合一个 SessionStart 钩子。

这套方式把技能直接复制进 ~/.claude/skills/(Claude Code 会自动发现该目录,技能以目录名直接加载),而不是使用上面的 wrapper。两种方式二选一即可。

先保留一个专用的稳定 clone(只用于同步技能,请不要在里面做开发提交):

mkdir -p ~/ai-skills
git clone https://github.com/Yuan1z0825/nature-skills.git ~/ai-skills/nature-skills

首次安装,把技能复制进 Claude Code 的技能目录:

~/ai-skills/nature-skills/scripts/autoupdate-skills.sh --force

然后在 ~/.claude/settings.json 里加一个 SessionStart 钩子(若已有 hooks,把这一项合并进去,不要整体覆盖):

{
  "hooks": {
    "SessionStart": [
      {
        "hooks": [
          {
            "type": "command",
            "command": "$HOME/ai-skills/nature-skills/scripts/autoupdate-skills.sh",
            "async": true,
            "timeout": 120
          }
        ]
      }
    ]
  }
}

async: true 让它在后台运行、不阻塞启动。脚本自带保护:默认 6 小时内不重复联网检查、断网或拉取失败自动跳过(exit 0,绝不卡住会话)、只有上游 HEAD 真正变化时才重新同步、并且拒绝在有未提交改动的 clone 上强行前进。拉到的新版会在下一次开启会话时生效(当前会话的技能已经加载完毕)。运行日志在 ~/.local/state/nature-skills/autoupdate.log。

目标目录与检查频率都可配置:

# 默认同步到 ~/.claude/skills;用 --dest 指到别处,例如 Codex:
~/ai-skills/nature-skills/scripts/autoupdate-skills.sh --dest ~/.codex/skills
# 只在最多每小时检查一次:
~/ai-skills/nature-skills/scripts/autoupdate-skills.sh --throttle 3600

5.3 Codex 安装方式

推荐使用仓库自带脚本安装或更新 Codex skills。脚本会同步 skills/ 下所有顶层技能目录,并在复制后做 diff 验证;它不会覆盖其他无关 Codex skills。

git clone https://github.com/Yuan1z0825/nature-skills.git
cd nature-skills
scripts/update-codex-skills.sh --pull

如果已经 clone 过仓库:

cd nature-skills
scripts/update-codex-skills.sh --pull

验证当前 Codex 安装是否和这个 checkout 一致:

scripts/update-codex-skills.sh --check

如果你长期用这个脚本更新,并希望清理上游已经删除的旧技能目录:

scripts/update-codex-skills.sh --pull --prune

--prune 只会删除以前由这个脚本记录过、但当前仓库已经不再包含的目录。第一次运行没有历史记录时,它不会猜测删除旧目录。

也可以把仓库链接交给 Codex,让 Codex 执行安装脚本。推荐提示词:

请从这个仓库安装 Codex skills:
https://github.com/Yuan1z0825/nature-skills.git

请 clone 仓库后运行 scripts/update-codex-skills.sh --pull。
安装后再运行 scripts/update-codex-skills.sh --check 验证。
请保留 skills/ 下的完整技能目录,不要只复制 SKILL.md。

如果只安装单个技能,请明确说明技能名:

只安装这个仓库里的 nature-reader:
https://github.com/Yuan1z0825/nature-skills.git

如果该技能需要共享文件,也请一并安装 skills/nature-shared。

关键规则:保留完整目录结构。请复制或引用整个技能文件夹,而不是只复制 SKILL.md,因为许多技能依赖 references/、static/、manifest.yaml、脚本、资产或共享文件。

安装脚本不会自动安装 Python 依赖。需要使用相关脚本或 MCP 服务时,再按需安装:

python -m pip install -r skills/nature-paper-to-patent/requirements.txt
python -m pip install -r skills/nature-paper-to-patent/scripts/disclosure/requirements-cnipa.txt  # 可选:国知局公布公告检索
python -m pip install -r skills/nature-academic-search/mcp-server/requirements.txt

如果启用 nature-paper-to-patent 的国知局公布公告检索,还需要执行 python -m playwright install chromium。

nature-academic-search 的 MCP 服务还需要单独配置 PUBMED_EMAIL,Scopus / ScienceDirect 等可选 provider 需要使用本机凭据配置,不要把 API key 写入仓库文件。

安装后,请开启一个新的 Codex 会话,然后自然描述任务,例如:

把这篇论文做成中英文对照的完整 Markdown reader。
把这篇论文做成中文PPT。

如果你使用 OpenClaw、OpenCode、Hermes 等开源 agent / 编程框架,请看 OpenClaw / OpenCode / Hermes 接入教程。

自动更新(可选)

Codex 支持全局 SessionStart hook。保留一个专用 clone 后,可以在每次启动或恢复 Codex 会话时检查更新,并把新版同步到 ~/.codex/skills/。

先创建专用 clone 并完成首次同步:

mkdir -p ~/.codex
git clone https://github.com/Yuan1z0825/nature-skills.git ~/.codex/.nature-skills-src
~/.codex/.nature-skills-src/scripts/autoupdate-skills.sh \
  --dest ~/.codex/skills --force

然后创建或合并 ~/.codex/hooks.json:

{
  "hooks": {
    "SessionStart": [
      {
        "matcher": "startup|resume",
        "hooks": [
          {
            "type": "command",
            "command": "/bin/bash \"$HOME/.codex/.nature-skills-src/scripts/autoupdate-skills.sh\" --dest \"$HOME/.codex/skills\"",
            "timeout": 75,
            "statusMessage": "Checking Nature Skills updates"
          }
        ]
      }
    ]
  }
}

若 hooks.json 中已有其他 hook,请合并 SessionStart 项,不要整体覆盖。首次启用或修改 hook 后,在 Codex 中运行 /hooks 检查并信任它。Codex 当前按同步方式执行 command hook,因此这里依靠脚本自带的 6 小时节流、60 秒网络保护和断网自动跳过,避免每次会话都重复联网或因更新失败阻断启动。

更新日志位于 ~/.local/state/nature-skills/autoupdate.log。拉取到的新技能通常在下一次会话中完整生效。

5.4 其他 Agent 场景

OpenClaw、OpenCode、Hermes 的具体接入方式见 OpenClaw / OpenCode / Hermes 接入教程。

用于其他 agent 时,建议保留一个稳定的仓库 clone,再创建轻量 subagent、slash command 或 custom prompt wrapper,指向真实的 skills/*/SKILL.md,并保留 skills/nature-shared/。

手动或其他 agent 使用时:

  1. 将完整技能目录复制到你的 prompt library 或项目中。
  2. 保留 SKILL.md、manifest.yaml、static/、references/、脚本、资产和需要的 skills/nature-shared/ 文件。
  3. 如目标 agent 有自己的格式要求,可调整 frontmatter 和正文结构。

6. 技能索引

当前 skills/ 下包含以下可触发技能;skills/nature-shared/ 是共享内容目录,不计入技能索引。点击技能名或“详情页”可以进入每个 skill 的单独说明页面。

技能状态用途触发词详情页
nature-figureStable面向 Nature / 高影响力期刊的 Python 或 R 投稿级科研图工作流,内置 figures4papers demo,并支持通过 OpenRouter GPT Image 2 生成论文示意图草稿“Nature figure”, “投稿级图片”, “publication plot”, “scientific figure”, “figures4papers”, “论文示意图”, “GPT Image 2”详情
nature-polishingStable将学术文本润色、重构或翻译为 Nature 风格英文“Nature style”, “润色”, “academic writing”, “论文英文”详情
nature-writingDraft起草 Nature 风格手稿章节,并重建论文论证“Nature writing”, “写摘要”, “写引言”, “manuscript draft”, “论文写作”详情
nature-reviewerDraft从审稿人视角模拟 Nature 风格评审,输出三份 reviewer reports 和综合意见“Nature reviewer”, “预投稿评审”, “reviewer report”, “审稿人视角评估”详情
nature-citationBeta检索严格限定在 Nature / CNS 系列的支撑文献,并导出 ENW、RIS 或 Zotero RDF“Nature citation”, “CNS citation”, “分段引用”, “支撑文献”, “Zotero RDF”详情
nature-dataDraft准备 Data Availability statement、数据仓储方案和 FAIR 检查“Data Availability”, “数据可用性”, “repository”, “FAIR metadata”详情
nature-statisticsDraft审查、改写或起草 Nature / 高影响力期刊投稿中的统计报告,覆盖样本量、独立分析单位、重复数、p 值、多重比较、效应量、置信区间、图注统计和审稿人统计意见“Nature statistics”, “统计审查”, “statistical analysis”, “p value”, “sample size”, “replicates”, “multiple comparisons”, “图注统计”, “统计分析小节”详情
nature-readerBeta生成带来源锚点、图文对应和中英文对照的全文 Markdown reader“nature reader”, “全文 Markdown”, “原文对照”, “图文对应”, “全文翻译”详情
nature-paper-cardBeta精读单篇论文并生成有来源约束的 01–16 节 Paper Card,覆盖方法逻辑、实验—结论证据链、结论边界、批判性分析和可检验研究想法“nature paper card”, “论文精读”, “Paper Card”, “证据链”, “结论边界”详情
nature-responseBeta解析返修邮件,起草、审查和修改返修 cover letter、逐点回复审稿人的 response letter、标红修改稿,并提供 LaTeX 模板“response to reviewers”, “rebuttal letter”, “cover letter”, “major revision”, “返修邮件”, “审稿意见回复”, “修回信”, “LaTeX 模板”详情
nature-paper2pptBeta从科研论文生成中文 PPTX 文献汇报 deck“paper PPT”, “journal club”, “paper to slides”, “论文汇报”详情
nature-paper-to-patentBeta从论文、技术报告或项目材料生成有证据约束的中国发明专利草稿,并支持专利点挖掘、查新和技术交底书迭代“paper to patent”, “Chinese patent”, “论文转专利”, “权利要求书”, “技术交底书”, “专利点”详情
nature-ref-verifierStable参考文献多源交叉验证:逐字段对比作者/标题/年份/卷期/页码,标记卷年冲突、作者编造、页码偏差等“verify refs”, “校验文献”, “check references”, “文献验证”, “ref check”详情
nature-academic-searchBeta多源文献检索、引用核验、严格他引审计、文章引用指标表、高影响力引用者画像和参考文献管理“search papers”, “find articles”, “literature search”, “查文献”, “verify DOI”, “严格他引”, “文章引用表”, “引用我的文章的人有没有大牛”详情
nature-downloaderBeta通过图书馆资源入口、Chrome 登录态和开放获取路径合法获取学术全文/PDF“download papers”, “图书馆下载文献”, “CARSI”, “Web of Science”, “PDF 下载”详情
nature-literature-pipelineStable自动化文献发现管线:多源检索、六维评分、精读推送和本地归档“literature pipeline”, “每日文献”, “文献推送”, “daily literature push”, “cron”详情
nature-experiment-logDraft标准化记录实验图片、语音和文字材料,生成带 YAML frontmatter 的 Obsidian 实验日志并归档原始材料“实验日志”, “记录实验”, “experiment log”, “Obsidian vault”, “飞书科研群”详情
nature-proposal-writerBetaproposal-first 科研写作状态机,先建立证据、论证和章节契约,再起草或审查文本“researchwrite”, “proposal”, “开题报告”, “研究方案”, “科研写作 QA”详情

7. 贡献与开发

7.1 共享设计原则

所有技能都遵守以下原则:

  1. 优先使用一手来源:规则基于已发表 Nature 内容、官方期刊指南或明确的本地来源,而不是泛泛审美偏好。
  2. 显式胜过隐式:每条规则都应说明理由,而不是只给断言。
  3. 感知章节与任务上下文:学术写作、图件、引用和回复都依赖上下文;不同论文部分使用不同逻辑。
  4. 输出优先:每个技能都应返回能直接使用的产物,例如可粘贴文本、.svg、.pptx、.docx 或具体建议。
  5. 可扩展:每个技能自包含在自己的目录中,新增技能不应要求修改既有技能。

7.2 仓库目录结构

skills/
├── nature-shared/              # 当技能引用 ../nature-shared 时需要保留
├── nature-<topic>/
│   ├── README.md
│   ├── README_EN.md
│   ├── SKILL.md
│   ├── manifest.yaml     # router-style 技能会包含
│   ├── static/           # router-style 技能会包含
│   └── references/...
└── nature-proposal-writer/
    ├── README.md
    ├── README_EN.md
    ├── SKILL.md
    ├── scripts/...
    ├── templates/...
    └── references/...

7.3 新增技能流程

向本仓库添加技能时,请按以下流程:

1. 创建技能目录

skills/nature-<topic>/

2. 添加必需文件

文件是否必需用途
SKILL.md必需frontmatter(name、description)+ 规则 + 工作流;触发后由 agent 加载
README.md必需面向人的中文说明文档
README_EN.md必需与中文详情页配套的英文说明文档
references/*.md复杂技能推荐模块化规则文件,例如 API、设计理论、教程、图表类型等

3. 编写中英文 README

每个新增技能都必须同时提供 README.md 和 README_EN.md。README 是面向人的技能入口页,不是 SKILL.md 的重复版,也不是安装手册。它的目标是让用户在 30 秒内判断:这个 skill 能不能解决我的问题、我要给它什么、它会产出什么、边界在哪里。

基本规则:

  • 中文 README 和英文 README 必须一一镜像:标题数量一致、顺序一致、信息点一致;英文页不要写成另一套独立模板。
  • 顶部固定为技能名、语言切换链接和一句定位说明。
  • 默认使用下面的基础结构;只有确实需要时才插入可选章节。
  • 不要在单个 skill README 中重复仓库安装教程、作者信息、变更日志、开发过程、完整文件树或大段内部实现细节。
  • 复杂规则、API 参数、长教程、脚本说明和模板索引应放进 references/、static/、scripts/ 或 SKILL.md,README 只保留路标。
  • 如果技能有视觉资产,可以放一个小型预览表;不要把 README 变成大型图库或长篇技术手册。

中文 README 基础结构:

# `nature-<topic>` 技能

[English](README_EN.md)

一句话说明这个技能的定位、主要任务和使用边界。

## 适合用它做什么
## 典型请求
## 你需要提供
## 产出
## 边界
## 相关技能

英文 README 必须对应为:

# `nature-<topic>` Skill

[中文说明](README.md)

One sentence describing the skill's role, main task, and usage boundary.

## What To Use It For
## Typical Requests
## What You Need To Provide
## Outputs
## Boundaries
## Related Skills

可选章节必须中英文同步插入,并保持相同顺序。常见可选章节包括:

中文章节英文章节使用场景
## 工作方式## Workflow需要解释核心流程或路由方式
## 运行和依赖## Runtime and Dependencies有脚本、MCP、API key、本地配置或外部依赖
## 示例预览## Example Preview有少量图件、截图或可视化资产值得展示
## 内置参考## Built-In References需要指向 references/、assets/ 或 demo
## 方法来源## Method Sources写作、审查或分析规则来自特定来源
## 三种模式## Three Modes技能有清晰的 compose/revise/hybrid 等模式
## 与 ... 的关系## Relationship With ...容易和另一个技能混淆,需要说明分工

提交前至少做这些 README 检查:

python scripts/validate-readmes.py
git diff --check
for d in skills/nature-*; do
  [ -f "$d/README.md" ] && [ -f "$d/README_EN.md" ] || continue
  rg -q '^\[English\]\(README_EN\.md\)$' "$d/README.md"
  rg -q '^\[中文说明\]\(README\.md\)$' "$d/README_EN.md"
  test "$(rg -c '^## ' "$d/README.md")" = "$(rg -c '^## ' "$d/README_EN.md")"
done

4. 录制使用教程

提交 PR 时,请同时录制一个简短的使用教程,说明这个 skill 解决什么问题、如何触发、需要什么输入,以及会产出什么结果。可以在 PR 描述中附上视频、录屏链接或可公开访问的教程地址。

5. 配置 SKILL.md frontmatter

---
name: nature-<topic>
description: >-
  用一句话说明这个技能做什么、什么时候触发、主要输出格式和核心使用场景。
---

6. 更新技能索引

在上方 技能索引 表格中添加一行:

| [`nature-<topic>`](skills/nature-<topic>/README.md) | Draft / Stable | 一句话用途 | 触发词 | [详情](skills/nature-<topic>/README.md) |

7. 设置状态标签

标签含义
Draft规则已定义,但尚未在真实案例上测试
Beta已在示例上测试,仍可能存在边界问题
Stable已在真实学术内容上验证,规则相对稳定

8. Star 历史

Star History Chart

研究与检索浏览器与自动化

高风险

  • 来源需自行核对维护者身份。
  • 包含脚本或命令调用,安装前请复核。
  • 可能需要外部 token、网络权限或第三方服务。
  • 存在潜在风险命令,请谨慎安装。
  • 扫描发现:5 条。

Codex — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/Yuan1z0825/nature-skills.git
  3. 将 "skills/nature-downloader" 文件夹复制到 Codex 的 skills 目录中。
  4. 重启 Codex 让新的 skill 生效。

Codex — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Codex 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Codex 让新的 skill 生效。

Claude Code — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/Yuan1z0825/nature-skills.git
  3. 将 "skills/nature-downloader" 文件夹复制到 Claude Code 的 skills 目录中。
  4. 重启 Claude Code 让新的 skill 生效。

Claude Code — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Claude Code 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Claude Code 让新的 skill 生效。

Cursor — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/Yuan1z0825/nature-skills.git
  3. 将 "skills/nature-downloader" 文件夹复制到 Cursor 的 skills 目录中。
  4. 重启 Cursor 让新的 skill 生效。

Cursor — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Cursor 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Cursor 让新的 skill 生效。

GitHub Copilot — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/Yuan1z0825/nature-skills.git
  3. 将 "skills/nature-downloader" 文件夹复制到 GitHub Copilot 的 skills 目录中。
  4. 重启 GitHub Copilot 让新的 skill 生效。

GitHub Copilot — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 GitHub Copilot 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 GitHub Copilot 让新的 skill 生效。

Windsurf — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/Yuan1z0825/nature-skills.git
  3. 将 "skills/nature-downloader" 文件夹复制到 Windsurf 的 skills 目录中。
  4. 重启 Windsurf 让新的 skill 生效。

Windsurf — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Windsurf 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Windsurf 让新的 skill 生效。
查看 SKILL.md 原文
name: nature-downloader
description: Use when a user needs lawful academic full text, CNKI institutional access, English OA retrieval, publisher API access, institutional browser fallback, or supporting information downloads.
metadata:
  compatibility: Requires Node.js 22+ and Python 3. CNKI, Web Access, and SI routes additionally require the user's authenticated Chrome session and remote debugging. Uses only lawful OA, publisher API, and user-authorized institutional access.

Nature Literature Downloader

This skill routes literature through lawful open-access, publisher-API, CNKI institutional, and browser-based institutional providers. scripts/batch_download.mjs is the orchestration entry point; school configuration, publisher credentials, metadata/OA resolution, provider downloads, content validation, and manifests are separate modules.

Verified routes are examples, not defaults. Every institution should start from the user's actual library resource URL, because resource portals, CAS callbacks, EZproxy, WebVPN, IP-authenticated database pages, and database detail pages reveal the live authorization path more reliably than a school name.

SI confirmation gate — do this first. Before downloading any PDF, CAJ, HTML, XML, archive, or attachment, ask whether the user wants Supporting Information. An explicit request for SI counts as yes; an explicit request for正文 only counts as no. Otherwise ask once for the whole batch. Run the downloader with exactly one of --si or --no-si. Without either flag the script returns si_confirmation_required and does not create the output directory.

Main workflow. Normalize the DOI/title and identify language and publisher before routing. Chinese literature always uses CNKI. For English Elsevier, Springer Nature, and IEEE articles with usable provider credentials, try the publisher API first and do not require an OA determination after a successful API download. If that API attempt fails, automatically check legitimate OA sources. Other English publishers check OA first, then use the institutional Web Access route when OA is unavailable.

规范化 DOI/题名并识别语言、出版商
├─ 中文文献:直接走 CNKI
└─ 英文文献
   ├─ Elsevier / Springer Nature / IEEE,且已配置有效 Key
   │  ├─ 优先通过出版商 API 下载
   │  ├─ API 下载成功:结束,不强制判断 OA
   │  └─ API 下载失败:检查文章级 OA,再走 PMC / Unpaywall / 合法仓储
   └─ 其他出版商
      ├─ 检查文章级 OA
      └─ OA 不可用:走 Web Access 机构授权

Chinese literature is CNKI-only. A Chinese title, zh metadata language, explicit CNKI source URL, or --route cnki must use CNKI even if another OA copy appears to exist. Reuse the user's current Chrome library/CNKI login state and prefer configured discovery.cnki_url. Never export cookies or collect the institutional password.

Publisher API fallback. A valid API key does not guarantee full-text entitlement. When an Elsevier, Springer Nature, or IEEE API attempt returns no entitlement or no usable full text, automatically try legitimate OA sources first. Return api_fallback_confirmation_required and ask once whether to use Web Access only after both the publisher API and OA routes fail. Do not switch to institutional Web Access automatically.

Browser-state principle. Authorized downloads depend on the exact browser profile where the user is logged in. If a proxy, CDP session, or browser automation tool opens a fresh profile or a different browser with no login state, do not treat the failure as missing library permission. Switch to a control path that reuses the user's active browser session, or ask the user to authenticate in the controlled browser instance.

Format principle. PDF, HTML full text, and database-native formats such as CAJ are different deliverables. If the user asks for PDF only, require a real PDF link or %PDF response and report no_authorized_pdf_found / pdf_fetch_failed when none exists. Do not save CAJ, HTML, or a login page as if it were a PDF.

Download Intake and First-Run Configuration

For every download request, first establish the paper list and ask:

是否同时下载这些文献的 Supporting Information(SI,补充材料)?

Do metadata lookup before this question only when needed to identify the requested papers. Do not download files until the answer is known. Configure a library only when the selected route is CNKI or Web Access. Configure a publisher API only when the selected English article belongs to Elsevier, Springer Nature, or IEEE; an OA determination is not required before trying a configured provider API.

Paid Library Resource Configuration

Ask for the library resource URL the user actually uses:

请发你平时进入图书馆电子资源/数据库的平台链接。
可以是资源门户、数据库列表、Web of Science 入口、某个数据库详情页,
或跳转到统一身份认证的登录链接。

Then infer the authorization route from the URL before saving config:

python3 scripts/configure_school.py infer "https://example.edu/library/resources"
python3 scripts/configure_school.py url "https://example.edu/library/resources"
python3 scripts/configure_school.py show
python3 scripts/configure_school.py health --force

The distributed skill contains no school presets. If the user cannot provide a resource URL, ask them to locate their institution's library/database entry instead of guessing a school-specific domain.

The default config path is:

~/.config/lit-dl/school.json

For tests or isolated profiles, set:

LIT_DL_CONFIG_DIR=/path/to/configdir

The downloader reads this config automatically. If discovery.web_of_science_url is present, scripts/batch_download.mjs uses it as the Web of Science entry; otherwise it falls back to https://www.webofscience.com/wos/woscc/basic-search.

For Chinese literature, the downloader also reads discovery.cnki_url when present. If absent, scripts/batch_download.mjs --title "<中文题名>" falls back to https://kns.cnki.net/kns8s/defaultresult/index.

API-First and Open-Access Fallback

For an English article, identify its publisher before deciding when to resolve article-level OA:

  1. Collect a DOI, PMID, exact title, article URL, or a definite paper list, then normalize its metadata and publisher.

  2. If it belongs to Elsevier, Springer Nature, or IEEE and usable provider credentials are configured, try that publisher API first. On success, record accessMode: publisher_api and oa_status: not_checked_api_first; do not run OA resolution only to label the article.

  3. If the publisher API fails, automatically search legitimate OA sources such as PMC, Unpaywall, publisher OA pages, arXiv, and other lawful repositories or clearly open PDF URLs. Preserve the failed API attempt in the manifest.

  4. For all other English publishers, search those legitimate OA sources before Web Access.

  5. For an exact title or an explicit OA-only request, prefer:

    node scripts/batch_download.mjs --title "<exact title>" --open-access --no-si --out "<project>"
    

    Use --pdf-url when the user supplies a known legitimate OA PDF URL.

  6. Verify the downloaded file and record the source. Mark a successful PDF as open_access_downloaded.

  7. If no lawful OA full text is found, mark oa_not_found. For a supported publisher whose API already failed, request confirmation before Web Access. For another publisher, continue to Web Access. If --route open_access was explicitly requested, stop after the OA result.

Publisher API Credentials

Configure credentials lazily, only when the route first needs them:

python3 scripts/configure_credentials.py set elsevier
python3 scripts/configure_credentials.py set springer_nature
python3 scripts/configure_credentials.py set ieee --fulltext-endpoint 'https://issued-endpoint.example/articles/{doi}'
python3 scripts/configure_credentials.py set elsevier --stdin
python3 scripts/configure_credentials.py show
python3 scripts/configure_credentials.py validate <provider>
python3 scripts/configure_credentials.py delete <provider>
python3 scripts/configure_credentials.py contact-email researcher@example.org

Give the user the official registration link: Elsevier https://dev.elsevier.com/, Springer Nature https://dev.springernature.com/docs/quick-start/api-access/, or IEEE https://developer.ieee.org/member/register.

Do not proactively ask the user to paste an API key into chat. If the user voluntarily sends a publisher API key, treat that as authorization to save that exact key: do not reject it, ask them to regenerate it, or repeat it back. Pass it to configure_credentials.py set <provider> --stdin, keep it out of command-line arguments, logs, replies, and manifests, then report only the masked confirmation and validation status. The local hidden prompt remains the preferred path when the key has not already been provided. IEEE Metadata API access is not paid full-text access; require the issued Full-Text Access endpoint/template before treating IEEE as downloadable through the API. Secrets are stored in ~/.config/lit-dl/credentials.json with mode 0600.

Resource URL Triage

Classify the user-provided URL before choosing an access path:

cas.* / /authserver/login        CAS / SSO login page; inspect service= callback, then return to the service portal
idp/shibboleth / carsi           CARSI / Shibboleth institutional route
ezproxy / libproxy               EZproxy remote-access proxy
webvpn / vpn                     WebVPN route
metaersp / metaauth / uas        Library resource aggregation portal
webofscience / sciencedirect     Database or publisher entry; check whether it was reached through a portal

If the URL is a login page with a service= parameter, treat the callback host as the resource service and do not make the login page the whole workflow. For example, https://login.university.example/authserver/login?service=https://resources.university.example/callback means the identity service returns to the user's resource portal after authentication.

Institution-Specific Domains

Confirm against what actually appears in the user's address bar; correct these for each institution instead of assuming a preset is complete.

Library home / aggregation:  library.example.edu, resources.example.edu
Discovery/database entry:    webofscience.com, clarivate.com, cnki.net, sciencedirect.com, provider.example.com
Unified identity / SSO:      sso.example.edu, cas.example.edu, idp.example.edu
Federation / WAYF:           ds.carsi.edu.cn, wayf.example.org, shibboleth/openathens hosts
Proxy / WebVPN:              ezproxy.example.edu, webvpn.example.edu

Treat configured institutional login, federation, proxy, and database-login hosts as sign-in stages. Do not treat reaching them as a final failure.

Boundaries

Use only the user's legitimate institutional access. Do not bypass paywalls, DRM, or two-factor authentication.

Verification-first rule: When a visible slider, checkbox, robot check, or simple verification control appears in the user's authenticated Chrome session, attempt it in the browser before asking the user to intervene. Keep the attempt bounded (at most two attempts on one tab), verify that the challenge disappeared, and continue from that same tab when successful.

  • Slider/drag challenges (including CNKI puzzle sliders): estimate the visible travel distance and simulate a gradual drag.
  • ScienceDirect robot checks, managed Turnstile, and reCAPTCHA checkbox stages: try the visible checkbox once.
  • Simple Continue, Verify, or equivalent visible controls: click once, then re-check the page state.

User handoff: Ask the user only after the bounded attempt fails, or immediately when the page requires secret or identity-bearing input such as an image-selection answer, QR approval, SMS/OTP, passkey, hardware key, or two-factor authentication. Keep the challenged tab open and never ask the user to paste credentials or codes into chat.

Avoid unbounded or indiscriminate downloading. Process only the definite paper list confirmed by the user, apply provider-friendly pacing, and leave a clear audit trail of what was downloaded, from where, and whether supporting information was found.

Do not ask the user to paste institutional passwords, database passwords, OTP codes, recovery codes, or session tokens into chat or terminal. If the user offers one of those identity-bearing secrets, decline and use the handoff-login workflow instead. Publisher API keys follow the separate save-on-receipt rule above.

Exception for saved institutional login pages: if the user explicitly says that the browser has already filled credentials and authorizes clicking the visible login/confirm button, the agent may click that button once on the expected institutional SSO / CAS / CARSI / Shibboleth page without reading, copying, or typing any credential. This exception does not apply to CAPTCHA, QR login, SMS/OTP, publisher bot checks, consent/security warnings, or any page outside the expected institutional login flow.

Do not inspect or export cookies, passwords, local storage, browser profiles, or session files. Use the browser's already-authenticated page context only.

Preconditions

Before attempting downloads, confirm the conditions that apply to the selected access branch.

For the OA-only branch, confirm the target paper identifier/list, output folder, Node.js 22+, and Python 3 when PDF verification needs it. Do not require a library configuration or institutional browser login.

For the paid-library branch, confirm these conditions:

  1. The browser that holds the user's library/database login state is open on the user's machine.
  2. The school configuration exists and is valid.
    • Run python3 scripts/configure_school.py show.
    • If missing, run python3 scripts/configure_school.py preset "<school name>" or guide the user through src/wizard.py.
  3. The user has personally logged in to their institution/library route in that same browser, and can reach the library aggregation service, target database, or discovery entry.
  4. The browser-control path can reuse that same logged-in browser profile.
    • For Chrome CDP, ask the user to open chrome://inspect/#remote-debugging and enable remote debugging for the current browser instance.
    • If CDP attaches to a stale browser, a temporary profile, or a different browser, use a browser-control channel that can reuse the user's active session instead of launching a new profile.
  5. The environment can run Node.js 22+.
    • Try node --version.
    • If node is not on PATH in Codex Desktop, try %LOCALAPPDATA%\OpenAI\Codex\bin\node.exe.
  6. The environment can run Python 3 for configuration and PDF text verification.
    • Try python3 --version.
    • Install Python helpers with pip install -r requirements.txt when needed.
  7. The web-access CDP proxy is available or can be started.
    • Typical Claude Code path: %USERPROFILE%\.claude\skills\web-access-main\scripts\check-deps.mjs.
    • Typical shared agent path: %USERPROFILE%\.agents\skills\web-access-main\scripts\check-deps.mjs.
    • In Codex-only setups also check %USERPROFILE%\.codex\skills\web-access-main\scripts\check-deps.mjs.
  8. The user has approved the target output folder.

If Claude Code says this skill is not installed, install or copy it to:

$env:USERPROFILE\.claude\skills\nature-downloader

Codex and other agent setups may instead use .codex\skills or .agents\skills; treat the three locations as install targets, not as different skill versions.

Batch Scope

Definite DOI/title/PMID lists are supported without a fixed per-batch paper-count recommendation.

Operational safeguards:

  • pace requests appropriately for each provider and maintain the manifest throughout the batch
  • attempt visible verification controls first; stop after at most two failed attempts, on institutional login expiry, or when an unusual/security-sensitive prompt appears

Do not turn a broad keyword search into unlimited automatic downloading. Do not download whole journal issues, volumes, or large result sets.

Status Categories

Classify every paper into one of these statuses, and keep the status in the manifest:

downloaded
downloaded_with_si
open_access_downloaded
full_text_html_available
available_not_downloaded
native_fulltext_downloaded
si_confirmation_required
credentials_missing
credentials_invalid
api_not_entitled
api_fulltext_unavailable
api_fallback_confirmation_required
oa_not_found
oa_resolution_inconclusive
metadata_ambiguous
carsi_waiting_user
carsi_resolved_retry_needed
publisher_verification_waiting_user
sciencedirect_robot_check
retry_after_user_verification
verification_auto_passed
verification_auto_failed
do_not_auto_retry
url_needs_repair
library_no_permission
no_full_text_link
publisher_blocked_waiting_user
no_authorized_pdf_found
failed_after_retry

Use verification_auto_passed when an automatic CAPTCHA/slider/robot check was successfully solved by the skill, and the download then proceeded normally.

Use verification_auto_failed when auto-verification was attempted but could not pass the challenge. This is a user-handoff status, not a final failure.

Use carsi_waiting_user only when the browser is visibly at an institutional SSO / CAS / CARSI-Shibboleth / OpenAthens / database authentication page. Do not treat this as a final failure.

Use publisher_verification_waiting_user or sciencedirect_robot_check when a publisher page shows a verification challenge but no automatic interaction was possible. When a bounded automatic attempt was made and failed, use verification_auto_failed instead. None of these is a final download failure.

Use open_access_downloaded when a legitimate open-access route such as PMC, the publisher's OA PDF, arXiv, or another lawful open PDF source provides the downloaded PDF without institutional authorization.

For a successful API-first download, record oa_status: not_checked_api_first; this means OA resolution was intentionally skipped, not that the article is non-OA. Use api_fallback_confirmation_required only after a supported publisher API attempt and its automatic OA fallback both fail.

Use full_text_html_available when the library/full-text resolver grants access to a readable HTML full text but no valid PDF link or %PDF response is available. This is a successful full-text access result, not a PDF download. Save the HTML/text if the user asked for the article, and explicitly tell the user that the PDF was not available through the current authorized route.

Use library_no_permission when the library portal, SFX/OpenURL resolver, database, or publisher page clearly says the user's institution has no full-text entitlement for the paper. Tell the user plainly that the current library resources do not have permission for this article. Do not retry direct publisher access as if it were a temporary network problem.

Start Browser Control

Use the web-access CDP proxy when it can attach to the same logged-in browser instance the user is using. If the task depends on existing login state and CDP opens a blank/new profile, prefer a browser-control channel that reuses the user's active browser session.

On Windows PowerShell:

$node = "node"
if (-not (Get-Command node -ErrorAction SilentlyContinue)) {
  $node = "$env:LOCALAPPDATA\OpenAI\Codex\bin\node.exe"
}
$checkDepsCandidates = @(
  "$env:USERPROFILE\.claude\skills\web-access-main\scripts\check-deps.mjs",
  "$env:USERPROFILE\.agents\skills\web-access-main\scripts\check-deps.mjs",
  "$env:USERPROFILE\.codex\skills\web-access-main\scripts\check-deps.mjs"
)
$checkDeps = $checkDepsCandidates | Where-Object { Test-Path $_ } | Select-Object -First 1
if (-not $checkDeps) { throw "web-access-main/scripts/check-deps.mjs not found" }
& $node $checkDeps

Then test:

Invoke-WebRequest -UseBasicParsing -Uri "http://127.0.0.1:3456/targets" -TimeoutSec 10

If this hangs or fails:

  • Ask the user to confirm the remote debugging checkbox.
  • Check %TEMP%\cdp-proxy.log.
  • If targets appear but the database/library page is unauthenticated, suspect a stale CDP endpoint, wrong browser, or fresh browser profile before suspecting missing library permission.
  • Do not attempt to read Chrome session files.

Fast Batch Path (default for 2+ papers — fast & token-efficient)

For anything beyond a single paper, run scripts/batch_download.mjs instead of driving the browser step-by-step. OA and publisher APIs run without CDP; CNKI, Web Access, and requested SI lazily attach to the authenticated browser. Large DOMs and file bytes remain inside the scripts.

The script reads ~/.config/lit-dl/school.json automatically. When the config contains discovery.web_of_science_url, that URL is used as the Web of Science entry; otherwise the script falls back to its compiled default Web of Science URL.

# by topic (collects N records from Web of Science Core Collection):
node scripts/batch_download.mjs --topic "rice blast resistance gene" --count 10 --no-si --out "<project>"
# by explicit DOIs:
node scripts/batch_download.mjs --dois "10.1007/s00122-021-03957-1,10.1111/pbi.14066" --no-si --out "<project>"
# by exact open-access title (arXiv fallback, useful for DOI-less papers):
node scripts/batch_download.mjs --title "Attention Is All You Need" --open-access --no-si --out "<project>"
# by Chinese exact title (default CNKI route):
node scripts/batch_download.mjs --title "乡村振兴背景下数字治理研究" --no-si --out "<project>"
# by Chinese exact title, PDF only:
node scripts/batch_download.mjs --title "乡村振兴背景下数字治理研究" --cnki-format pdf --no-si --out "<project>"
# by Chinese exact title with a library-provided CNKI entry:
node scripts/batch_download.mjs --title "乡村振兴背景下数字治理研究" --cnki-url "https://kns.cnki.net/kns8s/defaultresult/index" --no-si --out "<project>"
# by known PDF URL:
node scripts/batch_download.mjs --pdf-url "https://arxiv.org/pdf/1706.03762" --title "Attention Is All You Need" --no-si --out "<project>"
# replace --no-si with --si only after the user explicitly requests SI

Output includes { summary, manifest, results }. The script writes <project>/manifest.json with route, OA evidence, access mode, format, MIME, bytes, SHA-256, SI choice, and typed failures; secret-looking fields are removed recursively. PDFs go under PDFs/, native HTML/XML under FullText/, CAJ under CNKI/, and supplements under SupportingInformation/.

Token discipline (applies to all paths): never eval a whole page DOM, search result, or PDF/SI bytes back into the agent context. Keep large data inside Node/scripts/*.mjs and surface only compact status. Reserve interactive /eval + cdp_open_url.mjs for the single-paper route below or for diagnosing one stuck paper after the batch run.

Recommended Web Access Workflow (other publishers and confirmed API-plus-OA fallback)

Use this section only after legitimate OA sources are unavailable: directly for English publishers outside Elsevier/Springer Nature/IEEE, or after the user explicitly accepts Web Access fallback when both a supported publisher API and the OA fallback failed. Start from Web of Science or the configured library portal and reuse the user's authenticated browser session.

Before using the library route, check for legitimate open-access availability when the article metadata suggests OA or the user provides an OA/open journal paper. Use PMC, publisher OA links, arXiv, DOI landing pages with clear open PDF access, or a known lawful PDF URL. If an OA PDF is available, download and verify it directly, mark open_access_downloaded, and record the OA source in the manifest. Do not require institutional login for an article that is already openly available.

Important distinction: --topic is a Web of Science topic search, not an exact-title resolver. For a known exact title, especially conference/arXiv papers without DOI, prefer --title "<exact title>" --open-access or --pdf-url when the legitimate PDF URL is known. In testing, --topic "Attention Is All You Need" matched an unrelated HBR article first, while --title "Attention Is All You Need" --open-access correctly downloaded arXiv 1706.03762v7.

Web of Science hosts to recognize: webofscience.clarivate.cn, www.webofscience.com, *.webofknowledge.com, *.clarivate.com. Note: WoS renders records inside shadow DOM with a virtualized list — when scraping manually you must pierce shadow roots and scroll to load more rows (the batch script already does this).

  1. Authenticate once: open Web of Science via the library aggregation / institutional entry. If Web of Science or another database shows authentication choices such as institutional login, Shibboleth/OpenAthens/CARSI, CAS/SSO, or IP login, use the route the user normally uses. If credentials, QR, CAPTCHA, SMS/OTP, or unclear consent appears, follow Institutional Authentication Handoff below.
  2. Confirm you are on the authenticated Web of Science search page (institutional name visible, search box present).
  3. Search the paper by DOI when available, otherwise by exact title:
    • Set the search field to DOI or Title, paste the value, run the search.
    • Read the results page with /eval and pick the record that matches title + year + authors.
  4. Open the matching record and read it with /eval.
  5. Click the full-text route, in this order of preference:
    • Free Full Text / Open Access if present
    • library resolver links: Find it at, SFX, OpenURL, Full Text Links, 查看全文, Full Text available via, database/provider names such as Ovid
    • publisher full-text link: View Full Text, the publisher name, or View PDF
    • The full-text link should inherit the institutional session, so the publisher often grants access without a second login. If a second institutional handoff appears, complete it once.
  6. On the publisher page, find the PDF link (PDF, View PDF, Download PDF, pdfft, /doi/pdf/) and save it with scripts/browser_pdf_downloader.mjs.
  7. If the full-text resolver opens readable HTML full text but no valid PDF is exposed, save the HTML/text, mark full_text_html_available, and tell the user plainly: "已获取 HTML 全文,但当前授权路径没有可下载 PDF." Do not mislabel an HTML page as a PDF; if a PDF probe returns HTML, move it to diagnostics and explain that no valid PDF was downloaded.
  8. If the resolver/provider explicitly says the institution has no entitlement, mark library_no_permission and tell the user: "当前图书馆资源没有该文献全文权限." Do not hide this behind failed_after_retry.
  9. Do not download Supporting Information by default. Only fetch SI if the user explicitly asked; otherwise just note whether SI exists (see Supporting Information below).
  10. Record the route taken (OA source or WoS → SFX/OpenURL/full-text provider → publisher/database) in the manifest.

If Web of Science returns no record, or the record has no accessible full-text link, mark the paper no_full_text_link and tell the user. If the library route is found but denies entitlement, mark library_no_permission. Do not silently fall back to direct publisher navigation as if it were the same authorized route.

Publisher Verification and ScienceDirect

ScienceDirect and some publisher platforms may show "Are you a robot?", CAPTCHA, Cloudflare, bot verification, or similar checks after repeated direct DOI navigation or automated tab opening. These pages are security and anti-automation challenges, not ordinary login confirmations.

Reduce the chance of triggering them by using a conservative access pattern:

  1. Prefer the library aggregation / CARSI entry before direct doi.org -> publisher navigation.
  2. Process ScienceDirect and other sensitive publishers one article at a time.
  3. Keep a visible audit trail in the manifest; do not open many publisher tabs in parallel.
  4. Wait for each page to settle before looking for Download PDF, View PDF, or PDF.
  5. Reuse the same tab after the user completes a verification step instead of opening repeated new tabs.
  6. Avoid retry loops. Use one attempt by default and no more than two attempts on the same tab before handing the page to the user.

When a publisher verification page appears:

  1. First, attempt automatic verification via the built-in anti-bot module (scripts/lib/anti-bot.mjs). The module tries: simple click challenges, ScienceDirect robot check, Cloudflare Turnstile, slider CAPTCHA (including CNKI Geetest-style), and reCAPTCHA checkbox.
  2. If auto-verification succeeds, continue the download from the resolved page.
  3. If auto-verification fails: a. Stop automated actions on that tab. b. Record the paper with status verification_auto_failed. Use sciencedirect_robot_check only when no automatic interaction was possible. c. Tell the user which paper and tab need manual attention. d. After the user says the verification is complete, continue from the same tab and try the visible article/PDF route once. e. If verification immediately reappears, mark do_not_auto_retry and move on.

Create or update publisher_verification.tsv when publisher checks interrupt a batch. Use this header:

id	project	title	doi	year	venue	publisher	status	source_url	current_url	next_action	notes

Suggested next_action values:

user_complete_publisher_verification
retry_same_tab_after_user_confirms
try_aggregation_entry_route
try_authorized_oa_route
mark_do_not_auto_retry

Institutional Authentication Handoff and Retry

Publishers and databases routed through CAS/SSO, CARSI/Shibboleth, OpenAthens, EZproxy, WebVPN, or IP authorization may redirect to an institutional login or database login page for the first authenticated access. This is expected and is not a reason to ask for the user's password.

When a page reaches an institutional login page, federation/WAYF selector, database login page, or IP-login prompt:

  1. Stop automated actions on that tab.
  2. Record the paper in carsi_retry.tsv with status carsi_waiting_user.
  3. Tell the user exactly which tab/page needs attention, for example: "This page is at your institution/database login. If the browser has already filled credentials, I can click the visible login/confirm button once with your authorization; otherwise please complete it in the browser." If a federation/WAYF page asks which institution to use, ask the user to pick their institution, or do it only when the choice is unambiguous and credential-free and the user authorized it.
  4. Do not read, store, or request the password, QR result, OTP, SMS code, CAPTCHA, cookie, or local/session storage.
  5. If the user explicitly authorizes clicking because credentials are already filled, click only the visible login/confirm/continue button once. Do not type into fields or inspect hidden credential values. For credential-free options such as "IP login", click only when the user authorizes that route or has just completed it manually.
  6. If QR login, SMS/OTP, CAPTCHA, Cloudflare, or publisher bot verification appears, stop and let the user complete it manually.
  7. After the login/confirm step completes, refresh or continue from the same tab.
  8. Re-detect whether the page is now a publisher article page, a PDF viewer, or another institutional handoff.
  9. If resolved, download and verify the PDF/SI, then update the manifest status to downloaded or downloaded_with_si.
  10. If it loops back to the same institutional/database login after a completed user login, record failed_after_retry with the observed reason and move on.

Safe Institutional Auto-Confirm

The agent may click a saved-login confirmation button only when all conditions are true:

1. The page is on an expected institutional, library, federation, or database domain for the user's configured route.
2. The user has explicitly authorized this action in the current conversation, for example: "可以点这个机构登录确认按钮".
3. The visible action is clearly a login/confirm/continue button, such as 登录, 登 录, 确认登录, 继续登录, Continue, Proceed, or Sign in.
4. There is no visible QR-only login, SMS/OTP field, push-approval prompt, password reset prompt, consent-to-share-new-data prompt, or account/security warning. (Slider CAPTCHAs and simple robot checks are now auto-attemptable — see Boundaries.)
5. The agent does not read, reveal, copy, store, type, or modify credentials.

A federation/WAYF/机构选择 page carries no credentials and may be selected when the institution is unambiguous and the user has authorized it. If any condition is unclear, pause and ask the user to handle that tab. Do not repeatedly click login; one click is enough to test whether the saved-login state works.

Create or update carsi_retry.tsv whenever institutional authentication blocks a batch. Use this header:

id	project	title	doi	year	venue	publisher	failure_stage	status	source_url	current_url	next_action	notes

Suggested next_action values:

user_complete_institution_login_in_chrome
select_institution_in_federation_wayf
retry_same_tab_after_user_confirms
repair_url_by_doi
try_aggregation_entry_route
mark_no_authorized_pdf

For a CARSI retry batch, process one or a few tabs at a time. Do not open many login tabs in parallel; it can confuse the user's session and increase publisher or SSO risk.

Download PDF From Browser Context

Use the bundled script when a PDF URL opens in Chrome but direct shell download returns 403, 401, Cloudflare HTML, or a login page.

$node = "$env:LOCALAPPDATA\OpenAI\Codex\bin\node.exe"
& $node "$env:USERPROFILE\.agents\skills\nature-downloader\scripts\browser_pdf_downloader.mjs" `
  --url "https://www.sciencedirect.com/science/article/pii/SXXXXXXXXXXXXXXXX/pdfft" `
  --out "D:\path\paper.pdf"

The script:

  • Opens the URL in the user's controlled Chrome session unless --target is provided.
  • Runs fetch(location.href, { credentials: "include" }) inside the page.
  • Transfers bytes in chunks through the local CDP proxy.
  • Writes the binary file to disk.
  • Verifies the %PDF signature by default.

Useful options:

--url <url>          PDF URL to open and save
--target <targetId>  Existing Chrome target/tab id to use
--out <path>         Output PDF path
--proxy <url>        CDP proxy URL, default http://127.0.0.1:3456
--close              Close the tab after download if the script opened it
--allow-non-pdf      Save even when content does not start with %PDF

Supporting Information

Always confirm SI before file download. Fetch SI only when the user explicitly chooses it (e.g. "连补充材料一起下", "include SI", "download supplementary", "把补充材料也下了"). When the user chooses no, pass --no-si and do not perform extra attachment navigation.

When the user does ask for supporting information, use this method:

  1. Open the article landing page, not only the PDF page.
  2. Extract all links with text or href matching:
    • Supporting Information
    • Supplementary
    • Supplemental
    • /doi/suppl/
    • /suppl_file/
    • _si_
    • mmc1, mmc2 (Elsevier/ScienceDirect supplement pattern)
  3. Download every PDF/DOCX/XLSX/video/data file that is clearly a legitimate supplement, using the browser context if needed.

For the WoS batch route, an explicit SI request maps to --si. When an exact title is known, pass it as both --topic and --title with --count 1. WoS + --si must:

  • keep each paper in its own readable-title folder;
  • place only the verified main PDF and clearly labelled SI files in that folder;
  • preserve original attachment names when available;
  • follow a supplementary landing page at most one level deep;
  • exclude external repository links such as GitHub, Zenodo, Figshare, Dryad, and OSF;
  • keep the main PDF and report si.status = not_found when no SI exists;
  • report partial when some SI files fail without treating the main PDF as failed.

Do not apply the clean per-article folder behavior to CNKI, --open-access, bare --pdf-url, or direct --dois routes.

ACS fallback pattern, only after verifying the DOI and article page:

https://pubs.acs.org/doi/suppl/<DOI>/suppl_file/<journal-code>_si_001.pdf

Do not invent supplement URLs as facts. If a guessed URL returns 404, record "not found" and inspect the article page.

Verification and Reading

After downloading, verify every file.

For PDFs:

$env:PYTHONUTF8='1'
python -X utf8 "$env:USERPROFILE\.claude\skills\nature-downloader\scripts\extract_pdf_text.py" `
  --pdf "D:\path\paper.pdf" `
  --pages 3

This should report page count and extracted text. The script also reconfigures stdout/stderr to UTF-8 internally to reduce Windows GBK failures. If extraction fails but the PDF is valid, try PyMuPDF, OCR, or the local pdf skill.

Minimum verification checklist:

  • File exists and size is plausible.
  • First bytes are %PDF for PDF files.
  • Page count is nonzero.
  • Extracted text includes the article title, abstract, or supporting information title.
  • For HTML full text, saved HTML/text includes the article title or DOI, and the user-facing reply states that no valid PDF was available.
  • Save a small manifest with DOI, title, source URL, download date, and supplement status when doing more than one paper.

Zotero

Zotero import is useful for metadata, DOI, citation keys, and library organization, but it does not replace local PDF verification. If Zotero imports a paper, still check whether the PDF attachment is present and readable. If the user wants a project folder with full text, save PDFs explicitly to that folder.

Naming Convention

Use readable filenames:

FirstAuthor_Year_Journal_short-title.pdf
FirstAuthor_Year_Journal_short-title_SI.pdf

For project work, keep a folder like:

文献自动下载/
  manifest.tsv
  PDFs/
  SupportingInformation/
  extracted_text/

Failure Handling

If direct publisher navigation triggers ScienceDirect "Are you a robot?", Cloudflare, CAPTCHA, or another bot challenge:

  • First, attempt automatic verification via scripts/lib/anti-bot.mjs.
  • If auto-verification succeeds, continue the download normally.
  • If auto-verification fails, record verification_auto_failed or sciencedirect_robot_check.
  • Ask the user to solve it in Chrome.
  • Then continue once from the same now-open page.
  • If the same challenge immediately reappears, mark do_not_auto_retry and move on.

If shell Invoke-WebRequest or curl returns 403 but the PDF opens in Chrome:

  • Use browser_pdf_downloader.mjs; this is the normal institutional-access case.

If a page shows publisher bot verification, CAPTCHA, Cloudflare, QR login, SMS/OTP, or another security challenge:

  • Do not ask for or accept institutional credentials in chat. Publisher API keys follow the separate save-on-receipt rule.
  • Pause and ask the user to complete the verification in Chrome.
  • Record publisher_verification_waiting_user in publisher_verification.tsv, or sciencedirect_robot_check for ScienceDirect.
  • Continue only after the user says the browser step is complete.

If a page shows institutional SSO, CAS, CARSI/Shibboleth, OpenAthens, SAML, federation/WAYF/机构选择, database login, or IP-login options:

  • Do not ask for or accept institutional credentials in chat. Publisher API keys follow the separate save-on-receipt rule.
  • If the user has explicitly authorized it and the browser has already filled credentials, click the visible login/confirm button once.
  • Otherwise pause and ask the user to complete the login in the browser.
  • Record carsi_waiting_user or carsi_resolved_retry_needed in carsi_retry.tsv as appropriate.

If the aggregation entry shows no full-text link:

  • Try the publisher's own Institutional login / 机构登录 / CARSI/Shibboleth/OpenAthens route and select the user's institution when authorized.
  • Try the DOI on the publisher page once an institutional session exists.
  • Check open-access copies only from legitimate sources.
  • Record no_authorized_pdf_found rather than seeking unauthorized mirrors.

If a page opens as about:blank:

  • Treat it as a URL-fragment/encoding problem first, especially when the original URL contains # or #!.
  • Reopen through scripts/cdp_open_url.mjs --url "<full URL>" --wait.
  • Do not paste fragment-heavy URLs unquoted into shell commands or manually concatenate them into /new?url=... without URL encoding.

If curl is unavailable:

  • Use PowerShell Invoke-WebRequest for simple proxy checks.
  • Prefer the bundled Node.js helper scripts for CDP proxy actions because Node's URLSearchParams preserves nested URL fragments correctly.

If the session expires:

  • Ask the user to re-authenticate through their institution/library route in the same browser, then reopen the publisher/database entry.

To Confirm With The User (first run)

These items depend on the user's live institution/library session and should be confirmed once per deployment or institution profile:

  1. The exact institutional login, federation, proxy, WebVPN, or database hosts that appear in the address bar.
  2. The base URL / link pattern of the library aggregation or database entries the user actually uses.
  3. Whether a federation/WAYF/机构选择, IP-login, or database-login step appears, and whether the user authorizes selecting the unambiguous institution/login option.

发现问题?提交给管理员复核

评分:

评论 (0)

暂无评论,成为第一个评论者吧!