SkillAtlasSkill 详情

ref-downloader

Stop losing an afternoon to chasing dozens of reference PDFs by hand.

审核状态:已审核Quality 72Security 70

复制安装命令

用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。

复制前请先查看来源、License 和安全提示。

项目 README

来源文件:README.md

抓取于 2026年9月12日

ref-downloader

Stop losing an afternoon to chasing dozens of reference PDFs by hand. One DOI in, every reference PDF out — using your existing institutional access.

Version: 0.4.1 Status: beta License: MIT Python 3.11+ Verified on Windows + Edge

中文完整文档 / Full Chinese version

Status: beta (v0.4.1). Windows + Microsoft Edge verified path. macOS / Linux / Chromium untested. Expect rough edges around supplementary downloads and publisher-site changes. PR-worthy issues welcome.

Heads up — not a paywall bypass. ref-downloader uses your institutional access. If your university or organization subscribes to a journal, those refs work. If they don't, those refs become manual_pending for you to follow up on by hand.

Demo (30-second console preview)

$ python run_ref_downloader.py 10.1021/jacs.5c05017

=== Ref Downloader Wrapper ===
DOI:         10.1021/jacs.5c05017
PROJECT:     jacs.5c05017
Config:      config.example.toml + config.local.toml

>>> extract_refs.py
  Title: Designing Natural Cell-Inspired Heme-Spurred Membrane...
  References found: 38

>>> validate_refs.py
  Total: 38  Verified: 38  Failed: 0  No DOI: 0

>>> download_refs.py
  [ 1] downloaded (842 KB)        Lee2016_NatEnergy.pdf
  [ 2] downloaded (1.2 MB)        Wang2018_AdvMater.pdf
  [ 3] manual_pending (auth_redirect)
  [ 4] downloaded (655 KB)        Chen2019_JACS.pdf
  [ 5] failed (challenge_timeout)
  [ 6] ignored (ignored_institution_access)
  ... 31 more refs processed ...
  [38] downloaded (956 KB)        Park2024_JElectrochemSoc.pdf

========== Download report ==========
Total references:  38
Main PDFs:         33 downloaded · 3 manual_pending · 1 failed · 1 ignored
SI files:          12 captured
PDFs land in:      ./jacs.5c05017_refs/jacs.5c05017/
=====================================

Contents

What you get

  • Paywalled refs work without setup. Drives your real Microsoft Edge profile, so any institutional login already in your browser carries through. No API keys, no proxies, no reverse engineering.
  • One DOI in, every reference PDF out. Crossref-driven extraction + 17+ publisher-specific download paths (Wiley PDFDirect, Elsevier viewer, AIP loading-page wait — see per-publisher reliability tier), not generic scraping.
  • Two routing modes from an agent (v0.4.1+). Mode A: hand the agent one paper, get all of its references. Mode B: hand the agent a custom list — DOIs / titles / arXiv-PMID IDs / BibTeX / abstract queries like "Wang 2024 Nature Energy papers" — and it resolves to DOIs and downloads them. See SKILL.md.
  • You always know which refs failed and why. download_report.csv gives every ref a status + reason (manual_pending (auth_redirect), failed (challenge_timeout), ignored); events.jsonl keeps the per-ref event trace.
  • Pick up where you left off after a VPN drop, browser crash, or Ctrl+C. State persists per project; rerunning skips already-downloaded refs and retries only the failures.

Why not Zotero, scihub, or generic scrapers?

  • vs. Zotero's Find Available PDF — walks one paper at a time and silently gives up at SSO redirects. ref-downloader walks the whole reference list at once and treats SSO as a configurable step instead of a dead end.
  • vs. scihub-style tools — don't carry your institutional license, so paywalled refs you legitimately have access to just fail. ref-downloader uses your authenticated browser session, so subscriptions you already pay for actually count.
  • vs. generic web scrapers — don't know Wiley needs PDFDirect, Elsevier needs a viewer click, or AIP serves a Chinese loading page first. ref-downloader has 17+ publisher-specific paths plus Elsevier popup state machine + --auto mode retry queue (manual-pending refs get a second async attempt 60s later, hot-session preserved).
  • vs. raw Playwright — gets blocked on Cloudflare / Radware / Turnstile-heavy sites. Set REF_DOWNLOADER_BROWSER=cloak to swap in cloakbrowser's stealth Chromium with humanized input — no code changes, same pipeline. See Configuration.

Quick start

The skill is self-contained under skills/ref-downloader/. Pick the install path for your agent framework:

git clone https://github.com/ltczding-gif/ref-downloader.git

# Pick ONE install destination for your agent framework:
#   Claude Code:        cp -r ref-downloader/skills/ref-downloader ~/.claude/skills/
#   Codex CLI:          cp -r ref-downloader/skills/ref-downloader ~/.codex/skills/
#   Copilot CLI / VSC:  cp -r ref-downloader/skills/ref-downloader .github/skills/
#   Project-local:      cp -r ref-downloader/skills/ref-downloader .agents/skills/

cd ~/.claude/skills/ref-downloader     # or wherever you copied it
pip install playwright pymupdf
playwright install msedge
cp config.example.toml config.local.toml      # then set [crossref].mailto

# In your agent: just describe the task; the skill triggers via its description.
# Direct CLI for testing: python scripts/run_ref_downloader.py 10.1021/jacs.5c05017

What you'll see: 30–80 refs discovered for a typical chemistry/physics paper, then a mix of downloaded (refs your institution covers), manual_pending (SSO bounce or paywall), and occasional failed (publisher quirk). Run on a DOI from a journal your institution actually subscribes to for the highest hit rate. Details below.

Requirements

  • OS: Windows 10/11 (verified). macOS / Linux untested — PRs welcome.
  • Browser: Microsoft Edge (Stable channel). The script claims your persistent Edge profile, so close all Edge windows before running.
  • Python: 3.11 or newer (uses stdlib tomllib).
  • Optional: A Zotero installation (auto-detects DOI from a PDF's filename via Zotero's SQLite database — much faster than text extraction).
  • Optional: PyMuPDF (pip install pymupdf) for DOI extraction from PDF text when Zotero lookup is unavailable.

Install

As an agent skill (recommended)

Pick the install path for your agent framework:

FrameworkInstall command
Claude Codecp -r skills/ref-downloader ~/.claude/skills/
Claude Agent SDKsame (auto-discovers ~/.claude/skills/)
Codex CLIcp -r skills/ref-downloader ~/.codex/skills/
Copilot CLI / VS Code agentcp -r skills/ref-downloader .github/skills/
Any framework (project-local)cp -r skills/ref-downloader .agents/skills/

Then install Python prereqs INSIDE the copied skill folder (the skill protocol doesn't manage Python deps):

cd ~/.claude/skills/ref-downloader            # or wherever you copied it
pip install playwright pymupdf                # or use the source's requirements.txt
playwright install msedge

cp config.example.toml config.local.toml
# Edit config.local.toml — at minimum set [crossref].mailto.
# Windows: notepad config.local.toml
# macOS / Linux: $EDITOR config.local.toml   (or vim / nano / code / ...)

As a Python tool (for developers)

If you want to hack on the code, the skill folder is a runnable Python project:

git clone https://github.com/ltczding-gif/ref-downloader.git
cd ref-downloader

pip install -r requirements.txt -r requirements-dev.txt
playwright install msedge

cp skills/ref-downloader/config.example.toml skills/ref-downloader/config.local.toml
# Edit config.local.toml — at minimum set [crossref].mailto.

# Run the offline test suite
python -m pytest tests/ -v

# Run the tool directly
python skills/ref-downloader/scripts/run_ref_downloader.py 10.1021/jacs.5c05017

Usage examples

(After install — paths assume the skill is at <SKILL_DIR>, e.g. ~/.claude/skills/ref-downloader/. In source, <SKILL_DIR> = skills/ref-downloader/.)

Input: a DOI

python <SKILL_DIR>/scripts/run_ref_downloader.py 10.1021/jacs.5c05017

Default output: <cwd>/jacs.5c05017_refs/jacs.5c05017/

Input: a local PDF (with DOI in metadata or in PDF text)

python <SKILL_DIR>/scripts/run_ref_downloader.py "C:\path\to\your_paper.pdf"

Default output: <pdf_dir>/your_paper_refs/<doi-derived-name>/

Custom output directory

python <SKILL_DIR>/scripts/run_ref_downloader.py 10.1021/jacs.5c05017 --output-dir refs/

Non-interactive (CI / batch)

python <SKILL_DIR>/scripts/run_ref_downloader.py 10.1021/jacs.5c05017 --yes --auto

Alternate config file

python <SKILL_DIR>/scripts/run_ref_downloader.py 10.1021/jacs.5c05017 --config ./alt.toml

Configuration

All configuration lives in config.local.toml (gitignored). Copy config.example.toml to bootstrap.

SectionKeyPurpose
[crossref]mailtoYour email — entry into Crossref polite pool
[zotero]db_pathOptional path to zotero.sqlite for DOI lookup from PDF filename
[browser]edge_profile_dirEdge profile directory; empty = OS default
[browser]disable_extensionsSet true to launch with --disable-extensions
[institution]auth_hostsHostnames that mean "you got bounced to SSO" (e.g. ["sso.your-uni.edu"])
[institution]auth_url_fragmentsURL substrings indicating SSO (e.g. ["oauth", "saml"])
[institution]auth_page_titles<title> text for SSO pages (catches HTML served as PDF)
[institution]auth_loading_titlesLoading-page titles (also reused for AIP/AVS publisher loading detection)
[institution]ignored_access_doisDOIs you know are paywalled at your institution; skipped without retry

Environment variables override file values:

VariableMaps to
REF_DOWNLOADER_MAILTOcrossref.mailto
REF_DOWNLOADER_ZOTERO_DBzotero.db_path
REF_DOWNLOADER_EDGE_PROFILEbrowser.edge_profile_dir
REF_DOWNLOADER_DISABLE_EXTENSIONSbrowser.disable_extensions (1/true to enable)
REF_DOWNLOADER_CONFIGPath to alternate TOML file

See skills/ref-downloader/config.example.toml for full documentation.

Alternative backend: CloakBrowser (optional, for Cloudflare-heavy sites)

What it is. CloakBrowser is a third-party Python package by CloakHQ (MIT-licensed, available on PyPI as cloakbrowser). It ships a patched Chromium build with source-level anti-fingerprint changes designed to look like a normal browser to common bot-detection layers (Cloudflare Turnstile, Radware, DataDome, FingerprintJS, etc). Its launch_persistent_context_async() API is intentionally compatible with Playwright's — that's what lets ref-downloader swap backends with a single env var instead of rewriting the download flow.

Not a dependency of ref-downloader. If you don't run pip install cloakbrowser it's never imported. The default Edge backend is unchanged. When CloakBrowser IS the active backend, ref-downloader uses Chromium under a separate persistent profile at ~/.local/cloakbrowser/profiles/ref-downloader (or REF_DOWNLOADER_CLOAK_PROFILE), so your Edge profile is not touched — Edge does NOT need to be closed.

When to use it. Sites you'd reach for it on: CCS Chemistry (10.31635, Cloudflare-protected), some Elsevier paths gated by Radware, anything where the Edge backend keeps producing manual_pending (radware_bot_manager) or failed (challenge_timeout). Don't reach for it as a default — the Edge backend is more reliable when your institutional access is the actual bottleneck, because Edge carries your authenticated cookies.

Caveats. CloakBrowser is beta third-party software; install + use at your own discretion (review its repo before pulling it). It is not a captcha solver — interactive challenges still need you. It also does not carry your institutional cookies (separate profile), so it's most useful for open-Cloudflare sites, less useful for paywalled-but-license-covered refs.

pip install cloakbrowser                              # one-time, separate from ref-downloader
$env:REF_DOWNLOADER_BROWSER = "cloak"
$env:REF_DOWNLOADER_CLOAK_HUMAN_PRESET = "careful"    # optional: slower mouse/scroll
python skills/ref-downloader/scripts/run_ref_downloader.py 10.31635/ccsorg...

CloakBrowser env vars (all optional):

VariableDefaultPurpose
REF_DOWNLOADER_BROWSERedgeSet to cloak (or cloakbrowser) to switch backend
REF_DOWNLOADER_CLOAK_PROFILE~/.local/cloakbrowser/profiles/ref-downloaderPersistent Chromium profile path
REF_DOWNLOADER_CLOAK_HUMANIZE10/false to disable humanized input
REF_DOWNLOADER_CLOAK_HUMAN_PRESETdefaultdefault or careful (slower)
REF_DOWNLOADER_CLOAK_PROXYunsetHTTP/SOCKS proxy URL
REF_DOWNLOADER_CLOAK_GEOIPauto1 to force GeoIP rerouting (auto when proxy is set)
CLOAKBROWSER_PYTHONPATHunsetsys.path hint for a local cloakbrowser source checkout

Notes:

  • Edge does not need to be closed when using the cloak backend — it uses its own Chromium.
  • A fresh cloak profile may still hit Cloudflare/security pages on first visit — warm it manually with that profile before batch downloads.
  • human_preset=careful reduces behavior-based detection but is not a captcha solver.
  • cloakbrowser is NOT a hard dependency of ref-downloader. If you never set REF_DOWNLOADER_BROWSER=cloak, it's not imported.

Architecture

Three-stage pipeline + a wrapper:

skills/ref-downloader/
├── SKILL.md                            agent runbook (slim entry)
├── references/agent-runbook.md         extended manual flow + DOI fallback
├── config.example.toml                 config schema (copy to config.local.toml)
└── scripts/
    ├── run_ref_downloader.py           entry — config + DOI resolution + sequencing
    │     └─> extract_refs.py    (1) Crossref API: fetch parent's reference list
    │     └─> validate_refs.py   (2) Crossref API: per-ref metadata + publisher classify
    │     └─> download_refs.py   (3) Playwright/Edge: download main PDF + SI per publisher
    └── _config.py                      TOML + env-var loader

You can also run the three scripts manually for debugging or partial restarts. See the agent runbook in skills/ref-downloader/references/agent-runbook.md for the manual flow.

Agent users can install or inspect the packaged skill at skills/ref-downloader/SKILL.md. The repository root remains the human-facing Python project; the skill bundle is kept separate so Codex does not treat README, changelog, tests, and source files as always-associated skill context.

Supported publishers

ACS, Nature, Science, Elsevier, Wiley, RSC, Springer, PNAS, ECS, IOP, AIP, AVS, IEEE, OSA, KPS, Beilstein, APS, Annual Reviews, Taylor & Francis, CCS Chemistry. Maturity varies — see docs/SUPPORTED_PUBLISHERS.md for the per-publisher tier table and known issues. CCS Chemistry sits behind Cloudflare; pair it with REF_DOWNLOADER_BROWSER=cloak for reliable access.

Known limitations

  • Windows + Microsoft Edge only: that's the verified path. macOS / Linux / Chromium support has not been tested. If you try, please open an issue with results.
  • Headed mode required: empirically, headless=True yields empty results for Wiley / ACS supplementary downloads. The default is headed.
  • Edge must be fully closed before running: Playwright needs exclusive access to the persistent profile. Check Task Manager for any background msedge.exe processes.
  • SSO redirects are detected, not solved: when the script bounces to your institution's SSO, the ref becomes manual_pending so you can sign in interactively. Configure [institution] to teach it which redirects to recognize.
  • SI download is the most fragile path: main PDFs are reliable; SI lookup varies by publisher and is the area most likely to need a tweak when a publisher updates their site.
  • Paywalled content needs institutional access: this is not a bypass tool.
  • Crossref dependency: papers with no reference list deposited at Crossref can't be processed automatically.

Contributing

See CONTRIBUTING.md for guidance on:

  • Adding a new publisher (DOI prefix → strategy)
  • Adding institutional SSO patterns
  • Reporting download failures with useful logs

Security

This tool launches your real Edge profile, with all your cookies and saved sessions. Read SECURITY.md before running it against a profile you also use for daily browsing.

License

MIT — see LICENSE.

其他

中风险

  • 来源需自行核对维护者身份。
  • 包含脚本或命令调用,安装前请复核。
  • 可能需要外部 token、网络权限或第三方服务。
  • 未检测到高风险命令。
  • 扫描发现:4 条。

Codex — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/ltczding-gif/ref-downloader.git
  3. 将 "skills/ref-downloader" 文件夹复制到 Codex 的 skills 目录中。
  4. 重启 Codex 让新的 skill 生效。

Codex — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Codex 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Codex 让新的 skill 生效。

Claude Code — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/ltczding-gif/ref-downloader.git
  3. 将 "skills/ref-downloader" 文件夹复制到 Claude Code 的 skills 目录中。
  4. 重启 Claude Code 让新的 skill 生效。

Claude Code — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Claude Code 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Claude Code 让新的 skill 生效。

Cursor — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/ltczding-gif/ref-downloader.git
  3. 将 "skills/ref-downloader" 文件夹复制到 Cursor 的 skills 目录中。
  4. 重启 Cursor 让新的 skill 生效。

Cursor — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Cursor 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Cursor 让新的 skill 生效。

GitHub Copilot — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/ltczding-gif/ref-downloader.git
  3. 将 "skills/ref-downloader" 文件夹复制到 GitHub Copilot 的 skills 目录中。
  4. 重启 GitHub Copilot 让新的 skill 生效。

GitHub Copilot — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 GitHub Copilot 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 GitHub Copilot 让新的 skill 生效。

Windsurf — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/ltczding-gif/ref-downloader.git
  3. 将 "skills/ref-downloader" 文件夹复制到 Windsurf 的 skills 目录中。
  4. 重启 Windsurf 让新的 skill 生效。

Windsurf — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Windsurf 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Windsurf 让新的 skill 生效。
查看 SKILL.md 原文
name: ref-downloader
description: >
  Use when the user asks to batch-download academic PDFs with
  ref-downloader — either ALL references of one paper (Mode A: DOI or
  PDF input), OR a custom batch of papers (Mode B: DOI/title/arXiv-PMID
  list, or abstract query like "Author X's recent papers"). Not for
  one-off PDFs, paper search, or Zotero import.

Ref Downloader — Agent Runbook

Slim entry for agent mode. The full 8-step manual runbook with code snippets for Mode A debug + PUBLISHER_MAP extension procedure lives in references/agent-runbook.md. Human users see ../../README.md.

<SKILL_DIR> = this folder (skills/ref-downloader in the source repo, or wherever the user copied this skill — e.g. ~/.claude/skills/ref-downloader/). Python scripts live in <SKILL_DIR>/scripts/; config files (config.example.toml, config.local.toml) live at <SKILL_DIR>/.

Mode router

This skill handles two flows. Pick before running.

  • Mode A — Reference-list download (original use case). User provides ONE paper (DOI or local PDF) and wants "all of its references". Pipeline: extract_refs.py → validate_refs.py → download_refs.py.

  • Mode B — Custom batch download. User provides their own batch of papers — DOIs, paper titles, non-DOI identifiers (arXiv / PMID / Semantic Scholar IDs), OR an abstract query ("Smith 在 Google Scholar 上的文章" / "Nature Energy 2023 papers"). The agent resolves whatever was given to DOIs, then runs validate_refs.py → download_refs.py directly. Skip the wrapper — run_ref_downloader.py assumes a parent DOI and will fail.

Both modes share install, config, per-publisher strategies, failure modes, output layout, and the CloakBrowser opt-in backend.

Trigger family

User input shapeModeSub-flow
One DOI/PDF + "all refs of" / "全部参考文献" / "把这篇引用都下了"A—
≥2 DOIs in input (any wrapping: bare / {} / https://doi.org/… / dx.doi.org/…; ASCII or full-width slashes)BB.1 (after canonicalize)
Non-DOI IDs only: arXiv: / PMID: / S2: / corpusId:BB.0 normalize → B.1
Title list ("下载这几篇:title1, title2, …")BB.2
Abstract query (author / topic / journal+year / "Google Scholar 上 …")BB.3
Mixed (DOIs + titles + IDs + queries)Brun each, merge
Single DOI without "of refs" qualifierBB.1 single-item
Title + author + year for ONE paper ("Smith 2024 Nature paper on X")BB.2 (specific paper, lookup)
Open-ended query for a corpus ("Smith 2024 之后所有的 Nature 文章")BB.3 (discovery)
Insufficient resolvable content ("上次给你的那 5 篇" / pure pronouns)—Ask user to repaste / attach file; do NOT guess
Genuinely ambiguous A vs B—ask user

Key disambiguators:

  • "refs OF a paper" (Mode A) vs "these papers themselves" (Mode B).
  • Identifies a specific paper (B.2) vs describes a set to discover (B.3). If user names a specific paper with enough metadata to uniquely identify it (title + author + year), it's B.2 (lookup exact); if they describe a class of papers ("all Smith's 2024 Nature papers"), it's B.3 (discover then filter).

Mode A — Reference-list download

When to invoke

Trigger phrases:

  • "帮我下载 [paper / DOI] 的参考文献" / "批量下载引用文献" / "把这篇论文的所有引用下载下来"
  • "Download all refs of [paper / DOI]" / "Batch-download every reference"
  • User provides a DOI (10.x/y form) or local PDF path and asks for "all references" / "全部参考文献"

Don't invoke for:

  • Downloading one arbitrary PDF (not a reference list) → user wants one paper itself, route to Mode B as a single-item B.1
  • Generic web scraping
  • Paper search / Zotero import — different tools

Primary entry

python "<SKILL_DIR>/scripts/run_ref_downloader.py" <DOI_OR_PDF_PATH>

The wrapper handles DOI resolution (Zotero → fitz fallback), output-dir layout, sequential 3-stage pipeline (extract_refs.py → validate_refs.py → download_refs.py), and end-of-run cleanup.

Useful flags:

  • --yes — non-interactive (CI/batch), overwrite prompts default-yes
  • --auto — forwarded to download_refs.py: skip "press Enter" confirm + shorter challenge wait + async retry queue for manual_pending refs (60s delay, single retry, max 3 concurrent). Use for CI / overnight runs; not for sessions where you want to drive captchas yourself.
  • --fail-fast — terminate after first actionable unresolved ref (useful in CI to surface real failures fast)
  • --output-dir <path> — override default output location
  • --config <path> — alternate TOML config (overrides config.local.toml)

Pre-flight checklist (confirm before running)

  1. DOI correct? Echo back to user: 即将下载参考文献:DOI=<doi>
  2. Edge fully closed? All msedge.exe processes killed (Task Manager check). The script claims the user's persistent Edge profile and needs exclusive access. (Cloak backend skips this.)
  3. Config set? First-run users need <SKILL_DIR>/config.local.toml with [crossref].mailto. Missing config → wrapper prints a WARNING but continues with placeholder defaults.
  4. Output location agreed? Default for DOI input: <cwd>/<project_name>_refs/. For PDF input: <pdf_dir>/<pdf_stem>_refs/. Override with --output-dir.

Mode B — Custom batch download

Input variability is the point. Don't refuse — route. The agent handles whatever shape the user gave (paste, file, prose, BibTeX, RIS, abstract query) and resolves it to a clean DOI list before handing off to the pipeline.

Step 0 — Canonicalize input (always runs first)

Before any extraction or routing, normalize the input string:

  1. Unwrap delimiters around DOIs: strip leading/trailing {}, <>, (), [], and quote marks.
  2. Strip URL prefixes: https://doi.org/, http://doi.org/, https://dx.doi.org/, http://dx.doi.org/ → bare DOI.
  3. Full-width / Unicode punctuation to ASCII:
    • / (U+FF0F) → /
    • : (U+FF1A) → :
    • . (U+FF0E) → .
    • Smart quotes "" '' → straight " '
  4. Trim trailing punctuation: .,;)}"' (note }).

Apply Step 0 BEFORE the regex pass in B.1 AND before the canonical dedupe compare in step 5 of the main flow.

B.0 — Normalize non-DOI identifiers

For each non-DOI identifier the user gave:

InputAction
arXiv:2401.12345 or bare arXiv IDUse 10.48550/arXiv.<id> (canonical), but prefer a journal DOI if the agent can discover one via Crossref query.bibliographic=<arXiv_id>
PMID:12345678 or pubmed.ncbi.nlm.nih.gov/12345678Hit eutils.ncbi.nlm.nih.gov/entrez/eutils/esummary.fcgi?db=pubmed&id=<pmid> → grab articleids[type=doi]
Semantic Scholar paper ID (S2:abc... / corpusId:N)Hit api.semanticscholar.org/graph/v1/paper/<id>?fields=externalIds → grab externalIds.DOI
Anything else non-DOI shapedLeave for B.1 regex pass to ignore; if it survives B.1+B.2 unresolved, drop with skipped (unresolvable_identifier) in the confirm table — do NOT auto-fire requests on garbage

Network preflight: before firing B.0 lookups, do one cheap sanity probe (e.g. HEAD https://api.crossref.org/). If it fails, tell the user "no network — Mode B can't resolve non-DOI identifiers or do discovery; only direct DOI extraction will work" and let them decide to proceed with B.1 only.

Semantic Scholar rate limit: unauthenticated ≤ 1 req/sec. Pace batches; on 429 back off 30s then retry once; on second 429, drop the entry with skipped (ss_rate_limited).

B.1 — DOI extraction

After Step 0 canonicalization:

  1. Regex 10\.\d{4,9}/[^\s,;<>"'{}]+ on the canonicalized input (or file contents). The character class explicitly excludes {} so wrappers like BibTeX doi = {10.x/y} don't leak braces.
  2. Strip trailing punctuation .,;)}"'.
  3. Dedupe via canonical lowercase compare (see flow step 5).
  4. Fallback rule: if regex yields DOIs for less than 50% of the apparent entries (e.g. BibTeX with 20 @article entries but only 5 have ASCII DOIs), send the remaining entries to B.2 title lookup rather than silently dropping them.

B.2 — Title → DOI lookup

For each title:

  1. Query Crossref: GET https://api.crossref.org/works?query.title=<urlencode>&rows=5&mailto=<config crossref.mailto>
  2. Pick top by score. Note: Crossref score is unbounded relevance, NOT 0–100 — absolute thresholds across queries don't compare. Use relative + content rules:
    • If top1.score / top2.score < 1.5 → ambiguous; show user top 3.
    • Always sanity-check matched_title vs input_title by token overlap (or Levenshtein). If overlap < 50%, mark low-confidence regardless of score ratio.
  3. If user gave an author or year ("Smith 2024"), cross-check against the candidate's author[].family and issued.date-parts.
  4. Each B.2 result enters the confirm table as one of:
    • confidence=high (top1, ratio ≥ 1.5, overlap ≥ 50%, any author/year match consistent) — included by default.
    • confidence=low (any of: ratio < 1.5, overlap < 50%, no author match) — excluded by default; user must explicitly pick.
    • unresolved (0 candidates or all rejected) — dropped with skipped (no_match).

B.3 — Discovery from abstract description

Triggered by queries like "Smith 在 Google Scholar 上的文章", "topic Y top 20", "Nature Energy 2023". Do NOT scrape Google Scholar (anti-bot + ToS). Interpret "Scholar" semantically and use the ladder below.

Tool ladder (try in order, use what's available):

  1. Crossref (api.crossref.org/works?query.author= / query.bibliographic= / query.container-title=) — always available, no auth.
  2. OpenAlex (api.openalex.org/works?search= or ?filter=author.id:A...) — free, no auth, broader coverage than Crossref author search, returns DOIs directly.
  3. Semantic Scholar (api.semanticscholar.org/graph/v1/paper/search?query=...) — better for topic / abstract search. Rate limit ~1 req/sec unauthenticated; pace requests, on 429 back off 30s then drop the query on a second 429.
  4. PubMed (bio-research:pubmed MCP) — if biomedical AND the MCP is loaded in the host framework.
  5. WebSearch / Tavily / Exa — last-resort if a web-search-* skill is loaded. Highest hallucination risk; agent MUST round-trip every candidate through Crossref or OpenAlex to verify the DOI exists before accepting.

After discovery: present candidates as a numbered list with title + first-author + year + DOI + source (which API found it). Default top 20; ask if user wants more. User strikes out / picks subset → final list locked.

Open-ended-query clarifier: if the agent's discovery would return more than 50 candidates (e.g. user said "Smith 的所有文章" and the author has 200+ publications), confirm scope with user BEFORE returning — "found 200+; you want all of them, top 20 most cited, or filter by year?".

Mode B flow (after sub-flow resolution)

  1. Resolve input via Step 0 → B.0 → B.1 → B.2 (leftovers + B.1 fallback) → B.3 (for queries).
  2. Consolidate + dedupe. Canonicalize EVERY DOI (lowercase, strip wrappers, no URL prefix, ASCII slash) BEFORE comparing. 10.X/ABC} and 10.x/abc must collapse to one entry.
  3. Confirm with user — FULL TABLE, not just first 5:
    找到 N 个唯一 DOI(去重后)。完整列表:
    
      [ 1] doi=10.xxxx/yyy  source=B.1   confidence=high
           title= ...  author= ...  year= ...
      [ 2] doi=10.zzzz/www  source=B.2   confidence=high
           matched_title= ...  (input: "...")
      [ 3] doi=10.aaaa/bbb  source=B.3   confidence=high
           via=Crossref  (query: "...")
      [ 4] doi=10.cccc/ddd  source=B.2   confidence=LOW
           matched_title= ...  (input: "...")  ← excluded; pick to include
      [ 5] (unresolvable)   source=B.0   from: "PMID:99999"
           ← dropped
      ...
    
    开始下载吗?(y=accept all high-confidence / n=cancel /
    include 4 / exclude 1,3 / show <N> / ...)
    
    Default: download confidence=high rows only. confidence=low excluded unless user explicitly includes. Unresolvables dropped.
  4. Propose project name (context-aware: "组会文献" → groupmtg_<date>; "Smith 综述补充" → smith_review_extras; nothing topical → custom_<date>). Ask user confirm.
  5. Append vs new — if <OUTPUT_DIR>/<project_name>/refs_raw.json exists, ask append / new / rename (default: ask again on any other input — DO NOT default-append). Append rules:
    • Canonicalize new DOIs (Step 0) BEFORE compare.
    • Drop new DOIs already present in existing entries (case-insensitive canonical-form compare).
    • New ids start at id = max(existing_ids) + 1.
    • Never renumber existing entries — validate_refs.py keys its incremental skip on id, renumbering re-assigns prior verified metadata to the wrong DOI. Only verified rows are skipped; failed/pending rows revalidate on re-run.
  6. Build refs_raw.json (heredoc the agent runs):
    import json
    from datetime import datetime
    dois = [...]  # finalized canonical-lowercase list
    start_id = 1  # or max(existing_ids)+1 in append mode
    data = {
      "parent_doi": "",  # empty string for clean report labels;
                         # validate_refs.py reads as raw JSON, null
                         # would also work but "" is preferred.
      "parent_title": f"Custom batch — {user_label}",
      "extracted_at": datetime.now().isoformat(timespec="seconds"),
      "total": len(dois),
      "with_doi": len(dois),
      "without_doi": 0,
      "references": [
        {"id": i, "doi": d,
         "key": "", "unstructured": "",
         "author": "", "year": "", "journal": "",
         "volume": "", "first_page": ""}
        for i, d in enumerate(dois, start=start_id)
      ],
    }
    with open(f"{project_name}/refs_raw.json", "w", encoding="utf-8") as f:
        json.dump(data, f, ensure_ascii=False, indent=2)
    
    All metadata fields stay empty — validate_refs.py fills them from Crossref per DOI on success. Rows whose DOI Crossref can't resolve become status=failed with empty metadata (not partially enriched).
  7. Run pipeline (skip the wrapper):
    cd <OUTPUT_DIR>
    python <SKILL_DIR>/scripts/validate_refs.py <project_name>
    python <SKILL_DIR>/scripts/download_refs.py <project_name> [--auto] [--fail-fast]
    

Mode B pre-flight checklist

  1. Full confirm table approved? User saw EVERY row, not just top 5.
  2. Total count + estimated runtime? Roughly: 0.35s × N for validate, ~30-90s per ref for download (publisher-dependent).
  3. Project name + new/append decided? First-write = new project; second-write = append OR rename.
  4. Edge fully closed? (Edge backend only — cloak skips.)
  5. Config set? [crossref].mailto for polite-pool latency.
  6. Network reachable? (api.crossref.org HEAD probe).

Compatibility with v0.4 flags

  • --auto: works with Mode B (manual_pending refs go to async retry queue same as Mode A).
  • --fail-fast: works with Mode B (stops on first actionable unresolved ref).
  • [user].verified_no_si_dois: works with Mode B (matches by lowercase DOI — independent of how refs_raw.json was produced).
  • CloakBrowser backend (REF_DOWNLOADER_BROWSER=cloak): works with Mode B (browser backend is decided by env var, independent of input mode).

Install prerequisites (before first invocation)

The skill protocol can't manage Python deps. If python -c "import playwright" fails, the user needs:

cd "<SKILL_DIR>"
pip install playwright pymupdf
playwright install msedge          # downloads Edge driver
cp config.example.toml config.local.toml   # then user edits [crossref].mailto

If the user is developing from the source repo instead of an installed skill copy, they can also install from the repo root with pip install -r requirements.txt -r requirements-dev.txt.

Alternative backend: CloakBrowser (Cloudflare-heavy sites)

Default backend is Microsoft Edge. For sites that keep blocking ordinary Playwright (Cloudflare Turnstile, Radware, persistent Just a moment / 安全验证 pages), switch to the CloakBrowser stealth Chromium backend.

What CloakBrowser is. Third-party MIT-licensed Python package by CloakHQ (github.com/CloakHQ/CloakBrowser, pypi:cloakbrowser). Ships a patched Chromium build with anti-fingerprint changes. Its launch_persistent_context_async() is Playwright-API-compatible, which is why ref-downloader can swap it in with one env var. NOT a dependency of ref-downloader — if the user doesn't pip install cloakbrowser, it's never imported and the default Edge path runs as normal. Beta software; user installs it at their own discretion.

# One-time setup (separate from ref-downloader's `pip install playwright pymupdf`)
pip install cloakbrowser

# Switch backend (env vars; no CLI flag changes)
$env:REF_DOWNLOADER_BROWSER = "cloak"
$env:REF_DOWNLOADER_CLOAK_HUMAN_PRESET = "careful"   # optional: slower mouse/scroll
# Optional overrides:
# $env:REF_DOWNLOADER_CLOAK_PROFILE = "<custom path>"  # default: ~/.local/cloakbrowser/profiles/ref-downloader
# $env:REF_DOWNLOADER_CLOAK_PROXY = "http://..."
# $env:REF_DOWNLOADER_CLOAK_GEOIP = "1"
# $env:CLOAKBROWSER_PYTHONPATH = "<dev source>"      # sys.path hint if cloakbrowser is checked out, not pip-installed

python "<SKILL_DIR>/scripts/download_refs.py" <PROJECT_NAME>

Caveats:

  • cloakbrowser uses its own Chromium with a separate profile — Edge does NOT need to be closed.
  • The separate profile means your institutional cookies are NOT carried — best for open-Cloudflare sites, less useful for paywalled refs your institution licenses.
  • A fresh cloak profile may still show Cloudflare/security-verification pages on first visit; warm it manually by opening the target site once with the same REF_DOWNLOADER_CLOAK_PROFILE and finishing any verification before running the downloader.
  • human_preset=careful lowers behavior-detection trigger rates but is not a captcha solver.

Output layout

Same for both modes:

<OUTPUT_DIR>/
├── <PROJECT_NAME>/
│   ├── refs_raw.json           # extract_refs.py output (Mode A) or
│   │                           # hand-built JSON (Mode B)
│   ├── refs_validated.json     # validate_refs.py output
│   ├── download_report.csv     # per-ref status (only on graceful
│   │                           # completion; OVERWRITTEN each run —
│   │                           # NOT historical truth)
│   ├── *.pdf                   # reference PDFs
│   └── *_SI.pdf                # supplementary files (where supported)
└── runs/<timestamp>-round-03/
    └── events.jsonl            # full event trace per ref
                                # (append-only across runs;
                                #  THIS is the authoritative history)

Interruption note: if the run is interrupted (Ctrl+C / Edge crash / VPN drop), the root download_report.csv may be stale. Trust the latest runs/<timestamp>/events.jsonl + actual files in <PROJECT_NAME>/.

Common failure modes

Status / symptomMeaningAction
manual_pending (auth_redirect)Bounced to institution SSOUser signs in via live Edge tab; re-run (incremental skips done refs)
manual_pending (challenge_timeout)Cloudflare / publisher challenge unsolved in timeRe-run interactively; solve captcha when prompted
manual_pending (elsevier_crasolve_shell)Elsevier viewer stuck in transitionIn --auto mode the async retry queue picks it up ~60s later; in interactive mode the hot-session retry usually catches it, else manual click in live page
failed (auto)Generic auto path failedCheck events.jsonl for that ref; may need a publisher-specific patch
ignored (ignored_institution_access)DOI listed in [institution].ignored_access_doisSkip-by-design; remove from config to retry
Edge won't launchBackground msedge.exe still holding profileKill all msedge.exe in Task Manager, re-run (cloak backend skips this)
ModuleNotFoundError: playwrightInstall prereqs not doneSee "Install prerequisites" section above
WARNING: crossref.mailto is the placeholderFirst-run config uncustomizedEdit <SKILL_DIR>/config.local.toml → set [crossref].mailto to a real email (Crossref polite pool)
Mode B: Step 0 left full-width slash unconvertedCanonicalization bugVerify Step 0 ran before regex; flag for design fix
Mode B: BibTeX doi = {10.x/y} left trailing } in refs_raw.jsonStep 0 bypassedStep 0 MUST run before B.1 regex
Mode B: Crossref title query 0 hits for a B.2 rowTitle couldn't matchDrop entry as unresolved; suggest user provide author/year/journal
Mode B: B.3 abstract query returns 0 across all ladder stepsNo matches foundSuggest user narrow (add author / year / journal); or accept that no papers match
Mode B: B.3 discovery returns 200+ candidatesQuery too broadAsk user to scope (year range / top-N by citations / specific journal) BEFORE listing
Mode B: run_ref_downloader.py invoked accidentallyIt assumes parent DOI — will failDirect validate_refs.py + download_refs.py invocation only
Mode B: Semantic Scholar 429Unauthenticated rate limit hitBack off 30s, retry once; on second 429, drop with skipped (ss_rate_limited)
Mode B: PMID lookup returns no articleids[type=doi]PubMed has no DOI for this entryEntry has no DOI; tell user, suggest alternative identifier
Mode B: Input is purely conversational ("上次那 5 篇")No resolvable contentRefuse; ask user to repaste / attach file
Mode B: No network reachableB.0 / B.2 / B.3 all need networkTell user before starting; only B.1 viable

Manual / debug mode

If the wrapper fails partway (Mode A), run the 3 scripts standalone for partial re-execution:

python <SKILL_DIR>/scripts/extract_refs.py <DOI>           # → refs_raw.json
python <SKILL_DIR>/scripts/validate_refs.py <PROJECT>      # → refs_validated.json
python <SKILL_DIR>/scripts/download_refs.py <PROJECT>      # → PDFs + download_report.csv

Mode B uses the same standalone invocation, just skipping extract_refs.py (the agent builds refs_raw.json directly).

Full 8-step manual flow with code snippets, DOI-resolution fallback chain (Zotero query → fitz text → user prompt), and the procedure for extending PUBLISHER_MAP when encountering an unknown DOI prefix → references/agent-runbook.md.

See also

发现问题?提交给管理员复核

评分:

评论 (0)

暂无评论,成为第一个评论者吧!