SkillAtlasSkill 详情

sn-search-academic

The SenseNova model family plugs directly into agent runtimes such as OpenClaw and hermes-agent...

审核状态:已审核Quality 72Security 70

复制安装命令

用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。

复制前请先查看来源、License 和安全提示。

项目 README

来源文件:README.md

抓取于 2026年8月3日

SenseNova-Skills

English | 简体中文

Website Raccoon Token Plan SenseNova U1 SenseNova 6.7

The SenseNova model family plugs directly into agent runtimes such as OpenClaw and hermes-agent, with the skills in this repository extending the models with concrete, end-to-end office capabilities.

In this repository each skill lives in its own directory and declares triggers, capabilities, and execution flow through a SKILL.md file, following the Agent Skills convention.

The skills cover image generation & visualization, slide-deck (PPT) generation, Excel data analysis, and deep research — usable standalone or composed into end-to-end workflows.

🎨 Want to see what it can do? Check out our sn-infographic Gallery to explore nearly 100 stunning generation cases and steal their prompt designs !

🦝 Available out-of-the-box in Raccoon

The latest SenseNova models and the full Cowork-Skill suite in this repo are bundled into Raccoon, with enterprise-grade security and a zero-setup experience — if you'd rather not provision env, API keys, and runtimes yourself, you can use these capabilities directly through Raccoon. Free trial available — no payment required to get started.

Raccoon now ships a full upgrade across product capability and client experience:

  • Three core office capabilities, strengthened: powered by SenseNova 6.7 Flash + Cowork-Skill, data analysis, PPT generation, and task planning each take a step up — covering the full loop from multi-file cleaning/analysis to formal report decks, industry/competitive research, and investment memos.
  • New: infographic generation: built on the SenseNova U1 model, compresses complex data, long reports, and business insights into dense, structured, visual infographics that are easier to digest and share.
  • New client + local Agent OS: the cloud model handles heavy reasoning and multimodal understanding; the local Agent OS sits next to your files, work context, and personal habits — delivering a more personalized, local, and secure AI-native office experience.
  • Proven at scale: chosen by 15M+ individual users and thousands of enterprise customers.

👉 Try it: xiaohuanxiong.com

How to Use

These skills are designed to run inside an Agent Skills-compatible agent.

Recommended: let the agent install the skills for you. Hand it the repo URL and ask it to clone and drop the skills into the right directory — for example:

"Please install SenseNova-Skills from https://github.com/OpenSenseNova/SenseNova-Skills into your skills directory."

After it finishes, you may need to manually restart the agent service before the new skills are picked up.

AgentTarget directory
OpenClaw~/.openclaw/skills/
hermes-agent~/.hermes/skills/
Prefer to install manually?

Clone this repository, then copy the subdirectories under skills/ into the target directory yourself:

git clone https://github.com/OpenSenseNova/SenseNova-Skills.git --depth=1
mkdir -p ~/.openclaw/skills
cp -r SenseNova-Skills/skills/* ~/.openclaw/skills/

For Hermes, swap the target to ~/.hermes/skills/.

Per-category Python dependencies, API keys, and invocation examples are documented in the 📖 Full guide for each section.

Skills List

🎨 Image & Visualization

📖 Full guide: docs/sn-image-generate_en.md (prerequisites, Quick Start, API config, and invocation samples).

NameLabelDescription
sn-image-doctorEnvironment DoctorValidates the SenseNova-Skills environment — checks sn-image-base install, Python deps, and required env vars; interactively fills missing values into .env.
sn-image-baseImage Base Layer (Tier 0)Low-level tools — text-to-image (sn-image-generate), image recognition (sn-image-recognize), and text optimization (sn-text-optimize) — exposed through a unified sn_agent_runner.py, designed to be called by upper-layer skills.
sn-infographicInfographic Generation (Tier 1)Auto prompt-quality scoring, layout/style selection (87 layouts / 66 styles), multi-round generation with VLM review and quality ranking, producing publication-ready infographics.
sn-image-imitateImage Imitation (Tier 1)Given one reference image and a target content prompt, generates a new image that imitates the reference.
sn-image-resumeResume Image Generation (Tier 1)Given resume information, generates a resume image.

📊 Presentations (PPT)

📖 Full guide: docs/sn-ppt-generate.md (prerequisites, Quick Start, API config, and invocation samples).

NameLabelDescription
sn-ppt-entryPPT Entry PointUnified entry point for PPT generation. Asks the user to choose fast, standard, or creative mode, then collects role / audience / scenario / page count. For standard mode, also asks about image sourcing (AI, web search, or none) and chart rendering (U1 infographics or ECharts). Parses uploaded pdf / docx / md / txt, emits task_pack.json + info_pack.json, and dispatches to the chosen mode.
sn-ppt-doctorPPT Environment DoctorEnvironment check for the PPT pipeline — validates sn-image-base, API keys, the Node runtime, and optional deps; writes missing required vars into .env.
sn-ppt-creativePPT Creative ModeOne full-page 16:9 PNG per slide, generated via sn-image-generate with a per-page composed prompt. Falls back to web image search when T2I generation fails.
sn-ppt-standardPPT Standard & Faststyle_spec → outline → asset plan + per-slot images + VLM QC → per-page HTML → per-page review → PPTX export. Fast mode builds a complete draft immediately with autonomous decisions, then provides structured refinement suggestions. Supports AI-generated infographics (U1) for diagrams and web image search (Serper) for real photos.

📈 Data Analysis (DA)

📖 Full guide: docs/sn-data-analysis.md (prerequisites, Quick Start, API config, and invocation samples).

NameLabelDescription
sn-da-excel-workflowExcel Analysis OrchestrationEnd-to-end Excel pipeline — multi-sheet read, large-file detection (≥10k rows triggers Parquet), cleaning, conditional filtering, cross-sheet aggregation, and Excel/CSV export.
sn-da-image-captionImage Understanding & Data ExtractionFor image-first inputs — table OCR, chart understanding, screenshot/UI description; parses captions into DataFrames, recreates visualizations, exports Excel/CSV.
sn-da-large-file-analysisHigh-Performance Large-File AnalysisStreaming reads for ≥10k-row Excel datasets (openpyxl read_only + iter_rows), Parquet conversion, memory optimization, chunked processing, large-file writes.

🔬 Deep Research

📖 Full guide: docs/sn-deep-research.md (prerequisites, web_search precheck, Quick Start, and per-stage invocation).

NameLabelDescription
sn-deep-researchDeep Research Entry PointUnified deep-research orchestrator with true-dependency DAGs, reusable source snapshots, and evidence-informed content units, producing final report.md.
sn-research-reportFinal Report Writing & EditingRenders the judgment layer into the final report.md; also handles targeted rewrites — restructuring, polishing, table-augmentation — for an existing draft.
sn-report-format-discoveryPresentation-Format DiscoveryCompares final forms such as a research report, academic paper, table-first analysis, decision memo, or a custom Markdown form; scout uses it before research and user confirmation.
sn-prepare-citationsCitation RenderingPost-processes [^source_id] footnotes into numbered citations and appends references from evidence sources.
sn-md-to-html-reportMarkdown → HTML ReportConverts the research report.md (or any Markdown doc) into a clean, single-file HTML reading view that opens offline — embedded images, side-panel TOC, responsive tables, and table-delimiter repair.

🔍 Search

📖 Search skills are documented together with deep research: docs/sn-deep-research.md (includes per-platform API keys, invocation, and unified JSON output).

NameLabelDescription
sn-search-academicAcademic SearchArXiv (with section-level HTML reading) / Semantic Scholar (with citation counts) / PubMed (with PMC open-access full text) / Wikipedia, in one aggregated interface.
sn-search-codeDeveloper SearchGitHub (repo / code / issue) / Stack Overflow / Hacker News / HuggingFace (models / datasets / spaces), aggregated.
sn-search-social-cnChinese Social SearchBilibili / Zhihu / Douyin search; some platforms require cookie auth.
sn-search-social-enEnglish Social SearchReddit / Twitter (X) / YouTube search.

Sample Outputs

🎨 Infographic (sn-infographic)

A few sn-infographic outputs (more in docs/sn-infographic-examples.md).

sn-infographic sample outputs

🧩 Memory price analysis — insight → analysis → presentation → end-to-end workflow

examples/memory-price-end2end-analysis. Starting from a raw quote CSV, the agent profiles fields, normalizes categories and timestamps, then attacks the rally from three angles — overall trend, top movers per category, and the gap between server-grade and consumer-grade SKUs — locating a late-February inflection along the way. Treating those findings as the research question, it switches to deep research: planning per-dimension web searches over supply contraction, AI-server demand, and vendor output discipline, then triaging and cross-checking evidence across sources before committing it to the report. The data and research conclusions are then handed to PPT generation, which lays out a 16-page outline, plans per-slot imagery, renders per-page HTML, runs VLM review, and finally composites screenshots into the PPTX. The result is a clear three-step storyline: prices are rising → here is why → here is what to do. This is the only example that exercises the full data analysis → deep research → PPT chain end-to-end.

📊 Employee performance analysis — data analysis

examples/employee-performance-analysis. The agent reads 10 separate monthly review xlsx files, aligns column schemas across months and joins them into one longitudinal table. From that table it produces aggregate views — monthly average trend, score-distribution boxplots, grade mix change, and a 38-role ranking — and individual views — top performers, needs-attention, and consistently-improving cohorts plus per-employee year trends. The findings are written up with explicit improvement suggestions tied to specific roles and individuals, backed by 8 supporting charts. The same content is delivered as a Word doc (for distribution) and a visualized HTML report (for browsing). The example shows how sn-da-excel-workflow handles "many small spreadsheets that should be one analysis" rather than a single big file.

🔬 Embodied AI industry research — deep research

examples/embodied-ai-deep-research. Given only an industry name, the agent first commits to a research plan — market size, vendor share, financing, cost structure, development roadmap — instead of jumping straight into search. For each dimension it runs targeted web searches, fetches and reads source pages, and extracts both numeric and qualitative evidence; conflicting figures across sources are explicitly reconciled before being trusted. A synthesis stage organizes per-dimension evidence into a traceable, reader-oriented information structure rather than a stack of disconnected bullets. The output is an illustrated report (Markdown + visualized HTML) with 5 dimension-specific charts. The example shows how sn-deep-research turns "go research X" into a structured plan-then-execute loop with traceable evidence.

🎯 Property fee pricing — PPT generation

examples/property-fee-pricing-ppt. The agent takes a free-form brief — topic (property fee pricing), audience (property staff + committee), 26 pages, black-and-white warm style — and first commits to an outline plus a per-page asset plan that conforms to the style spec. Each slide is then built as semantic per-page HTML rather than free-form image generation: copy, layout, illustrations, icons, and any data charts are reasoned about per slot. Imagery is produced or selected per slot and VLM-checked against the page's intent; each rendered page goes through a review pass with optional rewrite for coherence and copy quality. Final pages are screenshotted and composited into the PPTX, with the per-page HTML kept alongside for direct browser preview or re-editing. The example demonstrates sn-ppt-standard style consistency on a long, prose-heavy deck where every slide must obey the same audience and palette constraints.

FAQ

Common setup and runtime questions (400/401 errors, rate limits, PPT timeouts, infographic quality, model names) are answered in docs/faq.md.

Contributing

Feel free to use the skills here as templates for your own OpenClaw skills. The qualities that make a skill good:

  • Clear triggers: state in description exactly when the skill should and should not run, so the agent recognizes it accurately
  • Focused scope: each skill does one thing well; complex workflows compose multiple skills
  • Solid documentation: examples, artifact contracts, edge cases, failure handling
  • Supporting resources: use references/, scripts/, prompts/ to provide additional context

Join the Community

Join our growing community to share feedback, get support, and stay updated on the latest developments. Scan the QR code below to hop into the chat — we'd love to hear from you!

DiscordLark Group

License

MIT — see LICENSE.

研究与检索

中风险

  • 来源需自行核对维护者身份。
  • 包含脚本或命令调用,安装前请复核。
  • 可能需要外部 token、网络权限或第三方服务。
  • 未检测到高风险命令。
  • 扫描发现:4 条。

Codex — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/OpenSenseNova/SenseNova-Skills.git
  3. 将 "skills/sn-search-academic" 文件夹复制到 Codex 的 skills 目录中。
  4. 重启 Codex 让新的 skill 生效。

Codex — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Codex 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Codex 让新的 skill 生效。

Claude Code — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/OpenSenseNova/SenseNova-Skills.git
  3. 将 "skills/sn-search-academic" 文件夹复制到 Claude Code 的 skills 目录中。
  4. 重启 Claude Code 让新的 skill 生效。

Claude Code — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Claude Code 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Claude Code 让新的 skill 生效。

Cursor — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/OpenSenseNova/SenseNova-Skills.git
  3. 将 "skills/sn-search-academic" 文件夹复制到 Cursor 的 skills 目录中。
  4. 重启 Cursor 让新的 skill 生效。

Cursor — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Cursor 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Cursor 让新的 skill 生效。

GitHub Copilot — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/OpenSenseNova/SenseNova-Skills.git
  3. 将 "skills/sn-search-academic" 文件夹复制到 GitHub Copilot 的 skills 目录中。
  4. 重启 GitHub Copilot 让新的 skill 生效。

GitHub Copilot — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 GitHub Copilot 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 GitHub Copilot 让新的 skill 生效。

Windsurf — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/OpenSenseNova/SenseNova-Skills.git
  3. 将 "skills/sn-search-academic" 文件夹复制到 Windsurf 的 skills 目录中。
  4. 重启 Windsurf 让新的 skill 生效。

Windsurf — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Windsurf 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Windsurf 让新的 skill 生效。
查看 SKILL.md 原文
name: sn-search-academic
description: 用于学术调研、论文精读、相关工作梳理、百科知识查询和引用链追溯。

sn-search-academic - 学术搜索

凭证配置

API key、token 与 cookie 统一建议写在仓库根目录 .env(参考 .env.example),并由 runtime 或用户在执行前加载为同名环境变量。脚本仍只从环境变量或显式 CLI 参数读取凭证;不要把真实密钥写入 skill payload、报告、日志或提交。

使用三个统一入口完成学术调研:

  • search.py:搜索论文和百科条目
  • paper.py:列出论文章节,读取论文全文或指定章节
  • refTree.py:查询论文的 references 和 citations

不要直接调用历史 provider 脚本;它们只是统一入口的内部实现细节。 需要 provider 回退链、参数分发或完整输出字段时,按需读取 references/search.md、references/paper.md、references/refTree.md。

可用脚本

脚本用途主要输入主要输出
scripts/search.py搜索论文/百科query,可选 --source、--limit、--category、--lang按 source 分组的论文/百科条目,位于 source_results[*].items
scripts/paper.py列出章节,读取论文全文或章节论文 ID,可选 --source、--list_section、--section章节列表位于 sections;全文或章节正文位于 content
scripts/refTree.py查询引用树--paper_id、--title,可选 --direction参考文献与被引论文,位于 source_results[*].references / source_results[*].citations

执行约定

本技能的 scripts/...、requirements.txt、references/... 路径均相对本 skill 目录;若当前工作目录不同,先解析为绝对路径,不要依赖 ${SKILL_DIR} 运行时变量。

调用约定:

  • 不要并行启动多个本技能脚本;search.py 和 refTree.py 内部已经处理并发、超时和 provider 回退链。
  • 长结果优先加 --output <path> 写入文件,再读取必要字段,避免终端输出过长。
  • --provider-timeout 表示单个 provider 超时;默认使用脚本内置超时。

依赖

首次运行或脚本提示缺库时,使用本技能的依赖清单安装到当前 Python 环境:

python3 -m pip install -r requirements.txt

不要在脚本内部自动安装依赖。若安装失败、网络不可用或包不可用,停止使用对应脚本并改用 WebSearch/browser-use,说明缺少依赖。

Crawler 回退还需要额外运行时环境:

python3 -m playwright install firefox

arxiv_crawler_search.py 和 semantic_scholar_crawler_refTree.py 还需要 Node.js,以及某个当前目录或祖先目录中已安装 camoufox-js 的 node_modules。缺少这些环境时,不要尝试绕过;改用非 crawler provider 或网页搜索。

参数说明

search.py

统一搜索入口。默认搜索所有支持的 source,并按 source 分组返回结果。

python3 scripts/search.py <query> [选项]
参数说明默认值
query搜索关键词,必填位置参数-
--source, --sources, -s搜索源;支持重复传参或逗号分隔all
--limit, -n每个 source 返回数量10
--category, -cArXiv 分类过滤,只传给支持分类的 source-
--lang, -l语言提示,只传给支持语言参数的 source-
--output, -o将最终 JSON 写入文件-
--provider-timeout每个 provider 的超时时间,单位秒;0 表示不限制60

支持的 --source:

  • all
  • arxiv
  • semantic
  • google_scholar
  • pubmed
  • wikipedia

示例:

python3 scripts/search.py "retrieval augmented generation" --limit 5
python3 scripts/search.py "diffusion model" --source arxiv,semantic --category cs.CV --limit 5
python3 scripts/search.py "阿尔茨海默病 多模态诊断" --source pubmed,wikipedia --lang zh --limit 5
python3 scripts/search.py "agentic memory" --source all --limit 8 --output results/search.json

paper.py

统一论文阅读入口。默认按 arXiv 论文读取;读取 PMC 论文时显式传 --source pmc。不确定章节名时先用 --list_section 列出可用章节,再用 --section 精读。

python3 scripts/paper.py <id> [选项]
参数说明默认值
id论文 ID。arXiv 支持原始 ID、arXiv: 前缀、abs/pdf URL;PMC 支持 PMC11119143、11119143、PMC URL-
--source论文来源:arxiv 或 pmcarxiv
--section, -s读取指定章节;不填则读取全文-
--list_section, --list-section列出论文可用章节,不返回正文;不能和 --section 同时使用false
--output, -o将最终 JSON 写入文件-

示例:

python3 scripts/paper.py 2603.00729
python3 scripts/paper.py 2603.00729 --list_section
python3 scripts/paper.py arXiv:2603.00729 --section introduction
python3 scripts/paper.py 2603.00729 --section method --output results/paper-method.json
python3 scripts/paper.py PMC11119143 --source pmc
python3 scripts/paper.py PMC11119143 --source pmc --list-section
python3 scripts/paper.py PMC11119143 --source pmc --section results

refTree.py

统一引用树入口。--paper_id 和 --title 都必填;标题用于回退时精确匹配。

python3 scripts/refTree.py --paper_id <paper_id> --title <title> [选项]
参数说明默认值
--paper_id论文 ID:Semantic Scholar ID、DOI、ArXiv ID、PMID 等-
--title论文标题,必填-
--direction查询方向:references 或 citations;不填则两者都查-
--source, --sources, -s引用树 source;当前支持 all、semanticall
--limit, -n每个 source、每个 direction 返回数量10
--api-keySemantic Scholar API 密钥,可选-
--provider-timeout每个 provider 的超时时间,单位秒;0 表示不限制60
--output, -o将最终 JSON 写入文件-

注意:参数名是 --paper_id,不是 --paper-id;paper_id 不支持位置参数。

示例:

python3 scripts/refTree.py --paper_id "2309.16609" --title "Qwen Technical Report"
python3 scripts/refTree.py --paper_id "2309.16609" --title "Qwen Technical Report" --direction references --limit 20
python3 scripts/refTree.py --paper_id "10.1038/s41586-024-07487-w" --title "AlphaFold 3" --direction citations
python3 scripts/refTree.py --paper_id "2309.16609" --title "Qwen Technical Report" --output results/refTree.json

输出格式

所有脚本都输出 JSON。先看顶层 success;失败时读取 error、errors 和 attempts 判断是无结果、超时还是 provider 失败。

search.py 输出

CLI 输出的顶层不包含 items,论文条目在 source_results[*].items 中:

{
  "success": true,
  "query": "retrieval augmented generation",
  "provider": "search.py",
  "sources": ["arxiv", "semantic"],
  "source_results": [
    {
      "source": "arxiv",
      "success": true,
      "provider": "arxiv_official",
      "items": [
        {
          "source": "arxiv",
          "provider": "arxiv_official",
          "title": "Example title",
          "abstract": "Example abstract",
          "citation_count": null,
          "arxiv_id": "2301.00001",
          "url": "https://arxiv.org/abs/2301.00001"
        }
      ],
      "attempts": [],
      "error": null
    }
  ],
  "errors": [],
  "error": null
}

常用 item 字段:

  • 通用:title、abstract、snippet、url、citation_count、doi
  • arXiv:arxiv_id、pdf_url、categories
  • Semantic Scholar:paper_id、venue、year
  • PubMed:pmid、pmc_id、journal、pub_date
  • Wikipedia:page_id、word_count、section_title

paper.py 输出

默认读取全文;指定 --section 时读取章节。正文在顶层 content:

{
  "success": true,
  "source": "arxiv",
  "provider": "arxiv_html",
  "arxiv_id": "2603.00729",
  "section": "introduction",
  "content": "<全文或章节正文>",
  "char_count": 12345,
  "attempts": [],
  "error": null
}

指定 --list_section 时只返回章节结构,不返回 content:

{
  "success": true,
  "source": "arxiv",
  "provider": "arxiv_html",
  "arxiv_id": "2603.00729",
  "section_count": 2,
  "sections": [
    {"name": "Abstract", "level": 0},
    {"name": "1 Introduction", "level": 1}
  ],
  "attempts": [],
  "error": null
}

常用字段:

  • arXiv:arxiv_id、title、abs_url、html_url、pdf_url、section_count、sections
  • PMC:pmc_id、pmid、title、pmc_url、section_count、sections
  • 指定 --list_section 时返回 sections 和 section_count,不包含 content
  • 指定 --section 时会包含 section;不指定 --list_section / --section 时读取全文

refTree.py 输出

引用树结果在 source_results[*].references 和 source_results[*].citations:

{
  "success": true,
  "id": "2309.16609",
  "title": "Qwen Technical Report",
  "provider": "refTree.py",
  "direction": "all",
  "source_results": [
    {
      "source": "semantic",
      "success": true,
      "provider": "semantic_official",
      "references": [
        {
          "title": "Example reference",
          "abstract": "Example abstract",
          "citation_count": 128,
          "paper_id": "example-reference-id",
          "arxiv_id": "2301.00001"
        }
      ],
      "citations": [
        {
          "title": "Example citing paper",
          "abstract": "Example abstract",
          "citation_count": 42,
          "paper_id": "example-citing-id",
          "doi": "10.1234/example"
        }
      ],
      "attempts": [],
      "error": null
    }
  ],
  "errors": [],
  "error": null
}

如果使用 --output,三个脚本都会在 JSON 中额外加入 output_path。

并发与限流约定

这些脚本会访问外部学术服务,必须控制请求频率。 执行本技能脚本时:

  • 不要并发运行多个搜索脚本。
  • 不要使用并行工具同时调用多个 python3 scripts/... 命令。
  • 一次只运行一个脚本命令,等待结果返回后再运行下一个。
  • 批量查询时,优先使用脚本自带的 --limit、--id-list 等参数,而不是启动多个进程。
  • 如果需要连续调用,按顺序执行,并在必要时等待数秒。

全文阅读工作流

搜索结果只有摘要时,用 paper.py 先列章节,再补充全文或关键章节。

  1. 先用 search.py 搜索,优先从 source_results[*].items 里记录 title、arxiv_id、pmc_id、paper_id、doi、citation_count。
  2. 如果条目有 arxiv_id,先用 python3 scripts/paper.py <arxiv_id> --source arxiv --list_section 查看章节;再用 --section <section> 精读。
  3. 如果条目有 pmc_id,先用 python3 scripts/paper.py <pmc_id> --source pmc --list_section 查看章节;再按需读 --section <section> 。
  4. 如果需要整体理解,再不带 --section / --list_section 读取全文。
  5. 全文很长时使用 --output results/paper.json,再读取 content、sections、char_count 等字段。

引用追溯工作流

通过论文的引用关系发现关键词搜索覆盖不到的相关工作。

通过 references 找奠基工作,通过 citations 找后续进展。refTree.py 需要同时传论文 ID 和标题。

后向追溯(找奠基工作):

  1. 关键词搜索找到高相关论文 → 取其 paper_id 或 arxiv_id 和 title
  2. refTree.py --paper_id "<id>" --title "<title>" --direction references --limit 20 → 找到高引参考文献
  3. 筛选与研究问题相关的条目 → 用 paper.py深入阅读

前向追踪(找后续进展):

  1. 找到领域奠基论文或关键论文 → 取其 ID
  2. refTree.py --paper_id "<id>" --title "<title>" --direction citations --limit 20 → 找到近期高引跟进工作
  3. 筛选与研究问题相关的条目 → 用 paper.py深入阅读

引用链:构建演化路径

  1. 从种子论文 A 出发 → backward 找到 A 的关键参考文献 B
  2. 从 B 出发 → forward 找到引用 B 的后续工作(可能发现 A 没引用的相关论文 C)
  3. 形成 B → A → ... 和 B → C → ... 的知识脉络

主工作流

严格遵循本工作流去执行学术搜索的全流程

  1. 在提供的学术平台选择所有可能的平台搜索学术文献
  2. 如果摘要不足或论文高度相关时,列出章节,尝试读取论文章节或全文,判断论文和搜索需求的相关性
  3. 选择相关性高的论文,搜索它的参考文献和被引(使用引用追溯工作流)。
  4. 选择引用树中 高引用的文献,执行步骤 2、步骤 3。
  5. 重复以上步骤,进行多轮搜索,尽可能多的进行搜索。
  6. 当文献数量、引用链和全文证据足够支撑回答时停止搜索,并在结论中说明主要依据。

ArXiv 分类速查

顶层领域可直接用(如 --category cs),子分类更精确(如 --category cs.AI)。

领域分类代码说明
计算机科学cs.AI人工智能
cs.LG机器学习
cs.CL计算语言学 / NLP
cs.CV计算机视觉
cs.IR信息检索
cs.RO机器人
cs.SE软件工程
cs.DC分布式/并行计算
cs.NI网络与互联网
cs.CR密码学与安全
cs.DB数据库
cs.HC人机交互
统计stat.ML统计机器学习
stat.AP应用统计
stat.ME统计方法论
数学math.OC优化与控制
math.ST统计理论
math.CO组合数学
物理physics物理(全类)
cond-mat凝聚态物理
quant-ph量子物理
hep-th高能理论物理
经济/金融econ.GN经济学综合
q-fin.CP计算金融
q-fin.ST统计金融
生物/医学q-bio.NC神经科学
q-bio.GN基因组学
q-bio.QM定量方法

发现问题?提交给管理员复核

评分:

评论 (0)

暂无评论,成为第一个评论者吧!