SkillAtlasSkill 详情

paper-orchestra

A pluggable skill pack that lets any coding agent in Claude Code, Cursor,

审核状态:已审核Quality 80Security 80

复制安装命令

用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。

复制前请先查看来源、License 和安全提示。

项目 README

来源文件:README.md

抓取于 2026年8月25日

PaperOrchestra

A pluggable skill pack that lets any coding agent in Claude Code, Cursor, Antigravity, Cline, Aider, OpenCode, etc. which can run the PaperOrchestra multi-agent pipeline for turning unstructured research materials into a submission-ready LaTeX paper.

Song, Y., Song, Y., Pfister, T., Yoon, J. PaperOrchestra: A Multi-Agent Framework for Automated AI Research Paper Writing. arXiv:2604.05018, 2026. https://arxiv.org/pdf/2604.05018

PaperOrchestra paper — first page preview
Click to read the paper on arXiv

Why this exists

The paper defines a five-agent pipeline

  • Outline
  • Plotting
  • Literature Review
  • Section Writing
  • Content Refinement

that substantially outperforms single-agent and tree-search baselines on the PaperWritingBench benchmark (50–68% absolute win margin on literature review quality; 14–38% on overall quality). The paper ships the exact prompts for every agent in Appendix F.

This repo turns those prompts, schemas, halt rules, and verification pipelines into a set of host-agent-executable skills. There are no API keys, no SDK dependencies, no embedded LLM calls. The skills are instruction documents plus deterministic helpers; your coding agent does all LLM reasoning and web search using its own tools.

image

How skills work here

Each skill is:

  • SKILL.md — a dense instruction document the host agent reads and follows.
  • references/ — reference material: verbatim paper prompts (Appendix F), JSON schemas, rubrics, halt rules, example outputs.
  • scripts/ — purely deterministic local helpers: JSON schema validation, Levenshtein fuzzy matching, BibTeX formatting, dedup, LaTeX sanity checks, coverage gates. No network, no LLM, no API keys.

Everything else (LLM reasoning, web search, Semantic Scholar lookups, LaTeX compilation) is delegated to the host agent by instruction. See skills/paper-orchestra/references/host-integration.md for per-host invocation (Claude Code, Cursor, Antigravity, Cline, Aider).

The seven skills

SkillPaper step# LLM callsRole
paper-orchestraorchestrator—Top-level driver. Coordinates the other six.
outline-agentStep 11Idea + log + template + guidelines → structured outline JSON (plotting plan, lit review plan, section plan).
plotting-agentStep 2~20–30Execute plotting plan; render plots & conceptual diagrams; optional VLM-critique refinement loop; caption everything.
literature-review-agentStep 3~20–30Web-search candidates; Semantic Scholar verify (Levenshtein > 70, cutoff, dedup); draft Intro + Related Work with ≥90% citation integration.
section-writing-agentStep 41One single multimodal call: draft remaining sections, build tables from experimental log, splice figures.
content-refinement-agentStep 5~5–7Simulated peer review; accept/revert per strict halt rules; safety constraints prevent gaming the evaluator.
paper-writing-bench§3—Reverse-engineer raw materials (Sparse/Dense idea, experimental log) from an existing paper to build benchmark cases.
paper-autoratersApp. F.3—Run the paper's own autoraters: Citation F1 (P0/P1), LitReview quality (6-axis), SxS paper quality, SxS litreview quality.

Steps 2 and 3 run in parallel (see skills/paper-orchestra/references/pipeline.md).

agent-research-aggregator (optional)

A pre-pipeline skill that bridges the gap between scattered AI coding-agent history and the structured (idea.md, experimental_log.md) inputs that PaperOrchestra expects. If you have been running experiments through Claude Code, Cursor, Antigravity, or OpenClaw — but never wrote up a clean experiment log — this skill does that extraction for you.

It is optional. If workspace/inputs/idea.md and workspace/inputs/experimental_log.md already exist, the skill skips itself and the pipeline proceeds directly. It only runs when the inputs are missing or when you explicitly point an agent at a directory.

The simplest way to use it: just tell your agent the folder. If you have a directory (a project root, an agent cache, any folder with research notes), the aggregator figures out what's inside and structures it for PaperOrchestra. The first thing it does is aggregate — scanning, extracting, and synthesising — so even if the data is scattered across multiple files and formats, it produces clean, reviewable inputs before anything gets written.

Run it before paper-orchestra (or let paper-orchestra call it automatically when inputs are missing).

What it does

[.claude/]  [.cursor/]  [.antigravity/]  [.openclaw/]
      │            │              │               │
      └────────────┴──────────────┴───────────────┘
                        │
                Phase 1: Discovery  (deterministic)
                        │
                Phase 2: Extraction (LLM — per batch)
                        │
                Phase 3: Synthesis  (LLM — one call)
                        │
                Phase 4: Formatting (deterministic)
                        │
             ┌──────────┴──────────┐
      workspace/inputs/      workspace/ara/
        idea.md                aggregation_report.md
        experimental_log.md    discovered_logs.json
                               raw_experiments.json
                               synthesis.json

The four phases are:

PhaseToolWhat happens
1 Discoverydiscover_logs.pyWalks --search-roots to catalog every relevant log file across all agent caches. Prints a summary for user review before anything is read.
2 ExtractionLLM (per ~50 KB batch)Applies references/extraction-prompt.md to each batch; produces raw_experiments.json. PII is stripped; unverified numbers are flagged [UNVERIFIED].
3 SynthesisLLM (one call)Merges possibly-redundant experiment records into a single research narrative (synthesis.json). Detects multiple disconnected projects and pauses to ask the user.
4 Formattingformat_po_inputs.pyConverts synthesis.json into idea.md (Sparse Idea format, §3.1) and experimental_log.md (App. D.3), ready for paper-orchestra.

Integration

Install — no extra dependencies beyond the base requirements.txt.

Symlink the skill into your host's skill directory alongside the others:

ln -sf ~/paper-orchestra/skills/agent-research-aggregator \
       ~/.claude/skills/agent-research-aggregator

For Cursor / Antigravity / Cline / Aider, follow the same per-host instructions in skills/paper-orchestra/references/host-integration.md.

Invoke by telling your coding agent:

"Aggregate my agent logs for paper writing" — or — "Prepare PaperOrchestra inputs from my cache" — or — "Turn my agent logs into a paper"

The trigger phrases are listed in the description field of skills/agent-research-aggregator/SKILL.md.

Parameters

FlagDefaultDescription
--search-rootscwd, ~Directories to scan for agent caches
--agentsallSubset: claude,cursor,antigravity,openclaw
--workspace./workspacePaperOrchestra workspace root
--depth4Max scan depth (prevents runaway traversal)
--since—Only logs modified after this date (ISO 8601)

Example workflows

From Claude Code memory + CLAUDE.md only:

python skills/agent-research-aggregator/scripts/discover_logs.py \
    --search-roots . \
    --agents claude \
    --out workspace/ara/discovered_logs.json
# → finds .claude/projects/<hash>/memory/*.md and CLAUDE.md

From a Cursor project (chat history + rules):

python skills/agent-research-aggregator/scripts/discover_logs.py \
    --search-roots ~/my-project \
    --agents cursor \
    --out workspace/ara/discovered_logs.json
# → finds .cursor/chat/chatHistory.json and .cursorrules

From Antigravity worker logs, restricted to the last 60 days:

python skills/agent-research-aggregator/scripts/discover_logs.py \
    --search-roots ~/my-project \
    --agents antigravity \
    --since 2026-02-09 \
    --out workspace/ara/discovered_logs.json
# → finds .antigravity/workers/<id>/log.jsonl and output.md

From OpenClaw sessions + run metrics:

python skills/agent-research-aggregator/scripts/discover_logs.py \
    --search-roots ~/my-project \
    --agents openclaw \
    --out workspace/ara/discovered_logs.json
# → finds .openclaw/sessions/*/conversation.md and runs/*/metrics.json

Full run across all caches:

# Phase 1 — discovery
python skills/agent-research-aggregator/scripts/discover_logs.py \
    --search-roots . ~ --out workspace/ara/discovered_logs.json

# Phase 2 — LLM extraction (your agent handles this; validate afterward)
python skills/agent-research-aggregator/scripts/extract_experiments.py \
    --discovered workspace/ara/discovered_logs.json \
    --out workspace/ara/raw_experiments.json --validate-only

# Phase 3 — LLM synthesis (your agent handles this)

# Phase 4 — format + audit report
python skills/agent-research-aggregator/scripts/format_po_inputs.py \
    --synthesis workspace/ara/synthesis.json \
    --out workspace/inputs/ \
    --report workspace/ara/aggregation_report.md

After Phase 4, the workspace is ready for paper-orchestra. You still need to supply workspace/inputs/template.tex (your conference LaTeX template) and workspace/inputs/conference_guidelines.md (page limit, deadline, formatting rules).

Reference docs

Install

git clone <this repo> ~/paper-orchestra
cd ~/paper-orchestra
pip install -r requirements.txt   # deterministic helpers only

Then symlink the skills you want into your host's skill directory:

# Claude Code
mkdir -p ~/.claude/skills
for s in paper-orchestra outline-agent plotting-agent literature-review-agent \
         section-writing-agent content-refinement-agent paper-writing-bench \
         paper-autoraters agent-research-aggregator; do
  ln -sf ~/paper-orchestra/skills/$s ~/.claude/skills/$s
done

# Or for ~/.all-skills/
mkdir -p ~/.all-skills
for s in paper-orchestra outline-agent plotting-agent literature-review-agent \
         section-writing-agent content-refinement-agent paper-writing-bench \
         paper-autoraters agent-research-aggregator; do
  ln -sf ~/paper-orchestra/skills/$s ~/.all-skills/$s
done

For Cursor / Antigravity / Cline / Aider, see skills/paper-orchestra/references/host-integration.md.

Optional integrations

The pipeline requires zero API keys to run under any host with a native web search tool. Two optional integrations improve throughput or coverage:

  • Semantic Scholar API key — Phase 2 (citation verification) uses the public unauthenticated Semantic Scholar endpoint by default (≤1 QPS). A free API key raises the rate limit and reduces 429 back-off during large runs. The bundled scripts/s2_search.py reads SEMANTIC_SCHOLAR_API_KEY from the environment automatically — if the variable is absent it silently falls back to unauthenticated mode. The repo never commits a key.

    export SEMANTIC_SCHOLAR_API_KEY="your-key-here"   # https://api.semanticscholar.org/
    # verify it's picked up:
    python skills/literature-review-agent/scripts/s2_search.py --check-key
    

    See skills/literature-review-agent/references/s2-api-cookbook.md for endpoint details, field reference, and error-handling notes.

  • PaperBanana (Zhu et al., 2026) — the figure-generation backbone used by PaperOrchestra for Step 2. Runs a Retriever → Planner → Stylist → Visualizer → Critic loop that produces publication-quality diagrams grounded in real paper examples. Requires one API key — fill at least one, you don't need both:

    git clone https://github.com/dwzhu-pku/PaperBanana
    cd PaperBanana
    pip install -r requirements.txt
    cp configs/model_config.template.yaml configs/model_config.yaml
    # open model_config.yaml — paste your Gemini key into api_keys.google_api_key
    #                        OR your OpenRouter key into api_keys.openrouter_api_key
    export PAPERBANANA_PATH="/path/to/PaperBanana"
    

    That's it. Set PAPERBANANA_PATH and the plotting-agent uses PaperBanana automatically for diagram figures; falls back to matplotlib if unset. See skills/plotting-agent/references/paperbanana-cookbook.md for details.

  • Exa — research-paper-focused search engine. The literature-review-agent can use it as a Phase 1 candidate-discovery backend via skills/literature-review-agent/scripts/exa_search.py. Set EXA_API_KEY in your environment (the repo never commits a key) and the helper queries Exa with category: "research paper", returning 10–20 candidates per query in the format the rest of the pipeline expects. See skills/literature-review-agent/references/exa-search-cookbook.md for the full recipe, query patterns, cost (~$0.007/query), and security notes.

    export EXA_API_KEY="your-key-here"   # https://dashboard.exa.ai/
    python skills/literature-review-agent/scripts/exa_search.py \
        --query "Sparse attention long context" --num-results 15
    

    Skip Exa entirely if your host (Claude Code, Cursor, Antigravity) already has a native web search tool — the agent will use that instead.

Quickstart

Option A — you already have structured inputs

# 1. scaffold a workspace next to your raw materials
python skills/paper-orchestra/scripts/init_workspace.py --out workspace/

# 2. drop your inputs into workspace/inputs/
#    (idea.md, experimental_log.md, template.tex, conference_guidelines.md;
#     optional pre-existing figures go in workspace/inputs/figures/)

# 3. ask your coding agent:
#    "Run the paper-orchestra pipeline on ./workspace"

Option B — your research is scattered across a directory or agent caches

If you have a project folder and haven't written up a clean experiment log yet, just tell your coding agent the folder. The aggregator runs first — automatically — and produces idea.md and experimental_log.md before handing off to the pipeline:

"Write a paper from my work in ~/my-project"
"Turn my experiments in ~/lord into a paper"
"Aggregate ~/market-crispony and write a conference submission"

The agent will:

  1. Scan the directory for agent caches (.claude/, .cursor/, .antigravity/, .openclaw/) and any research notes it finds there.
  2. Extract and synthesize them into workspace/inputs/idea.md and workspace/inputs/experimental_log.md.
  3. Ask you to review both files, then run the full paper-orchestra pipeline.

You can also point it at any arbitrary directory — not just known agent caches:

# Phase 1: discover what's in the folder
python skills/agent-research-aggregator/scripts/discover_logs.py \
    --search-roots ~/my-project \
    --out workspace/ara/discovered_logs.json

# Then let your agent handle the rest ("Run paper-orchestra on ./workspace")

The aggregator is optional. If workspace/inputs/idea.md and workspace/inputs/experimental_log.md already exist, it is skipped entirely.

A ready-to-run toy case lives at examples/minimal/.

Repo layout

paper-orchestra/
├── README.md, LICENSE, CITATION.cff, requirements.txt
├── skills/                  # 7 skills + orchestrator
├── examples/minimal/        # toy end-to-end example
└── docs/
    ├── architecture.md      # deep-dive on the pipeline
    ├── paper-fidelity.md    # design-decision → paper page map
    └── coding-agent-integration.md  # per-host setup

Fidelity to the paper

Every agent prompt in skills/*/references/prompt.md is reproduced verbatim from Appendix F of arXiv:2604.05018, with a header pointing to the page number. See docs/paper-fidelity.md for a design-decision → paper-page map.

On top of the paper, this repo adds a few deterministic hardening scripts (orphan-citation gate, anti-leakage grep, worklog-based rollback, provenance snapshots). These are clearly marked as out-of-paper improvements in docs/paper-fidelity.md.

Citation

If you use this skill pack, please cite the PaperOrchestra paper. If you use the PaperBanana plotting backbone, cite that too:

@article{song2026paperorchestra,
  title={{PaperOrchestra}: A Multi-Agent Framework for Automated {AI} Research Paper Writing},
  author={Song, Yiwen and Song, Yale and Pfister, Tomas and Yoon, Jinsung},
  journal={arXiv preprint arXiv:2604.05018},
  year={2026},
  url={https://arxiv.org/abs/2604.05018}
}

@article{zhu2026paperbanana,
  title={{PaperBanana}: Automating Academic Illustration for {AI} Scientists},
  author={Zhu, Dawei and Meng, Rui and Song, Yale and Wei, Xiyu and Li, Sujian and Pfister, Tomas and Yoon, Jinsung},
  journal={arXiv preprint arXiv:2601.23265},
  year={2026},
  url={https://arxiv.org/abs/2601.23265}
}
It would have been fun if the repo wrote the paper.

License

MIT — see LICENSE.

文档与办公研究与检索内容与创作Agent / MCP / Skill 创作

中风险

  • 来源需自行核对维护者身份。
  • 包含脚本或命令调用,安装前请复核。
  • 未检测到明显外部权限要求。
  • 未检测到高风险命令。
  • 扫描发现:2 条。

Codex — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/Ar9av/PaperOrchestra.git
  3. 将 "skills/paper-orchestra" 文件夹复制到 Codex 的 skills 目录中。
  4. 重启 Codex 让新的 skill 生效。

Codex — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Codex 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Codex 让新的 skill 生效。

Claude Code — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/Ar9av/PaperOrchestra.git
  3. 将 "skills/paper-orchestra" 文件夹复制到 Claude Code 的 skills 目录中。
  4. 重启 Claude Code 让新的 skill 生效。

Claude Code — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Claude Code 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Claude Code 让新的 skill 生效。

Cursor — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/Ar9av/PaperOrchestra.git
  3. 将 "skills/paper-orchestra" 文件夹复制到 Cursor 的 skills 目录中。
  4. 重启 Cursor 让新的 skill 生效。

Cursor — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Cursor 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Cursor 让新的 skill 生效。

GitHub Copilot — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/Ar9av/PaperOrchestra.git
  3. 将 "skills/paper-orchestra" 文件夹复制到 GitHub Copilot 的 skills 目录中。
  4. 重启 GitHub Copilot 让新的 skill 生效。

GitHub Copilot — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 GitHub Copilot 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 GitHub Copilot 让新的 skill 生效。

Windsurf — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/Ar9av/PaperOrchestra.git
  3. 将 "skills/paper-orchestra" 文件夹复制到 Windsurf 的 skills 目录中。
  4. 重启 Windsurf 让新的 skill 生效。

Windsurf — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Windsurf 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Windsurf 让新的 skill 生效。
查看 SKILL.md 原文
name: paper-orchestra
description: Orchestrate the full PaperOrchestra (Song et al., 2026, arXiv:2604.05018) five-agent pipeline to turn unstructured research materials (idea, experimental log, LaTeX template, conference guidelines, optional figures) into a submission-ready LaTeX manuscript and compiled PDF. TRIGGER when the user asks to "write a paper from my experiments", "turn this idea and these results into a paper", "generate a conference submission", "run paper-orchestra on X", or otherwise wants the end-to-end paper-writing pipeline. Coordinates the outline-agent, plotting-agent, literature-review-agent, section-writing-agent, and content-refinement-agent skills.
data_access_level: raw

paper-orchestra (Orchestrator)

Top-level driver for the PaperOrchestra pipeline. Read this document and follow the steps below. The detailed prompts and rules live in each sub-skill's SKILL.md and references/ directories — you (the host agent) will load them as you go.

Source paper: Song et al., PaperOrchestra: A Multi-Agent Framework for Automated AI Research Paper Writing, arXiv:2604.05018, 2026. https://arxiv.org/pdf/2604.05018

What this skill produces

A complete submission package P = (paper.tex, paper.pdf) written into workspace/final/, plus a full audit trail under workspace/ (outline, figures, refs, drafts, refinement worklog, provenance snapshot).

Inputs (the (I, E, T, G, F) tuple from the paper)

The workspace MUST contain:

FileSymbolRequiredDescription
workspace/inputs/idea.mdIyesIdea Summary (Sparse or Dense variant — see references/io-contract.md)
workspace/inputs/experimental_log.mdEyesExperimental Log: setup, raw numeric data, qualitative observations
workspace/inputs/template.texTyesLaTeX template for the target conference (with \section{...} commands)
workspace/inputs/conference_guidelines.mdGyesFormatting rules, page limit, mandatory sections
workspace/inputs/figures/FnoOptional pre-existing figures. If empty, the plotting agent generates everything.

scripts/init_workspace.py will scaffold this layout. scripts/validate_inputs.py will check it before the pipeline runs.

Pipeline (read references/pipeline.md for the full diagram)

Step 1: Outline           ──▶  outline.json                       (1 LLM call)
Step 2: Plotting     ─┐
                      ├──▶  figures/*.png + captions.json         (~20-30 calls)
Step 3: Lit Review   ─┘                                           (~20-30 calls)
                          intro_relwork.tex + refs.bib

Step 4: Section Writing  ──▶  drafts/paper.tex                    (1 LLM call)
Step 5: Content Refine   ──▶  final/paper.tex + final/paper.pdf   (~5-7 calls, ~3 iters)

Step 2 and Step 3 are independent and MUST run in parallel when your host supports parallel sub-agents. If not, run Step 3 first (it has the longer wall time due to Semantic Scholar rate limits) and Step 2 second.

Critical pre-instruction (read once, apply always)

Before any LLM call that writes paper content (outline, intro/related work, section writing, refinement), you MUST prepend the Anti-Leakage Prompt at references/anti-leakage-prompt.md to your system prompt. This is verbatim from Appendix D.4 of the paper and prevents pre-training-data leakage. The paper applies it uniformly across all baselines for fair comparison; we apply it for fidelity and to keep generated papers grounded in the user's inputs.

Step-by-step execution

0. Pre-flight Checks

Before running the pipeline, perform the following quality gates in order:

# 1. Scaffold the workspace
python skills/paper-orchestra/scripts/init_workspace.py --out workspace/
# user drops their inputs into workspace/inputs/

# 2. Validate required files are present and well-formed
python skills/paper-orchestra/scripts/validate_inputs.py --workspace workspace/

# 3. Check input density — idea and experimental log must meet minimum thresholds
python skills/paper-orchestra/scripts/check_idea_density.py \
    --idea workspace/inputs/idea.md \
    --log workspace/inputs/experimental_log.md

# 4. Cross-validate consistency between idea and experimental log
python skills/paper-orchestra/scripts/validate_consistency.py \
    --idea workspace/inputs/idea.md \
    --log workspace/inputs/experimental_log.md

If validate_inputs.py or check_idea_density.py fail (exit code 1 or 2), stop and tell the user what's missing or below threshold — do not proceed until fixed.

validate_consistency.py produces warnings only (exit code 1 = WARN, non-blocking); report warnings to the user but continue.

Before failing on missing inputs, check whether aggregation can supply them:

Inputs stateAction
idea.md and experimental_log.md both present and non-emptyContinue to Step 1.
Either is missing/empty, and the user mentioned a directoryLoad and run agent-research-aggregator with that directory as --search-roots, then re-validate.
Either is missing/empty, no directory mentionedAsk the user: "Your workspace is missing idea.md / experimental_log.md. Do you have a folder with research notes or agent history I can aggregate from? If so, tell me the path — or drop the files manually into workspace/inputs/."

If validation still fails after aggregation (e.g. template.tex or conference_guidelines.md are missing), stop and tell the user exactly which files remain outstanding.

Also probe the TeX installation (once per workspace, result cached):

python skills/paper-orchestra/scripts/check_tex_packages.py \
    --out workspace/tex_profile.json

The Section Writing Agent reads tex_profile.json to decide which LaTeX patterns to use (e.g., Figure~\ref{} vs \cref{}, whether to include \usepackage{microtype}, etc.). This eliminates compile-time package failures that previously required iterative manual edits.

1. Outline (Step 1 — 1 LLM call)

Load skills/outline-agent/SKILL.md and follow it. Output: workspace/outline.json. Validate with python skills/outline-agent/scripts/validate_outline.py workspace/outline.json. Halt the pipeline if validation fails — every downstream agent depends on the schema.

2 ∥ 3. Plotting and Literature Review (in parallel)

Parse outline.json. Extract:

  • outline.plotting_plan → drives Step 2
  • outline.intro_related_work_plan → drives Step 3

If your host supports parallel sub-agents (Claude Code's Agent tool with multiple concurrent calls; Cursor's parallel agents; Antigravity's worker pool), spawn two concurrent sub-tasks:

  • Sub-task A: load skills/plotting-agent/SKILL.md, execute the plotting plan, produce workspace/figures/<figure_id>.png for every entry, plus workspace/figures/captions.json.
  • Sub-task B: load skills/literature-review-agent/SKILL.md, execute the research strategy, produce workspace/drafts/intro_relwork.tex and workspace/refs.bib.

If your host does not support parallel sub-agents, run Sub-task B first (it has slower wall-clock due to Semantic Scholar QPS limits) then Sub-task A. The artifacts are independent, so order doesn't affect correctness.

3.5. Outline Reconciliation (after Step 3 completes, before Step 4)

Once Step 3 (Literature Review) has produced citation_pool.json and cross_verification_report.json, run the reconciliation step.

Load references/outline-reconciliation.md and follow its prompt. Output: workspace/outline_reconciled.json.

Validate and diff:

python skills/outline-agent/scripts/validate_outline.py workspace/outline_reconciled.json
python skills/paper-orchestra/scripts/diff_outlines.py \
    --original   workspace/outline.json \
    --reconciled workspace/outline_reconciled.json \
    --summary    workspace/reconciliation_summary.md

If validation fails, fall back to outline.json for Step 4 and warn the user. Show the user the reconciliation_summary.md (even if no changes — it confirms the outline matched the actual literature).

Skip conditions: citation pool empty, Step 3 failed, or Step 2 is still running and the host cannot issue another call concurrently. See references/outline-reconciliation.md for full skip conditions.

4. Section Writing (Step 4 — ONE single multimodal LLM call)

Load skills/section-writing-agent/SKILL.md and follow it. This is one single call in the paper (App. B: "Section Writing Agent (1 call)") — do not split it per section. The agent receives:

  • outline_reconciled.json (use this if it exists; fall back to outline.json)
  • idea.md, experimental_log.md
  • intro_relwork.tex (already-filled from Step 3 — preserve verbatim)
  • refs.bib (the citation map)
  • conference_guidelines.md
  • research_brief.md (if it exists — read §1–§3 for accumulated pipeline context)
  • The actual figure image files from workspace/figures/ (multimodal input)

Output: workspace/drafts/paper.tex (a complete LaTeX document).

Then run the deterministic gates:

python skills/section-writing-agent/scripts/orphan_cite_gate.py workspace/drafts/paper.tex workspace/refs.bib
python skills/section-writing-agent/scripts/latex_sanity.py workspace/drafts/paper.tex
python skills/paper-orchestra/scripts/anti_leakage_check.py workspace/drafts/paper.tex
python skills/paper-orchestra/scripts/claim_evidence_gate.py \
    --paper workspace/drafts/paper.tex \
    --log   workspace/inputs/experimental_log.md \
    --out   workspace/claim_evidence_report.json

claim_evidence_gate.py is a WARN gate (exit 1 = warnings, not a hard stop). Report the count of unsupported claims to the user. The content-refinement agent will address them in Step 5.

If any gate fails, the host agent must fix the issue (re-prompting the writing step with the gate's error report) before proceeding.

5. Content Refinement (Step 5 — ~3 iterations, ~5-7 calls)

Load skills/content-refinement-agent/SKILL.md and follow it. The skill implements the loop with strict halt rules from halt-rules.md. Maintain workspace/refinement/worklog.json and snapshot each iteration into workspace/refinement/iter<N>/.

Halt conditions (any one triggers the loop to stop and accept the current best snapshot):

  1. Iteration count reaches the cap (default 3, see halt-rules.md).
  2. Overall score from the simulated reviewer decreases vs the previous iteration → revert to previous snapshot, halt.
  3. Overall score ties but at least one sub-axis decreases while none gain compensatingly (negative net sub-axis change) → revert, halt.
  4. Reviewer issues no new actionable weaknesses.

The accepted snapshot is copied to workspace/final/paper.tex.

6. Compile and finalize

cd workspace/final && latexmk -pdf paper.tex

Then write workspace/provenance.json capturing input file hashes, outline hash, refs hash, figure hashes, and final tex/pdf hashes (helper: scripts/snapshot.py in the orchestrator scripts dir if you want a one-shot; otherwise the host agent computes hashes inline).

Report to the user: the path to workspace/final/paper.pdf, a brief summary of which sections were drafted, citation count, refinement iterations completed, and any gates that failed mid-pipeline.

Workspace layout

See references/io-contract.md. Summary:

workspace/
├── inputs/                          # user-provided
│   ├── idea.md
│   ├── experimental_log.md
│   ├── template.tex
│   ├── conference_guidelines.md
│   └── figures/                     # optional pre-existing figures
├── outline.json                     # Step 1 output
├── figures/                         # Step 2 output
│   ├── <figure_id>.png
│   └── captions.json
├── refs.bib                         # Step 3 output
├── drafts/                          # Step 3 + Step 4 output
│   ├── intro_relwork.tex
│   └── paper.tex
├── refinement/                      # Step 5 working dir
│   ├── worklog.json
│   ├── iter1/
│   ├── iter2/
│   └── iter3/
├── final/                           # accepted snapshot + compiled PDF
│   ├── paper.tex
│   └── paper.pdf
└── provenance.json                  # input/output hashes for reproducibility

Cost budget (from paper App. B)

Total: ~60–70 LLM calls per paper, ~40 minutes wall-time on the paper's setup. Budget breakdown:

StepCalls
Outline1
Plotting~20–30
Literature Review~20–30
Section Writing1
Content Refinement~5–7

Host integration

See references/host-integration.md for per-host invocation details (Claude Code, Cursor, Antigravity, Cline, Aider, OpenCode).

Resources

  • references/pipeline.md — full step-by-step flow + parallelism rules + halt rules
  • references/io-contract.md — workspace layout, input file schemas
  • references/anti-leakage-prompt.md — verbatim from App. D.4, prepend to every writing call
  • references/paper-summary.md — 1-page distillation of arXiv:2604.05018
  • references/host-integration.md — per-host invocation guide
  • references/outline-reconciliation.md — NEW Step 3.5 outline reconciliation protocol (AutoSci-inspired)
  • scripts/init_workspace.py — scaffold workspace dir tree
  • scripts/validate_inputs.py — verify (I, E, T, G) before running
  • scripts/anti_leakage_check.py — grep draft for leaked author names/emails/affils
  • scripts/claim_evidence_gate.py — NEW WARN gate: verify numeric claims in draft are grounded in experimental_log.md
  • scripts/diff_outlines.py — NEW diff original vs reconciled outline; writes reconciliation_summary.md
  • skills/shared/research_brief_template.md — NEW schema for workspace/research_brief.md (accumulated cross-agent context)

发现问题?提交给管理员复核

评分:

评论 (0)

暂无评论,成为第一个评论者吧!