SkillAtlasSkill 详情

content-refinement-agent

A pluggable skill pack that lets any coding agent in Claude Code, Cursor,

审核状态:已审核Quality 80Security 80

复制安装命令

用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。

复制前请先查看来源、License 和安全提示。

项目 README

来源文件:README.md

抓取于 2026年8月25日

PaperOrchestra

A pluggable skill pack that lets any coding agent in Claude Code, Cursor, Antigravity, Cline, Aider, OpenCode, etc. which can run the PaperOrchestra multi-agent pipeline for turning unstructured research materials into a submission-ready LaTeX paper.

Song, Y., Song, Y., Pfister, T., Yoon, J. PaperOrchestra: A Multi-Agent Framework for Automated AI Research Paper Writing. arXiv:2604.05018, 2026. https://arxiv.org/pdf/2604.05018

PaperOrchestra paper — first page preview
Click to read the paper on arXiv

Why this exists

The paper defines a five-agent pipeline

  • Outline
  • Plotting
  • Literature Review
  • Section Writing
  • Content Refinement

that substantially outperforms single-agent and tree-search baselines on the PaperWritingBench benchmark (50–68% absolute win margin on literature review quality; 14–38% on overall quality). The paper ships the exact prompts for every agent in Appendix F.

This repo turns those prompts, schemas, halt rules, and verification pipelines into a set of host-agent-executable skills. There are no API keys, no SDK dependencies, no embedded LLM calls. The skills are instruction documents plus deterministic helpers; your coding agent does all LLM reasoning and web search using its own tools.

image

How skills work here

Each skill is:

  • SKILL.md — a dense instruction document the host agent reads and follows.
  • references/ — reference material: verbatim paper prompts (Appendix F), JSON schemas, rubrics, halt rules, example outputs.
  • scripts/ — purely deterministic local helpers: JSON schema validation, Levenshtein fuzzy matching, BibTeX formatting, dedup, LaTeX sanity checks, coverage gates. No network, no LLM, no API keys.

Everything else (LLM reasoning, web search, Semantic Scholar lookups, LaTeX compilation) is delegated to the host agent by instruction. See skills/paper-orchestra/references/host-integration.md for per-host invocation (Claude Code, Cursor, Antigravity, Cline, Aider).

The seven skills

SkillPaper step# LLM callsRole
paper-orchestraorchestrator—Top-level driver. Coordinates the other six.
outline-agentStep 11Idea + log + template + guidelines → structured outline JSON (plotting plan, lit review plan, section plan).
plotting-agentStep 2~20–30Execute plotting plan; render plots & conceptual diagrams; optional VLM-critique refinement loop; caption everything.
literature-review-agentStep 3~20–30Web-search candidates; Semantic Scholar verify (Levenshtein > 70, cutoff, dedup); draft Intro + Related Work with ≥90% citation integration.
section-writing-agentStep 41One single multimodal call: draft remaining sections, build tables from experimental log, splice figures.
content-refinement-agentStep 5~5–7Simulated peer review; accept/revert per strict halt rules; safety constraints prevent gaming the evaluator.
paper-writing-bench§3—Reverse-engineer raw materials (Sparse/Dense idea, experimental log) from an existing paper to build benchmark cases.
paper-autoratersApp. F.3—Run the paper's own autoraters: Citation F1 (P0/P1), LitReview quality (6-axis), SxS paper quality, SxS litreview quality.

Steps 2 and 3 run in parallel (see skills/paper-orchestra/references/pipeline.md).

agent-research-aggregator (optional)

A pre-pipeline skill that bridges the gap between scattered AI coding-agent history and the structured (idea.md, experimental_log.md) inputs that PaperOrchestra expects. If you have been running experiments through Claude Code, Cursor, Antigravity, or OpenClaw — but never wrote up a clean experiment log — this skill does that extraction for you.

It is optional. If workspace/inputs/idea.md and workspace/inputs/experimental_log.md already exist, the skill skips itself and the pipeline proceeds directly. It only runs when the inputs are missing or when you explicitly point an agent at a directory.

The simplest way to use it: just tell your agent the folder. If you have a directory (a project root, an agent cache, any folder with research notes), the aggregator figures out what's inside and structures it for PaperOrchestra. The first thing it does is aggregate — scanning, extracting, and synthesising — so even if the data is scattered across multiple files and formats, it produces clean, reviewable inputs before anything gets written.

Run it before paper-orchestra (or let paper-orchestra call it automatically when inputs are missing).

What it does

[.claude/]  [.cursor/]  [.antigravity/]  [.openclaw/]
      │            │              │               │
      └────────────┴──────────────┴───────────────┘
                        │
                Phase 1: Discovery  (deterministic)
                        │
                Phase 2: Extraction (LLM — per batch)
                        │
                Phase 3: Synthesis  (LLM — one call)
                        │
                Phase 4: Formatting (deterministic)
                        │
             ┌──────────┴──────────┐
      workspace/inputs/      workspace/ara/
        idea.md                aggregation_report.md
        experimental_log.md    discovered_logs.json
                               raw_experiments.json
                               synthesis.json

The four phases are:

PhaseToolWhat happens
1 Discoverydiscover_logs.pyWalks --search-roots to catalog every relevant log file across all agent caches. Prints a summary for user review before anything is read.
2 ExtractionLLM (per ~50 KB batch)Applies references/extraction-prompt.md to each batch; produces raw_experiments.json. PII is stripped; unverified numbers are flagged [UNVERIFIED].
3 SynthesisLLM (one call)Merges possibly-redundant experiment records into a single research narrative (synthesis.json). Detects multiple disconnected projects and pauses to ask the user.
4 Formattingformat_po_inputs.pyConverts synthesis.json into idea.md (Sparse Idea format, §3.1) and experimental_log.md (App. D.3), ready for paper-orchestra.

Integration

Install — no extra dependencies beyond the base requirements.txt.

Symlink the skill into your host's skill directory alongside the others:

ln -sf ~/paper-orchestra/skills/agent-research-aggregator \
       ~/.claude/skills/agent-research-aggregator

For Cursor / Antigravity / Cline / Aider, follow the same per-host instructions in skills/paper-orchestra/references/host-integration.md.

Invoke by telling your coding agent:

"Aggregate my agent logs for paper writing" — or — "Prepare PaperOrchestra inputs from my cache" — or — "Turn my agent logs into a paper"

The trigger phrases are listed in the description field of skills/agent-research-aggregator/SKILL.md.

Parameters

FlagDefaultDescription
--search-rootscwd, ~Directories to scan for agent caches
--agentsallSubset: claude,cursor,antigravity,openclaw
--workspace./workspacePaperOrchestra workspace root
--depth4Max scan depth (prevents runaway traversal)
--since—Only logs modified after this date (ISO 8601)

Example workflows

From Claude Code memory + CLAUDE.md only:

python skills/agent-research-aggregator/scripts/discover_logs.py \
    --search-roots . \
    --agents claude \
    --out workspace/ara/discovered_logs.json
# → finds .claude/projects/<hash>/memory/*.md and CLAUDE.md

From a Cursor project (chat history + rules):

python skills/agent-research-aggregator/scripts/discover_logs.py \
    --search-roots ~/my-project \
    --agents cursor \
    --out workspace/ara/discovered_logs.json
# → finds .cursor/chat/chatHistory.json and .cursorrules

From Antigravity worker logs, restricted to the last 60 days:

python skills/agent-research-aggregator/scripts/discover_logs.py \
    --search-roots ~/my-project \
    --agents antigravity \
    --since 2026-02-09 \
    --out workspace/ara/discovered_logs.json
# → finds .antigravity/workers/<id>/log.jsonl and output.md

From OpenClaw sessions + run metrics:

python skills/agent-research-aggregator/scripts/discover_logs.py \
    --search-roots ~/my-project \
    --agents openclaw \
    --out workspace/ara/discovered_logs.json
# → finds .openclaw/sessions/*/conversation.md and runs/*/metrics.json

Full run across all caches:

# Phase 1 — discovery
python skills/agent-research-aggregator/scripts/discover_logs.py \
    --search-roots . ~ --out workspace/ara/discovered_logs.json

# Phase 2 — LLM extraction (your agent handles this; validate afterward)
python skills/agent-research-aggregator/scripts/extract_experiments.py \
    --discovered workspace/ara/discovered_logs.json \
    --out workspace/ara/raw_experiments.json --validate-only

# Phase 3 — LLM synthesis (your agent handles this)

# Phase 4 — format + audit report
python skills/agent-research-aggregator/scripts/format_po_inputs.py \
    --synthesis workspace/ara/synthesis.json \
    --out workspace/inputs/ \
    --report workspace/ara/aggregation_report.md

After Phase 4, the workspace is ready for paper-orchestra. You still need to supply workspace/inputs/template.tex (your conference LaTeX template) and workspace/inputs/conference_guidelines.md (page limit, deadline, formatting rules).

Reference docs

Install

git clone <this repo> ~/paper-orchestra
cd ~/paper-orchestra
pip install -r requirements.txt   # deterministic helpers only

Then symlink the skills you want into your host's skill directory:

# Claude Code
mkdir -p ~/.claude/skills
for s in paper-orchestra outline-agent plotting-agent literature-review-agent \
         section-writing-agent content-refinement-agent paper-writing-bench \
         paper-autoraters agent-research-aggregator; do
  ln -sf ~/paper-orchestra/skills/$s ~/.claude/skills/$s
done

# Or for ~/.all-skills/
mkdir -p ~/.all-skills
for s in paper-orchestra outline-agent plotting-agent literature-review-agent \
         section-writing-agent content-refinement-agent paper-writing-bench \
         paper-autoraters agent-research-aggregator; do
  ln -sf ~/paper-orchestra/skills/$s ~/.all-skills/$s
done

For Cursor / Antigravity / Cline / Aider, see skills/paper-orchestra/references/host-integration.md.

Optional integrations

The pipeline requires zero API keys to run under any host with a native web search tool. Two optional integrations improve throughput or coverage:

  • Semantic Scholar API key — Phase 2 (citation verification) uses the public unauthenticated Semantic Scholar endpoint by default (≤1 QPS). A free API key raises the rate limit and reduces 429 back-off during large runs. The bundled scripts/s2_search.py reads SEMANTIC_SCHOLAR_API_KEY from the environment automatically — if the variable is absent it silently falls back to unauthenticated mode. The repo never commits a key.

    export SEMANTIC_SCHOLAR_API_KEY="your-key-here"   # https://api.semanticscholar.org/
    # verify it's picked up:
    python skills/literature-review-agent/scripts/s2_search.py --check-key
    

    See skills/literature-review-agent/references/s2-api-cookbook.md for endpoint details, field reference, and error-handling notes.

  • PaperBanana (Zhu et al., 2026) — the figure-generation backbone used by PaperOrchestra for Step 2. Runs a Retriever → Planner → Stylist → Visualizer → Critic loop that produces publication-quality diagrams grounded in real paper examples. Requires one API key — fill at least one, you don't need both:

    git clone https://github.com/dwzhu-pku/PaperBanana
    cd PaperBanana
    pip install -r requirements.txt
    cp configs/model_config.template.yaml configs/model_config.yaml
    # open model_config.yaml — paste your Gemini key into api_keys.google_api_key
    #                        OR your OpenRouter key into api_keys.openrouter_api_key
    export PAPERBANANA_PATH="/path/to/PaperBanana"
    

    That's it. Set PAPERBANANA_PATH and the plotting-agent uses PaperBanana automatically for diagram figures; falls back to matplotlib if unset. See skills/plotting-agent/references/paperbanana-cookbook.md for details.

  • Exa — research-paper-focused search engine. The literature-review-agent can use it as a Phase 1 candidate-discovery backend via skills/literature-review-agent/scripts/exa_search.py. Set EXA_API_KEY in your environment (the repo never commits a key) and the helper queries Exa with category: "research paper", returning 10–20 candidates per query in the format the rest of the pipeline expects. See skills/literature-review-agent/references/exa-search-cookbook.md for the full recipe, query patterns, cost (~$0.007/query), and security notes.

    export EXA_API_KEY="your-key-here"   # https://dashboard.exa.ai/
    python skills/literature-review-agent/scripts/exa_search.py \
        --query "Sparse attention long context" --num-results 15
    

    Skip Exa entirely if your host (Claude Code, Cursor, Antigravity) already has a native web search tool — the agent will use that instead.

Quickstart

Option A — you already have structured inputs

# 1. scaffold a workspace next to your raw materials
python skills/paper-orchestra/scripts/init_workspace.py --out workspace/

# 2. drop your inputs into workspace/inputs/
#    (idea.md, experimental_log.md, template.tex, conference_guidelines.md;
#     optional pre-existing figures go in workspace/inputs/figures/)

# 3. ask your coding agent:
#    "Run the paper-orchestra pipeline on ./workspace"

Option B — your research is scattered across a directory or agent caches

If you have a project folder and haven't written up a clean experiment log yet, just tell your coding agent the folder. The aggregator runs first — automatically — and produces idea.md and experimental_log.md before handing off to the pipeline:

"Write a paper from my work in ~/my-project"
"Turn my experiments in ~/lord into a paper"
"Aggregate ~/market-crispony and write a conference submission"

The agent will:

  1. Scan the directory for agent caches (.claude/, .cursor/, .antigravity/, .openclaw/) and any research notes it finds there.
  2. Extract and synthesize them into workspace/inputs/idea.md and workspace/inputs/experimental_log.md.
  3. Ask you to review both files, then run the full paper-orchestra pipeline.

You can also point it at any arbitrary directory — not just known agent caches:

# Phase 1: discover what's in the folder
python skills/agent-research-aggregator/scripts/discover_logs.py \
    --search-roots ~/my-project \
    --out workspace/ara/discovered_logs.json

# Then let your agent handle the rest ("Run paper-orchestra on ./workspace")

The aggregator is optional. If workspace/inputs/idea.md and workspace/inputs/experimental_log.md already exist, it is skipped entirely.

A ready-to-run toy case lives at examples/minimal/.

Repo layout

paper-orchestra/
├── README.md, LICENSE, CITATION.cff, requirements.txt
├── skills/                  # 7 skills + orchestrator
├── examples/minimal/        # toy end-to-end example
└── docs/
    ├── architecture.md      # deep-dive on the pipeline
    ├── paper-fidelity.md    # design-decision → paper page map
    └── coding-agent-integration.md  # per-host setup

Fidelity to the paper

Every agent prompt in skills/*/references/prompt.md is reproduced verbatim from Appendix F of arXiv:2604.05018, with a header pointing to the page number. See docs/paper-fidelity.md for a design-decision → paper-page map.

On top of the paper, this repo adds a few deterministic hardening scripts (orphan-citation gate, anti-leakage grep, worklog-based rollback, provenance snapshots). These are clearly marked as out-of-paper improvements in docs/paper-fidelity.md.

Citation

If you use this skill pack, please cite the PaperOrchestra paper. If you use the PaperBanana plotting backbone, cite that too:

@article{song2026paperorchestra,
  title={{PaperOrchestra}: A Multi-Agent Framework for Automated {AI} Research Paper Writing},
  author={Song, Yiwen and Song, Yale and Pfister, Tomas and Yoon, Jinsung},
  journal={arXiv preprint arXiv:2604.05018},
  year={2026},
  url={https://arxiv.org/abs/2604.05018}
}

@article{zhu2026paperbanana,
  title={{PaperBanana}: Automating Academic Illustration for {AI} Scientists},
  author={Zhu, Dawei and Meng, Rui and Song, Yale and Wei, Xiyu and Li, Sujian and Pfister, Tomas and Yoon, Jinsung},
  journal={arXiv preprint arXiv:2601.23265},
  year={2026},
  url={https://arxiv.org/abs/2601.23265}
}
It would have been fun if the repo wrote the paper.

License

MIT — see LICENSE.

内容与创作Agent / MCP / Skill 创作

中风险

  • 来源需自行核对维护者身份。
  • 包含脚本或命令调用,安装前请复核。
  • 未检测到明显外部权限要求。
  • 未检测到高风险命令。
  • 扫描发现:2 条。

Codex — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/Ar9av/PaperOrchestra.git
  3. 将 "skills/content-refinement-agent" 文件夹复制到 Codex 的 skills 目录中。
  4. 重启 Codex 让新的 skill 生效。

Codex — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Codex 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Codex 让新的 skill 生效。

Claude Code — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/Ar9av/PaperOrchestra.git
  3. 将 "skills/content-refinement-agent" 文件夹复制到 Claude Code 的 skills 目录中。
  4. 重启 Claude Code 让新的 skill 生效。

Claude Code — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Claude Code 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Claude Code 让新的 skill 生效。

Cursor — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/Ar9av/PaperOrchestra.git
  3. 将 "skills/content-refinement-agent" 文件夹复制到 Cursor 的 skills 目录中。
  4. 重启 Cursor 让新的 skill 生效。

Cursor — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Cursor 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Cursor 让新的 skill 生效。

GitHub Copilot — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/Ar9av/PaperOrchestra.git
  3. 将 "skills/content-refinement-agent" 文件夹复制到 GitHub Copilot 的 skills 目录中。
  4. 重启 GitHub Copilot 让新的 skill 生效。

GitHub Copilot — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 GitHub Copilot 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 GitHub Copilot 让新的 skill 生效。

Windsurf — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/Ar9av/PaperOrchestra.git
  3. 将 "skills/content-refinement-agent" 文件夹复制到 Windsurf 的 skills 目录中。
  4. 重启 Windsurf 让新的 skill 生效。

Windsurf — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Windsurf 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Windsurf 让新的 skill 生效。
查看 SKILL.md 原文
name: content-refinement-agent
description: Step 5 of the PaperOrchestra pipeline (arXiv:2604.05018). Iteratively refine drafts/paper.tex by simulating peer review and applying targeted revisions, with strict accept/revert halt rules, deterministic 0-100 decision bands (Accept/Minor/Major/Reject) that drive a target-met early stop, and a Devil's Advocate concession-threshold guard that blocks acceptance on unresolved critical findings. Maintains a worklog and snapshots each iteration so revert is real, not symbolic. TRIGGER when the orchestrator delegates Step 5 or when the user asks to "refine the draft", "iterate on the paper", or "run peer review on this paper".
data_access_level: verified_only

Content Refinement Agent (Step 5)

Faithful implementation of the Content Refinement Agent from PaperOrchestra (Song et al., 2026, arXiv:2604.05018, §4 Step 5, App. F.1 pp. 49–51).

Cost: ~5–7 LLM calls (App. B), typically ~3 refinement iterations, each consisting of one reviewer call and one revision call.

The paper highlights this step as one of the largest contributors to overall quality: refinement alone accounts for +19% (CVPR) and +22% (ICLR) absolute acceptance-rate improvement (Fig. 4). Get this step right.

Inputs

  • workspace/drafts/paper.tex — output of Step 4
  • workspace/inputs/conference_guidelines.md
  • workspace/inputs/experimental_log.md — used as ground truth for the hallucination check
  • workspace/citation_pool.json / workspace/refs.bib — the allowed bibliography

Outputs

  • workspace/refinement/iter1/, iter2/, iter3/ — per-iteration snapshots containing paper.tex, paper.pdf, review.json, score.json
  • workspace/refinement/worklog.json — append-only history of decisions
  • workspace/final/paper.tex and workspace/final/paper.pdf — copy of the best accepted snapshot

The refinement loop

prev_score = score(paper.tex)                  # baseline from initial draft
snapshot iter0/

for iter in 1..ITER_CAP (default 3):
    1. simulate_review(paper.tex) → review.json
       (uses `references/reviewer-rubric.md` rubric)

    2. apply_revision(paper.tex, review.json) → new_paper.tex
       (uses verbatim Refinement Agent prompt at `references/prompt.md`)

    3. snapshot iter<N>/ with new_paper.tex, review.json
       latexmk -pdf new_paper.tex → iter<N>/paper.pdf

    4. score(new_paper.tex) → curr_score

    5. decide via score_delta.py:
       - if curr.overall > prev.overall:                       ACCEPT
       - elif curr.overall == prev.overall and net_subaxis ≥0: ACCEPT
       - else:                                                 REVERT

    6. apply_worklog.py to append the decision

    7. if REVERT or no actionable weaknesses or iter == ITER_CAP: HALT

    paper.tex ← new_paper.tex   (only on ACCEPT)
    prev_score ← curr_score

cp <best iter>/paper.tex → workspace/final/paper.tex

The "best" snapshot at HALT is the one with the highest accepted overall score. On a REVERT halt, the best is the iteration immediately before the revert.

Step-by-step

0. Pre-refinement integrity gate

Before snapshotting or scoring the initial draft, run two gates in order:

Gate A — AI failure modes (load references/ai-failure-modes.md, runs once):

Load references/ai-failure-modes.md (which points to skills/shared/ai_failure_modes.md). Run all 7 checks against the draft and the inputs. This gate runs once only, at the start of iteration 1.

  • CONFIRMED failure → write HALT entry to worklog.json, report to user, stop.
  • SUSPECTED failure → add WARNING comment to paper.tex, log in worklog.json, continue.
  • No failures → proceed.

Gate B — Claim-evidence provenance (runs once, WARN gate):

python skills/paper-orchestra/scripts/claim_evidence_gate.py \
    --paper workspace/drafts/paper.tex \
    --log   workspace/inputs/experimental_log.md \
    --out   workspace/claim_evidence_report.json

Exit 0 → PASS, proceed normally. Exit 1 → WARN: unsupported numeric claims found. Log in worklog.json as: {gate: "claim_evidence", status: "WARN", unsupported_count: N, report: "workspace/claim_evidence_report.json"} Pass the unsupported list from the report to the revision agent in Step 3 as an additional instruction: "The following numeric values appear in the paper but cannot be corroborated in experimental_log.md — verify or remove them: ..." Do NOT halt on Gate B warnings; the revision agent will address them.

Gate C — Read research brief (every run, no exit code):

If workspace/research_brief.md exists, read it before all reviewer calls. Pass the "Sections where evidence was thin" list from §4 as additional context to the Devil's Advocate reviewer. This surfaces the highest-risk sections for CRITICAL scrutiny.

0b. Snapshot the initial draft

python skills/content-refinement-agent/scripts/snapshot.py \
    --src workspace/drafts/paper.tex \
    --dst workspace/refinement/iter0/

This creates iter0/paper.tex. Then compile to iter0/paper.pdf:

cd workspace/refinement/iter0/ && latexmk -pdf -interaction=nonstopmode paper.tex

Score it (see Step 1 below) → iter0/score.json.

1. Simulate peer review

For each iteration N starting from 1:

Writing quality pre-check (start of every iteration): Load references/writing-quality-check.md and run the 5-category checklist (Categories A–E) against the current draft. Note violations and add them to the revision agenda.

Update critique memory before the reviewer call (iter N ≥ 2 only — skip for iter 1):

python skills/content-refinement-agent/scripts/update_critique_memory.py \
    --worklog workspace/refinement/worklog.json \
    --review  workspace/refinement/iter<N-1>/review.json \
    --iter    <N> \
    --out     workspace/refinement/critique_memory.json

This produces critique_memory.json with focus_on (persistent unresolved issues) and do_not_reflag (already-resolved issues). Inject both lists into the reviewer system prompt verbatim:

CRITIQUE MEMORY — you must honour this before reviewing:

FOCUS ON (flagged in prior iterations, not yet resolved — prioritise these):
<critique_memory.focus_on items, one per line>

DO NOT RE-FLAG (already addressed in prior iterations):
<critique_memory.do_not_reflag items, one per line>

This prevents the reviewer from re-discovering already-fixed issues and from missing genuinely stuck problems.

Load references/reviewer-rubric.md as the system prompt for the simulated reviewer call. The reviewer reads iter<N-1>/paper.pdf (or paper.tex if your host LLM lacks PDF input) and produces a JSON of strengths, weaknesses, questions, and per-axis scores.

The rubric is structured to mimic AgentReview (Jin et al., 2024) — the paper's chosen evaluator. We ship a faithful rubric in the references directory; the host agent's LLM does the actual reviewing.

Devil's Advocate reviewer: One simulated reviewer must be designated the DA following references/da-reviewer.md. The DA challenges core claims from first principles (causal overclaiming, ablation coverage, baseline fairness, generalization claims, novelty inflation) rather than surface polish. If the DA issues a CRITICAL finding that remains unaddressed after all reviewers weigh in, that finding blocks the "refinement accepted" decision regardless of rubric scores. Log DA CRITICAL findings in worklog.json: {da_critical: true, finding: "..."}.

Record the DA's per-round findings and concession decisions in workspace/refinement/da_concessions.json (schema in references/da-reviewer.md) and enforce the concession-threshold protocol deterministically — this stops the simulated DA from sycophantically caving:

python skills/content-refinement-agent/scripts/concession_guard.py \
    --log workspace/refinement/da_concessions.json \
    --out workspace/refinement/iter<N>/da_guard.json
# exit 0 = clear; exit 1 = standing CRITICAL → force REVERT this iteration;
# exit 2 = a concession was rejected (caving/consecutive) → DA must restate;
# exit 3 = schema error.

The guard rejects any concession made at rebuttal_score < 4 or in a round immediately following another concession, and restores the affected finding to "standing". A standing CRITICAL (exit 1) overrides an ACCEPT into a REVERT.

Save to workspace/refinement/iter<N>/review.json.

2. Score the draft

The reviewer call produces both qualitative feedback and a per-axis score:

{
  "axis_scores": {
    "scientific_depth":     {"score": 65, "justification": "..."},
    "technical_execution":  {"score": 70, "justification": "..."},
    "logical_flow":         {"score": 60, "justification": "..."},
    "writing_clarity":      {"score": 55, "justification": "..."},
    "evidence_presentation":{"score": 72, "justification": "..."},
    "academic_style":       {"score": 68, "justification": "..."}
  },
  "overall_score": 64.5,
  "decision_band": "Major Revision",
  "strengths": [...],
  "weaknesses": [...],
  "questions": [...]
}

Save to iter<N>/score.json. (Combined with review.json if your host emits one document; the schemas overlap.)

decision_band is derived deterministically from overall_score — Accept (≥80) / Minor Revision (65–79) / Major Revision (50–64) / Reject (<50). Fill it in with python skills/content-refinement-agent/scripts/decision_band.py --score-json iter<N>/score.json rather than by hand, so it can never disagree with the number. The bands drive the target-met halt in Step 5.

3. Apply revision

Load the verbatim Content Refinement Agent prompt at references/prompt.md. Prepend the Anti-Leakage Prompt. Inputs:

  • paper.tex — current draft
  • paper.pdf — compiled PDF (multimodal context if available)
  • conference_guidelines.md
  • experimental_log.md — ground truth for numeric claims
  • worklog.json — history of previous changes
  • citation_pool.json — the allowed bibliography
  • reviewer_feedback — the JSON from Step 1

The prompt instructs the model to address weaknesses, integrate question answers, and emit two output blocks:

  1. A worklog JSON {addressed_weaknesses[], integrated_answers[], actions_taken[]}
  2. The full revised LaTeX code

Save the revised LaTeX as iter<N>/paper.tex. Append the worklog JSON to workspace/refinement/worklog.json via apply_worklog.py.

4. Compile and re-score

cd workspace/refinement/iter<N>/ && latexmk -pdf -interaction=nonstopmode paper.tex

Then re-run the simulated review on the new draft → updated score.json for the new iteration. (This is the "re-score after revision" call.)

5. Apply the accept/revert decision

The calling loop must track CONSECUTIVE_SMALL (starts at 0) and pass it on each call so score_delta.py can detect the plateau:

python skills/content-refinement-agent/scripts/score_delta.py \
    --prev workspace/refinement/iter<N-1>/score.json \
    --curr workspace/refinement/iter<N>/score.json \
    --plateau-threshold 1.0 \
    --plateau-streak 3 \
    --accept-threshold 80 \
    --consecutive-small $CONSECUTIVE_SMALL \
    > workspace/refinement/iter<N>/delta.json

EXIT=$?
# Update streak for next iteration:
CONSECUTIVE_SMALL=$(python3 -c "
import json
d = json.load(open('workspace/refinement/iter<N>/delta.json'))
print(d['consecutive_small'])
")

Exit codes:

  • 0 — ACCEPT (overall improved or tied with non-negative net sub-axis, below the Accept band, no plateau)
  • 1 — REVERT (overall decreased)
  • 2 — REVERT (tied overall, but net sub-axis change negative)
  • 4 — HALT_PLATEAU (accepted but N consecutive iterations below threshold — stop early)
  • 5 — HALT_TARGET_MET (accepted AND reached the Accept band, overall ≥ 80 — stop)

Behavior:

  • ACCEPT (exit 0): keep iter<N>/paper.tex as the new best. Continue to iter N+1.
  • REVERT (exit 1 or 2): copy iter<N-1>/paper.tex back as canonical, halt.
  • HALT_PLATEAU (exit 4): keep current (it was accepted), but stop — further iterations are unlikely to yield meaningful gains. In practice ~85% of refinement gain comes in iteration 1; the plateau fires when subsequent iterations improve by less than 1 point for 3 consecutive rounds.
  • HALT_TARGET_MET (exit 5): keep current (it was accepted), but stop — the paper has reached the Accept band (overall ≥ 80), so there is no reason to keep iterating and risk a regression. The delta.json carries decision_band_prev / decision_band_curr for the run report.

Override — DA CRITICAL. If concession_guard.py (Step 1) returned exit 1 for this iteration, treat the outcome as REVERT even when score_delta.py says ACCEPT: roll back to iter<N-1>/paper.tex and require the next revision to address the standing CRITICAL finding.

Always log the decision via apply_worklog.py --decision ....

6. Halt rules

Halt the loop when ANY of these is true:

  1. Iteration count reaches ITER_CAP (default 3).
  2. score_delta.py returned exit code 1 or 2 (REVERT), OR concession_guard.py returned exit 1 (standing DA CRITICAL → forced REVERT).
  3. The simulated reviewer's weaknesses list is empty (no actionable feedback to apply).
  4. score_delta.py returned exit code 4 (HALT_PLATEAU — plateau early-stop).
  5. score_delta.py returned exit code 5 (HALT_TARGET_MET — reached the Accept band, overall ≥ 80; promote the current draft).

7. Promote the best snapshot

Identify the iteration with the highest accepted overall_score (this may be the latest accepted iteration, OR an earlier one if a later iteration was reverted). Copy:

cp workspace/refinement/iter<best>/paper.tex workspace/final/paper.tex
cp workspace/refinement/iter<best>/paper.pdf workspace/final/paper.pdf

Then in the final report, tell the user:

  • How many iterations were run
  • The final overall score and its decision band (Accept / Minor / Major / Reject)
  • The score trajectory with bands (e.g., "iter0 58.0 Major → iter1 67.3 Minor (accept) → iter2 81.0 Accept (halt: target met)")
  • Which iteration was promoted, and the halt reason (revert / plateau / target met / iter cap / DA critical)

Critical safety constraints (App. F.1 page 50–51)

The paper explicitly notes that early versions of the Refinement Agent "exploited the automated reviewer's scoring function by superficially listing missing baselines as limitations to artificially inflate acceptance scores." The verbatim prompt forbids this. You must honor it:

  • [IRON RULE] Halt on score regression. If score_delta.py returns exit code 1 or 2 (REVERT), immediately revert to the previous snapshot and halt. No further revision attempts are permitted after a regression.
  • [IRON RULE] No new experiments in revision. Ignore reviewer requests for new experiments, ablations, or baselines. The Refinement Agent's job is presentation, not new science. If the reviewer asks for missing data, simply skip those points — do NOT add fabricated experiments, do NOT add a "future work" item promising them.
  • [IRON RULE] All numeric claims must match experimental_log.md. The agent cannot introduce new numbers, only re-present existing ones. Any number in the revised paper that does not appear in experimental_log.md is a hallucination.
  • Never explicitly state a limitation. The phrase "we acknowledge as a limitation that..." is forbidden. The model can address weaknesses through clearer explanation, but must not game the evaluator by listing them defensively.

These rules prevent reward hacking and keep the refinement loop honest.

Resources

  • references/prompt.md — verbatim Content Refinement Agent prompt from App. F.1
  • references/reviewer-rubric.md — AgentReview-style scoring rubric (6 axes)
  • references/halt-rules.md — accept/revert/halt logic in formal pseudocode
  • references/safe-revision-rules.md — anti-reward-hack constraints
  • references/writing-quality-check.md — 5-category anti-AI-prose checklist (pointer to shared)
  • references/ai-failure-modes.md — 7-mode integrity gate run before first iteration (pointer to shared)
  • references/da-reviewer.md — Devil's Advocate reviewer protocol and concession rules
  • scripts/score_delta.py — accept/revert/halt decision from two score JSONs; emits decision bands + target-met halt (exit 5)
  • scripts/decision_band.py — map an overall score to a canonical decision band (Accept/Minor/Major/Reject)
  • scripts/concession_guard.py — enforce the DA concession-threshold protocol; blocks accept on a standing CRITICAL
  • scripts/score_trajectory.py — per-dimension score history, regression and plateau detection
  • scripts/apply_worklog.py — append iteration entries to worklog.json
  • scripts/snapshot.py — copy paper.tex/paper.pdf into iter/ for rollback
  • scripts/update_critique_memory.py — NEW build/update critique_memory.json from worklog + review (AutoSci-inspired reviewer memory)
  • skills/shared/writing_quality_check.md — full anti-AI-prose checklist (5 categories)
  • skills/shared/ai_failure_modes.md — full AI research failure modes gate (7 modes)
  • skills/shared/handoff_schemas.md — formal data contracts between all pipeline steps
  • skills/shared/research_brief_template.md — NEW research brief schema (read §1–§4 before first reviewer call)

发现问题?提交给管理员复核

评分:

评论 (0)

暂无评论,成为第一个评论者吧!