SkillAtlasSkill 详情

sn-ppt-entry

The SenseNova model family plugs directly into agent runtimes such as OpenClaw and hermes-agent...

审核状态:已审核Quality 80Security 62

复制安装命令

用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。

复制前请先查看来源、License 和安全提示。

项目 README

来源文件:README.md

抓取于 2026年8月3日

SenseNova-Skills

English | 简体中文

Website Raccoon Token Plan SenseNova U1 SenseNova 6.7

The SenseNova model family plugs directly into agent runtimes such as OpenClaw and hermes-agent, with the skills in this repository extending the models with concrete, end-to-end office capabilities.

In this repository each skill lives in its own directory and declares triggers, capabilities, and execution flow through a SKILL.md file, following the Agent Skills convention.

The skills cover image generation & visualization, slide-deck (PPT) generation, Excel data analysis, and deep research — usable standalone or composed into end-to-end workflows.

🎨 Want to see what it can do? Check out our sn-infographic Gallery to explore nearly 100 stunning generation cases and steal their prompt designs !

🦝 Available out-of-the-box in Raccoon

The latest SenseNova models and the full Cowork-Skill suite in this repo are bundled into Raccoon, with enterprise-grade security and a zero-setup experience — if you'd rather not provision env, API keys, and runtimes yourself, you can use these capabilities directly through Raccoon. Free trial available — no payment required to get started.

Raccoon now ships a full upgrade across product capability and client experience:

  • Three core office capabilities, strengthened: powered by SenseNova 6.7 Flash + Cowork-Skill, data analysis, PPT generation, and task planning each take a step up — covering the full loop from multi-file cleaning/analysis to formal report decks, industry/competitive research, and investment memos.
  • New: infographic generation: built on the SenseNova U1 model, compresses complex data, long reports, and business insights into dense, structured, visual infographics that are easier to digest and share.
  • New client + local Agent OS: the cloud model handles heavy reasoning and multimodal understanding; the local Agent OS sits next to your files, work context, and personal habits — delivering a more personalized, local, and secure AI-native office experience.
  • Proven at scale: chosen by 15M+ individual users and thousands of enterprise customers.

👉 Try it: xiaohuanxiong.com

How to Use

These skills are designed to run inside an Agent Skills-compatible agent.

Recommended: let the agent install the skills for you. Hand it the repo URL and ask it to clone and drop the skills into the right directory — for example:

"Please install SenseNova-Skills from https://github.com/OpenSenseNova/SenseNova-Skills into your skills directory."

After it finishes, you may need to manually restart the agent service before the new skills are picked up.

AgentTarget directory
OpenClaw~/.openclaw/skills/
hermes-agent~/.hermes/skills/
Prefer to install manually?

Clone this repository, then copy the subdirectories under skills/ into the target directory yourself:

git clone https://github.com/OpenSenseNova/SenseNova-Skills.git --depth=1
mkdir -p ~/.openclaw/skills
cp -r SenseNova-Skills/skills/* ~/.openclaw/skills/

For Hermes, swap the target to ~/.hermes/skills/.

Per-category Python dependencies, API keys, and invocation examples are documented in the 📖 Full guide for each section.

Skills List

🎨 Image & Visualization

📖 Full guide: docs/sn-image-generate_en.md (prerequisites, Quick Start, API config, and invocation samples).

NameLabelDescription
sn-image-doctorEnvironment DoctorValidates the SenseNova-Skills environment — checks sn-image-base install, Python deps, and required env vars; interactively fills missing values into .env.
sn-image-baseImage Base Layer (Tier 0)Low-level tools — text-to-image (sn-image-generate), image recognition (sn-image-recognize), and text optimization (sn-text-optimize) — exposed through a unified sn_agent_runner.py, designed to be called by upper-layer skills.
sn-infographicInfographic Generation (Tier 1)Auto prompt-quality scoring, layout/style selection (87 layouts / 66 styles), multi-round generation with VLM review and quality ranking, producing publication-ready infographics.
sn-image-imitateImage Imitation (Tier 1)Given one reference image and a target content prompt, generates a new image that imitates the reference.
sn-image-resumeResume Image Generation (Tier 1)Given resume information, generates a resume image.

📊 Presentations (PPT)

📖 Full guide: docs/sn-ppt-generate.md (prerequisites, Quick Start, API config, and invocation samples).

NameLabelDescription
sn-ppt-entryPPT Entry PointUnified entry point for PPT generation. Asks the user to choose fast, standard, or creative mode, then collects role / audience / scenario / page count. For standard mode, also asks about image sourcing (AI, web search, or none) and chart rendering (U1 infographics or ECharts). Parses uploaded pdf / docx / md / txt, emits task_pack.json + info_pack.json, and dispatches to the chosen mode.
sn-ppt-doctorPPT Environment DoctorEnvironment check for the PPT pipeline — validates sn-image-base, API keys, the Node runtime, and optional deps; writes missing required vars into .env.
sn-ppt-creativePPT Creative ModeOne full-page 16:9 PNG per slide, generated via sn-image-generate with a per-page composed prompt. Falls back to web image search when T2I generation fails.
sn-ppt-standardPPT Standard & Faststyle_spec → outline → asset plan + per-slot images + VLM QC → per-page HTML → per-page review → PPTX export. Fast mode builds a complete draft immediately with autonomous decisions, then provides structured refinement suggestions. Supports AI-generated infographics (U1) for diagrams and web image search (Serper) for real photos.

📈 Data Analysis (DA)

📖 Full guide: docs/sn-data-analysis.md (prerequisites, Quick Start, API config, and invocation samples).

NameLabelDescription
sn-da-excel-workflowExcel Analysis OrchestrationEnd-to-end Excel pipeline — multi-sheet read, large-file detection (≥10k rows triggers Parquet), cleaning, conditional filtering, cross-sheet aggregation, and Excel/CSV export.
sn-da-image-captionImage Understanding & Data ExtractionFor image-first inputs — table OCR, chart understanding, screenshot/UI description; parses captions into DataFrames, recreates visualizations, exports Excel/CSV.
sn-da-large-file-analysisHigh-Performance Large-File AnalysisStreaming reads for ≥10k-row Excel datasets (openpyxl read_only + iter_rows), Parquet conversion, memory optimization, chunked processing, large-file writes.

🔬 Deep Research

📖 Full guide: docs/sn-deep-research.md (prerequisites, web_search precheck, Quick Start, and per-stage invocation).

NameLabelDescription
sn-deep-researchDeep Research Entry PointUnified deep-research orchestrator with true-dependency DAGs, reusable source snapshots, and evidence-informed content units, producing final report.md.
sn-research-reportFinal Report Writing & EditingRenders the judgment layer into the final report.md; also handles targeted rewrites — restructuring, polishing, table-augmentation — for an existing draft.
sn-report-format-discoveryPresentation-Format DiscoveryCompares final forms such as a research report, academic paper, table-first analysis, decision memo, or a custom Markdown form; scout uses it before research and user confirmation.
sn-prepare-citationsCitation RenderingPost-processes [^source_id] footnotes into numbered citations and appends references from evidence sources.
sn-md-to-html-reportMarkdown → HTML ReportConverts the research report.md (or any Markdown doc) into a clean, single-file HTML reading view that opens offline — embedded images, side-panel TOC, responsive tables, and table-delimiter repair.

🔍 Search

📖 Search skills are documented together with deep research: docs/sn-deep-research.md (includes per-platform API keys, invocation, and unified JSON output).

NameLabelDescription
sn-search-academicAcademic SearchArXiv (with section-level HTML reading) / Semantic Scholar (with citation counts) / PubMed (with PMC open-access full text) / Wikipedia, in one aggregated interface.
sn-search-codeDeveloper SearchGitHub (repo / code / issue) / Stack Overflow / Hacker News / HuggingFace (models / datasets / spaces), aggregated.
sn-search-social-cnChinese Social SearchBilibili / Zhihu / Douyin search; some platforms require cookie auth.
sn-search-social-enEnglish Social SearchReddit / Twitter (X) / YouTube search.

Sample Outputs

🎨 Infographic (sn-infographic)

A few sn-infographic outputs (more in docs/sn-infographic-examples.md).

sn-infographic sample outputs

🧩 Memory price analysis — insight → analysis → presentation → end-to-end workflow

examples/memory-price-end2end-analysis. Starting from a raw quote CSV, the agent profiles fields, normalizes categories and timestamps, then attacks the rally from three angles — overall trend, top movers per category, and the gap between server-grade and consumer-grade SKUs — locating a late-February inflection along the way. Treating those findings as the research question, it switches to deep research: planning per-dimension web searches over supply contraction, AI-server demand, and vendor output discipline, then triaging and cross-checking evidence across sources before committing it to the report. The data and research conclusions are then handed to PPT generation, which lays out a 16-page outline, plans per-slot imagery, renders per-page HTML, runs VLM review, and finally composites screenshots into the PPTX. The result is a clear three-step storyline: prices are rising → here is why → here is what to do. This is the only example that exercises the full data analysis → deep research → PPT chain end-to-end.

📊 Employee performance analysis — data analysis

examples/employee-performance-analysis. The agent reads 10 separate monthly review xlsx files, aligns column schemas across months and joins them into one longitudinal table. From that table it produces aggregate views — monthly average trend, score-distribution boxplots, grade mix change, and a 38-role ranking — and individual views — top performers, needs-attention, and consistently-improving cohorts plus per-employee year trends. The findings are written up with explicit improvement suggestions tied to specific roles and individuals, backed by 8 supporting charts. The same content is delivered as a Word doc (for distribution) and a visualized HTML report (for browsing). The example shows how sn-da-excel-workflow handles "many small spreadsheets that should be one analysis" rather than a single big file.

🔬 Embodied AI industry research — deep research

examples/embodied-ai-deep-research. Given only an industry name, the agent first commits to a research plan — market size, vendor share, financing, cost structure, development roadmap — instead of jumping straight into search. For each dimension it runs targeted web searches, fetches and reads source pages, and extracts both numeric and qualitative evidence; conflicting figures across sources are explicitly reconciled before being trusted. A synthesis stage organizes per-dimension evidence into a traceable, reader-oriented information structure rather than a stack of disconnected bullets. The output is an illustrated report (Markdown + visualized HTML) with 5 dimension-specific charts. The example shows how sn-deep-research turns "go research X" into a structured plan-then-execute loop with traceable evidence.

🎯 Property fee pricing — PPT generation

examples/property-fee-pricing-ppt. The agent takes a free-form brief — topic (property fee pricing), audience (property staff + committee), 26 pages, black-and-white warm style — and first commits to an outline plus a per-page asset plan that conforms to the style spec. Each slide is then built as semantic per-page HTML rather than free-form image generation: copy, layout, illustrations, icons, and any data charts are reasoned about per slot. Imagery is produced or selected per slot and VLM-checked against the page's intent; each rendered page goes through a review pass with optional rewrite for coherence and copy quality. Final pages are screenshotted and composited into the PPTX, with the per-page HTML kept alongside for direct browser preview or re-editing. The example demonstrates sn-ppt-standard style consistency on a long, prose-heavy deck where every slide must obey the same audience and palette constraints.

FAQ

Common setup and runtime questions (400/401 errors, rate limits, PPT timeouts, infographic quality, model names) are answered in docs/faq.md.

Contributing

Feel free to use the skills here as templates for your own OpenClaw skills. The qualities that make a skill good:

  • Clear triggers: state in description exactly when the skill should and should not run, so the agent recognizes it accurately
  • Focused scope: each skill does one thing well; complex workflows compose multiple skills
  • Solid documentation: examples, artifact contracts, edge cases, failure handling
  • Supporting resources: use references/, scripts/, prompts/ to provide additional context

Join the Community

Join our growing community to share feedback, get support, and stay updated on the latest developments. Scan the QR code below to hop into the chat — we'd love to hear from you!

DiscordLark Group

License

MIT — see LICENSE.

文档与办公

高风险

  • 来源需自行核对维护者身份。
  • 包含脚本或命令调用,安装前请复核。
  • 未检测到明显外部权限要求。
  • 存在潜在风险命令,请谨慎安装。
  • 扫描发现:5 条。

Codex — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/OpenSenseNova/SenseNova-Skills.git
  3. 将 "skills/sn-ppt-entry" 文件夹复制到 Codex 的 skills 目录中。
  4. 重启 Codex 让新的 skill 生效。

Codex — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Codex 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Codex 让新的 skill 生效。

Claude Code — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/OpenSenseNova/SenseNova-Skills.git
  3. 将 "skills/sn-ppt-entry" 文件夹复制到 Claude Code 的 skills 目录中。
  4. 重启 Claude Code 让新的 skill 生效。

Claude Code — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Claude Code 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Claude Code 让新的 skill 生效。

Cursor — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/OpenSenseNova/SenseNova-Skills.git
  3. 将 "skills/sn-ppt-entry" 文件夹复制到 Cursor 的 skills 目录中。
  4. 重启 Cursor 让新的 skill 生效。

Cursor — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Cursor 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Cursor 让新的 skill 生效。

GitHub Copilot — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/OpenSenseNova/SenseNova-Skills.git
  3. 将 "skills/sn-ppt-entry" 文件夹复制到 GitHub Copilot 的 skills 目录中。
  4. 重启 GitHub Copilot 让新的 skill 生效。

GitHub Copilot — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 GitHub Copilot 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 GitHub Copilot 让新的 skill 生效。

Windsurf — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/OpenSenseNova/SenseNova-Skills.git
  3. 将 "skills/sn-ppt-entry" 文件夹复制到 Windsurf 的 skills 目录中。
  4. 重启 Windsurf 让新的 skill 生效。

Windsurf — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Windsurf 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Windsurf 让新的 skill 生效。
查看 SKILL.md 原文
name: sn-ppt-entry
description: |
  Entry point for PPT generation. Asks the user to choose a mode (fast,
  standard, or creative), then collects role / audience / scene / page_count
  as needed. For standard mode, also asks how images should be sourced (AI
  generation, web search, or none), whether charts should use AI-generated
  infographics or ECharts, and whether the final deliverable should be PPTX or
  PDF. Parses uploaded pdf/docx/md/txt files, produces
  task_pack.json + info_pack.json in a new deck_dir, then dispatches to
  sn-ppt-creative or sn-ppt-standard. Fast mode skips optional questions and
  gets straight to building. Use when the user asks to make a PPT /
  presentation / 演示 / PPT. If the user asks to open, preview, inspect, or
  edit previously generated HTML slides in the WebUI/workbench without
  regenerating, dispatch to sn-ppt-workbench instead of this generation entry.
metadata:
  project: SenseNova-Skills
  tier: 1
  category: scene
  user_visible: true
triggers:
  - "生成 PPT"
  - "做一套 PPT"
  - "做一份演示"
  - "sn-ppt-entry"

sn-ppt-entry

Hard preconditions

Run sn-ppt-doctor hard checks (SN_API_KEY or capability-specific API keys / node / sn-image-base) at the start of this skill. If any fails, stop and tell the user to run /skill sn-ppt-doctor.

If the user is not asking to generate a new deck and only wants to open an existing/generated deck in the WebUI, do not run these generation preconditions. Dispatch directly to /skill sn-ppt-workbench.

Flow

  1. Extract parameters from the user's message:

    • role (speaker identity)
    • audience
    • scene (where the deck will be used)
    • page_count
    • language — detect from the user's query: zh-Hans (Simplified Chinese), zh-Hant (Traditional Chinese), or en (English). Do NOT ask the user; just infer and record it. If unsure, use zh-Hans.
  2. If the user asks only to open/preview/edit an existing generated deck in the WebUI, dispatch to /skill sn-ppt-workbench deck_dir=<abs-or-user-provided-path> and stop. Do not ask mode questions.

  3. If task_pack.json + info_pack.json already exist in a deck_dir the user refers to and the user asks to continue generation, read them and jump to step 10 (see "Resume" below).

  4. Always ask the user which mode to use first. Call ask_user:

    Question — Mode: "Which generation mode should I use?"

    • "Fast mode — build the slides now so you can review and iterate"
    • "Standard mode — plan the style and content thoroughly first, then build"
    • "Creative mode — full-page AI-generated images per slide"

    Store as ppt_mode in task_pack.

  5. Only ask standard-mode option questions for standard mode. Fast mode and creative mode have fixed defaults — asking extra questions defeats the purpose of "fast."

    If ppt_mode == "standard", ask three more questions:

    Question — Normal images (decorative / conceptual): "Should I include images, and how should they be sourced?"

    • "AI generation — create images from scratch"
    • "Web search — pull real photos from the web (requires Serper API key)"
    • "No images — use text, charts, and CSS visuals only"

    If the user picks web search and SERPER_API_KEY is not set, tell them how to get a free key at https://serper.dev. Store as image_source in task_pack.params.

    Question — Infographics (charts, flowcharts, diagrams): "For charts and diagrams, should I use AI-generated infographics or ECharts?"

    • "AI-generated infographics — U1 creates custom diagram images"
    • "ECharts — rendered as interactive charts in the HTML"

    Store as infographic_source in task_pack.params ("ai-gen" or "echarts").

    Question — Final output: "Which final file format should I generate?"

    • "PPTX — editable PowerPoint deck"
    • "PDF — fixed-layout presentation file"

    Store as output_format in task_pack.params ("pptx" or "pdf"). Default to "pptx" only if resuming an older task_pack.json that lacks this field; do not silently default during a new standard-mode run.

    If ppt_mode == "fast": skip image/output questions. Default to image_source = "ai-gen", infographic_source = "echarts", and output_format = "pptx". Also skip role/audience/scene/page_count questions — infer reasonable defaults from the user's query and move directly to building slides. Fast mode means fewer questions, faster start. If the user didn't explicitly state these, make your best guess and proceed.

    If ppt_mode == "creative": skip image/output questions. Default to image_source = "ai-gen" (full-page T2I rendering) and output_format = "pptx". Infographics are not applicable. Skip role/audience/scene/page_count unless explicitly stated.

  6. Collect role -> audience -> scene -> page_count — for standard mode only. Use the wording in references/ask_user_templates.md. 2-3 options per question; do not write "其他". For fast/creative modes, infer from the query and move on.

  7. Create deck_dir — location is FIXED, do not guess:

    • Parent: always $(pwd)/ppt_decks/. In OpenClaw, cwd at skill-invocation time is the agent's workspace directory (e.g. ~/.openclaw/workspace/). Do NOT use /tmp, the home directory, the repo root, or $SKILL_DIR as the parent. Do NOT honor $PPT_DECK_ROOT either — it's been removed to avoid drift.
    • Parent directory must be created if missing: mkdir -p $(pwd)/ppt_decks.
    • Deck name: <topic_concise>_<YYYYMMDD_HHMMSS>.
    • Full deck_dir path: $(pwd)/ppt_decks/<topic_concise>_<YYYYMMDD_HHMMSS>/.
    • Immediately resolve to absolute (realpath / Path.resolve()) before writing it into task_pack.json — downstream must see an absolute path.
    • Create subdirs: pages/ always; images/ if ppt_mode in {standard, fast}.
    • If $(pwd)/ppt_decks/ cannot be created (permission denied) → abort, tell the user to check workspace permissions.
  8. If user attached reference_docs (pdf/docx/md/txt):

    • Run $SKILL_DIR/scripts/parse_user_docs.py --files <paths...> --output <deck_dir>/raw_documents.json. The --output flag tells the script to write the JSON itself (recommended — works reliably even on agents that don't handle shell redirection well). The script prints a single-line JSON status {"status":"ok","output":"...","documents":N,"errors":M} to stdout when --output is used.
    • Call the LLM with $SKILL_DIR/prompts/document_digest.md as system prompt + (user_query + concatenated document text) as user prompt. See "Invoking the LLM" below.
    • On success: write document_digest JSON into info_pack.document_digest.
    • On failure: degrade — set info_pack.document_digest = null, continue (do NOT abort entry).
  9. Write task_pack.json + info_pack.json to deck_dir (see "Schemas" below). All path-bearing fields absolute.

  10. Start the generation progress WebUI (best-effort, non-blocking):

  • First write the initial progress event:
    python3 $PPT_STANDARD_DIR/scripts/progress_event.py --deck-dir <deck_dir> --stage entry --status ok --artifact "task_pack.json / info_pack.json" --label "task_pack.json / info_pack.json 已写入"
    
  • Then start/reuse the WebUI and immediately echo the returned generation_url:
    python3 $PPT_STANDARD_DIR/scripts/launch_workbench.py --deck-dir <deck_dir> --source-session-id "${HERMES_SESSION_KEY:-}" --agent-managed 1
    
  • On native Windows Hermes installs where python3 is unavailable, use python.
  • If the helper returns {"status":"ok",...}, tell the user 生成进度工作台已启动:<generation_url> before continuing. The returned generation_url is the progress page at /progress; the editor is a separate /editor URL exposed as editor_url.
  • If it returns {"status":"skipped","reason":"nodejs_missing",...}, ask the user whether to install NodeJS/dependencies. If they decline, say generation will continue without the WebUI and proceed. If they agree, first use any approved dependency-install skill/tool exposed by the active environment; otherwise use platform install means only after explicit dangerous-operation confirmation.
  • If it returns any other skipped/failed status, echo a short reason and continue generation without the WebUI.
  • Default bind behavior is handled by the helper: explicit host/env wins; otherwise Docker/WSL binds 0.0.0.0, native hosts bind localhost unless the user requested another IP/host.
  1. Caption every image once with VLM (mandatory, idempotent — runs after info_pack.json is written so both pools are visible):
PPT_STANDARD_DIR="$(dirname "$SKILL_DIR")/sn-ppt-standard" python3 $SKILL_DIR/scripts/caption_images.py --deck-dir <deck_dir>

⚠️ Set PPT_STANDARD_DIR — caption_images.py imports from sn-ppt-standard/lib/model_client.py and resolves it via $PPT_STANDARD_DIR. Without this env var, the script fails with FileNotFoundError: ppt-standard/lib/model_client.py not found. On Windows, use python instead of python3 and set the env var inline:

set PPT_STANDARD_DIR=C:\Users\...\Repository\ppt-editor\skills\sn-ppt-standard && python %SKILL_DIR%\scripts\caption_images.py --deck-dir <deck_dir>

Safe to skip when there are no attachment images (user_assets.reference_images empty and no doc-embedded images) — the script is a no-op with no images, and the failure is harmless. This script is the single source of truth for image-content descriptions:

  • Pool A — doc-embedded images (raw_documents.json documents[*].inherited_images[*]): caption written into the same JSON as vlm_caption.
  • Pool B — standalone uploads (info_pack.user_assets.reference_images): caption written into a sister field info_pack.user_assets.reference_image_captions: {abs_path: caption}.
  • Already-captioned images are skipped silently, so re-running is cheap and safe. Only newly added images incur a VLM call.
  • Failures don't abort: the script reports them in the JSON status; downstream stages fall back to filename / alt / digest hint when a caption is missing. Downstream (sn-ppt-standard cmd_page_html) reads these cached captions and never re-captions — that's the "single source of truth" rule. If you change image files in a deck, delete their vlm_caption (or reference_image_captions[path]) entry and re-run this script to refresh.
  1. Dispatch to sn-ppt-creative or sn-ppt-standard based on task_pack.ppt_mode.

ask_user boundary conditions

  • User answers multiple params in one turn -> extract all with a single sn-text-optimize call; skip asked-already params.
  • User's answer isn't in the 2-3 options -> record verbatim; don't force into the enumeration.
  • Session interrupted before task_pack.json written -> discard temp params; next entry starts over.
  • task_pack.json already exists -> skip param collection, go straight to dispatch.

Invoking the LLM for document_digest

parse_user_docs.py --output <deck_dir>/raw_documents.json already creates the file. Then call the LLM with a user prompt that gives only counts + indices of tables/images (not row contents) so the LLM can't accidentally paraphrase numbers:

python3 -c "
import sys, json, pathlib
sys.path.insert(0, '$PPT_STANDARD_DIR/lib')
from model_client import llm

raw = json.loads(pathlib.Path('<deck_dir>/raw_documents.json').read_text())

# Build the digest-safe view: strip tables[] and image paths, keep text + indices
docs_view = []
for d in raw.get('documents', []):
    docs_view.append({
        'doc_index': d['doc_index'],
        'type': d['type'],
        'text': d.get('text',''),
        'tables_count': len(d.get('tables') or []),
        'images_count': len(d.get('inherited_images') or []),
    })

user_prompt = json.dumps({
    'user_query': '<the user's original query>',
    'documents': docs_view,
}, ensure_ascii=False)

sys_prompt = open('$SKILL_DIR/prompts/document_digest.md').read()

out = llm(sys_prompt, user_prompt)
# Parse JSON; if it fails, degrade digest to null (not abort entry)
try:
    digest = json.loads(out)
except Exception:
    digest = None
pathlib.Path('<deck_dir>/digest_tmp.json').write_text(json.dumps(digest, ensure_ascii=False))
"

The digest JSON then merges into info_pack.document_digest. Downstream stages (outline, page_html) read both info_pack.document_digest (structured summary + inherited_tables/images index lists) AND raw_documents.json (actual table rows + image paths).

Substitute $PPT_STANDARD_DIR with the sn-ppt-standard skill install dir.

Schemas

task_pack.json:

{
  "deck_id": "AI产品发布会_20260318_154500",
  "deck_dir": "/abs/path/ppt_decks/AI产品发布会_20260318_154500",
  "ppt_mode": "standard",
  "params": {
    "role": "...",
    "audience": "...",
    "scene": "...",
    "page_count": 10,
    "language": "zh",
    "image_source": "ai-gen",
    "infographic_source": "ai-gen",
    "output_format": "pptx"
  },
  "created_at": "2026-04-21T15:45:00+08:00",
  "skill_version": "0.1.0"
}

info_pack.json:

{
  "user_query": "...",
  "user_assets": {
    "reference_images": ["/abs/..."],
    "reference_docs": ["/abs/..."],
    "reference_docs_failed": []
  },
  "document_digest": {
    "topic_summary": "...",
    "key_sections": [],
    "key_points": [],
    "data_highlights": [],
    "inherited_tables": [{"doc_index": 0, "table_index": 2, "title_hint": "..."}],
    "inherited_images": [{"doc_index": 0, "image_index": 0, "caption_hint": "..."}]
  },
  "raw_document_excerpts": {
    "enabled": true,
    "path": "/abs/.../raw_documents.json"
  }
}

🚫 Hard rules

  1. Do NOT use python-pptx, pptxgenjs, or any alternative PPTX builder. PPTX is produced by the downstream mode skills (sn-ppt-standard / sn-ppt-creative) through their designated scripts. Never pip install python-pptx or write Node scripts that import pptxgenjs.
  2. Wait for ask_user responses. When you ask the user a question, do NOT proceed until they reply. Never continue with assumed or default values.
  3. Validate paths before writing. Always ls or pwd to verify the current working directory before creating files. The only valid output location is $(pwd)/ppt_decks/<deck_dir>/. Never write to /workspace/, /tmp/, ~/, or any hallucinated path. If a path doesn't start with the verified $(pwd), it's wrong.

Failure handling

  • Missing required env var -> stop, tell user /skill sn-ppt-doctor.
  • $(pwd)/ppt_decks/ not creatable / not writable -> stop, tell user to check workspace permissions.
  • Per-file doc parse failure -> record in reference_docs_failed, continue.
  • document_digest LLM failure -> set to null, continue.

Progress echo — MANDATORY

Emit a short chat reply at each boundary. Silence between ask_user rounds and mode dispatch is a bug.

WhenExample
Right after entering sn-ppt-entry已进入 sn-ppt-entry,开始收集参数...
Missing a param缺少参数:<role>,马上问你 (then ask_user)
All params collected参数齐备:mode=standard, image_source=ai-gen, output_format=pptx, role=...。开始创建 deck_dir...
Before doc parse检测到 2 个附件,开始解析...
After doc parse解析完成:sample.pdf (12 页) / sample.docx (45 段)
Before digest[LLM] 正在汇总文档要点...
After digest文档摘要已入 info_pack.json
task_pack / info_pack writtentask_pack.json / info_pack.json 已写入 <deck_dir>
After WebUI launch生成进度工作台已启动:<url> or NodeJS 不可用;将继续生成但不启动 WebUI
Dispatching分发到 sn-ppt-creative(deck_dir=...)

Output and handoff

Final message includes a short summary:

准备就绪:
- 模式: <creative | standard>
- 页数: <n>
- deck_dir: <abs path>
即将进入<创意 | 标准>模式...

Then dispatch:

  • ppt_mode=creative -> invoke /skill sn-ppt-creative deck_dir=<abs>
  • ppt_mode=standard -> invoke /skill sn-ppt-standard deck_dir=<abs>

Does NOT

  • Do not generate any style / outline / page content (that's the mode skill's job).
  • Do not run any image generation.

发现问题?提交给管理员复核

评分:

评论 (0)

暂无评论,成为第一个评论者吧!