SkillAtlasSkill 详情

sn-ppt-creative

The SenseNova model family plugs directly into agent runtimes such as OpenClaw and hermes-agent...

审核状态:已审核Quality 80Security 62

复制安装命令

用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。

复制前请先查看来源、License 和安全提示。

项目 README

来源文件:README.md

抓取于 2026年8月3日

SenseNova-Skills

English | 简体中文

Website Raccoon Token Plan SenseNova U1 SenseNova 6.7

The SenseNova model family plugs directly into agent runtimes such as OpenClaw and hermes-agent, with the skills in this repository extending the models with concrete, end-to-end office capabilities.

In this repository each skill lives in its own directory and declares triggers, capabilities, and execution flow through a SKILL.md file, following the Agent Skills convention.

The skills cover image generation & visualization, slide-deck (PPT) generation, Excel data analysis, and deep research — usable standalone or composed into end-to-end workflows.

🎨 Want to see what it can do? Check out our sn-infographic Gallery to explore nearly 100 stunning generation cases and steal their prompt designs !

🦝 Available out-of-the-box in Raccoon

The latest SenseNova models and the full Cowork-Skill suite in this repo are bundled into Raccoon, with enterprise-grade security and a zero-setup experience — if you'd rather not provision env, API keys, and runtimes yourself, you can use these capabilities directly through Raccoon. Free trial available — no payment required to get started.

Raccoon now ships a full upgrade across product capability and client experience:

  • Three core office capabilities, strengthened: powered by SenseNova 6.7 Flash + Cowork-Skill, data analysis, PPT generation, and task planning each take a step up — covering the full loop from multi-file cleaning/analysis to formal report decks, industry/competitive research, and investment memos.
  • New: infographic generation: built on the SenseNova U1 model, compresses complex data, long reports, and business insights into dense, structured, visual infographics that are easier to digest and share.
  • New client + local Agent OS: the cloud model handles heavy reasoning and multimodal understanding; the local Agent OS sits next to your files, work context, and personal habits — delivering a more personalized, local, and secure AI-native office experience.
  • Proven at scale: chosen by 15M+ individual users and thousands of enterprise customers.

👉 Try it: xiaohuanxiong.com

How to Use

These skills are designed to run inside an Agent Skills-compatible agent.

Recommended: let the agent install the skills for you. Hand it the repo URL and ask it to clone and drop the skills into the right directory — for example:

"Please install SenseNova-Skills from https://github.com/OpenSenseNova/SenseNova-Skills into your skills directory."

After it finishes, you may need to manually restart the agent service before the new skills are picked up.

AgentTarget directory
OpenClaw~/.openclaw/skills/
hermes-agent~/.hermes/skills/
Prefer to install manually?

Clone this repository, then copy the subdirectories under skills/ into the target directory yourself:

git clone https://github.com/OpenSenseNova/SenseNova-Skills.git --depth=1
mkdir -p ~/.openclaw/skills
cp -r SenseNova-Skills/skills/* ~/.openclaw/skills/

For Hermes, swap the target to ~/.hermes/skills/.

Per-category Python dependencies, API keys, and invocation examples are documented in the 📖 Full guide for each section.

Skills List

🎨 Image & Visualization

📖 Full guide: docs/sn-image-generate_en.md (prerequisites, Quick Start, API config, and invocation samples).

NameLabelDescription
sn-image-doctorEnvironment DoctorValidates the SenseNova-Skills environment — checks sn-image-base install, Python deps, and required env vars; interactively fills missing values into .env.
sn-image-baseImage Base Layer (Tier 0)Low-level tools — text-to-image (sn-image-generate), image recognition (sn-image-recognize), and text optimization (sn-text-optimize) — exposed through a unified sn_agent_runner.py, designed to be called by upper-layer skills.
sn-infographicInfographic Generation (Tier 1)Auto prompt-quality scoring, layout/style selection (87 layouts / 66 styles), multi-round generation with VLM review and quality ranking, producing publication-ready infographics.
sn-image-imitateImage Imitation (Tier 1)Given one reference image and a target content prompt, generates a new image that imitates the reference.
sn-image-resumeResume Image Generation (Tier 1)Given resume information, generates a resume image.

📊 Presentations (PPT)

📖 Full guide: docs/sn-ppt-generate.md (prerequisites, Quick Start, API config, and invocation samples).

NameLabelDescription
sn-ppt-entryPPT Entry PointUnified entry point for PPT generation. Asks the user to choose fast, standard, or creative mode, then collects role / audience / scenario / page count. For standard mode, also asks about image sourcing (AI, web search, or none) and chart rendering (U1 infographics or ECharts). Parses uploaded pdf / docx / md / txt, emits task_pack.json + info_pack.json, and dispatches to the chosen mode.
sn-ppt-doctorPPT Environment DoctorEnvironment check for the PPT pipeline — validates sn-image-base, API keys, the Node runtime, and optional deps; writes missing required vars into .env.
sn-ppt-creativePPT Creative ModeOne full-page 16:9 PNG per slide, generated via sn-image-generate with a per-page composed prompt. Falls back to web image search when T2I generation fails.
sn-ppt-standardPPT Standard & Faststyle_spec → outline → asset plan + per-slot images + VLM QC → per-page HTML → per-page review → PPTX export. Fast mode builds a complete draft immediately with autonomous decisions, then provides structured refinement suggestions. Supports AI-generated infographics (U1) for diagrams and web image search (Serper) for real photos.

📈 Data Analysis (DA)

📖 Full guide: docs/sn-data-analysis.md (prerequisites, Quick Start, API config, and invocation samples).

NameLabelDescription
sn-da-excel-workflowExcel Analysis OrchestrationEnd-to-end Excel pipeline — multi-sheet read, large-file detection (≥10k rows triggers Parquet), cleaning, conditional filtering, cross-sheet aggregation, and Excel/CSV export.
sn-da-image-captionImage Understanding & Data ExtractionFor image-first inputs — table OCR, chart understanding, screenshot/UI description; parses captions into DataFrames, recreates visualizations, exports Excel/CSV.
sn-da-large-file-analysisHigh-Performance Large-File AnalysisStreaming reads for ≥10k-row Excel datasets (openpyxl read_only + iter_rows), Parquet conversion, memory optimization, chunked processing, large-file writes.

🔬 Deep Research

📖 Full guide: docs/sn-deep-research.md (prerequisites, web_search precheck, Quick Start, and per-stage invocation).

NameLabelDescription
sn-deep-researchDeep Research Entry PointUnified deep-research orchestrator with true-dependency DAGs, reusable source snapshots, and evidence-informed content units, producing final report.md.
sn-research-reportFinal Report Writing & EditingRenders the judgment layer into the final report.md; also handles targeted rewrites — restructuring, polishing, table-augmentation — for an existing draft.
sn-report-format-discoveryPresentation-Format DiscoveryCompares final forms such as a research report, academic paper, table-first analysis, decision memo, or a custom Markdown form; scout uses it before research and user confirmation.
sn-prepare-citationsCitation RenderingPost-processes [^source_id] footnotes into numbered citations and appends references from evidence sources.
sn-md-to-html-reportMarkdown → HTML ReportConverts the research report.md (or any Markdown doc) into a clean, single-file HTML reading view that opens offline — embedded images, side-panel TOC, responsive tables, and table-delimiter repair.

🔍 Search

📖 Search skills are documented together with deep research: docs/sn-deep-research.md (includes per-platform API keys, invocation, and unified JSON output).

NameLabelDescription
sn-search-academicAcademic SearchArXiv (with section-level HTML reading) / Semantic Scholar (with citation counts) / PubMed (with PMC open-access full text) / Wikipedia, in one aggregated interface.
sn-search-codeDeveloper SearchGitHub (repo / code / issue) / Stack Overflow / Hacker News / HuggingFace (models / datasets / spaces), aggregated.
sn-search-social-cnChinese Social SearchBilibili / Zhihu / Douyin search; some platforms require cookie auth.
sn-search-social-enEnglish Social SearchReddit / Twitter (X) / YouTube search.

Sample Outputs

🎨 Infographic (sn-infographic)

A few sn-infographic outputs (more in docs/sn-infographic-examples.md).

sn-infographic sample outputs

🧩 Memory price analysis — insight → analysis → presentation → end-to-end workflow

examples/memory-price-end2end-analysis. Starting from a raw quote CSV, the agent profiles fields, normalizes categories and timestamps, then attacks the rally from three angles — overall trend, top movers per category, and the gap between server-grade and consumer-grade SKUs — locating a late-February inflection along the way. Treating those findings as the research question, it switches to deep research: planning per-dimension web searches over supply contraction, AI-server demand, and vendor output discipline, then triaging and cross-checking evidence across sources before committing it to the report. The data and research conclusions are then handed to PPT generation, which lays out a 16-page outline, plans per-slot imagery, renders per-page HTML, runs VLM review, and finally composites screenshots into the PPTX. The result is a clear three-step storyline: prices are rising → here is why → here is what to do. This is the only example that exercises the full data analysis → deep research → PPT chain end-to-end.

📊 Employee performance analysis — data analysis

examples/employee-performance-analysis. The agent reads 10 separate monthly review xlsx files, aligns column schemas across months and joins them into one longitudinal table. From that table it produces aggregate views — monthly average trend, score-distribution boxplots, grade mix change, and a 38-role ranking — and individual views — top performers, needs-attention, and consistently-improving cohorts plus per-employee year trends. The findings are written up with explicit improvement suggestions tied to specific roles and individuals, backed by 8 supporting charts. The same content is delivered as a Word doc (for distribution) and a visualized HTML report (for browsing). The example shows how sn-da-excel-workflow handles "many small spreadsheets that should be one analysis" rather than a single big file.

🔬 Embodied AI industry research — deep research

examples/embodied-ai-deep-research. Given only an industry name, the agent first commits to a research plan — market size, vendor share, financing, cost structure, development roadmap — instead of jumping straight into search. For each dimension it runs targeted web searches, fetches and reads source pages, and extracts both numeric and qualitative evidence; conflicting figures across sources are explicitly reconciled before being trusted. A synthesis stage organizes per-dimension evidence into a traceable, reader-oriented information structure rather than a stack of disconnected bullets. The output is an illustrated report (Markdown + visualized HTML) with 5 dimension-specific charts. The example shows how sn-deep-research turns "go research X" into a structured plan-then-execute loop with traceable evidence.

🎯 Property fee pricing — PPT generation

examples/property-fee-pricing-ppt. The agent takes a free-form brief — topic (property fee pricing), audience (property staff + committee), 26 pages, black-and-white warm style — and first commits to an outline plus a per-page asset plan that conforms to the style spec. Each slide is then built as semantic per-page HTML rather than free-form image generation: copy, layout, illustrations, icons, and any data charts are reasoned about per slot. Imagery is produced or selected per slot and VLM-checked against the page's intent; each rendered page goes through a review pass with optional rewrite for coherence and copy quality. Final pages are screenshotted and composited into the PPTX, with the per-page HTML kept alongside for direct browser preview or re-editing. The example demonstrates sn-ppt-standard style consistency on a long, prose-heavy deck where every slide must obey the same audience and palette constraints.

FAQ

Common setup and runtime questions (400/401 errors, rate limits, PPT timeouts, infographic quality, model names) are answered in docs/faq.md.

Contributing

Feel free to use the skills here as templates for your own OpenClaw skills. The qualities that make a skill good:

  • Clear triggers: state in description exactly when the skill should and should not run, so the agent recognizes it accurately
  • Focused scope: each skill does one thing well; complex workflows compose multiple skills
  • Solid documentation: examples, artifact contracts, edge cases, failure handling
  • Supporting resources: use references/, scripts/, prompts/ to provide additional context

Join the Community

Join our growing community to share feedback, get support, and stay updated on the latest developments. Scan the QR code below to hop into the chat — we'd love to hear from you!

DiscordLark Group

License

MIT — see LICENSE.

文档与办公内容与创作

高风险

  • 来源需自行核对维护者身份。
  • 包含脚本或命令调用,安装前请复核。
  • 未检测到明显外部权限要求。
  • 存在潜在风险命令,请谨慎安装。
  • 扫描发现:4 条。

Codex — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/OpenSenseNova/SenseNova-Skills.git
  3. 将 "skills/sn-ppt-creative" 文件夹复制到 Codex 的 skills 目录中。
  4. 重启 Codex 让新的 skill 生效。

Codex — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Codex 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Codex 让新的 skill 生效。

Claude Code — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/OpenSenseNova/SenseNova-Skills.git
  3. 将 "skills/sn-ppt-creative" 文件夹复制到 Claude Code 的 skills 目录中。
  4. 重启 Claude Code 让新的 skill 生效。

Claude Code — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Claude Code 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Claude Code 让新的 skill 生效。

Cursor — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/OpenSenseNova/SenseNova-Skills.git
  3. 将 "skills/sn-ppt-creative" 文件夹复制到 Cursor 的 skills 目录中。
  4. 重启 Cursor 让新的 skill 生效。

Cursor — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Cursor 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Cursor 让新的 skill 生效。

GitHub Copilot — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/OpenSenseNova/SenseNova-Skills.git
  3. 将 "skills/sn-ppt-creative" 文件夹复制到 GitHub Copilot 的 skills 目录中。
  4. 重启 GitHub Copilot 让新的 skill 生效。

GitHub Copilot — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 GitHub Copilot 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 GitHub Copilot 让新的 skill 生效。

Windsurf — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/OpenSenseNova/SenseNova-Skills.git
  3. 将 "skills/sn-ppt-creative" 文件夹复制到 Windsurf 的 skills 目录中。
  4. 重启 Windsurf 让新的 skill 生效。

Windsurf — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Windsurf 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Windsurf 让新的 skill 生效。
查看 SKILL.md 原文
name: sn-ppt-creative
description: |
  Creative-mode PPT pipeline. One full-page 16:9 PNG per slide.
  LLM / VLM calls go through sn-ppt-standard/lib/model_client.py (shared thin
  client). Text-to-image (the actual png rendering) goes through
  sn-image-base/scripts/sn_agent_runner.py. Falls back to web image search
  when T2I generation fails. Expects task_pack.json + info_pack.json already
  written by sn-ppt-entry.
metadata:
  project: SenseNova-Skills
  tier: 1
  category: scene
  user_visible: false
triggers:
  - "sn-ppt-creative"

sn-ppt-creative

⚠️ This skill must be invoked through /skill sn-ppt-entry. Never start here directly — the entry skill collects parameters and writes task_pack.json + info_pack.json that this skill requires. If you arrived here without those files, stop and tell the user to enter via /skill sn-ppt-entry or "生成 PPT".

Call-routing policy

KindBackend
LLM (text)$PPT_STANDARD_DIR/lib/model_client.py → llm(sys, user)
VLM (image understanding)$PPT_STANDARD_DIR/lib/model_client.py → vlm(sys, user, images)
T2I (image generation)$SN_IMAGE_BASE/scripts/sn_agent_runner.py sn-image-generate

Never mix — LLM / VLM through sn-image-base, or T2I through model_client — both violate policy.

Visual asset priority

  • Creative mode renders each slide as a generated full-page PNG, so image generation is the first-priority visual path.
  • If image generation fails for a page, use web search (sn-search-image) as a fallback to find a real image that fits the page's topic. Each search result includes the image URL, source page, title, and domain for traceability.
  • Do not create placeholders. If generation and search both fail, record the page failure and continue; never write fake PNGs, grey boxes, broken-image icons, or "image pending" text.
  • Do not mention the search provider name in prompts, visible slide text, progress, or summaries.

Preconditions

  • <deck_dir>/task_pack.json exists and ppt_mode == "creative"
  • <deck_dir>/info_pack.json exists
  • <deck_dir>/pages/ exists
  • $SN_IMAGE_BASE env var (OpenClaw-injected) points at the sn-image-base skill root
  • $PPT_STANDARD_DIR env var points at the sn-ppt-standard skill root (so we can import model_client)

Any missing → stop and tell user to enter via /skill sn-ppt-entry.

Generation progress WebUI

sn-ppt-entry starts the generation progress WebUI after task_pack.json / info_pack.json are written. During creative-mode generation, publish progress with the shared writer from sn-ppt-standard:

P="python3 $PPT_STANDARD_DIR/scripts/progress_event.py"
$P --deck-dir <deck_dir> --stage creative-style --status running
$P --deck-dir <deck_dir> --stage creative-style --status ok --artifact style_spec.md
$P --deck-dir <deck_dir> --stage creative-outline --status running
$P --deck-dir <deck_dir> --stage creative-outline --status ok --artifact outline.json
$P --deck-dir <deck_dir> --stage creative-prompt --page N --status running
$P --deck-dir <deck_dir> --stage creative-prompt --page N --status ok
$P --deck-dir <deck_dir> --stage creative-render --page N --status running
$P --deck-dir <deck_dir> --stage creative-render --page N --status ok
$P --deck-dir <deck_dir> --stage export --status running
$P --deck-dir <deck_dir> --stage export --status ok

On failure, write the same stage with --status failed --error "<short reason>" before moving on or aborting. On native Windows, use python if python3 is unavailable.

Resume

python3 $SKILL_DIR/scripts/resume_scan.py --deck-dir <deck_dir>
# => {"style_spec_done": bool, "outline_done": bool, "pptx_done": bool,
#     "pages": [{"page_no": 1, "action": "skip|render_only|full"}, ...]}

Dispatch:

ManifestDo
style_spec_done == falseRun Stage 2
outline_done == falseRun Stage 3
per-page action == "full"Run Stage 4.1 + 4.2
per-page action == "render_only"Run Stage 4.2 only (prompt.txt already on disk)
per-page action == "skip"Skip
pptx_done == false (all pages done or failed)Run Stage 5

Stage 2 — style_spec.md (LLM or VLM via model_client)

One independent exec tool_call. Two branches based on reference images.

Branch A (no ref images, or all missing on disk) — use model_client.llm:

python3 -c "
import sys, pathlib, json
sys.path.insert(0, '$PPT_STANDARD_DIR/lib')
from model_client import llm

deck = pathlib.Path('<deck_dir>')
tp = json.loads((deck / 'task_pack.json').read_text())
ip = json.loads((deck / 'info_pack.json').read_text())

sys_prompt = open('$SKILL_DIR/prompts/style_from_query.md').read()
user_prompt = json.dumps({
    'params': tp['params'],
    'query': ip.get('user_query'),
    'digest': ip.get('document_digest'),
}, ensure_ascii=False)

md = llm(sys_prompt, user_prompt)
(deck / 'style_spec.md').write_text(md, encoding='utf-8')
print('style_spec.md ok')
"

Branch B (≥1 reference image on disk) — use model_client.vlm:

python3 -c "
import sys, pathlib, json
sys.path.insert(0, '$PPT_STANDARD_DIR/lib')
from model_client import vlm

deck = pathlib.Path('<deck_dir>')
ip = json.loads((deck / 'info_pack.json').read_text())
tp = json.loads((deck / 'task_pack.json').read_text())

refs = [p for p in (ip.get('user_assets') or {}).get('reference_images', []) if pathlib.Path(p).exists()]

sys_prompt = open('$SKILL_DIR/prompts/style_from_image.md').read()
user_prompt = f'PPT 主题/参数: {json.dumps(tp[\"params\"], ensure_ascii=False)}\nuser_query: {ip.get(\"user_query\") or \"\"}'

md = vlm(sys_prompt, user_prompt, images=refs)
(deck / 'style_spec.md').write_text(md, encoding='utf-8')
print(f'style_spec.md ok (from {len(refs)} ref images)')
"

If user_assets.reference_images is non-empty but all paths missing on disk: fall through to Branch A and prepend a line reference_images_missing: <original paths> at the top of style_spec.md.

Stage 3 — outline.json (LLM via model_client)

python3 -c "
import sys, pathlib, json
sys.path.insert(0, '$PPT_STANDARD_DIR/lib')
from model_client import llm

deck = pathlib.Path('<deck_dir>')
tp = json.loads((deck / 'task_pack.json').read_text())
ip = json.loads((deck / 'info_pack.json').read_text())
style = (deck / 'style_spec.md').read_text()

sys_prompt = open('$SKILL_DIR/prompts/outline.md').read()
user_prompt = json.dumps({
    'style_spec_markdown': style,
    'params': tp['params'],
    'query': ip.get('user_query'),
    'digest': ip.get('document_digest'),
}, ensure_ascii=False)

raw = llm(sys_prompt, user_prompt).strip()
if raw.startswith('\`\`\`'):
    raw = raw.split('\n', 1)[1].rsplit('\`\`\`', 1)[0]
data = json.loads(raw)
assert len(data['pages']) == tp['params']['page_count'], 'page_count mismatch'
(deck / 'outline.json').write_text(json.dumps(data, ensure_ascii=False, indent=2))
print(f'outline ok, {len(data[\"pages\"])} pages')
"

On failure (non-JSON / length mismatch): abort.

Stage 4 — per-page: one independent exec per page

4.1 Compose prompt (LLM via model_client) — skip if action == "render_only"

python3 -c "
import sys, pathlib, json
sys.path.insert(0, '$PPT_STANDARD_DIR/lib')
from model_client import llm

deck = pathlib.Path('<deck_dir>')
N = <NNN>
style = (deck / 'style_spec.md').read_text()
outline = json.loads((deck / 'outline.json').read_text())
page = next(p for p in outline['pages'] if int(p['page_no']) == N)

sys_prompt = open('$SKILL_DIR/prompts/page_prompt.md').read()
user_prompt = json.dumps({'style_spec_markdown': style, 'page': page}, ensure_ascii=False)

txt = llm(sys_prompt, user_prompt)
(deck / 'pages' / f'page_{N:03d}.prompt.txt').write_text(txt, encoding='utf-8')
print(f'prompt page {N} ok')
"

# sanitize the written prompt in-place: strip hex/rgb/hsl/CSS/px/em/rem etc
# to prevent T2I server-side prompt-enhance from baking them into the image.
# Silent: no chat-facing notification; removals go to stderr only.
python3 $SKILL_DIR/scripts/sanitize_prompt.py --path <deck_dir>/pages/page_<NNN>.prompt.txt

4.2 Generate image (T2I via sn-image-base)

--negative-prompt 是针对可能带自身 prompt-enhance 的 T2I 后端的最后一道防线: 即使前面的 sanitize 没拦住、或后端重写时引入了新的样式元数据,也通过反向约束压制模型把它们画出来。这段字符串在所有页上都一致。

python $SN_IMAGE_BASE/scripts/sn_agent_runner.py sn-image-generate \
  --prompt "$(cat <deck_dir>/pages/page_<NNN>.prompt.txt)" \
  --negative-prompt "hex color code, #RRGGBB, rgb(), rgba(), hsl(), hsla(), css, json, yaml, code snippet, pixel values, px, em, rem, pt, color palette text, typography label, design spec, style guide, font stack, hex code, layout annotation, dimensional callout, figma-style spec sheet, wireframe annotation, swatch with numbers" \
  --aspect-ratio 16:9 \
  --image-size 2k \
  --save-path <deck_dir>/pages/page_<NNN>.png \
  --output-format json

4.3 Failure handling

  • 4.1 failure (model timeout / empty / malformed): record page_no into failed_pages, echo failure line, continue.
  • 4.2 failure: same — record, echo, continue.
  • No retries. No placeholder PNG. Don't write 1x1 transparent PNGs to fake success.
  • .prompt.txt may remain on disk for a later manual re-run of 4.2 only.

Stage 5 — pptx 打包(一次独立 exec)

所有页图生成后(含部分失败的情况),把 pages/page_*.png 平铺打包成 16:9 整册 PPTX,每张图满版一页。由 scripts/build_pptx.py 完成,模型只负责执行脚本。

python3 $SKILL_DIR/scripts/build_pptx.py --deck-dir <deck_dir>
# => {"deck_id": "...", "output": "<deck_dir>/<deck_id>.pptx",
#     "total_slides": N, "included_pages": [...], "missing_pages": [...]}

行为约定:

  • 输出路径默认 <deck_dir>/<deck_id>.pptx;可用 --output 覆盖。
  • 页序按 outline.json 的 page_no 排;缺失 outline.json 时按 page_001..page_NNN 走。
  • 缺失的 PNG 会插入空白页并在 stderr 记录一行,不中止;这样跟 Stage 4 的"失败跳过"语义一致。
  • 脚本失败(依赖缺失 / 写盘失败):echo 失败原因,不中止整个 skill,仍进入 Stage 6 收尾;PNG 已在磁盘上。 如果 python-pptx 缺失导致失败:🚫 不要尝试 pip install python-pptx 或任何替代方案。PNG 页面已经是最终交付物,直接进入 Stage 6。

Stage 6 — closing

Emit:

创意模式已完成。

📁 输出目录:<deck_dir>
📄 结果文件:
  - style_spec.md
  - outline.json
  - pages/page_001.png ~ page_NNN.png(失败 M 页:page_..., page_...)
  - <deck_id>.pptx(整册,缺失页插入空白)

⚠️ 未完成:
  - page_007:生图返回超时,已跳过(pptx 中为空白页)

下一步:
  - 可直接打开 <deck_id>.pptx 查看整册
  - 或在 pages/ 目录查看 PNG

Progress echo — MANDATORY

StageExample
After resume_scan已进入 sn-ppt-creative,共 N 页
After each progress write.workbench/progress.json 已更新:<stage> <status>
After Stage 2[1] style_spec.md ✓
After Stage 3[2] outline.json ✓(N 页)
Per page-prompt (4.1)[prompt 3/10] ✓
Per page-image (4.2)[图 3/10] page_003.png ✓ or [图 3/10] ✗ 超时
After Stage 5[pptx] <deck_id>.pptx ✓(N 页,缺失 M 页) or [pptx] ✗ <reason>
Closingfull summary above
  • Each echo is a chat reply, not a log write.
  • Per-page echo is the heartbeat for Stage 4.
  • On failure, echo failure line with reason before moving on.

🚫 Hard rules

  1. Do NOT loop inside a single exec. One page = one tool_call.
  2. Do NOT fake images. Failed T2I → record failed, move on. No 1x1 placeholder PNGs.
  3. Do NOT use model_client.t2i — T2I must go through sn-image-base. model_client handles only LLM / VLM.
  4. Do NOT use sn-text-optimize or sn-image-recognize from sn-image-base — those must go through model_client.llm / model_client.vlm.
  5. Do NOT retry on first failure. If the same stage fails twice in a row with the same error, treat it as permanent and move on.
  6. Do NOT generate editable JSON from PNG (out of scope).
  7. Language integrity. All user-visible text MUST match the user's query language. A single English slide in a Chinese deck is a regression.
  8. Do NOT use python-pptx, pptxgenjs, or any alternative PPTX builder. scripts/build_pptx.py is the ONLY way to produce a PPTX. Never pip install python-pptx or write Node scripts that import pptxgenjs. If PPTX build fails, the PNG pages are the final deliverable.
  9. Do NOT fabricate data. All numbers and factual claims MUST come from the user's documents or web search. Use qualitative descriptions if no data source is available.
  10. Wait for responses. If you ask the user a question, do NOT proceed until they reply. Never assume default values.
  11. Multi-round edits: regenerate. When the user requests changes, re-run the affected pipeline stages. Do NOT sed/perl/patch files in-place.
  12. Validate paths before writing. All output goes under <deck_dir>/ — the absolute path written in task_pack.json. Before writing any file, verify the parent directory exists. Never write to /workspace/, /tmp/, ~/, ./, or any hallucinated path.

发现问题?提交给管理员复核

评分:

评论 (0)

暂无评论,成为第一个评论者吧!