复制安装命令
用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。
复制前请先查看来源、License 和安全提示。
The SenseNova model family plugs directly into agent runtimes such as OpenClaw and hermes-agent...
用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。
复制前请先查看来源、License 和安全提示。
来源文件:README.md
English | 简体中文
The SenseNova model family plugs directly into agent runtimes such as OpenClaw and hermes-agent, with the skills in this repository extending the models with concrete, end-to-end office capabilities.
In this repository each skill lives in its own directory and declares triggers, capabilities, and execution flow through a SKILL.md file, following the Agent Skills convention.
The skills cover image generation & visualization, slide-deck (PPT) generation, Excel data analysis, and deep research — usable standalone or composed into end-to-end workflows.
🎨 Want to see what it can do? Check out our sn-infographic Gallery to explore nearly 100 stunning generation cases and steal their prompt designs !
The latest SenseNova models and the full Cowork-Skill suite in this repo are bundled into Raccoon, with enterprise-grade security and a zero-setup experience — if you'd rather not provision env, API keys, and runtimes yourself, you can use these capabilities directly through Raccoon. Free trial available — no payment required to get started.
Raccoon now ships a full upgrade across product capability and client experience:
👉 Try it: xiaohuanxiong.com
These skills are designed to run inside an Agent Skills-compatible agent.
INSTALL.md.Recommended: let the agent install the skills for you. Hand it the repo URL and ask it to clone and drop the skills into the right directory — for example:
"Please install SenseNova-Skills from https://github.com/OpenSenseNova/SenseNova-Skills into your skills directory."
After it finishes, you may need to manually restart the agent service before the new skills are picked up.
| Agent | Target directory |
|---|---|
| OpenClaw | ~/.openclaw/skills/ |
| hermes-agent | ~/.hermes/skills/ |
Clone this repository, then copy the subdirectories under skills/ into the target directory yourself:
git clone https://github.com/OpenSenseNova/SenseNova-Skills.git --depth=1
mkdir -p ~/.openclaw/skills
cp -r SenseNova-Skills/skills/* ~/.openclaw/skills/
For Hermes, swap the target to ~/.hermes/skills/.
Per-category Python dependencies, API keys, and invocation examples are documented in the 📖 Full guide for each section.
📖 Full guide: docs/sn-image-generate_en.md (prerequisites, Quick Start, API config, and invocation samples).
| Name | Label | Description |
|---|---|---|
sn-image-doctor | Environment Doctor | Validates the SenseNova-Skills environment — checks sn-image-base install, Python deps, and required env vars; interactively fills missing values into .env. |
sn-image-base | Image Base Layer (Tier 0) | Low-level tools — text-to-image (sn-image-generate), image recognition (sn-image-recognize), and text optimization (sn-text-optimize) — exposed through a unified sn_agent_runner.py, designed to be called by upper-layer skills. |
sn-infographic | Infographic Generation (Tier 1) | Auto prompt-quality scoring, layout/style selection (87 layouts / 66 styles), multi-round generation with VLM review and quality ranking, producing publication-ready infographics. |
sn-image-imitate | Image Imitation (Tier 1) | Given one reference image and a target content prompt, generates a new image that imitates the reference. |
sn-image-resume | Resume Image Generation (Tier 1) | Given resume information, generates a resume image. |
📖 Full guide: docs/sn-ppt-generate.md (prerequisites, Quick Start, API config, and invocation samples).
| Name | Label | Description |
|---|---|---|
sn-ppt-entry | PPT Entry Point | Unified entry point for PPT generation. Asks the user to choose fast, standard, or creative mode, then collects role / audience / scenario / page count. For standard mode, also asks about image sourcing (AI, web search, or none) and chart rendering (U1 infographics or ECharts). Parses uploaded pdf / docx / md / txt, emits task_pack.json + info_pack.json, and dispatches to the chosen mode. |
sn-ppt-doctor | PPT Environment Doctor | Environment check for the PPT pipeline — validates sn-image-base, API keys, the Node runtime, and optional deps; writes missing required vars into .env. |
sn-ppt-creative | PPT Creative Mode | One full-page 16:9 PNG per slide, generated via sn-image-generate with a per-page composed prompt. Falls back to web image search when T2I generation fails. |
sn-ppt-standard | PPT Standard & Fast | style_spec → outline → asset plan + per-slot images + VLM QC → per-page HTML → per-page review → PPTX export. Fast mode builds a complete draft immediately with autonomous decisions, then provides structured refinement suggestions. Supports AI-generated infographics (U1) for diagrams and web image search (Serper) for real photos. |
📖 Full guide: docs/sn-data-analysis.md (prerequisites, Quick Start, API config, and invocation samples).
| Name | Label | Description |
|---|---|---|
sn-da-excel-workflow | Excel Analysis Orchestration | End-to-end Excel pipeline — multi-sheet read, large-file detection (≥10k rows triggers Parquet), cleaning, conditional filtering, cross-sheet aggregation, and Excel/CSV export. |
sn-da-image-caption | Image Understanding & Data Extraction | For image-first inputs — table OCR, chart understanding, screenshot/UI description; parses captions into DataFrames, recreates visualizations, exports Excel/CSV. |
sn-da-large-file-analysis | High-Performance Large-File Analysis | Streaming reads for ≥10k-row Excel datasets (openpyxl read_only + iter_rows), Parquet conversion, memory optimization, chunked processing, large-file writes. |
📖 Full guide: docs/sn-deep-research.md (prerequisites, web_search precheck, Quick Start, and per-stage invocation).
| Name | Label | Description |
|---|---|---|
sn-deep-research | Deep Research Entry Point | Unified deep-research orchestrator with true-dependency DAGs, reusable source snapshots, and evidence-informed content units, producing final report.md. |
sn-research-report | Final Report Writing & Editing | Renders the judgment layer into the final report.md; also handles targeted rewrites — restructuring, polishing, table-augmentation — for an existing draft. |
sn-report-format-discovery | Presentation-Format Discovery | Compares final forms such as a research report, academic paper, table-first analysis, decision memo, or a custom Markdown form; scout uses it before research and user confirmation. |
sn-prepare-citations | Citation Rendering | Post-processes [^source_id] footnotes into numbered citations and appends references from evidence sources. |
sn-md-to-html-report | Markdown → HTML Report | Converts the research report.md (or any Markdown doc) into a clean, single-file HTML reading view that opens offline — embedded images, side-panel TOC, responsive tables, and table-delimiter repair. |
📖 Search skills are documented together with deep research: docs/sn-deep-research.md (includes per-platform API keys, invocation, and unified JSON output).
| Name | Label | Description |
|---|---|---|
sn-search-academic | Academic Search | ArXiv (with section-level HTML reading) / Semantic Scholar (with citation counts) / PubMed (with PMC open-access full text) / Wikipedia, in one aggregated interface. |
sn-search-code | Developer Search | GitHub (repo / code / issue) / Stack Overflow / Hacker News / HuggingFace (models / datasets / spaces), aggregated. |
sn-search-social-cn | Chinese Social Search | Bilibili / Zhihu / Douyin search; some platforms require cookie auth. |
sn-search-social-en | English Social Search | Reddit / Twitter (X) / YouTube search. |
A few sn-infographic outputs (more in docs/sn-infographic-examples.md).
examples/memory-price-end2end-analysis. Starting from a raw quote CSV, the agent profiles fields, normalizes categories and timestamps, then attacks the rally from three angles — overall trend, top movers per category, and the gap between server-grade and consumer-grade SKUs — locating a late-February inflection along the way. Treating those findings as the research question, it switches to deep research: planning per-dimension web searches over supply contraction, AI-server demand, and vendor output discipline, then triaging and cross-checking evidence across sources before committing it to the report. The data and research conclusions are then handed to PPT generation, which lays out a 16-page outline, plans per-slot imagery, renders per-page HTML, runs VLM review, and finally composites screenshots into the PPTX. The result is a clear three-step storyline: prices are rising → here is why → here is what to do. This is the only example that exercises the full data analysis → deep research → PPT chain end-to-end.
sn-da-excel-workflow, sn-deep-research, sn-ppt-entry, sn-ppt-standard, sn-md-to-html-reportexamples/employee-performance-analysis. The agent reads 10 separate monthly review xlsx files, aligns column schemas across months and joins them into one longitudinal table. From that table it produces aggregate views — monthly average trend, score-distribution boxplots, grade mix change, and a 38-role ranking — and individual views — top performers, needs-attention, and consistently-improving cohorts plus per-employee year trends. The findings are written up with explicit improvement suggestions tied to specific roles and individuals, backed by 8 supporting charts. The same content is delivered as a Word doc (for distribution) and a visualized HTML report (for browsing). The example shows how sn-da-excel-workflow handles "many small spreadsheets that should be one analysis" rather than a single big file.
sn-da-excel-workflowexamples/embodied-ai-deep-research. Given only an industry name, the agent first commits to a research plan — market size, vendor share, financing, cost structure, development roadmap — instead of jumping straight into search. For each dimension it runs targeted web searches, fetches and reads source pages, and extracts both numeric and qualitative evidence; conflicting figures across sources are explicitly reconciled before being trusted. A synthesis stage organizes per-dimension evidence into a traceable, reader-oriented information structure rather than a stack of disconnected bullets. The output is an illustrated report (Markdown + visualized HTML) with 5 dimension-specific charts. The example shows how sn-deep-research turns "go research X" into a structured plan-then-execute loop with traceable evidence.
sn-deep-researchexamples/property-fee-pricing-ppt. The agent takes a free-form brief — topic (property fee pricing), audience (property staff + committee), 26 pages, black-and-white warm style — and first commits to an outline plus a per-page asset plan that conforms to the style spec. Each slide is then built as semantic per-page HTML rather than free-form image generation: copy, layout, illustrations, icons, and any data charts are reasoned about per slot. Imagery is produced or selected per slot and VLM-checked against the page's intent; each rendered page goes through a review pass with optional rewrite for coherence and copy quality. Final pages are screenshotted and composited into the PPTX, with the per-page HTML kept alongside for direct browser preview or re-editing. The example demonstrates sn-ppt-standard style consistency on a long, prose-heavy deck where every slide must obey the same audience and palette constraints.
sn-ppt-entry, sn-ppt-standardCommon setup and runtime questions (400/401 errors, rate limits, PPT timeouts, infographic quality, model names) are answered in docs/faq.md.
Feel free to use the skills here as templates for your own OpenClaw skills. The qualities that make a skill good:
description exactly when the skill should and should not run, so the agent recognizes it accuratelyreferences/, scripts/, prompts/ to provide additional contextJoin our growing community to share feedback, get support, and stay updated on the latest developments. Scan the QR code below to hop into the chat — we'd love to hear from you!
| Discord | Lark Group |
![]() | ![]() |
MIT — see LICENSE.
name: sn-da-large-file-analysis
description: "万行以上 Excel 数据集的高性能分析引擎。提供 openpyxl read_only 流式读取(iter_rows 支持 10 万行以上)、Parquet 转换加速、内存优化、分块处理和大文件写入模式。**遇到以下任一情况就主动使用本 skill**:①数据行数 ≥ 10k(由 sn-da-excel-workflow 的行数评估步骤触发);②用户出现触发词:大文件 / 大数据量 / 性能优化 / 内存不足 / OOM / 百万行 / 十万行 / 流式读取 / Parquet / 分块处理 / large file / big data / streaming read / chunked processing;③直接使用 pd.read_excel() 导致超时或内存溢出;④用户明确要求对大规模数据集进行高性能处理。仅不用于:小于 10k 行的常规 Excel 分析(使用 sn-da-excel-workflow 即可)。"When total rows >= 10,000, you MUST use the methods in this skill.
| Data Scale | Read Strategy | Reason |
|---|---|---|
| < 10k rows | pd.read_excel() directly | No memory pressure |
| 10k–100k rows | pd.read_excel() → convert to Parquet → pd.read_parquet() for analysis | Avoid repeated slow reads |
| 100k–1M rows | openpyxl read_only + iter_rows streaming → Parquet | pd.read_excel() will OOM or timeout |
| > 1M rows | Streaming read + multi-sheet split (Excel max 1,048,576 rows per sheet) | Must chunk |
Prohibited:
pd.read_excel() to fully load 100k+ row filesfc-list, find ... fonts, or install packages with pip installdf.iterrows() on large DataFrames (use itertuples() or vectorized ops)df.apply(lambda...) for operations that can be vectorizedimport pandas as pd
import numpy as np
import os
import gc
pd.options.mode.copy_on_write = True
# CJK font setup (fixed paths — do NOT search for fonts)
# ⚠️ Copy this block as-is. Do NOT use fc-list, find, subprocess, or glob to locate fonts.
import matplotlib
import matplotlib.pyplot as plt
import matplotlib.font_manager as fm
_FONT_PATHS = [
'/mnt/afs_agents/SimHei.ttf',
'/mnt/afs_agents/mnt/data/SimHei.ttf',
os.path.expanduser('~/.fonts/SimHei.ttf'),
'/usr/share/fonts/truetype/wqy/wqy-zenhei.ttc',
'/usr/share/fonts/SimHei.ttf',
]
for _p in _FONT_PATHS:
if os.path.exists(_p):
fm.fontManager.addfont(_p)
matplotlib.rcParams['font.family'] = fm.FontProperties(fname=_p).get_name()
break
matplotlib.rcParams['axes.unicode_minus'] = False
Before any operation on a large file, inspect sheets and row counts without loading data into memory:
import openpyxl
def inspect_excel(file_path):
"""Stream-inspect Excel structure. Returns {sheet_name: {rows, columns}}."""
wb = openpyxl.load_workbook(file_path, read_only=True, data_only=True)
info = {}
for name in wb.sheetnames:
ws = wb[name]
row_count = 0
header = None
for i, row in enumerate(ws.iter_rows(values_only=True)):
if i == 0:
header = [str(c) if c is not None else f"Col_{j}" for j, c in enumerate(row)]
else:
row_count += 1
info[name] = {"rows": row_count, "columns": header}
wb.close()
return info
# Usage
file_info = inspect_excel(file_path)
for sheet, meta in file_info.items():
print(f"Sheet '{sheet}': {meta['rows']} rows, {len(meta['columns'])} cols")
print(f" Columns: {meta['columns'][:10]}...")
total_rows = sum(m['rows'] for m in file_info.values())
print(f"Total rows: {total_rows}")
For 100k+ row files, never use pd.read_excel(). Use openpyxl streaming → Parquet:
import openpyxl
import pyarrow as pa
import pyarrow.parquet as pq
def stream_excel_to_parquet(excel_path, parquet_path, sheet_name=None, chunk_size=50000):
"""Stream Excel rows to Parquet with constant memory usage.
All columns are cast to string to avoid cross-chunk schema mismatches
(Excel mixed-type columns may be all-None in some chunks, causing PyArrow
to infer null type instead of string). Convert numeric columns after loading
Parquet with pd.to_numeric() as needed.
"""
wb = openpyxl.load_workbook(excel_path, read_only=True, data_only=True)
ws = wb[sheet_name] if sheet_name else wb.active
header = None
writer = None
chunk_rows = []
total_written = 0
def _flush(rows):
nonlocal writer
table = pa.table({
col: pa.array(
[str(r[idx]) if r[idx] is not None else None for r in rows],
type=pa.string(),
)
for idx, col in enumerate(header)
})
if writer is None:
writer = pq.ParquetWriter(parquet_path, table.schema)
writer.write_table(table)
for i, row in enumerate(ws.iter_rows(values_only=True)):
if i == 0:
header = [str(c) if c is not None else f"Col_{j}" for j, c in enumerate(row)]
continue
chunk_rows.append(list(row))
if len(chunk_rows) >= chunk_size:
_flush(chunk_rows)
total_written += len(chunk_rows)
print(f" Written {total_written:,} rows...")
chunk_rows = []
gc.collect()
if chunk_rows:
_flush(chunk_rows)
total_written += len(chunk_rows)
if writer:
writer.close()
wb.close()
print(f"Done: {total_written:,} rows -> {parquet_path}")
return total_written
For 10k–100k rows, pd.read_excel() won't OOM, but Parquet is much faster for repeated analysis:
def convert_excel_to_parquet(excel_path, parquet_path, sheet_name=0):
"""Medium file: pd.read_excel -> Parquet cache."""
if os.path.exists(parquet_path):
print(f"Cache exists: {parquet_path}")
return
df = pd.read_excel(excel_path, sheet_name=sheet_name)
df.columns = df.columns.astype(str)
df.to_parquet(parquet_path, engine='pyarrow', compression='snappy')
row_count = len(df)
del df
gc.collect()
print(f"Converted {row_count:,} rows -> {parquet_path}")
After loading Parquet, further reduce memory footprint:
def optimize_dtypes(df):
"""Auto-downcast numeric types + convert low-cardinality strings to Category.
Typically saves 50-80% memory."""
start_mb = df.memory_usage(deep=True).sum() / 1024**2
for col in df.select_dtypes(include=['int64', 'int32']).columns:
c_min, c_max = df[col].min(), df[col].max()
if c_min >= np.iinfo(np.int8).min and c_max <= np.iinfo(np.int8).max:
df[col] = df[col].astype(np.int8)
elif c_min >= np.iinfo(np.int16).min and c_max <= np.iinfo(np.int16).max:
df[col] = df[col].astype(np.int16)
elif c_min >= np.iinfo(np.int32).min and c_max <= np.iinfo(np.int32).max:
df[col] = df[col].astype(np.int32)
for col in df.select_dtypes(include=['float64']).columns:
df[col] = df[col].astype(np.float32)
for col in df.select_dtypes(include=['object', 'string']).columns:
if df[col].nunique() / max(len(df), 1) < 0.5:
df[col] = df[col].astype('category')
end_mb = df.memory_usage(deep=True).sum() / 1024**2
print(f"Memory: {start_mb:.1f} MB -> {end_mb:.1f} MB (saved {(1 - end_mb/start_mb)*100:.0f}%)")
return df
def write_large_excel(df, output_path, sheet_name="Sheet1"):
"""Auto-select write strategy based on data size."""
total_cells = len(df) * len(df.columns)
if len(df) > 1_000_000:
csv_path = output_path.rsplit('.', 1)[0] + '.csv'
df.to_csv(csv_path, index=False)
print(f"Over 1M rows — exported as CSV: {csv_path}")
return csv_path
if total_cells > 50_000:
from openpyxl import Workbook
from openpyxl.cell import WriteOnlyCell
wb = Workbook(write_only=True)
ws = wb.create_sheet(title=sheet_name)
ws.append(list(df.columns))
for idx, row in enumerate(df.itertuples(index=False)):
ws.append([None if pd.isna(v) else v for v in row])
if (idx + 1) % 100_000 == 0:
print(f" Written {idx + 1:,} rows...")
wb.save(output_path)
wb.close()
print(f"write_only mode: {len(df):,} rows -> {output_path}")
else:
df.to_excel(output_path, index=False, sheet_name=sheet_name)
print(f"Standard write: {len(df):,} rows -> {output_path}")
return output_path
Scenario: User has a 100k-row sales Excel file and wants regional sales distribution with a bar chart.
import pandas as pd
import os, gc
excel_path = "sales_100k.xlsx"
parquet_path = "sales_100k.parquet"
# === Step 1: Inspect structure ===
file_info = inspect_excel(excel_path)
total_rows = sum(m['rows'] for m in file_info.values())
print(f"Total rows: {total_rows}")
# === Step 2: Choose read strategy by row count ===
if total_rows >= 100_000:
stream_excel_to_parquet(excel_path, parquet_path)
else:
convert_excel_to_parquet(excel_path, parquet_path)
# === Step 3: Load Parquet + optimize memory ===
df = pd.read_parquet(parquet_path)
df = optimize_dtypes(df)
print(f"Shape: {df.shape}")
print(df.head(3))
# === Step 4: Analysis ===
region_sales = df.groupby('Region')['Sales'].sum().sort_values(ascending=False)
print(region_sales)
# === Step 5: Visualization ===
fig, ax = plt.subplots(figsize=(10, 6))
region_sales.plot(kind='bar', ax=ax, color='#4C72B0')
ax.set_title('Sales by Region')
ax.set_ylabel('Sales')
plt.tight_layout()
plt.savefig('region_sales.png', dpi=150, bbox_inches='tight')
plt.show()
# === Step 6: Cleanup ===
del df
gc.collect()
Scenario: User has a 1M-row transaction log and wants records with amount > 10,000 exported.
import pandas as pd
import os, gc
excel_path = "transactions_1m.xlsx"
parquet_path = "transactions_1m.parquet"
# === Step 1: Stream to Parquet (1M rows — MUST use streaming, never pd.read_excel) ===
stream_excel_to_parquet(excel_path, parquet_path, chunk_size=50000)
# === Step 2: Load only needed columns (saves memory) ===
df = pd.read_parquet(parquet_path, columns=['TransactionID', 'Amount', 'Date', 'Type'])
df = optimize_dtypes(df)
print(f"Shape: {df.shape}, Memory: {df.memory_usage(deep=True).sum()/1024**2:.1f} MB")
# === Step 3: Vectorized filtering (never use apply/iterrows) ===
mask = df['Amount'] > 10000
high_value = df[mask].copy()
print(f"Filtered: {len(high_value):,} / {len(df):,} rows")
# === Step 4: Export ===
output_path = write_large_excel(high_value, 'high_value_transactions.xlsx')
# === Step 5: Cleanup ===
del df, high_value
gc.collect()
On large files, never use slow operations — use vectorized alternatives:
| Slow (Prohibited) | Fast (Use This) |
|---|---|
df.apply(lambda x: x*2) | df['col'] * 2 |
df.iterrows() | df.itertuples(index=False) |
for i in range(len(df)): df.iloc[i] | Vectorized boolean indexing df[mask] |
df['a'].map(lambda x: 'Y' if x>0 else 'N') | np.where(df['a']>0, 'Y', 'N') |
df.groupby('a').apply(custom_func) | df.groupby('a').agg({'b':'sum','c':'mean'}) |
Estimate memory before loading to avoid OOM:
Estimated MB ≈ rows × cols × 8 / 1024² (numeric columns)
Estimated MB ≈ rows × cols × 50 / 1024² (with text columns)
| Rows | 20 cols (numeric) | 20 cols (with text) |
|---|---|---|
| 100k | ~15 MB | ~95 MB |
| 500k | ~76 MB | ~477 MB |
| 1M | ~153 MB | ~953 MB |
When estimated memory exceeds 80% of available RAM, use column-selective loading (pd.read_parquet(columns=[...])) or chunked processing.
read_only + iter_rows. Never pd.read_excel() for full load.del df; gc.collect() after every intermediate DataFrame.openpyxl Workbook(write_only=True).
评论 (0)
暂无评论,成为第一个评论者吧!