复制安装命令
用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。
复制前请先查看来源、License 和安全提示。
This is the open-source content repository behind
用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。
复制前请先查看来源、License 和安全提示。
来源文件:README.md
This is the open-source content repository behind Skill Store. It stores every approved Agent Skill, the records that go with it, and the automated security audits published with each skill.
This repo is a companion to the Skill Store platform, not the place to submit skills. Skills are added through skillstore.io — its review pipeline writes to this repo automatically. Please do not open a pull request here to add a skill; PRs adding skills will be closed. See Contributing a skill below.
The recommended way to install any skill is the skillstore CLI — one command works for both Claude Code and Codex:
npx skillstore add author/skill-name
For example:
npx skillstore add aiskillstore/code-review
It downloads the skill and drops it into the right skills/ directory for your tool. Claude Code auto-discovers it; for Codex, restart the session.
Prefer to do it by hand, or installing via Claude Web? See the full Installation Guides for every method (CLI, manual, and ZIP upload) and the scope directories (~/.agents/skills/, .claude/skills/, ~/.claude/skills/, .codex/skills/, …).
Submit through the platform — not through a pull request:
SKILL.md.SKILL.md — the skill definition (required, per the Agent Skills spec)LICENSE (recommended)Every submission is scanned automatically before it can be published. The audit flags things like:
eval, exec, raw system commands)Security analysis is report-only: findings inform maintainers and users, but a risk result does not automatically block an otherwise approved skill from being published. See our Security Trust Center for the methodology, limitations, and risk-level definitions.
Live Security Passport example:
.
├── skills/ # Approved, published skills (one folder each, with SKILL.md)
├── pending/ # Submissions awaiting review
├── packages/
│ ├── cli/ # The `skillstore` CLI (npx skillstore add …)
│ └── skillstore/
├── schemas/ # JSON schemas for skill records
├── scripts/ # Maintenance & scoring scripts
└── .github/workflows/ # Submission, audit, and sync automation
The contents of this repo are maintained by Skill Store's automated pipeline. Manual changes are limited to maintainers.
The marketplace catalog is MIT-licensed. Individual skills carry their own licenses — check each skill's LICENSE file.
name: pdf-page-extract
description: Extract rich data from PDF pages including text spans with metadata, rendered PNG images, and page mapping. Creates persistent artifacts for downstream processing.This skill extracts all necessary data from PDF pages to enable accurate AI-driven HTML generation. It produces three critical artifacts:
This is the deterministic, Python-based foundation for the entire pipeline. All extracted data is saved to persistent files for traceability and future processing.
Validate input parameters
Establish page mapping (if not already done)
python3 Calypso/tools/read_page_footers.pyanalysis/page_mapping.jsonExtract rich page data using PyMuPDF and pdfplumber
python3 Calypso/tools/rich_extractor.pyanalysis/chapter_XX/rich_extraction.jsonRender PDF page to PNG
output/chapter_XX/page_artifacts/page_YY/02_page_XX.pngExtract embedded images (if present)
python3 Calypso/tools/extract_images.pyoutput/chapter_XX/images/page_YY_image_*.pngpage_YY_images.jsonValidate extraction completeness
chapter: <int> - Chapter number (1-8)
start_page: <int> - Starting PDF index (0-based) or page range
end_page: <int> - Ending PDF index (optional if single page)
pdf_path: <str> - Path to PDF file (default: Calypso/PREP-AL 4th Ed 9-26-25.pdf)
output_base: <str> - Output directory (default: Calypso/output)
mapping_file: <str> - Page mapping file (default: Calypso/analysis/page_mapping.json)
Per-page artifacts (in output/chapter_XX/page_artifacts/page_YY/):
01_rich_extraction.json - Text spans with metadata02_page_XX.png - Rendered PDF page imagepage_mapping.json - Shared mapping file (symlink or copy)Extraction data (in analysis/chapter_XX/):
rich_extraction.json - Full extraction for all pages in chapterpage_6_pattern_analysis.json - (Optional) Pattern analysis for specific pagesImages (in output/chapter_XX/images/chapter_XX/):
page_XX_image_*.png - Embedded images from pagepage_XX_images.json - Metadata for embedded images{
"page_number": 16,
"pdf_index": 15,
"book_page": 17,
"chapter": 2,
"dimensions": {
"width": 612,
"height": 792
},
"text_spans": [
{
"text": "Rights in Real Estate",
"font": "Arial-BoldMT",
"size": 27.04,
"bold": true,
"italic": false,
"bbox": {
"x0": 72,
"y0": 150,
"x1": 400,
"y1": 177
},
"color": 0,
"sequence": 1
}
],
"analysis": {
"font_sizes": {
"27.04": 1,
"11.04": 45
},
"font_styles": {
"bold_27.04": 1,
"regular_11.04": 45
},
"likely_headings": [
{
"text": "Rights in Real Estate",
"level": 1,
"confidence": 0.95
}
],
"likely_paragraphs": [
{
"text": "Real property consists of...",
"type": "body_text"
}
]
},
"extraction_timestamp": "2025-11-08T14:30:00Z",
"extraction_tool": "rich_extractor.py v1.0"
}
cd Calypso/tools
python3 read_page_footers.py \
--start 15 \
--end 28 \
--pdf "../PREP-AL 4th Ed 9-26-25.pdf" \
--output "../analysis/page_mapping.json"
Success indicators:
cd Calypso/tools
python3 rich_extractor.py \
--pdf "../PREP-AL 4th Ed 9-26-25.pdf" \
--start 15 \
--end 28 \
--output "../analysis/chapter_02/rich_extraction.json"
Success indicators:
cd Calypso/tools
python3 -c "
import fitz
pdf = fitz.open('../PREP-AL 4th Ed 9-26-25.pdf')
for page_idx in range(15, 29):
page = pdf[page_idx]
pix = page.get_pixmap(matrix=fitz.Matrix(3, 3)) # 300% zoom for high-res
pix.save(f'../output/chapter_02/page_artifacts/page_{page_idx:02d}/02_page_{page_idx}.png')
pdf.close()
"
cd Calypso/tools
# For each page with images
python3 extract_images.py \
--page 17 \
--pdf "../PREP-AL 4th Ed 9-26-25.pdf" \
--output "../output" \
--mapping "../analysis/page_mapping.json"
Before declaring extraction complete:
File existence
01_rich_extraction.json exists02_page_XX.png exists and is validpage_mapping.json existsJSON validity
Data completeness
Image quality
If PDF file not found:
If page mapping fails:
If rich extraction produces no text:
"page_type": "image_only"If PNG rendering fails:
All artifacts include metadata:
This enables:
✓ All required files created in correct directories ✓ Rich extraction JSON is valid and complete ✓ PNG image renders correctly ✓ Page mapping is accurate ✓ All data persisted and ready for next skill ✓ No extraction errors or warnings
Once extraction completes successfully:
PDF won't open: Verify file path, ensure PDF is not corrupted No text extracted: Page may be image-only (OCR needed) Wrong page numbers: Check page_mapping.json for accuracy PNG images are blank: Try increasing zoom factor (3x = 300 DPI)
评论 (0)
暂无评论,成为第一个评论者吧!