复制安装命令
用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。
复制前请先查看来源、License 和安全提示。
Version 5.
用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。
复制前请先查看来源、License 和安全提示。
来源文件:README.md
Version 5.27.0 - Added walkthrough-script-agent: generate timed walkthrough video scripts for app features
Personal collection of agent skills using the open SKILL.md standard. Works with Claude Code and other AI assistants.
# Add the marketplace
/plugin marketplace add michaelboeding/skills
# Install the plugin
/plugin install skills@michaelboeding-skills
Copy the skills/ folder to your project or follow your tool's skill installation docs.
Many skills require Python packages. Run the install script:
# From the skills directory
./scripts/install.sh
Or install manually:
pip install -r requirements.txt
Requirements:
google-genai package)What gets installed:
| Package | Version | Used By |
|---|---|---|
google-genai | ≥1.0.0 | image-generation, video-generation, voice-generation, music-generation |
matplotlib | ≥3.7.0 | chart-generation |
numpy | ≥1.24.0 | chart-generation |
python-pptx | ≥0.6.21 | slide-generation |
Pillow | ≥10.0.0 | slide-generation, image processing |
rembg | ≥2.0.50 | background-remove, icon-generation |
Optional tools:
| Tool | Install | Used By |
|---|---|---|
ffmpeg | brew install ffmpeg | media-utils, audio/video processing |
Some skills require API keys to function. Copy the example environment file and add your keys:
# Copy to your config directory (recommended - keeps keys safe from git)
mkdir -p ~/.config/skills
cp env.example ~/.config/skills/.env
# Edit ~/.config/skills/.env with your keys
Then export the variables in your shell profile (~/.bashrc, ~/.zshrc, or ~/.bash_profile):
# Core APIs (used by multiple skills)
export OPENAI_API_KEY="sk-..." # DALL-E, Sora, TTS
export GOOGLE_API_KEY="..." # Imagen, Gemini (AI Studio)
export ELEVENLABS_API_KEY="..." # ElevenLabs TTS
# Music Generation
export SUNO_API_KEY="..." # Suno music
export UDIO_API_KEY="..." # Udio music
# Model Council (optional)
export ANTHROPIC_API_KEY="sk-ant-..." # Claude API
export XAI_API_KEY="..." # Grok API
Restart your terminal or run source ~/.bashrc (or equivalent) for changes to take effect.
Vertex AI is the default backend for all Google-powered skills with higher rate limits:
| Skill | AI Studio | Vertex AI |
|---|---|---|
| Video (Veo) | 10/day | 10/min |
| Voice (Gemini TTS) | Limited | Higher |
| Music (Lyria) | Limited | Higher |
| Image (Imagen) | Limited | Higher |
Setup Vertex AI (one-time):
# 1. Install Google Cloud SDK: https://cloud.google.com/sdk/docs/install
# 2. Login and set project
gcloud auth application-default login
gcloud config set project YOUR_PROJECT_ID
# 3. Enable Vertex AI API
gcloud services enable aiplatform.googleapis.com
# 4. Export project (add to .env or shell profile)
export GOOGLE_CLOUD_PROJECT="your-project-id"
export GOOGLE_CLOUD_LOCATION="us-central1" # or us-east4
The video generation scripts auto-detect and use Vertex AI when GOOGLE_CLOUD_PROJECT is set.
Where to get API keys:
| ✅ Do | ❌ Don't |
|---|---|
Store keys in ~/.config/skills/.env | Commit .env files to git |
Use gcloud auth for local dev | Hardcode keys in scripts |
| Use service accounts for CI/CD | Share API keys publicly |
| Rotate keys if exposed | Store keys in repo, even private |
For CI/CD / Production:
# Option 1: Service Account (recommended)
export GOOGLE_APPLICATION_CREDENTIALS="/path/to/service-account.json"
# Option 2: Workload Identity (GKE/Cloud Run)
# Automatically authenticated, no keys needed
Everything is a skill (has a SKILL.md file), but there are two types:
┌─────────────────────────────────────────────────────────────────────────────┐
│ AGENT SKILLS (Higher-Level) │
│ Skills that orchestrate other skills + have sub-agents │
│ │
│ ┌─────────────────────┐ ┌─────────────────────┐ ┌─────────────────────┐ │
│ │ patent-lawyer-agent │ │ product-engineer- │ │ video-producer- │ │
│ │ 5 sub-agents │ │ agent │ │ agent │ │
│ │ uses: image-gen │ │ 5 sub-agents │ │ uses: video-gen │ │
│ │ chart-gen │ │ uses: image-gen │ │ voice-gen │ │
│ └─────────────────────┘ └─────────────────────┘ └─────────────────────┘ │
│ │ calls │
├──────────────────────────────────────▼──────────────────────────────────────┤
│ BASE SKILLS (Single-Purpose) │
│ Do ONE thing well - can be used directly or by agents │
│ │
│ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ │
│ │ image-gen │ │ video-gen │ │ voice-gen │ │ music-gen │ │
│ │ Generate │ │ Generate │ │ Generate │ │ Generate │ │
│ │ images │ │ videos │ │ speech │ │ music │ │
│ └──────────────┘ └──────────────┘ └──────────────┘ └──────────────┘ │
│ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ │
│ │ chart-gen │ │ slide-gen │ │ media-utils │ │
│ │ Data charts │ │ PPTX slides │ │ Concat/mix │ │
│ └──────────────┘ └──────────────┘ └──────────────┘ │
└─────────────────────────────────────────────────────────────────────────────┘
Key Difference:
Note: Agents are still skills (they have SKILL.md files), but they're a higher-level type that combines other skills in their execution. Think of it as: agents are skills that use skills.
Base skills are focused tools that do one thing well. They can be used directly or called by agent skills.
| Skill | What It Does | API Keys |
|---|---|---|
| image-generation | Generate/edit images (Gemini, DALL-E) | GOOGLE_API_KEY or OPENAI_API_KEY |
| icon-generation | Generate app icons with transparent backgrounds | GOOGLE_API_KEY |
| background-remove | Remove backgrounds from images (AI-based) | None (pip install rembg) |
| video-generation | Generate videos (Veo, Sora) | GOOGLE_API_KEY or OPENAI_API_KEY |
| voice-generation | Text-to-speech (Gemini TTS, ElevenLabs, OpenAI) | GOOGLE_API_KEY, ELEVENLABS_API_KEY, or OPENAI_API_KEY |
| music-generation | Generate music (Lyria, Suno, Udio) | GOOGLE_API_KEY, SUNO_API_KEY, or UDIO_API_KEY |
| chart-generation | Data-driven charts (matplotlib) | None (pip install matplotlib) |
| slide-generation | PowerPoint slides from JSON | None (pip install python-pptx) |
| device-framer | Wrap screenshots/recordings in iPhone frames | None (pip install Pillow, brew install ffmpeg) |
| media-utils | Concat/mix audio/video (FFmpeg) | None (brew install ffmpeg) |
| docx | Create/edit Word documents (OOXML) | None |
| pptx | Create/edit PowerPoint (advanced) | None (npm install pptxgenjs) |
| xlsx | Create/edit Excel spreadsheets | None |
| PDF forms, extraction, validation | None |
Skills for development workflows (no API keys needed):
| Skill | What It Does |
|---|---|
| style-guide | Analyze codebase conventions, generate style guide |
| ios-to-android | Port iOS/Swift features to Android/Kotlin |
| android-to-ios | Port Android/Kotlin features to iOS/Swift |
| add-to-xcode | Auto-register new files with Xcode projects |
| sidequest | Spawn parallel Claude sessions in new terminal tabs |
| debug-council | Multi-agent debugging with majority voting |
| feature-council | Multi-agent feature implementation, synthesize best parts |
| parallel-builder | Decompose plans into parallel tasks |
| model-council | Get consensus from multiple AI models |
| auto-permissions-review | Per-session AI permission review using Claude Haiku |
Reduces permission prompt fatigue by auto-approving safe operations and sending ambiguous commands to Haiku for review. Per-session — each terminal enables independently.
| Tool | Default mode | Accept-edits mode (Shift+Tab) |
|---|---|---|
Read, Glob, Grep, LS, Agent | instant allow | instant allow |
Simple Bash (ls, cat, find, git status) | instant allow | instant allow |
| Complex Bash (pipes, substitution) | Haiku reviews | Haiku reviews |
Edit, Write | normal prompt (you decide) | Haiku reviews |
/auto-permissions-review-install # one-time setup
/auto-permissions-review-enable # turn on (this session)
/auto-permissions-review-disable # turn off (this session)
Agent skills are higher-level skills that:
All agent skills use the -agent suffix to indicate they orchestrate other skills.
Business analysis, research, and strategy:
| Agent | What It Does | Sub-Agents | Skills Used |
|---|---|---|---|
| cmo-agent | AI CMO: SEO audit, content, Reddit, HN, X growth | 6 (seo, geo, content-writer, reddit, hackernews, x) | site_audit.py, chart-generation |
| brand-research-agent | Analyze brands from websites | 5 (visual, voice, product, audience, competitive) | None |
| product-engineer-agent | Design products with specs + visuals | 5 (industrial, mechanical, user, manufacturing, innovation) | image-generation |
| market-researcher-agent | Research markets (TAM/SAM/SOM) | 4 (trend, consumer, industry, opportunity) | chart-generation |
| patent-lawyer-agent | Patent drafting + IP guidance | 5 (prior-art, patentability, claims, strategy, drafter) | image-generation |
| competitive-intel-agent | Analyze competitors | 4 (feature, pricing, positioning, market) | chart-generation, image-generation |
| copywriter-agent | Marketing copy | 4 (headlines, body, ads, CTA) | None |
| review-analyst-agent | Analyze product reviews | 4 (scraper, sentiment, issues, recommendations) | chart-generation |
| pitch-deck-agent | Create pitch decks | Workflow | slide-generation, chart-generation, image-generation |
Create complete media by combining multiple generation skills:
| Agent | What It Creates | Skills Used |
|---|---|---|
| walkthrough-script-agent | Walkthrough video scripts for app features | app-demo-agent, voice-gen |
| video-producer-agent | Complete videos with voiceover + music | video-gen, voice-gen, music-gen, media-utils |
| podcast-producer-agent | Podcast episodes, dialogues | voice-gen, music-gen, media-utils |
| audio-producer-agent | Audiobooks, ads, jingles | voice-gen, music-gen, media-utils |
| social-producer-agent | Multi-asset content packs | image-gen, video-gen, voice-gen |
| app-demo-agent | Polished demos from screen recordings | device-framer, voice-gen, music-gen, media-utils |
Example: patent-lawyer-agent workflow:
User: "Draft a patent for my self-watering planter"
│
▼
┌─────────────────────────────────────────────────────┐
│ patent-lawyer-agent │
│ │
│ 1. prior-art-searcher → Finds existing patents │
│ 2. patentability-analyst → Assesses novelty │
│ 3. claims-strategist → Drafts claims │
│ 4. ip-strategy-advisor → Recommends approach │
│ 5. patent-drafter → Writes full application │
│ │ │
│ ▼ calls │
│ ┌─────────────────────┐ │
│ │ image-generation │ → Patent figures │
│ └─────────────────────┘ │
│ ┌─────────────────────┐ │
│ │ chart-generation │ → Patent landscape │
│ └─────────────────────┘ │
└─────────────────────────────────────────────────────┘
│
▼
Output: Complete patent document + generated figures
USER: "Create a 30-second product video for my new wireless earbuds"
PRODUCER WORKFLOW:
1. Asks: Duration? Style? Have product images?
2. Plans: 5 scenes (reveal, features, lifestyle, CTA)
3. Generates:
- 5 video clips (Veo 3.1)
- Voiceover script (Gemini TTS)
- Background music (Lyria)
4. Assembles:
- Concat clips with transitions
- Mix voice + music (music ducks under voice)
- Merge audio with video
5. Delivers: final_product_video.mp4
OUTPUT: Professional video with VO, music, transitions
# FFmpeg for media assembly
brew install ffmpeg # macOS
apt install ffmpeg # Linux
# Python package for Google APIs
pip install google-genai
Combine professional agents with producer agents for complete workflows:
USER: "Analyze Nike's brand, then create a product video for my sneakers"
WORKFLOW:
1. brand-research-agent analyzes nike.com
→ Extracts colors, typography, voice, audience
→ Saves brand_profile.json
2. video-producer-agent uses brand_profile.json
→ Matches Nike's visual style
→ Uses appropriate music mood
→ Follows voice guidelines
RESULT: Video that feels "Nike-like"
USER: "Research the smart home market, design a new product, then create a pitch deck"
WORKFLOW:
1. market-researcher-agent → Market report with TAM/SAM/SOM
2. product-engineer-agent → Product spec with BOM
3. patent-lawyer-agent → IP assessment
4. pitch-deck-agent → Investor presentation
RESULT: Complete product launch package
10 debug solver agents focused on finding bugs:
| Agents | Purpose |
|---|---|
debug-solver-1 through debug-solver-10 | Independent bug finding and fixing |
Focus: Root cause analysis, finding the ONE correct fix, chain-of-thought debugging.
10 feature solver agents focused on building features:
| Agents | Purpose |
|---|---|
feature-solver-1 through feature-solver-10 | Independent feature implementation |
Focus: Codebase pattern matching, edge case coverage, comprehensive implementation.
10 builder solver agents focused on implementing assigned pieces:
| Agents | Purpose |
|---|---|
builder-solver-1 through builder-solver-10 | Implement assigned piece of decomposed plan |
Focus: File ownership, shared contracts, parallel execution, integration.
5 specialized analyzer agents, each focused on one aspect:
| Agent | Focus |
|---|---|
style-structure | Folder organization, file layout, module patterns |
style-naming | Naming conventions for files, variables, functions, classes |
style-patterns | Error handling, data access, logging, configuration |
style-testing | Test location, naming, structure, assertions |
style-frontend | Component patterns, styling, state (if applicable) |
Focus: Language-agnostic detection, real examples from codebase, structured output.
6 specialized marketing agents that work in parallel:
| Agent | Focus |
|---|---|
seo-analyst | Technical SEO audit with exact HTML fix snippets |
geo-analyst | AI search visibility (ChatGPT, Perplexity, Google AI Overview) |
content-writer | Full SEO articles (1500-3000 words) + 4-week content calendar |
reddit-scout | Active thread discovery + copy-paste-ready comments with risk assessment |
hackernews-scout | Show HN submission + founder comment + objection responses |
x-scout | Tweet threads + standalone tweets + 7-day calendar + influencer mapping |
Also includes site_audit.py: stdlib-only technical SEO crawler with 3-tier scoring (static analysis, PageSpeed Insights API, Lighthouse CLI).
5 specialized brand analysts that work in parallel:
| Agent | Focus |
|---|---|
visual-analyst | Colors, typography, logo, imagery style |
voice-analyst | Tone, messaging, taglines, copy patterns |
product-analyst | Offerings, features, USPs, pricing |
audience-analyst | Demographics, psychographics, pain points |
competitive-analyst | Market position, competitors, differentiation |
Focus: Web scraping, pattern extraction, structured brand profile output.
5 specialized engineering perspectives + visual generation:
| Agent | Focus |
|---|---|
industrial-designer | Form, ergonomics, aesthetics + generates concept renders |
mechanical-engineer | Mechanism, materials, assembly + generates exploded views |
user-researcher | User needs, pain points, usability |
manufacturing-advisor | Feasibility, costs, production |
innovation-scout | Existing solutions, patents, differentiation |
4 specialized market analysis perspectives:
| Agent | Focus |
|---|---|
trend-analyst | Market size, growth, trends, future outlook |
consumer-researcher | Customer segments, behavior, needs |
industry-analyst | Market structure, players, dynamics |
opportunity-finder | Gaps, opportunities, entry points |
5 specialized IP perspectives:
| Agent | Focus |
|---|---|
prior-art-searcher | Find existing patents, publications |
patentability-analyst | Assess novelty, non-obviousness |
claims-strategist | Draft claims, claim strategy |
ip-strategy-advisor | Protection strategy, timing, costs |
patent-drafter | Draft complete patent applications with generated figures |
4 specialized copywriting perspectives:
| Agent | Focus |
|---|---|
headlines-writer | Headlines, hooks, taglines |
body-copy-writer | Long-form persuasive copy |
ad-copy-writer | Platform-specific ad copy |
cta-specialist | Calls to action, conversion copy |
4 specialized competitive analysis perspectives:
| Agent | Focus |
|---|---|
feature-analyst | Product features, capabilities |
pricing-analyst | Pricing models, value comparison |
positioning-analyst | Brand positioning, messaging |
market-position-analyst | Market share, company health |
4 specialized review analysis perspectives:
| Agent | Focus |
|---|---|
review-scraper | Find and collect reviews from platforms |
sentiment-analyzer | Analyze sentiment, emotions, trends |
issue-identifier | Categorize complaints, find patterns |
improvement-recommender | Prioritize fixes, create action plans |
Both debug and feature agent types:
Builder agents are different:
Council skills will ask you how many agents to use (3-10), or specify directly:
| Mode | Agents | Use Case |
|---|---|---|
debug council of 3 | 3 | Fast, simple bugs |
debug council of 5 | 5 | Standard debugging |
debug council of 10 | 10 | Critical bugs |
feature council of 3 | 3 | Simple features |
feature council of 5 | 5 | Standard features |
feature council of 10 | 10 | Complex features |
Minimum 3 agents for councils - needed for meaningful voting/synthesis.
Parallel-builder uses as many agents as needed based on task decomposition (up to 10).
These agents are invoked automatically by their skills and should not be called directly.
Analyze a codebase to extract its conventions and patterns. Generates a reusable style guide:
style guide
generate style guide for this project
analyze codebase conventions
How it works:
.claude/codebase-style.mdOutput:
Use iOS/Swift code as reference to implement the equivalent Android feature:
ios to android: implement this feature for Android
convert this Swift code to Kotlin
port UserProfile from iOS to Android
How it works:
Key principle: Same behavior, same data shapes, but idiomatic for each platform.
Use Android/Kotlin code as reference to implement the equivalent iOS feature:
android to ios: implement this feature for iOS
convert this Kotlin code to Swift
port UserProfile from Android to iOS
Works the same as ios-to-android but in reverse direction.
Automatically register newly created source files with Xcode projects:
Create a new ProfileViewModel.swift in the ViewModels folder
What happens:
add_to_xcode.rb to register it with the .xcodeprojManual usage:
# After creating any source file in an Xcode project
ruby ${CLAUDE_PLUGIN_ROOT}/skills/add-to-xcode/scripts/add_to_xcode.rb Sources/MyNewFile.swift
Supported files: .swift, .m, .mm, .c, .cpp, .h
Requires: gem install xcodeproj
Spawn a new Claude Code session in a separate terminal to work on a different task:
/sidequest "Add a settings page with dark mode toggle"
/sidequest "Set up the database schema" --no-context
/sidequest # Interactive prompt for task description
What happens:
Use when: You're deep in a task but need to branch off for something else without losing your place.
macOS only (uses osascript for terminal control)
Research-aligned self-consistency for debugging. Each agent explores and debugs independently - no shared context:
debug council: fix this bug in my function
debug council of 5: important production issue
debug council of 10: critical bug, need maximum confidence
How it works (pure Wang et al., 2022):
Note: This is slower than shared-context approaches because each agent explores independently. Use for critical bugs where accuracy matters more than speed.
Multi-agent feature implementation. Each agent builds the feature independently, then synthesizes the best parts:
feature council: implement user authentication with OAuth
feature council of 5: add caching layer to the API
feature council of 10: complex payment integration
How it works:
Output shows:
Divide-and-conquer implementation from specs, PRDs, or plans. Decomposes into parallel tasks:
parallel-builder from docs/auth-prd.md
parallel-builder: full CRUD API for blog with posts, comments, users
parallel-builder something like src/features/users but for products
How it works:
Key differences from feature-council:
Where it shines (maximum speedup):
Falls back to sequential when:
Output shows:
Get consensus from multiple AI models (Claude, GPT, Gemini, Grok):
model council: review this architecture decision
model council with claude, gpt-4o: is this code secure?
model council all: critical decision, need all perspectives
Generate images with AI:
generate an image of a sunset over mountains
create a cyberpunk cityscape at night
make a watercolor painting of a cat
Generate app icons with transparent backgrounds:
generate an icon for a music app
create a flat style settings gear icon
make a 3D shopping cart icon for my e-commerce app
Remove backgrounds from images:
remove the background from this photo
make this image transparent
cut out the product from this image
Generate videos with AI:
generate a video of waves crashing on a beach at sunset
create a cinematic drone shot flying over mountains
make a video of a cat playing with yarn
Generate speech and audio:
read this text aloud: "Hello, welcome to my podcast"
generate a voiceover for this script
create narration for my video using a deep male voice
Generate music and songs:
create an upbeat pop song about summer
generate a cinematic orchestral soundtrack
make a lo-fi hip hop beat for studying
Create presentation slides:
create slides from this content: [paste JSON]
generate a PowerPoint presentation for my pitch
make slides for my market research report
Wrap screenshots and screen recordings in photorealistic iPhone frames:
frame this screenshot in an iPhone 16 Pro
wrap this screen recording in a device mockup
put this in an iPhone 17 Pro in cosmic orange on a dark background
Generate data-driven charts from data:
create a bar chart comparing our features to competitors
plot our monthly revenue: [100, 150, 220, 350]
generate a competitive positioning matrix
create a TAM/SAM/SOM chart: TAM $50B, SAM $5B, SOM $500M
make a pie chart showing use of funds
AI Chief Marketing Officer — enter a URL and get a full marketing team deployed:
be my AI CMO for https://mysite.com
run a full SEO audit on https://myapp.io and give me exact fixes
find Reddit and Hacker News opportunities for my product
write SEO articles for my site and create a content calendar
How it works:
site_audit.py crawls the site, scores SEO/Accessibility/Performance/Best PracticesOutput includes:
Optional API keys: GOOGLE_PSI_API_KEY for PageSpeed Insights scores, lighthouse CLI for full browser audit.
Analyze a brand from their website:
analyze the Nike brand from their website
research Apple's brand guidelines
what's the brand voice for Stripe?
Design new products with specs and visuals:
design a new portable phone charger
I have an idea for a smart water bottle, help me develop it
create a product spec for a pet feeding device
design a modular desk organizer and show me concept renders
create an exploded view of my product design
Research markets and opportunities:
what's the market size for smart home devices?
research the plant-based food market trends
is there an opportunity in sustainable packaging?
IP guidance and patent drafting (informational only):
is my invention patentable?
search for prior art on foldable drone designs
should I patent this or keep it as trade secret?
draft a full patent application for my invention
create a patent with figures for my self-watering planter
Create investor presentations:
create a pitch deck for my AI startup
build a seed round presentation
make investor slides for my SaaS company
Write marketing copy:
write headlines for our product launch
create ad copy for our Black Friday sale
write landing page copy for our new app
Analyze competitors:
analyze our competitors: Salesforce, HubSpot, Pipedrive
what are Notion's weaknesses?
create a competitive battlecard for sales
Analyze customer reviews:
analyze reviews for our product on Amazon
what are people complaining about with [competitor]?
find the top issues we should fix from customer feedback
Create complete videos with voiceover and music:
create a 30-second product video for my headphones
make a demo video for my SaaS app
create an explainer video about how our service works
Create podcast episodes and dialogues:
create a 5-minute podcast about AI with two hosts
make a fake interview between Einstein and Elon Musk
create an educational podcast episode about climate change
Create voiceovers, audiobooks, and audio ads:
create a 30-second radio ad for our coffee brand
generate an audiobook narration for this chapter
make a meditation audio with calming background music
Create social media content packs:
create a launch kit: 1 reel, 5 carousel images
make a week of social content for our product
create TikTok content for our new feature
Turn screen recordings into polished demo videos:
here's a screen recording of my app — turn it into a polished demo video
add voiceover to this screen recording: ~/Desktop/demo.mp4
take ~/Desktop/recording.mov, frame it in iPhone 17 Pro, add narration and music
If you see errors like:
API Error: Claude's response exceeded the 32000 output token maximum
Solution: Increase the max output tokens (only uses more when needed):
# Add to ~/.bashrc or ~/.zshrc
export CLAUDE_CODE_MAX_OUTPUT_TOKENS=64000
Then restart Claude Code.
This commonly happens with feature-council on complex features where agents generate complete implementations. The 64K limit allows full outputs without truncation.
If you see an error like:
OPENAI_API_KEY environment variable not set
Solution:
export OPENAI_API_KEY="sk-your-key-here"
~/.bashrc, ~/.zshrc, or ~/.bash_profile)source ~/.bashrcIf you hit rate limits:
Common causes:
If Claude doesn't use a skill when you expect it to:
/plugin list/plugin update skills@michaelboeding-skillsIf you update the plugin but Claude Code still uses an old version, or skills are missing:
Quick fix - run the update script:
# From the skills repo directory
./scripts/update-plugin.sh
Or manually clear the cache:
rm -rf ~/.claude/plugins/cache/michaelboeding-skills
rm -rf ~/.claude/plugins/cache/temp_local_*
Then in Claude Code:
/plugin update skills@michaelboeding-skills
Then restart Claude Code (quit and reopen - required for changes to take effect).
If a script fails to run:
python3 --versionecho $OPENAI_API_KEYpython3 ~/.claude/plugins/marketplaces/michaelboeding-skills/skills/image-generation/scripts/dalle.py --prompt "test"
If you see this error:
Architecture Mismatch Error
dlopen(...pydantic_core...incompatible architecture (have 'x86_64', need 'arm64'))
Cause: Pip installed x86_64 packages when running under Rosetta emulation.
Fix:
# Force arm64 architecture for pip installs
/usr/bin/arch -arm64 pip3 install --force-reinstall pydantic pydantic-core google-genai
Prevention:
./scripts/install.sh
All skills provide clear error messages when something goes wrong:
| Error | Meaning | Solution |
|---|---|---|
API_KEY environment variable not set | Missing API key | Export the required key (see Setup section) |
API error (401) | Invalid API key | Check your key is correct and active |
API error (429) | Rate limit exceeded | Wait and retry, or use different API |
API error (400) | Bad request | Check your prompt/parameters |
Content policy violation | Prompt rejected | Rephrase to be appropriate |
Text too long | Exceeded character limit | Shorten your text or split into parts |
MIT
name: image-generation
description: >
Use this skill for any image-related AI generation or editing task. Triggers include:
GENERATE: "generate image", "create image", "make picture", "draw", "visualize", "image of", "create art", "generate art"
EDIT: "edit image", "modify image", "change image", "update image", "fix image", "enhance image"
ADD/REMOVE: "add to image", "put in image", "remove from image", "delete from image", "add element"
STYLE: "style transfer", "make it look like", "convert style", "apply style", "in the style of"
PRODUCT: "product photo", "product placement", "place product", "mockup", "put product on"
COMPOSITE: "combine images", "merge images", "blend images", "create composite"
Supports text-to-image generation, image editing with references, product placement, style transfer, and multi-image composition using Google Gemini (Nano Banana Pro) or OpenAI DALL-E.Generate and edit images using AI (Google Gemini Nano Banana Pro, OpenAI DALL-E 3).
Capabilities:
Users can specify what they want:
| User Says | Mode | What Happens |
|---|---|---|
| "Generate an image of a sunset" | Generate | Text-to-image, no reference needed |
| "Create a logo for my coffee shop" | Generate | Text-to-image with text rendering |
| "Edit this image: add a hat to the cat" | Edit | User provides image, AI modifies it |
| "Remove the background from this photo" | Edit | User provides image, AI edits it |
| "Put this product on a kitchen counter" | Product | User provides product + optional scene |
| "Make this photo look like Van Gogh painted it" | Style | User provides photo, AI applies style |
| "Combine these photos into a group shot" | Composite | User provides multiple images |
Environment variables must be configured for the APIs to work. At least one API key is required:
OPENAI_API_KEY - For OpenAI DALL-E 3 image generationGOOGLE_API_KEY - For Google Gemini (Nano Banana / Nano Banana Pro)See the repository README for setup instructions.
gpt-image-1.5 (state of the art, best quality)gpt-image-1 (great quality, cost-effective)gpt-image-1-mini (fastest, most affordable)autoautoauto⚠️ Note: DALL-E 2 and DALL-E 3 are deprecated and will stop being supported on 05/12/2026.
gemini-2.5-flash-image): Fast, efficient, 1K resolution, up to 3 reference imagesgemini-3-pro-image-preview): Professional quality, up to 4K, thinking mode, up to 14 reference images (default)⚠️ Use interactive questioning — ask ONE question at a time.
⚠️ Use the AskUserQuestion tool for each question below. Do not just print questions in your response — use the tool to create interactive prompts with the options shown.
Q0: Model Selection
"Which image generation model would you like to use?
- Google Gemini (Nano Banana Pro) - Up to 4K, 14 reference images, style transfer, thinking mode (Recommended)
- OpenAI GPT Image 1.5 - State of the art, transparency, streaming, up to 16 input images
- OpenAI GPT Image 1 - Great quality, transparency, image editing
- OpenAI GPT Image 1 Mini - Fastest, most affordable"
Wait for response. If user doesn't have a preference, recommend Gemini for editing/reference tasks or GPT Image 1.5 for pure generation.
Q1: Reference
"I'll generate that image for you! First — do you have any reference images?
- Product photos to include
- Style references
- Images to edit
- No, generate from scratch"
Wait for response.
Q2: Aspect Ratio
"What aspect ratio?
- 1:1 (square)
- 16:9 (landscape/widescreen)
- 9:16 (portrait/vertical)
- 4:3 / 3:4 (classic)
- Other (2:3, 3:2, 4:5, 5:4, 21:9)
- Or specify"
Wait for response.
Q3: Resolution
"What resolution?
- 1K (fast)
- 2K (balanced)
- 4K (highest quality)"
Wait for response.
Q4: Style
"Any style preferences?
- Photorealistic
- Artistic/painterly
- Cartoon/illustration
- 3D render
- Or describe your own"
Wait for response.
| Question | Determines |
|---|---|
| Reference | Generation vs editing mode |
| Aspect Ratio | Image dimensions |
| Resolution | Quality level |
| Style | Prompt enhancement direction |
Parsing:
Transform the user request into an effective image generation prompt:
Example transformation:
Use the model selected by the user in Q0:
Check which API keys are configured in environment:
OPENAI_API_KEY → GPT Image models availableGOOGLE_API_KEY → Gemini (Nano Banana Pro) availableIf the user's selected model isn't available: Inform them and offer alternatives.
Model mapping from Q0:
gemini.py with gemini-3-pro-image-previewopenai_image.py with gpt-image-1.5openai_image.py with gpt-image-1openai_image.py with gpt-image-1-miniExecute the appropriate script from ${CLAUDE_PLUGIN_ROOT}/skills/image-generation/scripts/:
For OpenAI GPT Image - Text to Image:
python3 ${CLAUDE_PLUGIN_ROOT}/skills/image-generation/scripts/openai_image.py \
--prompt "your enhanced prompt" \
--model "gpt-image-1" \
--size "1024x1024" \
--quality "high" \
--output "/path/to/output.png"
For OpenAI GPT Image - With Transparent Background:
python3 ${CLAUDE_PLUGIN_ROOT}/skills/image-generation/scripts/openai_image.py \
--prompt "A product icon with no background" \
--model "gpt-image-1" \
--background "transparent" \
--quality "high" \
--output "/path/to/output.png"
For OpenAI GPT Image - Image Editing (with reference images):
python3 ${CLAUDE_PLUGIN_ROOT}/skills/image-generation/scripts/openai_image.py \
--prompt "Add a wizard hat to this cat" \
--model "gpt-image-1" \
--image "/path/to/cat.jpg" \
--input-fidelity "high" \
--output "/path/to/output.png"
For OpenAI GPT Image - Multiple Reference Images:
python3 ${CLAUDE_PLUGIN_ROOT}/skills/image-generation/scripts/openai_image.py \
--prompt "Create a gift basket containing these items" \
--model "gpt-image-1" \
--image "/path/to/item1.png" \
--image "/path/to/item2.png" \
--image "/path/to/item3.png" \
--output "/path/to/output.png"
For OpenAI GPT Image - With Mask (Inpainting):
python3 ${CLAUDE_PLUGIN_ROOT}/skills/image-generation/scripts/openai_image.py \
--prompt "Replace the pool with a garden" \
--model "gpt-image-1" \
--image "/path/to/scene.jpg" \
--mask "/path/to/mask.png" \
--output "/path/to/output.png"
For OpenAI GPT Image - Streaming with Partial Images:
python3 ${CLAUDE_PLUGIN_ROOT}/skills/image-generation/scripts/openai_image.py \
--prompt "A beautiful sunset over mountains" \
--model "gpt-image-1" \
--stream \
--partial-images 2 \
--output "/path/to/output.png"
For Google Gemini (Nano Banana Pro) - Text to Image:
python3 ${CLAUDE_PLUGIN_ROOT}/skills/image-generation/scripts/gemini.py \
--prompt "your enhanced prompt" \
--model "gemini-3-pro-image-preview" \
--aspect-ratio "1:1" \
--resolution "2K" \
--output "/path/to/output.png"
For Google Gemini - With Reference Images (editing, product placement, etc.):
python3 ${CLAUDE_PLUGIN_ROOT}/skills/image-generation/scripts/gemini.py \
--prompt "Add a wizard hat to this cat" \
--image "/path/to/cat.jpg" \
--aspect-ratio "1:1" \
--resolution "2K"
For Google Gemini - Multiple Reference Images (composition, style transfer):
python3 ${CLAUDE_PLUGIN_ROOT}/skills/image-generation/scripts/gemini.py \
--prompt "Place this product on the kitchen counter in this scene" \
--image "/path/to/product.png" \
--image "/path/to/kitchen.jpg" \
--aspect-ratio "16:9" \
--resolution "2K"
For Google Gemini (Nano Banana - faster, fewer features):
python3 ${CLAUDE_PLUGIN_ROOT}/skills/image-generation/scripts/gemini.py \
--prompt "your enhanced prompt" \
--model "gemini-2.5-flash-image" \
--aspect-ratio "1:1"
Missing API key: Inform the user which key is needed and how to set it up:
API rate limit: Suggest waiting or trying the other API.
Content policy violation: Rephrase the prompt to be more appropriate.
Generation failed: Retry with simplified prompt or different API.
Both OpenAI GPT Image and Google Gemini support reference images for advanced editing:
OpenAI GPT Image: Up to 16 input images, with input_fidelity: high for preserving faces/logos
Google Gemini: Nano Banana (up to 3), Nano Banana Pro (up to 14)
Tip: For best results with reference images, be specific about what you want to preserve vs. change.
| Feature | GPT Image 1.5 | GPT Image 1 | GPT Image 1 Mini | Nano Banana | Nano Banana Pro |
|---|---|---|---|---|---|
| Provider | OpenAI | OpenAI | OpenAI | ||
| Model ID | gpt-image-1.5 | gpt-image-1 | gpt-image-1-mini | gemini-2.5-flash-image | gemini-3-pro-image-preview |
| Best for | State of the art | Quality + value | Speed + cost | Fast generation | Professional assets |
| Sizes | 1024², 1536x1024, 1024x1536, auto | Same | Same | 1K only | Up to 4K |
| Quality options | low, medium, high, auto | Same | Same | N/A | N/A |
| Aspect ratios | 3 + auto | Same | Same | 10 options | 10 options |
| Reference images | Up to 16 | Up to 16 | Up to 16 | Up to 3 | Up to 14 |
| Image editing | Yes | Yes | Yes | Yes | Yes |
| Inpainting (mask) | Yes | Yes | Yes | Yes | Yes |
| Transparent background | Yes | Yes | Yes | No | No |
| Streaming | Yes | Yes | Yes | No | No |
| Input fidelity | high/low | high/low | low only | N/A | N/A |
| Output formats | png, jpeg, webp | Same | Same | png | png |
| Compression | 0-100% | Same | Same | No | No |
| Text rendering | Excellent | Excellent | Good | Good | Excellent |
| Thinking mode | No | No | No | No | Yes |
| Max prompt length | 32,000 chars | 32,000 chars | 32,000 chars | N/A | N/A |
| Speed | ~30-60s | ~20-40s | ~10-20s | ~10-20s | ~30-60s |
⚠️ DALL-E 2 and DALL-E 3 are deprecated and will stop being supported on 05/12/2026. Use GPT Image models instead.
评论 (0)
暂无评论,成为第一个评论者吧!