SkillAtlasSkill 详情

ima2

🌐 Live site: lidge-jun.github.io/ima2-gen · 한국어

审核状态:已审核Quality 72Security 52

复制安装命令

用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。

复制前请先查看来源、License 和安全提示。

项目 README

来源文件:README.md

抓取于 2026年8月13日

ima2-gen

ima2-gen logo

npm version Node.js License: MIT

🌐 Live site: lidge-jun.github.io/ima2-gen · 한국어

📖 Developer docs: Documentation site · 한국어

Read in other languages: 한국어 · 日本語 · 简体中文

ima2-gen is a local image generation studio for people who want the ChatGPT/Codex image workflow in a small desktop-like web app.

Install globally, sign in with ChatGPT OAuth or Grok OAuth, and start generating images and videos. Iterate with history, references, node branches, multimode batches, Canvas Mode cleanup, and Grok Video generation. Default OAuth paths need no API key; optional API-key providers (api, grok-api, gemini-api, agy) are also supported.

ima2-gen video playback with gallery sidebar showing generated images and videos.

Quick Start

npm install -g ima2-gen
ima2 setup
ima2 serve

Then open http://localhost:3333.

Docker

docker build -t ima2-gen .
docker run -d -p 3333:3333 -e IMA2_LAN_TOKEN=change-me -v ima2-data:/data ima2-gen

See docs/DOCKER.md for compose usage, required environment, and limitations.

To generate from the CLI, inspect the live lane catalog and choose explicit image/video defaults once:

ima2 models
ima2 defaults set image oauth/gpt-5.6-luna
ima2 defaults set video grok/grok-imagine-video-1.5
ima2 gen "a clean product photo of a red guitar pedal"
ima2 video "a cat playing piano" --duration 5 --resolution 720p
ima2 video "animate this scene" --ref photo.png --duration 10

ima2 gen and generate-mode ima2 video fail closed with NO_DEFAULT_MODEL until a CLI target is configured, unless that call passes --model <lane>/<model> or an explicit --provider <lane>. This prevents an upgrade from silently switching providers or billing lanes.

If 3333 is already occupied, ima2-gen binds the next available port and writes the actual URL to ~/.ima2/server.json. Use ima2 open or the URL printed in the terminal instead of assuming the port.

Using npx? See docs/NPX_QUICKSTART.md for the npx ima2-gen serve workflow.

One-Click Install (no npm required)

Don't have Node.js or npm? Use the platform install script — it detects your environment, installs Node LTS if needed, then installs ima2-gen.

macOS:

curl -fsSL https://lidge-jun.github.io/ima2-gen/install-mac.sh | bash

Windows (PowerShell):

irm https://lidge-jun.github.io/ima2-gen/install-windows.ps1 | iex

Linux / WSL:

curl -fsSL https://lidge-jun.github.io/ima2-gen/install-linux.sh | bash

Each script checks for nvm/fnm/brew/winget, installs Node LTS through the best available method, and handles stale process cleanup automatically.

Setup

ima2 setup offers four authentication choices:

  1. GPT OAuth — login with ChatGPT account (free, images only)
  2. Grok OAuth — login with xAI/Grok account (images + video)
  3. Both — GPT OAuth + Grok OAuth (full feature access)
  4. Web setup — configure everything in the web UI

Video generation requires Grok OAuth (option 2 or 3). Run ima2 grok login separately if you already have GPT OAuth configured and want to add video support; it defaults to the manual-paste flow.

Updating

Stop the running server with Ctrl+C, then:

npm install -g ima2-gen@latest

Ctrl+C now performs a clean shutdown — closing the database, stopping child processes, and releasing file locks. On older versions (< 1.1.22) or if you see EBUSY on Windows, use the install script which handles stale process cleanup automatically.

What It Does

  • Classic mode: generate, edit, reuse the current image, paste references, and continue from history.
  • Node mode: branch a good image into multiple directions without losing the original.
  • Multimode batches: launch several Classic outputs from one prompt, watch slot-by-slot progress, and continue from the best result.
  • Video generation: create short videos from text, a single image, or multiple reference images via Grok video models. SSE streaming shows planning → submitted → progress % → done. Video frame copy buttons (First/Mid/Last) let you extract and copy keyframes from generated videos.
  • Storyboard mode: toggle storyboard mode in the composer to maintain character and scene continuity across sequential frames. Works with both image and video generation — image keyframes are composed for video production, and video clips inherit character/environment lock rules.
  • Canvas Mode: zoom, pan, annotate, erase, clean backgrounds, keep transparent previews, and export either alpha or matte-backed versions.
  • Local gallery: keep generated assets on your machine with session-aware history. By default the gallery shows the current session and an All Images toggle reveals the full history; the default scope is sticky across sessions. Each image records its generation time and reasoning effort in the result metadata, so they persist across reloads.
  • Reference images: drag, drop, paste, and attach up to 5 references (images) or up to 7 references (video); large images are compressed before upload.
  • Prompt library imports: import local prompt packs, GitHub folders, and curated GPT-image prompt hints into the built-in prompt library.
  • Mobile shell: use the app bar, compose sheet, and compact settings toggle on smaller screens.
  • Observable jobs: active and recent jobs are tracked with safe logs and request IDs.

Agent Skills

ima2-gen ships three packaged skills for AI coding agents. These are Markdown instruction files that agents load to get structured workflows for image/video generation, frontend asset production, and design direction discovery.

SkillCommandWhat It Covers
Coreima2 skillCLI reference, prompting protocol, provider routing, Korean text, video workflows
Frontendima2 skill frontAsset pipeline (parallel gen, variant selection, provider routing), motion/video for web, responsive, a11y, anti-slop, 30+ reference files
UI/UX Designima2 skill uiuxImage-first design direction discovery, UX states, design-isms, product personalities, DESIGN.md workflow, 18 reference files
ima2 skill ls            # list available skills
ima2 skill front         # print the frontend skill
ima2 skill uiux          # print the design skill
ima2 skill front path    # print file path (for agents)
ima2 skill front --json  # JSON wrapper (for agents)
ima2 skill front refs    # list reference modules (35 files)
ima2 skill front ref motion        # load one reference module
ima2 skill install --dir <path>     # install skills to agent's skill dir
ima2 skill install --tmp            # install to temp dir (fallback)

The Frontend and UI/UX skills are production-grade design engineering guides adapted for the ima2 workflow. They cover typography, color systems, layout discipline, Korean UX patterns, motion choreography, and visual verification, with every asset generation step mapped to ima2 gen, ima2 video, and ima2 multimode commands.

SSE Multiplexing

The web UI uses a single GET /api/events Server-Sent Events connection for all generation progress. Multimode, node, and video requests are submitted as async POST (202 { requestId }) and progress events are multiplexed through a shared event bus. This eliminates the browser 6-connection limit that previously caused gallery hangs during concurrent generation. CLI clients that do not send async: true still receive per-request SSE streams for backward compatibility.

Provider Paths

Image generation can run through the local Codex/ChatGPT OAuth path, a configured OpenAI API key, the bundled Grok provider, or the Gemini provider via Antigravity CLI.

  • provider: "oauth" uses the local Codex OAuth proxy.
  • provider: "api" calls the OpenAI Responses API with the hosted image_generation tool.
  • provider: "grok" starts bundled progrok on 127.0.0.1:18645, runs mandatory xAI Web Search plus a planner pass (default: grok-4.5, configurable in settings or via --planner-model), then calls xAI Images API through the local proxy. grok-4.3 remains available as an explicit compatibility override.
  • provider: "grok-api" calls the xAI Images API directly with XAI_API_KEY (no bundled progrok OAuth proxy).
  • provider: "agy" spawns the Antigravity CLI (agy -p) to generate images via Google Gemini's default_api:generate_image tool (model: nano-banana-2). Output is fixed at 1024×1024 JPEG, max 3 reference images. No web search, quality, or size controls.
  • provider: "gemini-api" calls the Google Generative Language API directly. Supports two models: nano-banana-2 (Gemini 3.1 Flash Image) and nano-banana-pro (Gemini 3 Pro Image). Auth is via GEMINI_API_KEY env var, web UI key management, or a Vertex AI service account JSON (VERTEX_SERVICE_ACCOUNT_JSON). When both an API key and Vertex credentials are configured, Vertex takes priority. Supports variable aspect ratios (1:1 through 21:9) and four resolution tiers (512px, 1K, 2K, 4K); these controls are only honored on the direct API path — the Vertex AI endpoint ignores aspect/size because it does not accept the response_format field. Per-model cost differs: nano-banana-2 (Flash): 512=$0.001, 1K=$0.003, 2K=$0.004, 4K=$0.006; nano-banana-pro: 1K=$0.007, 2K=$0.007, 4K=$0.013. No web search or mask controls.
  • API-key generation supports classic generate, edit, mask-guided edit, multimode, and node generation.
  • Grok generation supports Classic, Node, and Agent flows. If a Classic reference, Node parent image, or Agent current image is present, ima2 switches the final Grok call to xAI image edit so image-to-image context is preserved.

If no provider is specified, the app keeps the current GPT OAuth/default behavior. GPT OAuth and API-key generation default to gpt-5.6-luna; the API-key path also defaults to low reasoning and 1024x1024 unless the request passes validated options. Grok image generation defaults to grok-imagine-image-quality.

Grok image generation exposes a model picker (grok-imagine-image / grok-imagine-image-quality) and a size picker (aspect ratio + 1k/2k resolution). The Settings page prefers the Grok Build weekly credits percentage and reset time from GET /v1/billing?format=credits; if that source is unavailable, it falls back to the legacy monthly billing window and $used/$limit. A Switch Account button starts a device-code OAuth flow (POST /api/auth/switch) for re-authenticating without leaving the app.

Grok video generation defaults to canonical grok-imagine-video-1.5; grok-imagine-video remains available for base-model-only Ref2V, V2V edit, and extension paths, and the legacy grok-imagine-video-1.5-preview string is accepted as an alias. Three modes are auto-detected from reference count: text-to-video (0 refs), image-to-video (1 ref), and reference-to-video (2-7 refs, max 10s duration). 1080p is available for grok-imagine-video-1.5 prompt-only text-to-video and single image/frame image-to-video; prompt-only 1.5 uses the internal white-canvas I2V shim before the upstream request. Video controls include duration (1-15s), resolution (480p, 720p, 1080p when supported), and aspect ratio (1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, auto).

Settings workspace showing GPT OAuth active and API key provider available.

Model Guidance

The app defaults to gpt-5.6-luna for image generation and Prompt Builder planning. Older supported models remain explicit compatibility choices.

  • gpt-5.6-luna — current image and Prompt Builder default.
  • gpt-5.6-terra / gpt-5.6-sol — current GPT-5.6 alternatives when your account exposes them.
  • gpt-5.5, gpt-5.4, gpt-5.4-mini — supported compatibility choices.

The app also exposes quality (low, medium, high) and moderation (auto, low) controls.

Workflows

Classic Mode

Use Classic when you want one strong result quickly.

  1. Write a prompt.
  2. Attach or paste references if needed.
  3. Pick model, quality, size, format, and moderation.
  4. Generate one image, or enable multimode to fan out several candidate slots from the same prompt.
  5. Copy, download, continue from the result, or send it into Canvas Mode.

For a control-by-control guide to Prompt Studio, multimode recipes, Direct mode, reasoning effort, and gallery favorite behavior, see the Prompt Studio manual.

Multimode sequence with four candidate slots generating from one prompt and active job history in the sidebar.

Node Mode

Use Node mode when you want to explore branches.

Node mode with connected generated cards and compact per-node metadata.

Each node keeps its own prompt and result. Root nodes can attach local references; child nodes use the parent image as their source. Completed jobs are matched back to nodes by request ID, so reloads and graph version conflicts can recover finished results.

Canvas Mode

Use Canvas Mode when a generated image is close but needs targeted cleanup before the next prompt.

  • Separate viewport panning from selection so you can move around a zoomed image without accidentally changing annotations.
  • Use annotation, eraser, multiselect, grouping, undo/redo, and sticky notes while keeping the original gallery image available.
  • Pick background-cleanup seeds, preview the mask, and save the cleanup as a canvas version.
  • Detect transparent images and show a checkerboard preview; export with preserved alpha or with a chosen matte color.
  • Saved canvas versions stay hidden from Gallery and HistoryStrip, but Canvas Mode can reuse them and attach a canvas version as the next reference.

Canvas Mode with zoom controls, annotation marks, a sticky note, and the canvas toolbar.

Prompt Library And Imports

The prompt library can now be filled from local files, GitHub folders, curated sources, and GPT-image hint packs. Imported prompts are indexed locally so search and ranking work without re-importing the same source every session.

Prompt import dialog for bringing prompts into the library, showing GitHub folder controls, curated sources, and searched prompt candidates before import.

Experimental Card News Mode

Card News is still dev-only and experimental. It is hidden in the default published runtime unless explicitly enabled for development, and it should not be treated as a stable public feature yet.

Settings

The settings workspace keeps account, model, appearance, and language controls away from the generation sidebar.

Settings workspace with account navigation and generation model controls.

CLI Commands

Server

CommandDescription
ima2 serve [--dev]Start the local web server; --dev enables verbose server diagnostics
ima2 setupReconfigure saved auth
ima2 statusShow config and OAuth status
ima2 doctorDiagnose Node, package, config, and auth
ima2 doctor image-probe [--json]Run sanitized image probes for no-image diagnostics
ima2 openOpen the web UI
ima2 resetRemove saved config

Client

These require a running ima2 serve. The CLI covers every server route. The most common ones are below — the full CLI reference lists everything (generation, history, sessions, prompt library, annotations, Card News, observability, config).

CommandDescription
ima2 models [--kind image|video] [--lane <lane>] [--json]List live lanes, status, model IDs, and capabilities
ima2 defaults set image|video <lane>/<model>Persist the fail-closed CLI target for image or video generation
ima2 defaults reset image|videoRemove a persisted CLI generation target
ima2 gen <prompt> [--model <lane>/<model>]Generate from the CLI; requires an explicit target or saved image default
ima2 edit <file> --prompt <text>Edit an existing image
ima2 multimode <prompt>Multi-image SSE generation
ima2 video <prompt> [--model <lane>/<model>]Generate video through a Grok or MCP lane; requires an explicit target or saved video default
ima2 ls [--session <id>] [--favorites]List recent history
ima2 show <name> [--metadata]Reveal a generated asset
ima2 prompt ls -q <search>Search the prompt library
ima2 inflight ls [--terminal]List active and recent jobs (alias of ps)
ima2 config set <key> <value>Write to ~/.ima2/config.json
ima2 pingHealth-check the running server

The server advertises its actual port at ~/.ima2/server.json. If 3333 is busy, the backend falls back to 3334+ and CLI commands follow the advertised URL. Override discovery with --server <url> or IMA2_SERVER=http://localhost:3333.

ima2 models --kind image
ima2 gen "poster" --model oauth/gpt-5.6-luna --reasoning-effort high
ima2 edit input.png --prompt "make it rainy" --web-search
ima2 multimode "two cats playing" -n 2
ima2 video "a cat playing piano" --model grok/grok-imagine-video-1.5 --duration 5 --resolution 720p
ima2 video "animate this" --model grok/grok-imagine-video-1.5 --ref photo.png --aspect-ratio 16:9
ima2 inflight ls --terminal
ima2 config set imageModels.reasoningEffort high

Full reference: docs/CLI.md.

Configuration

Config priority:

environment variables > ~/.ima2/config.json > built-in defaults
VariableDefaultDescription
IMA2_PORT / PORT3333Web server port
IMA2_HOST127.0.0.1Web server bind host
IMA2_OAUTH_PROXY_PORT / OAUTH_PORT10531OAuth proxy port
IMA2_SERVER—CLI target override
IMA2_CONFIG_DIR~/.ima2Config and SQLite location
IMA2_ADVERTISE_FILE~/.ima2/server.jsonRuntime discovery file
IMA2_GENERATED_DIR~/.ima2/generatedGenerated image directory
IMA2_IMAGE_MODEL_DEFAULTgpt-5.6-lunaServer fallback image model
IMA2_REASONING_EFFORTmediumDefault reasoning effort for the default (GPT OAuth) path; one of none, low, medium, high, xhigh
IMA2_NO_OAUTH_PROXY—Set 1 to disable the auto-started OAuth proxy
IMA2_LOG_LEVELinfoNormal serve defaults to info; dev mode defaults to debug; supports debug, info, warn, error, or silent
IMA2_INFLIGHT_TERMINAL_TTL_MS300000Recent terminal job retention for debug views
OPENAI_API_KEY—API key for the provider: "api" Responses API image path and auxiliary API-key features
XAI_API_KEY—API key for provider: "grok-api" direct xAI Images API path
IMA2_API_IMAGE_MODEL_DEFAULTgpt-5.6-lunaDefault image model for provider: "api"
IMA2_API_REASONING_EFFORTlowDefault reasoning effort for provider: "api"
IMA2_API_IMAGE_SIZE1024x1024Default size for provider: "api"
IMA2_API_ALLOW_WEB_SEARCHtrueToggle web search for provider: "api"
IMA2_GROK_PROXY_HOST127.0.0.1Host for the bundled progrok proxy
IMA2_GROK_PROXY_PORT18645Port for the bundled progrok proxy
IMA2_NO_GROK_PROXY—Set 1 to disable automatic progrok startup
IMA2_GROK_PLANNER_MODELgrok-4.5Grok search/planner model (also configurable via settings UI or --planner-model CLI flag)
IMA2_GROK_PLANNER_TIMEOUT_MS60000Timeout for Grok search and planner calls
IMA2_GROK_IMAGE_MODEL_DEFAULTgrok-imagine-image-qualityDefault final Grok image model
IMA2_GROK_VIDEO_MODEL_DEFAULTgrok-imagine-video-1.5Default Grok video model
IMA2_GROK_GENERATION_TIMEOUT_MS120000Timeout for the final Grok Images API call
IMA2_OAUTH_MASKED_EDIT_ENABLEDfalseOpt-in feature flag for masked-edit requests on the OAuth path (#31, groundwork only)
GEMINI_API_KEY—API key for provider: "gemini-api" direct Generative Language API path
VERTEX_SERVICE_ACCOUNT_JSON—Google service account JSON for Vertex AI auth with provider: "gemini-api"; takes priority over GEMINI_API_KEY when both are set
IMA2_AGY_BINagy on PATHExplicit path to the Antigravity CLI binary for provider: "agy"
IMA2_MAX_PARALLEL24Server-wide parallel generation cap

Logging modes

ima2 serve keeps terminal output intentionally quiet: startup URLs, warnings, and errors stay visible, while request/node/OAuth structured logs are hidden by default.

Use ima2 serve --dev, npm run dev, or IMA2_LOG_LEVEL=debug ima2 serve when you need request IDs, node generation phases, OAuth stream diagnostics, or inflight state transitions. Explicit IMA2_LOG_LEVEL and ~/.ima2/config.json values still override the built-in defaults.

API Reference

The endpoint list moved to docs/API.md so this README can stay focused on first-run use.

Useful references:

Troubleshooting

ima2 ping says the server is unreachable Start ima2 serve, then check ~/.ima2/server.json. You can also run ima2 ping --server http://localhost:3333.

GPT OAuth login does not work Re-run ima2 setup (option 1), confirm ima2 status, then restart ima2 serve.

fetch failed repeats on a proxy/VPN network Check that the local OAuth proxy is reachable. On networks that require a proxy, enable your proxy client's TUN/TURN-style mode, then retry openai-oauth --port 10531. If it still fails, set HTTP_PROXY and HTTPS_PROXY in the same terminal that runs ima2 serve or openai-oauth. On Windows, also check for auto-start network interception tools, including DNS/fragmentation bypass tools such as SecretDNS, because they can break OAuth or streaming image responses even when the browser appears connected.

Images fail with API_KEY_REQUIRED Set OPENAI_API_KEY or configure an API key before using provider: "api". The default GPT OAuth path still works without an API key.

Image generation returns EMPTY_RESPONSE or no image data Run ima2 doctor image-probe --json > ima2-image-probe.json and attach the safe JSON when opening an issue. For GPT OAuth cases, also capture ima2 gen "고양이" --model oauth/gpt-5.6-luna --no-web-search --json and ima2 gen "고양이" --model oauth/gpt-5.6-luna --json while ima2 serve is running. Do not share ChatGPT cookies, OAuth token files, API keys, raw upstream responses, prompt history, or generated base64. See the FAQ support bundle.

A large reference image fails The app compresses large JPEG/PNG references before upload. If a file still fails, convert it to JPEG or PNG at a lower resolution and try again. HEIC/HEIF files are not supported by the browser path.

Old gallery images are missing after updating Recent versions moved generated images from the installed package folder to ~/.ima2/generated. Run ima2 doctor and see Recover old images.

gpt-5.5 fails but other models work Update Codex CLI first, then retry. If it still fails, your account or backend route may not expose the same image capability or quota for gpt-5.5 yet; use gpt-5.4 as the stable fallback.

The app opened on a different port If the requested server port is busy, ima2-gen falls back to the next available port and records it in ~/.ima2/server.json. If the port is unexpectedly 3457, your shell may also have inherited PORT=3457 from another local tool. Run unset PORT or start with IMA2_PORT=3333 ima2 serve.

Port 10531 is already used on Windows Some Windows security tools, including AnySign4PC.exe, may occupy the default OAuth proxy port. Current builds track the actual fallback OAuth port. If you still need a manual override, start with IMA2_OAUTH_PROXY_PORT=11531 ima2 serve and check ima2 doctor.

For more beginner-friendly answers, see the FAQ.

Development

git clone https://github.com/lidge-jun/ima2-gen.git
cd ima2-gen
npm install
npm run dev
npm run typecheck
npm test
npm run build

npm run dev builds the UI and starts the TypeScript server entry with --watch and verbose server diagnostics. npm run typecheck, npm run build:server, and npm run build:cli verify the TypeScript migration and package emit path. Node mode and Canvas Mode are part of the packaged UI by default.

Contributors

License

MIT

其他

高风险

  • 来源需自行核对维护者身份。
  • 包含脚本或命令调用,安装前请复核。
  • 可能需要外部 token、网络权限或第三方服务。
  • 存在潜在风险命令,请谨慎安装。
  • 扫描发现:4 条。

Codex — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/lidge-jun/ima2-gen.git
  3. 将 "skills/ima2" 文件夹复制到 Codex 的 skills 目录中。
  4. 重启 Codex 让新的 skill 生效。

Codex — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Codex 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Codex 让新的 skill 生效。

Claude Code — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/lidge-jun/ima2-gen.git
  3. 将 "skills/ima2" 文件夹复制到 Claude Code 的 skills 目录中。
  4. 重启 Claude Code 让新的 skill 生效。

Claude Code — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Claude Code 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Claude Code 让新的 skill 生效。

Cursor — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/lidge-jun/ima2-gen.git
  3. 将 "skills/ima2" 文件夹复制到 Cursor 的 skills 目录中。
  4. 重启 Cursor 让新的 skill 生效。

Cursor — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Cursor 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Cursor 让新的 skill 生效。

GitHub Copilot — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/lidge-jun/ima2-gen.git
  3. 将 "skills/ima2" 文件夹复制到 GitHub Copilot 的 skills 目录中。
  4. 重启 GitHub Copilot 让新的 skill 生效。

GitHub Copilot — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 GitHub Copilot 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 GitHub Copilot 让新的 skill 生效。

Windsurf — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/lidge-jun/ima2-gen.git
  3. 将 "skills/ima2" 文件夹复制到 Windsurf 的 skills 目录中。
  4. 重启 Windsurf 让新的 skill 生效。

Windsurf — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Windsurf 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Windsurf 让新的 skill 生效。
查看 SKILL.md 原文
name: ima2
description: "Use the ima2-gen CLI/server to generate, edit, inspect, and manage local AI image generation jobs."

ima2 Skill

Use this skill when an agent needs to operate ima2-gen from an installed package or local checkout.

Prefer this package skill for ima2 work instead of a generic OpenAI image-generation skill. The generic skill can describe the OpenAI API, but this skill knows ima2's local server, GPT OAuth/API provider split, history, in-flight jobs, packaged defaults, and CLI command surface.

Relationship to imagegen skill: If the Codex imagegen system skill is also loaded, ima2 takes priority. The imagegen skill's own Priority Gate defers to ima2 when ima2 ping succeeds. Do not use both in the same generation task.

First Commands

Start by discovering the local package and running server state:

ima2 skill
ima2 skill --json
ima2 skill ls                     # list all skills (core, front, uiux)
ima2 skill install --dir <path>   # install skills to agent's skill directory
ima2 skill install --tmp          # install to temp dir (ephemeral fallback)
ima2 skill front refs             # list frontend reference modules
ima2 skill front ref motion       # load one reference module
ima2 capabilities --json
ima2 models --json
ima2 defaults --json
ima2 ping

If the server is not running:

ima2 serve
ima2 open

Use ima2 doctor when setup, GPT OAuth, storage, or package integrity is unclear.

Generate Images

List ready image lanes, choose a persistent CLI target, then generate:

ima2 models --kind image
ima2 defaults set image oauth/gpt-5.6-luna
ima2 gen "a clean product photo of a red guitar pedal"

Bare ima2 gen fails closed when no CLI image target is configured. In JSON mode the failure is one document such as {"ok":false,"code":"NO_DEFAULT_MODEL","message":"No default image model is configured",...} and exits 2. Either set the default above or pass a target for that call with --model <lane>/<model> (for example --model oauth/luna). Never rely on an implicit provider; --provider auto was removed.

Use high quality when output fidelity matters:

ima2 gen "a print-ready poster" --model oauth/luna --quality high

Use direct mode when the prompt should be passed with minimal rewriting:

ima2 gen "exact prompt text" --model oauth/luna --mode direct

--mode explained:

  • auto (default): the server may augment, restructure, or enrich the prompt before sending it to the image model. Good for casual or short prompts.
  • direct: the prompt is passed as-is with minimal server-side rewriting. Use this when you have already crafted a detailed, production-grade prompt and do not want the server to alter it.

Use request-level overrides only for that one call:

ima2 gen "cinematic mountain" --model oauth/gpt-5.5 --reasoning-effort high

Use Grok when the request should run through bundled progrok, mandatory xAI Web Search, planner pass (default: grok-4.3), and xAI Images API:

ima2 grok login
ima2 grok status
ima2 gen "cinematic neon city" --model grok/grok-imagine-image-quality

ima2 grok login defaults to the manual-paste flow.

Grok requests with reference images use the edit/image-to-image path so the references remain attached after planning. Keep Grok references to three total input images.

Prompting Guidance

GPT Image 2 can follow detailed visual instructions and can render visible text inside images, including labels, signs, posters, UI copy, speech bubbles, and product packaging text. Do not avoid text just because older image models were weak at it.

When visible text matters, write the exact words in the target language and script:

  • Good: A Korean poster with the exact headline "오늘 공연" and subtext "입장 무료".
  • Bad: A Korean poster with some Korean text.

Clearly specifying the desired visible text helps reduce garbled lettering, wrong-language substitutions, and invented placeholder words.

For dense or important text, specify:

  • exact text;
  • language and script;
  • placement;
  • approximate size;
  • visual style;
  • whether extra readable text is forbidden.

OpenAI's prompting guide additionally recommends: put literal text in quotes or ALL CAPS, state typography (font style, size, color, placement) as explicit constraints, and for exact copy demand it verbatim. The strongest official pattern is a dedicated text block:

Poster headline (EXACT, verbatim, no extra characters):
"Fresh and clean"
Typography: bold sans-serif, high contrast, centered, clean kerning.
Ensure the text appears once and is perfectly legible.

For tricky words such as brand names or uncommon spellings, spell them out letter-by-letter to improve character accuracy. Use medium or high quality whenever the image contains small text, dense panels, or multiple fonts. When localizing an existing image, translate the visible text verbatim, add no new words, and preserve everything else — layout, imagery, hierarchy — without reflowing the design.

GPT Image 2 can generate both stylized and realistic outputs. State the style directly, for example:

  • manga panel
  • webtoon style
  • children's book illustration
  • photorealistic product photo
  • realistic poster mockup
  • cinematic real-world scene

Text rendering is improved, but it is still not a typesetting engine. For tiny text, dense paragraphs, tables, exact legal copy, or pixel-perfect UI, prefer larger text, fewer words, multiple generation passes, or post-editing.

Agent Image Prompt Protocol

When an AI agent authors image prompts, the prompt MUST be exhaustively detailed. Vague one-liners produce generic, unusable output. Write every prompt as if you are briefing a senior photographer or illustrator who cannot ask follow-up questions. When using --mode auto, the server augments short prompts, but a detailed prompt still produces far better results than relying on auto-augmentation alone. For production assets, prefer --mode direct with a fully-specified prompt.

Structured Prompt Contract

Detailed is not enough — the prompt must be structured. OpenAI's official gpt-image prompting guide recommends composing prompts in a consistent field order — scene/background → subject → key details → constraints — and using labeled segments or line breaks instead of one long paragraph for complex requests. OpenAI's own showcase prompts use labeled blocks such as Context, Characters, and Composition. Apply these rules to every agent-authored prompt:

  • Write labeled sections, not a wall of prose. Long prompts are fine; an unstructured long prompt is not — it becomes impossible to iterate on.
  • Order fields by priority. Scene-first is the official default; lead with the subject when identity or product fidelity dominates. Field order is a priority signal to the model, not a fixed syntax.
  • Bind attributes locally. Keep each object's color, material, pose, count, and position in the same sentence as the object, and state spatial relationships explicitly (foreground/background, left/right, behind, facing, closest to camera).
  • Every sentence must change pixels. State aspect intent, exact hex colors, and transparent background needs directly; cut decorative filler words that describe nothing visible.
  • Do not wrap prompts in JSON. Structured fields are an authoring tool; render them as labeled natural-language sections. Vendors that support JSON prompts (e.g. FLUX) document that JSON and prose are understood equally well — JSON buys automation, not quality.

Required Spec Fields

Every agent-authored prompt MUST include all applicable fields. Omit a field only when it genuinely does not apply (e.g. no text in the image).

Use case: <slug: photorealistic-natural | product-mockup | ui-mockup | infographic-diagram | scientific-educational | ads-marketing | productivity-visual | logo-brand | illustration-story | stylized-concept | historical-scene>
Asset type: <where the asset will be used: hero, OG image, card, avatar, icon, texture, game sprite, etc.>
Primary request: <one clear sentence describing the desired image>
Scene/backdrop: <specific environment — not "nice background">
Subject: <main subject with identifying details: material, color, shape, posture, expression>
Style/medium: <exact style: editorial photography, flat illustration, 3D render, watercolor, etc.>
Composition/framing: <camera angle, crop, subject placement, negative space intent>
Lighting/mood: <light source, direction, color temperature, mood, time of day>
Color palette: <specific hex codes or named palette — not "modern colors">
Materials/textures: <surface details: matte plastic, brushed steel, linen, weathered wood, etc.>
Text (verbatim): "<exact text to render>" with font style, size, color, placement
Constraints: <must-keep invariants>
Avoid: <explicit negative constraints>

Specificity Rules

Bad (vague)Good (specific)
"a nice hero image""wide landscape product shot of a matte black thermos on a wet granite countertop, soft morning window light from the left, shallow depth of field, warm neutral tones, negative space on the right for headline overlay"
"modern background""soft radial gradient from #f8f9fa center to #e9ecef edges, subtle paper grain texture at 3% opacity, no objects, no patterns"
"Korean food photo""overhead flat-lay of budae-jjigae in a black stone pot, surrounded by small banchan dishes on a dark wood table, steam visible, warm tungsten lighting, editorial food photography style"
"logo on white""centered geometric mark: two interlocking triangles forming a hexagonal negative space, flat #1a1a2e on #ffffff, no gradients, strong silhouette at 32px, generous padding"
"a dashboard screenshot""realistic SaaS dashboard UI: top nav with avatar, left sidebar with 6 nav items, main area showing a line chart (3 series, 12 months) and a 4-column data table with 8 rows, light theme, Inter font, compact density"

Prompt Anti-Patterns

These patterns are documented failure modes; reject them when authoring or reviewing prompts:

Anti-patternWhy it failsDo instead
Keyword soup (beautiful, stunning, 8k, trending)Comma-separated tag piles are a documented anti-pattern for natural-language image modelsStructured narrative sentences: subject + attributes + relations
Unmotivated quality tokens (masterpiece, 8K, ultra-detailed)OpenAI's guide: lens, framing, and lighting language is more reliable for realism than generic quality tokensName the look: shallow depth of field, soft window light from the left, editorial photography
Trusting precision specs (85mm f/1.2, 5600K)Official guidance: detailed camera specs may be interpreted loosely — they are look cues, not optical simulationPrefer perceptual terms: medium close-up, eye level, warm tungsten mood; keep mm/Kelvin only as style hints
Contradictory constraints (minimalist + 12 required objects)Conflicting demands make the model silently drop some of themResolve conflicts before generating; one intent per field
Rewriting everything each iterationLoses working invariants, causes driftChange ONE variable per pass, restate invariants

Negative constraints are model-specific. For GPT Image, write exclusions as plain prose inside the prompt — No extra text, no logos, no watermark — this is the officially recommended form; there is no separate negative-prompt parameter. Do not copy diffusion-style negative lists (wall, frame) into GPT Image prompts; that syntax belongs to models with a dedicated negative field (e.g. Imagen), where instruction words like "no/don't" are in turn discouraged.

Quality and Size Selection

Asset PurposeQualitySizeNotes
Quick draft / iterationlow1024x1024Fastest; square
Final hero / product shothigh1536x1024 landscape, 1024x1536 portraitOr target aspect ratio
OG / social cardhigh1200x640Nearest 16px multiple of 1200x630
Mobile herohigh1024x1536Portrait
Print / 4Khigh3840x2160 or 2160x3840Max gpt-image-2 supports
Texture / tilemedium1024x1024Square, seamless edges
Icon / avatarmedium512x512 or 256x256Small canvas
Game environment concepthigh1792x1024 or 2048x1152Wide cinematic
Storyboard (for i2v)high1024x10243x3 grid, square

Cutout Assets and Background Strategy

GPT Image 2 does not reliably produce true transparent (alpha) backgrounds. Use the solid-background-then-remove strategy for cutout assets:

Generate on a pure solid background:

  • Black (#000000) for reflective/metallic/glass subjects
  • White (#ffffff) for dark/matte/opaque subjects
  • Brand color when the target page background is known

State the exact hex and ban AI additions: "PURE SOLID BLACK background hex #000000. No checkerboard, no transparency pattern, no gradient, no floor plane, no shadow, no vignette." Use --mode direct.

ima2 gen "3D chrome splash on PURE SOLID BLACK background hex #000000. \
  No gradient, no floor, no shadow, no vignette." \
  --quality high --size 1024x1024 --mode direct -o splash.png

Remove background after generation:

  • CSS mix-blend-mode: screen (black bg on light page)
  • CSS mix-blend-mode: multiply (white bg on dark page)
  • ima2 Canvas Mode background cleanup (export with alpha or matte)
  • ima2 edit asset.png --prompt "remove the background, keep only the subject"
  • Programmatic: sharp / ImageMagick / rembg

Anti-pattern: requesting "transparent background" or "PNG with alpha" in the prompt — the model often produces a fake checkerboard burned into the image.

Korean Text in Images

When generating images with Korean text:

  • Write the exact Korean string in quotes: "오늘의 추천", not "some Korean text"
  • Describe the scene in English and keep only the visible Hangul string in Korean: A clean summer poster with the exact Korean headline "여름 축제". Practitioner testing found all-Korean prompts produced garbled Hangul while English prompts with a quoted Korean string rendered correctly (heuristic, not a guarantee)
  • Start with short, label-like strings (a headline, a button) before attempting body copy; Hangul glyph complexity makes long dense text the most failure-prone case
  • Specify font style explicitly: 고딕체 (Gothic/Sans-serif) or 명조체 (Myeongjo/Serif)
  • Specify placement (top-center, bottom-left) and approximate size relative to the canvas
  • For mixed Korean + English, specify which script appears where and in what hierarchy
  • After generation, always inspect the result with view_image — garbled or substituted Hangul is common and must be caught before use
  • For critical Korean text, generate 2-4 candidates (-n 4) and pick the cleanest render
  • If a render is right except for the text, do a targeted ima2 edit pass that restates the exact string and changes only the text region; if spelling still will not stabilize after a couple of passes, stop retrying
  • For legally or commercially exact Korean copy (packaging, UI, contracts), the reproducible production path is: generate the image with a reserved empty text area (no text in that region), then composite real type with an actual Korean font in an editor or code. Korean text failure is a cross-model limitation, not an ima2-specific one

Multi-Candidate Strategy

For important visual assets (hero images, key illustrations, brand materials), generate multiple candidates and select the best:

# 4 candidates from one prompt
ima2 gen "<detailed prompt>" -n 4 -d ./candidates --quality high

# Or multimode for structurally different directions
ima2 multimode "<detailed prompt>" --max-images 4 -d ./candidates

After generation, inspect every candidate with view_image before selecting. Do not blindly use the first result.

Prompt Iteration

  • Start with one high-detail prompt. Inspect the result with view_image.
  • On the next pass, make ONE targeted change and re-specify all constraints. Do not rewrite the entire prompt from scratch.
  • Repeat invariants every iteration to prevent drift.
  • This mirrors the official guidance: start from a clean baseline, iterate with small single-variable follow-ups instead of overloading one prompt, and when a detail drifts, restate it explicitly — never assume it persists.
  • If the model consistently fails on a detail, try rephrasing, breaking the request into a base generation + ima2 edit pass, or switching --mode.

Frontend Asset Quick Recipes

Copy-paste starters for common frontend assets:

Hero image (landing page):

ima2 gen "Use case: product-mockup. Asset type: landing page hero. A premium wireless headphone floating at a slight angle against a soft warm-gray studio backdrop. Matte black finish with brushed aluminum accents. Soft three-point studio lighting, key light from upper-left. Shallow depth of field. Wide composition with generous negative space on the right for headline overlay. No text, no logos, no watermark." \
  --quality high --size 1536x1024 --mode direct -o hero.png

OG / social share image:

ima2 gen "Use case: ads-marketing. Asset type: social share card. Clean product flat-lay of a notebook, pen, and ceramic mug on a white marble desk. Overhead shot. Soft diffused daylight. Space in the upper third for title overlay. Warm neutral palette. No text, no logos, no watermark." \
  --quality high --size 1200x640 --mode direct -o og-image.png

App screenshot mockup background:

ima2 gen "Use case: stylized-concept. Asset type: hero background for device mockup. Soft abstract gradient from #f0f4f8 to #dbeafe with subtle geometric shapes at 5% opacity. Clean, modern, minimal. No objects, no patterns, no text." \
  --quality medium --size 1920x1088 --mode direct -o mockup-bg.png

Avatar / profile placeholder:

ima2 gen "Use case: stylized-concept. Asset type: user avatar. Friendly stylized portrait of a young professional, neutral expression, looking slightly left. Flat illustration style with subtle shadows. Solid #e5e7eb background. Circular crop safe. No text." \
  --quality medium --size 512x512 --mode direct -o avatar.png

Korean product hero:

ima2 gen "Use case: product-mockup. Asset type: Korean service landing hero. A modern smartphone at 15-degree tilt showing a clean fintech app UI. The screen displays a balance card with exact text \"잔액 1,234,500원\" in 고딕체, large centered. Soft gradient backdrop from #f8fafc to #e2e8f0. Studio lighting from upper-right. No other text, no logos, no watermark." \
  --quality high --size 1536x1024 --mode direct -o korean-hero.png

Game environment concept art:

ima2 gen "Use case: stylized-concept. Asset type: game environment concept art. A vast underground cavern with bioluminescent fungi on limestone walls. A narrow stone bridge crosses a dark chasm. Volumetric blue-green light from fungi clusters. Cinematic concept art style with industrial realism. Wide-angle, low camera, deep perspective. Mist rising from below. No characters, no text, no watermark." \
  --quality high --size 1792x1024 --mode direct -o cave-env.png

Reference / I2I Workflows

Reference generation:

ima2 gen "turn this into a clean product render" --ref input.png --quality high

Multimode reference workflow:

ima2 multimode "create four coherent variations" --ref input.png --max-images 4

Node-mode reference workflow:

ima2 node generate "continue this concept" --ref input.png

Image edit workflow:

ima2 edit input.png --prompt "make the object blue while preserving composition"

Do not use positional edit prompts. ima2 edit requires --prompt.

Structured Edit Brief

OpenAI's official edit pattern is "change only X" + "keep everything else the same" — an edit prompt does not need to re-describe the whole final image, but it must make the delta and the invariants explicit. Author every edit prompt as a brief:

Desired result: <one sentence describing the edited image's final state>
Change only: <the specific modification>
Preserve exactly: <named lock list: facial structure, pose, product
  silhouette, logo geometry, text spelling, framing, perspective, palette,
  lighting, shadows>
Do not add or remove: <protected elements>

"Keep everything else the same" alone is weak — name the fragile properties in the lock list, and repeat the same lock list on every iterative edit pass to prevent drift.

Annotated inputs. If the edit source or a reference image carries drawn markup (arrows, boxes, circled regions, sticky notes), the model tends to treat the markup as image content and reproduce it. Prefer sending the clean image plus text instructions derived from the markup. When the annotated image must be sent, state before and after the edit list that the markup is temporary editing instructions only — interpret it, apply the edits, then remove every trace of it from the output.

Removal edits. "Remove X" alone is weak. Pair the removal command with a positive description of what replaces it, then lock the rest: "Remove the sticky note. Show the continuous walnut desk surface where it was, matching the surrounding grain, lighting, and perspective — no residue, outline, or discoloration. Preserve every other object, the framing, and the color grading exactly." For stubborn removals, generate multiple candidates and re-edit only the residual region instead of enlarging the prompt.

Multi-Reference Rules

When passing multiple --ref images, label each reference by index and role inside the prompt, then state the relationships explicitly:

Image 1: base scene and composition.
Image 2: subject identity reference.
Image 3: style reference.

Place the subject from Image 2 into Image 1. Apply only Image 3's palette and
brushwork. Preserve Image 1's framing, background, perspective, and lighting.
  • Put the most identity-critical reference (face, logo, product) first: documented GPT Image behavior preserves the first input with the richest texture and detail.
  • When several faces must all stay recognizable, combine them into one composed reference image before generating instead of passing many separate portraits.
  • For compositing, specify the source element, its destination and location, the preserved context, and harmonization: scale, perspective, lighting, shadows.

Parallel Generation

There is no --parallel flag. For multiple candidates from the same prompt, prefer one server-side batch request:

ima2 gen "four poster candidates" -n 4 -d ./out --quality high
ima2 multimode "four different poster directions" --max-images 4

For truly different prompts, independent CLI jobs can run concurrently against the same server. Capture request IDs with JSON output, then monitor or cancel:

ima2 gen "variation 1" --quality high --json
ima2 gen "variation 2" --quality high --json
ima2 ps --json
ima2 cancel <requestId>

Treat capabilities.limits.maxParallel as advisory client-side queue guidance only. It is not a guaranteed server-side semaphore.

Agent Mode (web UI only)

Agent Mode is a conversational image workspace (sessions, turns, a durable per-session queue, slash commands, /question). It is served at /api/agent/* and lives in the web UI — there is no ima2 agent CLI command. From the CLI, drive generation with ima2 gen, ima2 edit, ima2 multimode, and ima2 node generate instead.

Watching Jobs

Use JSON when another agent needs to reason about active work:

ima2 inflight ls --json
ima2 inflight ls --kind multimode --terminal --json

Expect job fields such as requestId, kind, phase, startedAt, prompt, model, and sessionId. Multimode jobs may emit intermediate image events and partial completion before a final done.

Prompt Import

Build a structured image prompt from a message or transcript:

ima2 prompt build --message "make this product prompt clearer" --json
ima2 prompt build --messages @conversation.json --json

Preview a local markdown/text prompt source before committing:

ima2 prompt import preview ./prompts.md --json

Import a JSON export body:

ima2 prompt import json ./prompts-export.json --folder __root__

Import a raw image into history:

ima2 history import ./local-image.png

Defaults

Inspect the running server defaults, including defaults.cli.image and defaults.cli.video in JSON:

ima2 defaults --json

Inspect local effective defaults without contacting a server:

ima2 defaults --local --json

Discover live model IDs and lane status before choosing a CLI target:

ima2 models
ima2 models --kind image --lane oauth --json
ima2 models --kind video --json

ima2 models --json has the stable shape {"ok":true,"kinds":{"image":[],"video":[]}}. It requires the server; an unreachable server returns SERVER_UNREACHABLE and exits 3.

Persist the server-side model defaults shared by GPT OAuth and API provider paths:

The built-in OAuth image default is gpt-5.6-luna; Grok image and video code defaults are grok-imagine-image-quality and grok-imagine-video respectively. Use grok-imagine-video-1.5 explicitly when its quality or 1080p capabilities are needed.

ima2 defaults set model gpt-5.5

Persist the fail-closed CLI image and video targets separately:

ima2 defaults set image oauth/gpt-5.6-luna
ima2 defaults set video grok/grok-imagine-video
ima2 defaults reset image
ima2 defaults reset video

Setting a CLI target validates the live catalog. Unknown models and lanes are rejected, and locked/disconnected/key-missing lanes cannot become defaults.

Persist the default reasoning policy:

ima2 defaults set reasoning high

Restart a running server after changing persisted defaults:

ima2 serve

Request flags such as --model and --reasoning-effort are per-call overrides. They do not change persistent defaults.

Capability Values

Use ima2 capabilities --json as the source of truth for:

  • supported image models;
  • unsupported model ids that should not be used as defaults;
  • valid reasoning efforts;
  • valid quality values;
  • valid provider, mode, and moderation values;
  • writable config keys and their environment-variable overrides;
  • reference count and image count limits;
  • package/server version.

Use only models from:

valid.imageModels.supported

Do not pick models from:

valid.imageModels.unsupported

Discover writable configuration keys:

ima2 config keys --json

Safety Notes

  • Do not print API keys, OAuth tokens, config files, or .env values.
  • Use ima2 capabilities --json before guessing model names.
  • Use ima2 skill path when an agent needs the installed Markdown skill path.
  • Use ima2 skill <name> refs to discover reference modules for front/uiux skills.
  • Use ima2 skill <name> ref <refname> to load a specific reference module on demand.
  • Use ima2 skill install --dir <path> to install skills to the agent's skill directory.
  • Use ima2 inflight ls --json or ima2 ps --json to inspect active jobs.

Video Generation

Generate AI videos through a configured Grok or MCP lane. Grok OAuth requires a SuperGrok subscription; MCP lanes require their own connected subscription.

Quick Start

ima2 models --kind video
ima2 defaults set video grok/grok-imagine-video
ima2 video "a cat playing piano"                    # text-to-video, uses saved default
ima2 video "animate this" --model grok/grok-imagine-video --ref photo.png
ima2 video "cinematic" --model grok/grok-imagine-video --ref a.png --ref b.png

Targets use --model <lane>/<model>; a bare ID is accepted only when it is unique across lanes. Generate-mode video also accepts an explicit --provider <grok|grok-api|runway|higgsfield>. --provider auto is removed.

Runway and Higgsfield are MCP lanes. They submit POST /api/mcp/generate (202) and the CLI waits on SSE until completion. MCP generation supports -n 1 only, and --ref values must be generated gallery filenames, not arbitrary local paths. Core-only planning/session flags are rejected with FLAG_NOT_SUPPORTED. ima2 gen and ima2 video also accept --character <element-id|name> on MCP lanes: the element must be a character element in the assets workspace with a provider binding for the selected lane (Runway: stateless refs + optional @tag). Fail-closed envelopes: CHARACTER_ELEMENT_NOT_FOUND, CHARACTER_ELEMENT_AMBIGUOUS, CHARACTER_BINDING_MISSING, CAPABILITY_MISMATCH (core lane or model without image_references), BINDING_NOT_READY, and server-side CHARACTER_ELEMENT_CONFLICT / CHARACTER_REFS_EXCEED_PROVIDER_CAP. Runway is available when connected; Higgsfield remains locked until its paid lane is enabled. Inspect current state with ima2 models --kind video.

ima2 upscale <generated-file> upscales through the MCP media-action pipeline: images take --scale-factor 2|4|8|16 (above 2 requires --flavor sublime), --flavor, --sharpen, --smart-grain, --ultra-detail; videos take no parameters. Multishot generation is POST /api/mcp/multishot (CLI surface planned). Video edit is the 2-step edit-video-preview → edit-video-submit media action; stage-1 returns a synchronous keyframe preview.

Modes (auto-detected from --ref count)

RefsModeMax Duration
0text-to-video15s
1image-to-video15s
2-7reference-to-video10s

grok-imagine-video-1.5 supports image-to-video and supports 1080p for prompt-only text-to-video and single image/frame image-to-video. Prompt-only 1.5 text-to-video is implemented as an internal white-canvas image-to-video anchor because upstream 1.5 rejects raw T2V. The old grok-imagine-video-1.5-preview string is accepted as a compatibility alias. 1.5 does not support reference_images Ref2V, V2V edit, or extension. For 2+ references, use grok-imagine-video and keep duration at 10s or less. ima2 may auto-retry a rejected 1.5 Ref2V request with the base model; read effectiveModel and modelFallback from the final result before naming or reporting the output.

Parameters

FlagValuesDefault
--duration1–15 (seconds)5
--resolution480p, 720p, 1080p (1.5 T2V canvas shim or I2V)480p
--aspect-ratioauto, 1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3auto
--model<lane>/<model>; Grok base or 1.5 (preview alias accepted)grok-imagine-video after selecting the Grok lane
--topicany string(none)
--sessionsession ID(none)
-o, --outoutput file pathsaved under configured generated dir
--json(flag)false

Series Continuity (--topic)

--topic is legacy/best-effort series context. Prefer branch-local artifact continuity with ima2 video continue, Classic "Continue here", gallery video drag, or Node parent-video generation. Those flows use the previous generated video's last frame plus its stored revisedPrompt lineage.

Continuation is last-frame image-to-video, not true video-to-video. Only a single still frame carries over — motion, trajectory and camera movement are not preserved. Expect the next clip to start from that pose rather than to inherit momentum. Frames are pulled server-side with ffmpeg when the source is a generated file, falling back to in-browser canvas capture when ffmpeg is unavailable.

ima2 video "episode 1: morning routine" --topic "daily-vlog"
ima2 video "episode 2: commute" --model grok/grok-imagine-video --topic "daily-vlog"

Planning Layer

Prompts are NOT sent directly to the video model. A Grok planner rewrites your prompt with web search context for better results. The revisedPrompt in the response shows what was actually sent. Default planner model is grok-4.3 (configurable in settings UI).

Override the planner model per-request:

ima2 video "prompt" --model grok/grok-imagine-video --planner-model gpt-5.5
ima2 video "prompt" --model grok/grok-imagine-video --planner-model gpt-5.4

Grok 4.3 Prompt Surfaces

SurfaceFilesResponsibility
Image search/plannerlib/grokImageAdapter.tsWeb-search context and final image prompt for Grok image generation/editing.
Video plannerlib/grokVideoAdapter.ts, lib/grokVideoPlannerPrompt.tsFinal video prompt for T2V/I2V/Ref2V, duration pacing, and continuity lineage when present.
Video analyzerroutes/videoExtended.tsFirst/last-frame analysis prompt for recreating or continuing an existing generated video.
Agent/runtime prompt uselib/agentRuntime.ts, card/template planner modulesHigher-level orchestration surfaces that may create image/video prompt inputs but do not replace the video planner contract.

For video, the Grok 4.3 planner must produce one focused English prompt with: core subject, expected action/motion, camera/composition, environment/style, dialogue/audio intent, ending frame/continuity handoff, and constraints. If videoContinuity exists, the lineage is authoritative context: continue from the latest clip's final frame and final audio/dialogue state without restarting the scene. The planner also applies duration pacing: use the selected seconds as the full clip runtime, expand even short requests into a production-level sequence, and make the clip feel complete through composition, blocking, camera movement, motion rhythm, sound/dialogue timing, and an ending hold.

Active Video Prompt Requirement

Blank video prompts are blocked. Weak natural-language prompts are allowed, but agents should always write an active prompt that includes:

  • shot design: opening frame composition, one motivated reveal or change, settling final frame — not a checklist of elements
  • camera intent: choose the camera move that serves the scene (macro push-in for product, orbit for spatial VFX, handheld for documentary, crane for scale) — do not default to "slow dolly in"
  • production choices: concrete material/texture, motivated lighting source, depth layers (foreground/mid/background), lens framing — instead of generic "cinematic" or "volumetric lighting"
  • sound: music style, no music, room tone, or sound-effects-only
  • dialogue: exact line in original language or explicit no-dialogue
  • ending frame: final pose, camera state, last spoken words, and final sound cue — self-explanatory enough to serve as the first frame of a next clip
  • duration pacing: beat structure scales with length — 1-4s gets one action, 5-7s gets setup/turn/hold, 8-10s gets two connected beats, 11-15s gets a three-beat arc

The planner is model-aware: it adjusts for grok-imagine-video (simpler, bolder composition at 480p) vs grok-imagine-video-1.5 (finer detail, 1080p textures). For 1.5 text-to-video, the server uses a white-canvas shim internally; the planner automatically writes a fresh-scene prompt without referencing a source image.

Anti-slop: the planner rejects generic prestige phrases ("AAA trailer", "senior VFX artist", "shot on RED"), filler lighting ("volumetric", "neon glow"), and unmotivated dark/moody defaults. Write what the camera actually sees.

Prerequisites

ima2 grok login     # authenticate (manual-paste flow)
ima2 grok status    # verify connection
ima2 serve          # server must be running

Output

SSE streaming events: planning → submitted → progress (0-100%) → done. The submitted and done payloads include requestedModel, effectiveModel, and modelFallback so agents can report when a requested 1.5-preview Ref2V job actually ran on grok-imagine-video. CLI --json prints video.requestedModel, video.effectiveModel, and video.modelFallback; use path/filename for local chaining.

Discover Valid Parameters

ima2 capabilities --json | jq '.valid.videoModels'

Advanced Workflows

Image-First Video (best quality)

Generate a high-quality still image first, then animate it. This produces better results than text-to-video alone because the video model has a concrete visual anchor.

Critical rule for i2v: Compose ALL characters and the environment together in ONE image. Do NOT use individual portrait refs for i2v — the video model needs a single composed scene to animate from.

Keyframe image provider rule (MANDATORY):

  • Primary: GPT Image 2 (OpenAI, provider: oauth) with quality: high, maximum resolution matching the target video aspect ratio. For 16:9 video use 1792x1024. For 1:1 use 1024x1024. For 9:16 use 1024x1792.
  • Fallback: Grok (provider: grok, model grok-imagine-image-quality). Only aspect ratio must match — resolution does not matter because i2v accepts any resolution source image and internally rescales.
  • GPT Image 2 produces superior keyframes: better lighting coherence, character consistency, and fine detail that survives i2v animation. Always try GPT first.
  • The i2v model internally rescales the source image to its native resolution regardless of input size, so there is no benefit to upscaling a Grok fallback image.

ref2v vs i2v decision:

ScenarioUseWhy
Need 2+ character identity lock from separate refsref2v (grok-imagine-video, max 7 refs, max 10s)Refs lock character appearance
Single composed scene with all elementsi2v (1.5-preview or base, 1 ref)Better motion quality from composed start
Continue from previous videovideo continue (last frame as i2v ref)Lineage metadata preserved
# Multi-character scene: compose BOTH characters in one image first
# Primary: GPT Image 2 at high quality, max resolution, aspect ratio matching 16:9 video
ima2 gen "cinematic wide shot of Bruce Lee in yellow tracksuit facing Elon Musk in dark gi, underground fight arena, dramatic lighting, 16:9" --quality high --size 1792x1024 -o scene.png

# Fallback if GPT fails: Grok quality model, match aspect ratio only
# ima2 gen "same prompt" --provider grok --model grok-imagine-image-quality --size 1824x1024 -o scene.png

# Then animate from the composed scene
ima2 video "Bruce throws a rapid jeet kune do combination" --ref scene.png --duration 10 --resolution 720p --aspect-ratio 16:9

Multi-Shot Video (connected scenes)

Create a sequence of connected clips using --topic for narrative continuity. Each generation receives context from previous clips in the same topic.

# Scene 1: Establishing shot
ima2 video "wide establishing shot of a busy Tokyo street at night, neon signs" \
  --topic "tokyo-night" --duration 5

# Scene 2: Medium shot (planner sees Scene 1's revised prompt)
ima2 video "medium shot following a person walking through the crowd" \
  --topic "tokyo-night" --duration 5

# Scene 3: Close-up (planner sees Scenes 1+2)
ima2 video "close-up of rain drops on a neon sign reflection" \
  --topic "tokyo-night" --duration 5

The planner receives previous prompts from the same topic as continuity context. This is best-effort prompt guidance, not a guarantee that subjects, palette, or style will remain identical. For branch-local continuation, use ima2 video continue instead.

Storyboard-to-Video Chaining (9-panel storyboard → i2v loop)

The highest-quality video production workflow. Since Grok i2v accepts only one image input, pack the entire action sequence into a single 3×3 (9-panel) storyboard grid image. The i2v model reads the panels as a visual script and animates the progression.

Full workflow:

keyframe image (GPT high)
    → GPT i2i with reference → 9-panel storyboard grid
        → Grok i2v (reads panels, animates sequence)
            → extract last frame
                → GPT i2i with last frame → next 9-panel storyboard
                    → Grok i2v
                        → repeat

Step 1 — Opening keyframe (GPT Image 2, quality: high, max resolution matching target aspect ratio):

ima2 gen "cinematic wide shot of two fighters in a dojo, dramatic lighting" \
  --quality high --size 1792x1024 --storyboard

Fallback: Grok grok-imagine-image-quality, match aspect ratio only — resolution does not matter because i2v internally rescales.

Step 2 — 9-panel storyboard grid (GPT Image 2 with keyframe as reference):

# Use the keyframe as reference, prompt describes 9 sequential panels
ima2 gen "Using this scene as reference, create a 3x3 storyboard grid (9 panels, thin black borders) showing a 15-second action sequence. Panel 1 (0s): ... Panel 2 (2s): ... Panel 9 (15s): ... Maintain identical character designs across all panels." \
  --ref keyframe.png --quality high --size 1024x1024

9-panel storyboard rules:

  • Grid layout: 3×3, thin black borders between panels
  • Read order: left-to-right, top-to-bottom (panels 1-9)
  • Panel 1 (top-left) MUST be solid black — this is a lead-in frame, not content. The i2v model starts from Panel 1's pixels; a black frame ensures the video begins with a clean fade-in instead of showing the grid. The 1-second black lead-in is auto-trimmed by the server.
  • Panels 2-9 carry the action sequence (8 key moments with timestamps)
  • Character designs MUST be identical across all panels
  • Vary camera angle per panel for dynamic energy
  • Each panel should look like a film still, not a sketch
  • Do NOT add timestamp labels or text to panels — they burn into the video
  • Square format (1024×1024) works best — i2v rescales internally

Step 3 — Animate storyboard via i2v:

ima2 video "This is a 9-panel storyboard. Animate the full sequence as one continuous 15-second clip following panels left-to-right, top-to-bottom. Panel 1: ... Panel 9: ... Sound: [describe music, SFX, dialogue]. Camera: [describe movement per beat]." \
  --ref storyboard.png --duration 15 --resolution 720p --model grok-imagine-video-1.5

i2v prompt rules for storyboard input:

  • Explicitly state "This is a 9-panel storyboard" at the start
  • Reference each panel by number with its action description
  • Always include Sound/Music direction — never leave audio undefined
  • Include Camera direction per beat (wide, close-up, tracking, handheld, slow-mo)
  • Describe the end frame explicitly for continuation

Step 4 — Extract last frame and repeat:

# Extract last frame via ffmpeg
ffmpeg -sseof -0.1 -i clip.mp4 -frames:v 1 -q:v 2 -update 1 lastframe.jpg -y

# Generate next storyboard using last frame as reference
ima2 gen "Using this fight scene last frame as reference, create a 3x3 storyboard grid..." \
  --ref lastframe.jpg --quality high --size 1024x1024

# Animate next storyboard
ima2 video "This is a 9-panel storyboard..." --ref storyboard2.png --duration 15

Fallback: continueFromVideo — If a storyboard image triggers content moderation (common with intense action/fight scenes), fall back to video continue with a detailed text prompt instead:

ima2 video continue "detailed action description with sound and camera direction" \
  --video "$PREV_CLIP" --duration 15

Clip duration is flexible — use 15s for action-dense sequences with many beats, 10s for transitions, 5s for quick cuts. The 9-panel storyboard works best with 15s clips (each panel ≈ 1.5-2s of screen time).

Music and sound are MANDATORY in i2v prompts — describe the score (orchestral, percussion, taiko drums), sound effects (impacts, whooshes, crashes), dialogue lines, and audio transitions. "No music" or undefined audio produces flat, lifeless output.

Video Continuation (extend/sequel)

To continue from an existing video's last frame:

# Get the last generated video filename
LAST=$(ima2 ls -n 1 --json | jq -r '.items[0].filename')

# True extension keeps the original clip and appends new motion
ima2 video extend "the camera slowly pulls back revealing the full scene" --video "$LAST" --duration 6

# Branch-local sequel keeps revisedPrompt lineage and starts from the last frame
ima2 video continue "from the last frame, the camera slowly pulls back, no music, footsteps echo, end on a still wide shot" --video "$LAST"

Or in the UI: use "Continue here" on a video, drag a video from gallery/history to the prompt composer, or create a child from a video node. These flows attach the previous video's last frame and carry a branch-local videoContinuity lineage stack. The stack stores up to 4 revised prompts using keep-start-plus-latest-3: start clip is preserved, and the newest three clips stay in context.

ima2 video extend is xAI native extension: it returns original+extension as a combined artifact. ima2 video continue is ima2 branch continuation: it creates a new clip from the generated video's last frame and persists lineage metadata.

Marketing/Product Video

Generate a product showcase video from a product image:

# Step 1: Generate or provide product image
ima2 gen "clean product photo of wireless earbuds on white background" -o product.png

# Step 2: Create dynamic product video
ima2 video "sleek product reveal with rotating camera, premium feel, studio lighting" \
  --ref product.png --duration 10 --resolution 720p --aspect-ratio 16:9

Style-Consistent Series

For maintaining visual style across multiple videos (e.g., social media series):

# First video establishes the style
ima2 video "minimalist animation of a coffee cup, flat design, pastel colors" \
  --topic "coffee-series" --duration 5

# Subsequent videos inherit style via planner context
ima2 video "same style, now showing latte art being poured" \
  --topic "coffee-series" --duration 5

ima2 video "same style, steam rising from the cup" \
  --topic "coffee-series" --duration 5

Batch Generation (scripting)

#!/bin/bash
PROMPTS=("sunrise over ocean" "waves crashing" "seagulls flying" "sunset colors")
TOPIC="ocean-day"

for prompt in "${PROMPTS[@]}"; do
  ima2 video "$prompt" --topic "$TOPIC" --duration 5 --json >> results.jsonl
  sleep 2  # rate limiting
done

Limitations

  • Max 15 seconds per clip (extend adds 2-10s more)
  • Reference-to-video (2+ refs): max 10 seconds, max 7 refs, grok-imagine-video effective model
  • 1080p resolution is available for grok-imagine-video-1.5 prompt-only text-to-video via the white-canvas I2V shim, and for image-to-video with a single image/frame source
  • Video edit/extend: grok-imagine-video only (1.5 is not supported)
  • Video edit input: max 8.7 seconds
  • Video extend input: 2-15 seconds; extension duration: 2-10 seconds

Video Editing (V2V)

Edit an existing video with a text prompt. This uses xAI's real video edit endpoint and saves the result as a generated video artifact.

# Get the local video file from a previous generation
VIDEO_FILE=$(ima2 video "ocean waves" --json | jq -r '.path')

# Edit: change style
ima2 video edit "Make the water glow neon blue, bioluminescent" --video "$VIDEO_FILE"

# Edit: add object
ima2 video edit "Add a sailboat in the distance" --video "$VIDEO_FILE"

# Edit: change mood
ima2 video edit "Make it stormy with dark clouds" --video "$VIDEO_FILE"

Constraints: grok-imagine-video only, input mp4 <=8.7s. Use -o/--out if you also need a local copy outside the generated directory.

Video Extension (Continue from Last Frame)

Extend a video from its last frame using xAI's video extension endpoint. The output combines the source video and extension, but continuity quality is provider-dependent.

Constraints: grok-imagine-video only, extension duration 2-10s. 1.5-preview is not supported for extension.

# Generate initial clip
VIDEO_FILE=$(ima2 video "a bird takes flight from a branch" --duration 5 --json | jq -r '.path')

# Extend: add 5 more seconds
ima2 video extend "the bird soars higher into the clouds" --video "$VIDEO_FILE" --duration 5

# Chain extensions for longer videos
EXTENDED=$(ima2 video extend "camera follows the bird" --video "$VIDEO_FILE" --duration 5 --json | jq -r '.filename')
ima2 video extend "bird lands on a distant tree" --video "$EXTENDED" --duration 5

Video Frame Extraction

Extract frames from generated videos for use as references or analysis.

# Extract last frame
ima2 video frame 1780226256355_50252101.mp4 --last -o lastframe.png

# Extract frame at specific timestamp
ima2 video frame 1780226256355_50252101.mp4 --position 2.5 -o frame_2s.png

# Use extracted frame as reference for new generation
ima2 video "continue this scene" --ref lastframe.png

Video Analysis (Recreation Prompt)

Analyze first and last video frames with Grok 4.3 image understanding to get a structured recreation prompt. This infers motion from frames; it is not full temporal video understanding.

# Analyze a generated filename
ima2 video analyze 1780226256355_50252101.mp4

# Output: structured prompt with shot type, inferred camera movement, lighting, color, motion, mood

# Use the analysis to recreate with variations
ANALYSIS=$(ima2 video analyze 1780226256355_50252101.mp4 --json | jq -r '.analysis')
ima2 video "$ANALYSIS but in anime style" --ref reference.png

Audio in Video (Prompt-Controlled)

The API does not expose a separate audio on/off or audio-track control. Treat audio as prompt-compiled: describe dialogue, music, no-music, room tone, or sound-effects-only behavior in the video prompt. Output is provider-dependent, but the prompt must be explicit when audio matters.

# Explicit sound direction
ima2 video "ocean waves crashing on rocks with seagull calls and distant thunder"

# Music direction
ima2 video "timelapse of city at night, lo-fi hip hop background music"

# Dialogue
ima2 video "person speaking to camera: Hello world, welcome to my channel"

# No music / room tone
ima2 video "quiet forest scene, no background music, only subtle wind and leaves rustling"

# Sound effects only
ima2 video "no music, only footsteps, cloth movement, rain hits, and one radio click"

For continuity clips, always define the final audio state: whether dialogue finishes before the cut, music resolves or continues, or a sound effect carries into the next clip.

Structured Video Prompt Template

Use this structure for serious video generation, Ref2V, extension prompts, and multi-shot continuity. A static visual description is not enough. Write like a director calling a shot, not filling out a form.

Opening frame: composition, depth layers, spatial staging, material/texture.
Motivated movement: what changes and why — reveal, follow, discover, tension.
Camera intent: the specific move that serves this scene (macro push-in, orbit,
  lateral slider, rack focus, locked overhead, handheld, crane).
Visual turning point: a shift in focus, scale, light, or subject state.
Dialogue: speaker (by visual appearance, not name), exact line in original
  language, timing — or "no dialogue".
Sound: music style with swell/cut/resolve behavior, or "no background music,
  room tone only", or specific SFX (footsteps, rain, machine hum, impact).
Settling final frame: stable pose, camera angle, background, lighting, held
  audio state — self-explanatory for continuation.
Negative constraints: no visible subtitles/text unless requested, preserve
  identity/style.

When creating a sequence, write both motions explicitly: "A motion" for the first clip and "B motion" for the continuation. For last-frame Ref2V, use ref 1 as identity/style and ref 2 as current state/last frame.

Shot discipline (cross-vendor official guidance):

  • One camera move + one primary action per shot is the most reliable recipe; short clips follow instructions better than long ones. Write actions as observable, timed beats: "takes four steps to the window, pauses, pulls the curtain in the final second" — not abstract descriptions.
  • Split audio into explicit channels: Dialogue (speaker label + exact short line), Ambience, SFX, Music. Declare music policy explicitly — "diegetic only", "no score", or a concrete style. A 4-5s clip fits 1-2 short dialogue exchanges at most.
  • I2V prompts describe motion, not the image. When a reference image or last frame drives the clip, the image already fixes subject, composition, color, and lighting — do not re-describe them. Write only: subject motion, scene reaction, camera motion, motion style.
  • Reuse identical anchor phrases across clips. For multi-clip continuity, repeat the same character/wardrobe/palette wording verbatim in every prompt of the series.
  • Failure recovery ladder: freeze the camera, then simplify the action, then clear the background, then re-add one element per iteration.

Example — product reveal (10s, 1.5, 1080p):

A single continuous macro shot begins inches above a matte black desk surface,
tight on the brushed aluminum edge of wireless earbuds catching a narrow softbox
reflection. The camera glides laterally as focus racks from the charging case
texture to the earbud stem, revealing the full product silhouette against soft
warm-gray negative space. A gentle ambient hum, no music. The camera settles
into a medium close-up with the product centered, soft rim light from behind,
holding steady on the final composition.

End Frame Guidance (via Ref2V)

Guide the video toward a desired final scene using reference images:

# Start frame + end frame concept
ima2 video "smooth transition from day to night" \
  --ref sunrise.png --ref nightsky.png

The planner treats reference images as subject/style/composition guidance. This is best-effort guidance, not a guaranteed final-frame constraint.

Soul Character / Face Consistency (via Ref2V)

Guide character identity across multiple videos using reference photos:

# Provide face references for consistency
ima2 video "person walking through a park, smiling" \
  --ref face_front.png --ref face_side.png --ref face_smile.png

# Same character in different scenes
ima2 video "same person now sitting at a cafe" \
  --ref face_front.png --ref face_side.png --topic "character-series"

Marketing / Product Video

Turn a product image into a dynamic showcase video:

# Step 1: Generate or provide product image
ima2 gen "clean product photo of wireless earbuds on white background" -o product.png

# Step 2: Create product video
ima2 video "sleek product reveal, rotating camera, premium studio lighting" \
  --ref product.png --duration 10 --aspect-ratio 16:9

# Step 3: Extend with lifestyle shot
PRODUCT_VID=$(ima2 video "product reveal" --ref product.png --json | jq -r '.path')
ima2 video extend "person puts on the earbuds and smiles" --video "$PRODUCT_VID" --duration 5

MCP Provider Tool Contracts

Machine tool contracts (catalog sha256:863c92848fba)

Agents: run ima2 tools list --json for the live view; this section is the bundled-snapshot projection.

ima2 (5)

toolexecutable viadescription
ima2.generate_imageagent runtimeGenerate one or more images. Supports fanout: provide one prompt per variant.
ima2.generate_videoagent runtimeGenerate a single video with Grok Imagine. If the session has a last image, it is used as the image-to-video source automatically; prompt-on
ima2.get_generation_errorsagent runtimeRead-only lookup of the session's recent generation failures (failed queue jobs and error turns). Use when the user asks why a generation fa
ima2.get_image_contextagent runtimeLoad the session image context manifest (previous images, current image, locks). Runs automatically before image generation.
ima2.web_searchagent runtimeSearch the web for factual visual references before generating. Only available when web search is enabled for the session.

mcp.higgsfield (73)

toolexecutable viadescription
mcp.higgsfield.animation_actions—Read-only catalog of the 3D rig animation library (678 actions: locomotion, gestures, dancing, combat, daily actions). Search by name or bro
mcp.higgsfield.balance—Get the user's available credits and current subscription plan. For transaction history, call transactions instead.
mcp.higgsfield.cancel_trial_auto_renewal—Cancel the auto-renewal of the Higgsfield MCP 3-day free Plus trial. Call this when the user asks to cancel the trial, cancel auto-renewal,
mcp.higgsfield.confirm_billing_purchase—INTERNAL — invoked ONLY by the plans widget on an explicit user Confirm click. Do NOT call this tool yourself; it charges the user's real sa
mcp.higgsfield.confirm_trial_cancel—INTERNAL — invoked ONLY by the cancel-trial confirmation widget on an explicit user click of 'Cancel auto-renewal'. Do NOT call this tool yo
mcp.higgsfield.create_voice—Open the Create Voice Apps UI. Call this immediately when the user asks to create a voice, call the Create Voice tool, or needs a local brow
mcp.higgsfield.create_voice_from_confirmed_audio—Backend-only creation of a cloned voice from an already confirmed audio upload. Do not call this tool until audio_media_id and name are alre
mcp.higgsfield.create_website—Start a new full-stack website. Creates the website and a git repo: a React 19 + TanStack Start app, server-rendered, in ONE Cloudflare Work
mcp.higgsfield.deploy_game—Deploy a built browser game from an uploaded zip archive and get a shareable play URL. Deploying also lists the game in the Higgsfield marke
mcp.higgsfield.deploy_website—Build and deploy the website via CI, then return its live URL. Every deploy ships the live site at the website's public URL (there is no sep
mcp.higgsfield.dubbing—Dub a video into another language: translate the spoken audio, synthesize it in the target language, and lip-sync the result back onto the v
mcp.higgsfield.explainer_video—Assemble an explainer / narrated video from its per-block clips and voice takes: stitches two or more existing video clips into one MP4, in
mcp.higgsfield.generate_3d—Generate a 3D GLB mesh. Use models_explore(type:'3d') to pick a model and see its medias[].roles and parameters. Apps UI local file: c
mcp.higgsfield.generate_audio—Generate speech/voice audio (text-to-speech). DEFAULT model: seed_audio (Seed Audio 1.0 by ByteDance) — use it unless the user explicitly as
mcp.higgsfield.generate_imagePOST /api/mcp/generateGenerate an image. Apps UI local file media: call media_upload_widget; do not ask for Claude chat attachments because remote tools cannot
mcp.higgsfield.generate_videoPOST /api/mcp/generateGenerate a video. Apps UI local file: call media_upload_widget; do not ask for Claude chat attachments; remote tools cannot read them. Web
mcp.higgsfield.get_explainer_presets—Show the explainer video style presets (CMS-managed catalog). Returns preset ids, names, and preview media. When the user picks one, resolve
mcp.higgsfield.get_game_creation_bundle_file—Read a safe text file or directory from the game-generation resource folder. Use this after get_game_creation_instructions when the instruct
mcp.higgsfield.get_game_creation_instructions—REQUIRED before creating or editing any browser game. Reads the game-generation SKILL.md resource and returns the current list of files avai
mcp.higgsfield.get_website_creation_bundle_file—Read a safe text file or directory from the website-builder-flow resource folder. Use this after get_website_creation_instructions when the
mcp.higgsfield.get_website_creation_instructions—REQUIRED before creating or editing any website with the website tools (create_website / website_repo_access / deploy_website / website_db /
mcp.higgsfield.get_workflow_bundle_file—Read a safe text file or directory from a workflow's resource folder. Use this after get_workflow_instructions when the SKILL.md requires a
mcp.higgsfield.get_workflow_instructions—Discover and load multi-step content-generation workflows (each a bundled SKILL.md that orchestrates the generate_* tools). Call with NO arg
mcp.higgsfield.job_display—Show a single generation result in the UI widget by job ID. Pass exactly one job ID — to display multiple generations, call this tool once p
mcp.higgsfield.job_status—Check the status and results of an async job. Returns instantly. For non-terminal jobs the response includes poll_after_seconds — wait that
mcp.higgsfield.list_voices—List available voices for speech and voice tools. Returns built-in preset voices plus the user's own custom voices. Each voice has a voice_i
mcp.higgsfield.list_websites—List the websites you own — each with its id, name, slug, and live URL. Use this to find the id of a website you created earlier so you can
mcp.higgsfield.list_workspaces—List every workspace the user can access (their private workspace plus any shared/team workspaces). The is_selected field marks which work
mcp.higgsfield.media_confirm—Confirm file uploads after using the upload_url method. Call this after the curl uploads succeed. Supports confirming multiple uploads at on
mcp.higgsfield.media_import_url—Import an HTTPS image, video, or audio URL into Higgsfield storage and return a confirmed media_id. Use this before generate_image/generate_
mcp.higgsfield.media_upload—Upload media for use in generation, or general files (documents, archives, code) for sharing. Returns presigned URLs for clients that can up
mcp.higgsfield.media_upload_widget—Open the Higgsfield upload widget for a user-provided local image, video, or audio file. Call this immediately when the user says they have
mcp.higgsfield.models_explore—Find generation models. Use recommend with goal + input context; use get for model constraints.
mcp.higgsfield.motion_control—Animate an existing character image with the motion and camera movement from a reference video using Kling 3.0 Motion Control. Use this when
mcp.higgsfield.outpaint_image—Expand or uncrop an existing image by outpainting beyond the original frame while preserving the source content. Use this when the user asks
mcp.higgsfield.participate_in_contest—Enter the website in the current Higgsfield app contest, together with the social-media links promoting it. A website not yet PUBLISHED to t
mcp.higgsfield.personal_clipper_create—Turn YouTube videos into ready-to-share clips. This is a long-running job and can take up to 30+ minutes. Before starting, ask the user how
mcp.higgsfield.personal_clipper_jobs—Show recent clipping jobs.
mcp.higgsfield.personal_clipper_status—Check clip creation progress.
mcp.higgsfield.presets_show—Show available Higgsfield presets for image-to-video generation. Returns preset ids, names, previews, and descriptions.
mcp.higgsfield.publish_game—Publish a deployed game to the Higgsfield marketplace. This does not deploy anything: the game must already be live — use deploy_game first,
mcp.higgsfield.publish_website—Publish the website: lists the website's CURRENT LIVE production deploy on the Higgsfield community feed ('show in feed'), where other users
mcp.higgsfield.reframe—Expand or reframe an existing video to a new aspect ratio while preserving the source content. Use this when the user asks to make a video v
mcp.higgsfield.remove_background—Remove or cut out the background from an existing image or video. Use this when the user asks for background removal, a transparent backgrou
mcp.higgsfield.rename_website—Rename the website's SUBDOMAIN (the slug in its public URL). The site is re-deployed under the new subdomain and the OLD subdomain STOPS WOR
mcp.higgsfield.resolve_explainer_preset—Resolve an explainer video style preset (from get_explainer_presets) into a style reference media_id: the backend imports the preset's style
mcp.higgsfield.reveal_generation—Confirm the user has rights to the content of an ip_detected generation and flip its status to completed. Backend accepts only seedance-
mcp.higgsfield.select_workspace—Set or clear the active workspace — the one all subsequent MCP operations bill against and read from (generations, balance, transactions, up
mcp.higgsfield.shorts_studio_create—Start a Shorts Studio short: restyle one uploaded source video (4s–120s) into a set of AI-generated short-form clips using a style preset. P
mcp.higgsfield.shorts_studio_create_preset—Create a user-owned Shorts Studio style preset from reference media (videos + images). This just stores a STYLE — no generation, no credits.
mcp.higgsfield.shorts_studio_list_presets—Browse Shorts Studio style presets — the visual STYLE a short is restyled toward. Use this when the user wants to make a short and needs to
mcp.higgsfield.shorts_studio_list_sessions—List the caller's past Shorts Studio sessions (newest first) to find a session_id to poll with shorts_studio_status.
mcp.higgsfield.shorts_studio_status—Poll one Shorts Studio session. Returns {id, status, job_ids}. status='completed' means every clip job is terminal (not necessarily successf
mcp.higgsfield.show_characters—Soul Characters widget — reusable trained identity models. Actions: list (browse), train (needs name + 5-20 ref images, ~10 min, non-b
mcp.higgsfield.show_generations—Browse past completed non-Marketing Studio generations and render them directly in the widget. Returns generations with {id, type, status, m
mcp.higgsfield.show_marketing_studio—When replying to the user, do not say ms_image — refer to it as "DTC Ads".
mcp.higgsfield.show_marketing_studio_generations—Browse past completed Marketing Studio generations only. Returns Marketing Studio video and ad/image generations with {id, type, status, mod
mcp.higgsfield.show_medias—List your uploaded media files by type. Returns media IDs, URLs, and creation timestamps. Pass media IDs as value in the medias array of gen
mcp.higgsfield.show_plans_and_credits—Open the single combined pricing widget for everything billing-related. The widget has two tabs the user can switch between: *Upgrade Plan
mcp.higgsfield.show_reference_elements—Elements widget — reusable characters / environments / props per workspace. Actions:
mcp.higgsfield.sync_agents—Sync Agents — imports the user's user-authored Skills and a personality dump from the current host LLM into Higgsfield. One trigger, one upl
mcp.higgsfield.transactions—List the user's credit transactions (spend/refund/grant/deduct), newest first. Paginated: if next_cursor is not null, pass it as cursor to g
mcp.higgsfield.upscale_imagePOST /api/mcp/media-actionUpscale and enhance an existing image. Use this when the user asks to upscale, enhance, or increase the resolution of an image to 2K/4K. Thi
mcp.higgsfield.upscale_videoPOST /api/mcp/media-actionUpscale and enhance an existing video. Use this when the user asks to upscale, enhance, sharpen, denoise, restore, or convert a video to hig
mcp.higgsfield.video_analysis_create—Start a scene-by-scene analysis of a video. Provide EXACTLY ONE of: (a) video_input_id — UUID of a video the user has uploaded via media_upl
mcp.higgsfield.video_analysis_jobs—List the user's video analyses in the current workspace, newest first. Paginate by passing the previous response's cursor.
mcp.higgsfield.video_analysis_status—Get the status and result of a video analysis. Poll this after video_analysis_create until status='completed' (scenes populated) or 'failed'
mcp.higgsfield.virality_predictor—Virality Predictor predicts a video's virality potential, engagement, attention, audience response, retention risk, hook strength, and creat
mcp.higgsfield.voice_change—Replace the spoken voice in a video with a different voice while keeping the original timing and visuals, then re-merge the new audio onto t
mcp.higgsfield.website_db—Inspect the website's database (D1 / SQLite), READ-ONLY. The website has ONE database — the live site's real data. Pick an operation: 'table
mcp.higgsfield.website_repo_access—Get direct git access to a website's repo to edit it — THE way to get the website's code. Returns the repo URL, branch, slug, and a scoped t
mcp.higgsfield.website_secrets—Manage a website's SECRETS (environment variables: API keys, tokens). Set them HERE instead of hardcoding them in source. One tool, three op
mcp.higgsfield.website_status—Get the website's deploy status — the live URL and the status of the last deploy. Use to check a deploy that returned 'pending', or to fetch

mcp.runway (14)

toolexecutable viadescription
mcp.runway.complete_upload—Finalize a file upload and get the asset URL.
mcp.runway.edit_videoPOST /api/mcp/media-actionEdit an existing video with Runway's Aleph 2.0 in-context video editor: it changes ONLY what you ask for and preserves everything else — sub
mcp.runway.feedback—Call this when you (the AI agent) get stuck using Runway tools.
mcp.runway.generate_imagePOST /api/mcp/generateGenerate OR edit an image using a Runway-hosted image model. This is the only image tool — there is no separate "edit_image" tool. Pass the
mcp.runway.generate_multishot_video—Generate a multi-shot video — 3 to 5 connected scenes from a single story or per-shot prompts. Powered by Kling 3.0 (standard at 720p, pro a
mcp.runway.generate_product_marketing_video—Generate a polished creative product ad video from a product URL or product image plus a campaign idea.
mcp.runway.generate_videoPOST /api/mcp/generateGenerate OR edit a video using a Runway-hosted video model. Pass the source video as referenceVideo to edit/restyle it.
mcp.runway.get_task—Gets details for a Runway task by ID — used to check status and retrieve the result of a generation/edit task once it completes. Generation
mcp.runway.init_upload—Initialize a file upload to Runway. Returns temporary upload URLs for direct upload.
mcp.runway.list_recent—Lists recent uploaded and generated assets for the authenticated workspace. Returns asset IDs, media types, task IDs when available, and reu
mcp.runway.list_workspaces—Lists every Runway workspace the authenticated user belongs to, with role and a compact disabled-model summary per workspace. The MCP connec
mcp.runway.upscale_imagePOST /api/mcp/media-actionUpscale an existing image to a higher resolution (2x, 4x, 8x, or 16x) using Runway's AI image upscaler. Use this to sharpen, denoise, and in
mcp.runway.upscale_videoPOST /api/mcp/media-actionUpscale an existing video to a higher resolution (up to 4K) using Runway's AI video upscaler. Use this to sharpen, clean up, and increase th
mcp.runway.whoami—Returns the authenticated Runway user profile, the workspace this MCP connection is pinned to (chosen at sign-in), and the list of image/vid

发现问题?提交给管理员复核

评分:

评论 (0)

暂无评论,成为第一个评论者吧!