复制安装命令
用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。
复制前请先查看来源、License 和安全提示。
🌐 Live site: lidge-jun.github.io/ima2-gen · 한국어
用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。
复制前请先查看来源、License 和安全提示。
来源文件:README.md
🌐 Live site: lidge-jun.github.io/ima2-gen · 한국어
📖 Developer docs: Documentation site · 한국어
ima2-gen is a local image generation studio for people who want the ChatGPT/Codex image workflow in a small desktop-like web app.
Install globally, sign in with ChatGPT OAuth or Grok OAuth, and start generating images and videos. Iterate with history, references, node branches, multimode batches, Canvas Mode cleanup, and Grok Video generation. Default OAuth paths need no API key; optional API-key providers (api, grok-api, gemini-api, agy) are also supported.

npm install -g ima2-gen
ima2 setup
ima2 serve
Then open http://localhost:3333.
docker build -t ima2-gen .
docker run -d -p 3333:3333 -e IMA2_LAN_TOKEN=change-me -v ima2-data:/data ima2-gen
See docs/DOCKER.md for compose usage, required environment, and limitations.
To generate from the CLI, inspect the live lane catalog and choose explicit image/video defaults once:
ima2 models
ima2 defaults set image oauth/gpt-5.6-luna
ima2 defaults set video grok/grok-imagine-video-1.5
ima2 gen "a clean product photo of a red guitar pedal"
ima2 video "a cat playing piano" --duration 5 --resolution 720p
ima2 video "animate this scene" --ref photo.png --duration 10
ima2 gen and generate-mode ima2 video fail closed with NO_DEFAULT_MODEL until a CLI target is configured, unless that call passes --model <lane>/<model> or an explicit --provider <lane>. This prevents an upgrade from silently switching providers or billing lanes.
If 3333 is already occupied, ima2-gen binds the next available port and writes the actual URL to ~/.ima2/server.json. Use ima2 open or the URL printed in the terminal instead of assuming the port.
Using npx? See docs/NPX_QUICKSTART.md for the
npx ima2-gen serveworkflow.
Don't have Node.js or npm? Use the platform install script — it detects your environment, installs Node LTS if needed, then installs ima2-gen.
macOS:
curl -fsSL https://lidge-jun.github.io/ima2-gen/install-mac.sh | bash
Windows (PowerShell):
irm https://lidge-jun.github.io/ima2-gen/install-windows.ps1 | iex
Linux / WSL:
curl -fsSL https://lidge-jun.github.io/ima2-gen/install-linux.sh | bash
Each script checks for nvm/fnm/brew/winget, installs Node LTS through the best available method, and handles stale process cleanup automatically.
ima2 setup offers four authentication choices:
Video generation requires Grok OAuth (option 2 or 3). Run ima2 grok login separately if you already have GPT OAuth configured and want to add video support; it defaults to the manual-paste flow.
Stop the running server with Ctrl+C, then:
npm install -g ima2-gen@latest
Ctrl+C now performs a clean shutdown — closing the database, stopping child processes, and releasing file locks. On older versions (< 1.1.22) or if you see EBUSY on Windows, use the install script which handles stale process cleanup automatically.
ima2-gen ships three packaged skills for AI coding agents. These are Markdown instruction files that agents load to get structured workflows for image/video generation, frontend asset production, and design direction discovery.
| Skill | Command | What It Covers |
|---|---|---|
| Core | ima2 skill | CLI reference, prompting protocol, provider routing, Korean text, video workflows |
| Frontend | ima2 skill front | Asset pipeline (parallel gen, variant selection, provider routing), motion/video for web, responsive, a11y, anti-slop, 30+ reference files |
| UI/UX Design | ima2 skill uiux | Image-first design direction discovery, UX states, design-isms, product personalities, DESIGN.md workflow, 18 reference files |
ima2 skill ls # list available skills
ima2 skill front # print the frontend skill
ima2 skill uiux # print the design skill
ima2 skill front path # print file path (for agents)
ima2 skill front --json # JSON wrapper (for agents)
ima2 skill front refs # list reference modules (35 files)
ima2 skill front ref motion # load one reference module
ima2 skill install --dir <path> # install skills to agent's skill dir
ima2 skill install --tmp # install to temp dir (fallback)
The Frontend and UI/UX skills are production-grade design engineering guides
adapted for the ima2 workflow. They cover typography, color systems, layout
discipline, Korean UX patterns, motion choreography, and visual verification,
with every asset generation step mapped to ima2 gen, ima2 video, and
ima2 multimode commands.
The web UI uses a single GET /api/events Server-Sent Events connection for all generation progress. Multimode, node, and video requests are submitted as async POST (202 { requestId }) and progress events are multiplexed through a shared event bus. This eliminates the browser 6-connection limit that previously caused gallery hangs during concurrent generation. CLI clients that do not send async: true still receive per-request SSE streams for backward compatibility.
Image generation can run through the local Codex/ChatGPT OAuth path, a configured OpenAI API key, the bundled Grok provider, or the Gemini provider via Antigravity CLI.
provider: "oauth" uses the local Codex OAuth proxy.provider: "api" calls the OpenAI Responses API with the hosted image_generation tool.provider: "grok" starts bundled progrok on 127.0.0.1:18645, runs mandatory xAI Web Search plus a planner pass (default: grok-4.5, configurable in settings or via --planner-model), then calls xAI Images API through the local proxy. grok-4.3 remains available as an explicit compatibility override.provider: "grok-api" calls the xAI Images API directly with XAI_API_KEY (no bundled progrok OAuth proxy).provider: "agy" spawns the Antigravity CLI (agy -p) to generate images via Google Gemini's default_api:generate_image tool (model: nano-banana-2). Output is fixed at 1024×1024 JPEG, max 3 reference images. No web search, quality, or size controls.provider: "gemini-api" calls the Google Generative Language API directly. Supports two models: nano-banana-2 (Gemini 3.1 Flash Image) and nano-banana-pro (Gemini 3 Pro Image). Auth is via GEMINI_API_KEY env var, web UI key management, or a Vertex AI service account JSON (VERTEX_SERVICE_ACCOUNT_JSON). When both an API key and Vertex credentials are configured, Vertex takes priority. Supports variable aspect ratios (1:1 through 21:9) and four resolution tiers (512px, 1K, 2K, 4K); these controls are only honored on the direct API path — the Vertex AI endpoint ignores aspect/size because it does not accept the response_format field. Per-model cost differs: nano-banana-2 (Flash): 512=$0.001, 1K=$0.003, 2K=$0.004, 4K=$0.006; nano-banana-pro: 1K=$0.007, 2K=$0.007, 4K=$0.013. No web search or mask controls.If no provider is specified, the app keeps the current GPT OAuth/default behavior. GPT OAuth and API-key generation default to gpt-5.6-luna; the API-key path also defaults to low reasoning and 1024x1024 unless the request passes validated options. Grok image generation defaults to grok-imagine-image-quality.
Grok image generation exposes a model picker (grok-imagine-image / grok-imagine-image-quality) and a size picker (aspect ratio + 1k/2k resolution). The Settings page prefers the Grok Build weekly credits percentage and reset time from GET /v1/billing?format=credits; if that source is unavailable, it falls back to the legacy monthly billing window and $used/$limit. A Switch Account button starts a device-code OAuth flow (POST /api/auth/switch) for re-authenticating without leaving the app.
Grok video generation defaults to canonical grok-imagine-video-1.5; grok-imagine-video remains available for base-model-only Ref2V, V2V edit, and extension paths, and the legacy grok-imagine-video-1.5-preview string is accepted as an alias. Three modes are auto-detected from reference count: text-to-video (0 refs), image-to-video (1 ref), and reference-to-video (2-7 refs, max 10s duration). 1080p is available for grok-imagine-video-1.5 prompt-only text-to-video and single image/frame image-to-video; prompt-only 1.5 uses the internal white-canvas I2V shim before the upstream request. Video controls include duration (1-15s), resolution (480p, 720p, 1080p when supported), and aspect ratio (1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, auto).

The app defaults to gpt-5.6-luna for image generation and Prompt Builder planning. Older supported models remain explicit compatibility choices.
gpt-5.6-luna — current image and Prompt Builder default.gpt-5.6-terra / gpt-5.6-sol — current GPT-5.6 alternatives when your account exposes them.gpt-5.5, gpt-5.4, gpt-5.4-mini — supported compatibility choices.The app also exposes quality (low, medium, high) and moderation (auto, low) controls.
Use Classic when you want one strong result quickly.
For a control-by-control guide to Prompt Studio, multimode recipes, Direct mode, reasoning effort, and gallery favorite behavior, see the Prompt Studio manual.

Use Node mode when you want to explore branches.

Each node keeps its own prompt and result. Root nodes can attach local references; child nodes use the parent image as their source. Completed jobs are matched back to nodes by request ID, so reloads and graph version conflicts can recover finished results.
Use Canvas Mode when a generated image is close but needs targeted cleanup before the next prompt.

The prompt library can now be filled from local files, GitHub folders, curated sources, and GPT-image hint packs. Imported prompts are indexed locally so search and ranking work without re-importing the same source every session.

Card News is still dev-only and experimental. It is hidden in the default published runtime unless explicitly enabled for development, and it should not be treated as a stable public feature yet.
The settings workspace keeps account, model, appearance, and language controls away from the generation sidebar.

| Command | Description |
|---|---|
ima2 serve [--dev] | Start the local web server; --dev enables verbose server diagnostics |
ima2 setup | Reconfigure saved auth |
ima2 status | Show config and OAuth status |
ima2 doctor | Diagnose Node, package, config, and auth |
ima2 doctor image-probe [--json] | Run sanitized image probes for no-image diagnostics |
ima2 open | Open the web UI |
ima2 reset | Remove saved config |
These require a running ima2 serve. The CLI covers every server route. The most common ones are below — the full CLI reference lists everything (generation, history, sessions, prompt library, annotations, Card News, observability, config).
| Command | Description |
|---|---|
ima2 models [--kind image|video] [--lane <lane>] [--json] | List live lanes, status, model IDs, and capabilities |
ima2 defaults set image|video <lane>/<model> | Persist the fail-closed CLI target for image or video generation |
ima2 defaults reset image|video | Remove a persisted CLI generation target |
ima2 gen <prompt> [--model <lane>/<model>] | Generate from the CLI; requires an explicit target or saved image default |
ima2 edit <file> --prompt <text> | Edit an existing image |
ima2 multimode <prompt> | Multi-image SSE generation |
ima2 video <prompt> [--model <lane>/<model>] | Generate video through a Grok or MCP lane; requires an explicit target or saved video default |
ima2 ls [--session <id>] [--favorites] | List recent history |
ima2 show <name> [--metadata] | Reveal a generated asset |
ima2 prompt ls -q <search> | Search the prompt library |
ima2 inflight ls [--terminal] | List active and recent jobs (alias of ps) |
ima2 config set <key> <value> | Write to ~/.ima2/config.json |
ima2 ping | Health-check the running server |
The server advertises its actual port at ~/.ima2/server.json. If 3333 is busy, the backend falls back to 3334+ and CLI commands follow the advertised URL. Override discovery with --server <url> or IMA2_SERVER=http://localhost:3333.
ima2 models --kind image
ima2 gen "poster" --model oauth/gpt-5.6-luna --reasoning-effort high
ima2 edit input.png --prompt "make it rainy" --web-search
ima2 multimode "two cats playing" -n 2
ima2 video "a cat playing piano" --model grok/grok-imagine-video-1.5 --duration 5 --resolution 720p
ima2 video "animate this" --model grok/grok-imagine-video-1.5 --ref photo.png --aspect-ratio 16:9
ima2 inflight ls --terminal
ima2 config set imageModels.reasoningEffort high
Full reference: docs/CLI.md.
Config priority:
environment variables > ~/.ima2/config.json > built-in defaults
| Variable | Default | Description |
|---|---|---|
IMA2_PORT / PORT | 3333 | Web server port |
IMA2_HOST | 127.0.0.1 | Web server bind host |
IMA2_OAUTH_PROXY_PORT / OAUTH_PORT | 10531 | OAuth proxy port |
IMA2_SERVER | — | CLI target override |
IMA2_CONFIG_DIR | ~/.ima2 | Config and SQLite location |
IMA2_ADVERTISE_FILE | ~/.ima2/server.json | Runtime discovery file |
IMA2_GENERATED_DIR | ~/.ima2/generated | Generated image directory |
IMA2_IMAGE_MODEL_DEFAULT | gpt-5.6-luna | Server fallback image model |
IMA2_REASONING_EFFORT | medium | Default reasoning effort for the default (GPT OAuth) path; one of none, low, medium, high, xhigh |
IMA2_NO_OAUTH_PROXY | — | Set 1 to disable the auto-started OAuth proxy |
IMA2_LOG_LEVEL | info | Normal serve defaults to info; dev mode defaults to debug; supports debug, info, warn, error, or silent |
IMA2_INFLIGHT_TERMINAL_TTL_MS | 300000 | Recent terminal job retention for debug views |
OPENAI_API_KEY | — | API key for the provider: "api" Responses API image path and auxiliary API-key features |
XAI_API_KEY | — | API key for provider: "grok-api" direct xAI Images API path |
IMA2_API_IMAGE_MODEL_DEFAULT | gpt-5.6-luna | Default image model for provider: "api" |
IMA2_API_REASONING_EFFORT | low | Default reasoning effort for provider: "api" |
IMA2_API_IMAGE_SIZE | 1024x1024 | Default size for provider: "api" |
IMA2_API_ALLOW_WEB_SEARCH | true | Toggle web search for provider: "api" |
IMA2_GROK_PROXY_HOST | 127.0.0.1 | Host for the bundled progrok proxy |
IMA2_GROK_PROXY_PORT | 18645 | Port for the bundled progrok proxy |
IMA2_NO_GROK_PROXY | — | Set 1 to disable automatic progrok startup |
IMA2_GROK_PLANNER_MODEL | grok-4.5 | Grok search/planner model (also configurable via settings UI or --planner-model CLI flag) |
IMA2_GROK_PLANNER_TIMEOUT_MS | 60000 | Timeout for Grok search and planner calls |
IMA2_GROK_IMAGE_MODEL_DEFAULT | grok-imagine-image-quality | Default final Grok image model |
IMA2_GROK_VIDEO_MODEL_DEFAULT | grok-imagine-video-1.5 | Default Grok video model |
IMA2_GROK_GENERATION_TIMEOUT_MS | 120000 | Timeout for the final Grok Images API call |
IMA2_OAUTH_MASKED_EDIT_ENABLED | false | Opt-in feature flag for masked-edit requests on the OAuth path (#31, groundwork only) |
GEMINI_API_KEY | — | API key for provider: "gemini-api" direct Generative Language API path |
VERTEX_SERVICE_ACCOUNT_JSON | — | Google service account JSON for Vertex AI auth with provider: "gemini-api"; takes priority over GEMINI_API_KEY when both are set |
IMA2_AGY_BIN | agy on PATH | Explicit path to the Antigravity CLI binary for provider: "agy" |
IMA2_MAX_PARALLEL | 24 | Server-wide parallel generation cap |
ima2 serve keeps terminal output intentionally quiet: startup URLs, warnings, and errors stay visible, while request/node/OAuth structured logs are hidden by default.
Use ima2 serve --dev, npm run dev, or IMA2_LOG_LEVEL=debug ima2 serve when you need request IDs, node generation phases, OAuth stream diagnostics, or inflight state transitions. Explicit IMA2_LOG_LEVEL and ~/.ima2/config.json values still override the built-in defaults.
The endpoint list moved to docs/API.md so this README can stay focused on first-run use.
Useful references:
ima2 ping says the server is unreachable
Start ima2 serve, then check ~/.ima2/server.json. You can also run ima2 ping --server http://localhost:3333.
GPT OAuth login does not work
Re-run ima2 setup (option 1), confirm ima2 status, then restart ima2 serve.
fetch failed repeats on a proxy/VPN network
Check that the local OAuth proxy is reachable. On networks that require a proxy, enable your proxy client's TUN/TURN-style mode, then retry openai-oauth --port 10531. If it still fails, set HTTP_PROXY and HTTPS_PROXY in the same terminal that runs ima2 serve or openai-oauth. On Windows, also check for auto-start network interception tools, including DNS/fragmentation bypass tools such as SecretDNS, because they can break OAuth or streaming image responses even when the browser appears connected.
Images fail with API_KEY_REQUIRED
Set OPENAI_API_KEY or configure an API key before using provider: "api". The default GPT OAuth path still works without an API key.
Image generation returns EMPTY_RESPONSE or no image data
Run ima2 doctor image-probe --json > ima2-image-probe.json and attach the safe JSON when opening an issue. For GPT OAuth cases, also capture ima2 gen "고양이" --model oauth/gpt-5.6-luna --no-web-search --json and ima2 gen "고양이" --model oauth/gpt-5.6-luna --json while ima2 serve is running. Do not share ChatGPT cookies, OAuth token files, API keys, raw upstream responses, prompt history, or generated base64. See the FAQ support bundle.
A large reference image fails The app compresses large JPEG/PNG references before upload. If a file still fails, convert it to JPEG or PNG at a lower resolution and try again. HEIC/HEIF files are not supported by the browser path.
Old gallery images are missing after updating
Recent versions moved generated images from the installed package folder to ~/.ima2/generated. Run ima2 doctor and see Recover old images.
gpt-5.5 fails but other models work
Update Codex CLI first, then retry. If it still fails, your account or backend route may not expose the same image capability or quota for gpt-5.5 yet; use gpt-5.4 as the stable fallback.
The app opened on a different port
If the requested server port is busy, ima2-gen falls back to the next available port and records it in ~/.ima2/server.json. If the port is unexpectedly 3457, your shell may also have inherited PORT=3457 from another local tool. Run unset PORT or start with IMA2_PORT=3333 ima2 serve.
Port 10531 is already used on Windows
Some Windows security tools, including AnySign4PC.exe, may occupy the default OAuth proxy port. Current builds track the actual fallback OAuth port. If you still need a manual override, start with IMA2_OAUTH_PROXY_PORT=11531 ima2 serve and check ima2 doctor.
For more beginner-friendly answers, see the FAQ.
git clone https://github.com/lidge-jun/ima2-gen.git
cd ima2-gen
npm install
npm run dev
npm run typecheck
npm test
npm run build
npm run dev builds the UI and starts the TypeScript server entry with --watch and verbose server diagnostics. npm run typecheck, npm run build:server, and npm run build:cli verify the TypeScript migration and package emit path. Node mode and Canvas Mode are part of the packaged UI by default.
MIT
name: ima2
description: "Use the ima2-gen CLI/server to generate, edit, inspect, and manage local AI image generation jobs."Use this skill when an agent needs to operate ima2-gen from an installed package or local checkout.
Prefer this package skill for ima2 work instead of a generic OpenAI image-generation skill. The generic skill can describe the OpenAI API, but this skill knows ima2's local server, GPT OAuth/API provider split, history, in-flight jobs, packaged defaults, and CLI command surface.
Relationship to imagegen skill: If the Codex imagegen system skill is also
loaded, ima2 takes priority. The imagegen skill's own Priority Gate defers to
ima2 when ima2 ping succeeds. Do not use both in the same generation task.
Start by discovering the local package and running server state:
ima2 skill
ima2 skill --json
ima2 skill ls # list all skills (core, front, uiux)
ima2 skill install --dir <path> # install skills to agent's skill directory
ima2 skill install --tmp # install to temp dir (ephemeral fallback)
ima2 skill front refs # list frontend reference modules
ima2 skill front ref motion # load one reference module
ima2 capabilities --json
ima2 models --json
ima2 defaults --json
ima2 ping
If the server is not running:
ima2 serve
ima2 open
Use ima2 doctor when setup, GPT OAuth, storage, or package integrity is unclear.
List ready image lanes, choose a persistent CLI target, then generate:
ima2 models --kind image
ima2 defaults set image oauth/gpt-5.6-luna
ima2 gen "a clean product photo of a red guitar pedal"
Bare ima2 gen fails closed when no CLI image target is configured. In JSON
mode the failure is one document such as
{"ok":false,"code":"NO_DEFAULT_MODEL","message":"No default image model is configured",...}
and exits 2. Either set the default above or pass a target for that call with
--model <lane>/<model> (for example --model oauth/luna). Never rely on an
implicit provider; --provider auto was removed.
Use high quality when output fidelity matters:
ima2 gen "a print-ready poster" --model oauth/luna --quality high
Use direct mode when the prompt should be passed with minimal rewriting:
ima2 gen "exact prompt text" --model oauth/luna --mode direct
--mode explained:
auto (default): the server may augment, restructure, or enrich the prompt
before sending it to the image model. Good for casual or short prompts.direct: the prompt is passed as-is with minimal server-side rewriting. Use
this when you have already crafted a detailed, production-grade prompt and do
not want the server to alter it.Use request-level overrides only for that one call:
ima2 gen "cinematic mountain" --model oauth/gpt-5.5 --reasoning-effort high
Use Grok when the request should run through bundled progrok, mandatory xAI Web
Search, planner pass (default: grok-4.3), and xAI Images API:
ima2 grok login
ima2 grok status
ima2 gen "cinematic neon city" --model grok/grok-imagine-image-quality
ima2 grok login defaults to the manual-paste flow.
Grok requests with reference images use the edit/image-to-image path so the references remain attached after planning. Keep Grok references to three total input images.
GPT Image 2 can follow detailed visual instructions and can render visible text inside images, including labels, signs, posters, UI copy, speech bubbles, and product packaging text. Do not avoid text just because older image models were weak at it.
When visible text matters, write the exact words in the target language and script:
A Korean poster with the exact headline "오늘 공연" and subtext "입장 무료".A Korean poster with some Korean text.Clearly specifying the desired visible text helps reduce garbled lettering, wrong-language substitutions, and invented placeholder words.
For dense or important text, specify:
OpenAI's prompting guide additionally recommends: put literal text in quotes or ALL CAPS, state typography (font style, size, color, placement) as explicit constraints, and for exact copy demand it verbatim. The strongest official pattern is a dedicated text block:
Poster headline (EXACT, verbatim, no extra characters):
"Fresh and clean"
Typography: bold sans-serif, high contrast, centered, clean kerning.
Ensure the text appears once and is perfectly legible.
For tricky words such as brand names or uncommon spellings, spell them out
letter-by-letter to improve character accuracy. Use medium or high quality
whenever the image contains small text, dense panels, or multiple fonts. When
localizing an existing image, translate the visible text verbatim, add no new
words, and preserve everything else — layout, imagery, hierarchy — without
reflowing the design.
GPT Image 2 can generate both stylized and realistic outputs. State the style directly, for example:
manga panelwebtoon stylechildren's book illustrationphotorealistic product photorealistic poster mockupcinematic real-world sceneText rendering is improved, but it is still not a typesetting engine. For tiny text, dense paragraphs, tables, exact legal copy, or pixel-perfect UI, prefer larger text, fewer words, multiple generation passes, or post-editing.
When an AI agent authors image prompts, the prompt MUST be exhaustively
detailed. Vague one-liners produce generic, unusable output. Write every
prompt as if you are briefing a senior photographer or illustrator who cannot
ask follow-up questions. When using --mode auto, the server augments short
prompts, but a detailed prompt still produces far better results than relying
on auto-augmentation alone. For production assets, prefer --mode direct with
a fully-specified prompt.
Detailed is not enough — the prompt must be structured. OpenAI's official
gpt-image prompting guide recommends composing prompts in a consistent field
order — scene/background → subject → key details → constraints — and using
labeled segments or line breaks instead of one long paragraph for complex
requests. OpenAI's own showcase prompts use labeled blocks such as Context,
Characters, and Composition. Apply these rules to every agent-authored
prompt:
Every agent-authored prompt MUST include all applicable fields. Omit a field only when it genuinely does not apply (e.g. no text in the image).
Use case: <slug: photorealistic-natural | product-mockup | ui-mockup | infographic-diagram | scientific-educational | ads-marketing | productivity-visual | logo-brand | illustration-story | stylized-concept | historical-scene>
Asset type: <where the asset will be used: hero, OG image, card, avatar, icon, texture, game sprite, etc.>
Primary request: <one clear sentence describing the desired image>
Scene/backdrop: <specific environment — not "nice background">
Subject: <main subject with identifying details: material, color, shape, posture, expression>
Style/medium: <exact style: editorial photography, flat illustration, 3D render, watercolor, etc.>
Composition/framing: <camera angle, crop, subject placement, negative space intent>
Lighting/mood: <light source, direction, color temperature, mood, time of day>
Color palette: <specific hex codes or named palette — not "modern colors">
Materials/textures: <surface details: matte plastic, brushed steel, linen, weathered wood, etc.>
Text (verbatim): "<exact text to render>" with font style, size, color, placement
Constraints: <must-keep invariants>
Avoid: <explicit negative constraints>
| Bad (vague) | Good (specific) |
|---|---|
| "a nice hero image" | "wide landscape product shot of a matte black thermos on a wet granite countertop, soft morning window light from the left, shallow depth of field, warm neutral tones, negative space on the right for headline overlay" |
| "modern background" | "soft radial gradient from #f8f9fa center to #e9ecef edges, subtle paper grain texture at 3% opacity, no objects, no patterns" |
| "Korean food photo" | "overhead flat-lay of budae-jjigae in a black stone pot, surrounded by small banchan dishes on a dark wood table, steam visible, warm tungsten lighting, editorial food photography style" |
| "logo on white" | "centered geometric mark: two interlocking triangles forming a hexagonal negative space, flat #1a1a2e on #ffffff, no gradients, strong silhouette at 32px, generous padding" |
| "a dashboard screenshot" | "realistic SaaS dashboard UI: top nav with avatar, left sidebar with 6 nav items, main area showing a line chart (3 series, 12 months) and a 4-column data table with 8 rows, light theme, Inter font, compact density" |
These patterns are documented failure modes; reject them when authoring or reviewing prompts:
| Anti-pattern | Why it fails | Do instead |
|---|---|---|
Keyword soup (beautiful, stunning, 8k, trending) | Comma-separated tag piles are a documented anti-pattern for natural-language image models | Structured narrative sentences: subject + attributes + relations |
Unmotivated quality tokens (masterpiece, 8K, ultra-detailed) | OpenAI's guide: lens, framing, and lighting language is more reliable for realism than generic quality tokens | Name the look: shallow depth of field, soft window light from the left, editorial photography |
Trusting precision specs (85mm f/1.2, 5600K) | Official guidance: detailed camera specs may be interpreted loosely — they are look cues, not optical simulation | Prefer perceptual terms: medium close-up, eye level, warm tungsten mood; keep mm/Kelvin only as style hints |
Contradictory constraints (minimalist + 12 required objects) | Conflicting demands make the model silently drop some of them | Resolve conflicts before generating; one intent per field |
| Rewriting everything each iteration | Loses working invariants, causes drift | Change ONE variable per pass, restate invariants |
Negative constraints are model-specific. For GPT Image, write exclusions
as plain prose inside the prompt — No extra text, no logos, no watermark —
this is the officially recommended form; there is no separate negative-prompt
parameter. Do not copy diffusion-style negative lists (wall, frame) into
GPT Image prompts; that syntax belongs to models with a dedicated negative
field (e.g. Imagen), where instruction words like "no/don't" are in turn
discouraged.
| Asset Purpose | Quality | Size | Notes |
|---|---|---|---|
| Quick draft / iteration | low | 1024x1024 | Fastest; square |
| Final hero / product shot | high | 1536x1024 landscape, 1024x1536 portrait | Or target aspect ratio |
| OG / social card | high | 1200x640 | Nearest 16px multiple of 1200x630 |
| Mobile hero | high | 1024x1536 | Portrait |
| Print / 4K | high | 3840x2160 or 2160x3840 | Max gpt-image-2 supports |
| Texture / tile | medium | 1024x1024 | Square, seamless edges |
| Icon / avatar | medium | 512x512 or 256x256 | Small canvas |
| Game environment concept | high | 1792x1024 or 2048x1152 | Wide cinematic |
| Storyboard (for i2v) | high | 1024x1024 | 3x3 grid, square |
GPT Image 2 does not reliably produce true transparent (alpha) backgrounds. Use the solid-background-then-remove strategy for cutout assets:
Generate on a pure solid background:
#000000) for reflective/metallic/glass subjects#ffffff) for dark/matte/opaque subjectsState the exact hex and ban AI additions: "PURE SOLID BLACK background hex
#000000. No checkerboard, no transparency pattern, no gradient, no floor plane,
no shadow, no vignette." Use --mode direct.
ima2 gen "3D chrome splash on PURE SOLID BLACK background hex #000000. \
No gradient, no floor, no shadow, no vignette." \
--quality high --size 1024x1024 --mode direct -o splash.png
Remove background after generation:
mix-blend-mode: screen (black bg on light page)mix-blend-mode: multiply (white bg on dark page)ima2 edit asset.png --prompt "remove the background, keep only the subject"sharp / ImageMagick / rembgAnti-pattern: requesting "transparent background" or "PNG with alpha" in the prompt — the model often produces a fake checkerboard burned into the image.
When generating images with Korean text:
"오늘의 추천", not "some Korean text"A clean summer poster with the exact Korean headline "여름 축제".
Practitioner testing found all-Korean prompts produced garbled Hangul while
English prompts with a quoted Korean string rendered correctly (heuristic,
not a guarantee)고딕체 (Gothic/Sans-serif) or 명조체 (Myeongjo/Serif)view_image — garbled or
substituted Hangul is common and must be caught before use-n 4) and pick the cleanest renderima2 edit pass that
restates the exact string and changes only the text region; if spelling still
will not stabilize after a couple of passes, stop retryingno text in that region), then composite real type with an
actual Korean font in an editor or code. Korean text failure is a
cross-model limitation, not an ima2-specific oneFor important visual assets (hero images, key illustrations, brand materials), generate multiple candidates and select the best:
# 4 candidates from one prompt
ima2 gen "<detailed prompt>" -n 4 -d ./candidates --quality high
# Or multimode for structurally different directions
ima2 multimode "<detailed prompt>" --max-images 4 -d ./candidates
After generation, inspect every candidate with view_image before selecting.
Do not blindly use the first result.
view_image.ima2 edit pass, or switching --mode.Copy-paste starters for common frontend assets:
Hero image (landing page):
ima2 gen "Use case: product-mockup. Asset type: landing page hero. A premium wireless headphone floating at a slight angle against a soft warm-gray studio backdrop. Matte black finish with brushed aluminum accents. Soft three-point studio lighting, key light from upper-left. Shallow depth of field. Wide composition with generous negative space on the right for headline overlay. No text, no logos, no watermark." \
--quality high --size 1536x1024 --mode direct -o hero.png
OG / social share image:
ima2 gen "Use case: ads-marketing. Asset type: social share card. Clean product flat-lay of a notebook, pen, and ceramic mug on a white marble desk. Overhead shot. Soft diffused daylight. Space in the upper third for title overlay. Warm neutral palette. No text, no logos, no watermark." \
--quality high --size 1200x640 --mode direct -o og-image.png
App screenshot mockup background:
ima2 gen "Use case: stylized-concept. Asset type: hero background for device mockup. Soft abstract gradient from #f0f4f8 to #dbeafe with subtle geometric shapes at 5% opacity. Clean, modern, minimal. No objects, no patterns, no text." \
--quality medium --size 1920x1088 --mode direct -o mockup-bg.png
Avatar / profile placeholder:
ima2 gen "Use case: stylized-concept. Asset type: user avatar. Friendly stylized portrait of a young professional, neutral expression, looking slightly left. Flat illustration style with subtle shadows. Solid #e5e7eb background. Circular crop safe. No text." \
--quality medium --size 512x512 --mode direct -o avatar.png
Korean product hero:
ima2 gen "Use case: product-mockup. Asset type: Korean service landing hero. A modern smartphone at 15-degree tilt showing a clean fintech app UI. The screen displays a balance card with exact text \"잔액 1,234,500원\" in 고딕체, large centered. Soft gradient backdrop from #f8fafc to #e2e8f0. Studio lighting from upper-right. No other text, no logos, no watermark." \
--quality high --size 1536x1024 --mode direct -o korean-hero.png
Game environment concept art:
ima2 gen "Use case: stylized-concept. Asset type: game environment concept art. A vast underground cavern with bioluminescent fungi on limestone walls. A narrow stone bridge crosses a dark chasm. Volumetric blue-green light from fungi clusters. Cinematic concept art style with industrial realism. Wide-angle, low camera, deep perspective. Mist rising from below. No characters, no text, no watermark." \
--quality high --size 1792x1024 --mode direct -o cave-env.png
Reference generation:
ima2 gen "turn this into a clean product render" --ref input.png --quality high
Multimode reference workflow:
ima2 multimode "create four coherent variations" --ref input.png --max-images 4
Node-mode reference workflow:
ima2 node generate "continue this concept" --ref input.png
Image edit workflow:
ima2 edit input.png --prompt "make the object blue while preserving composition"
Do not use positional edit prompts. ima2 edit requires --prompt.
OpenAI's official edit pattern is "change only X" + "keep everything else the same" — an edit prompt does not need to re-describe the whole final
image, but it must make the delta and the invariants explicit. Author every
edit prompt as a brief:
Desired result: <one sentence describing the edited image's final state>
Change only: <the specific modification>
Preserve exactly: <named lock list: facial structure, pose, product
silhouette, logo geometry, text spelling, framing, perspective, palette,
lighting, shadows>
Do not add or remove: <protected elements>
"Keep everything else the same" alone is weak — name the fragile properties in the lock list, and repeat the same lock list on every iterative edit pass to prevent drift.
Annotated inputs. If the edit source or a reference image carries drawn markup (arrows, boxes, circled regions, sticky notes), the model tends to treat the markup as image content and reproduce it. Prefer sending the clean image plus text instructions derived from the markup. When the annotated image must be sent, state before and after the edit list that the markup is temporary editing instructions only — interpret it, apply the edits, then remove every trace of it from the output.
Removal edits. "Remove X" alone is weak. Pair the removal command with a positive description of what replaces it, then lock the rest: "Remove the sticky note. Show the continuous walnut desk surface where it was, matching the surrounding grain, lighting, and perspective — no residue, outline, or discoloration. Preserve every other object, the framing, and the color grading exactly." For stubborn removals, generate multiple candidates and re-edit only the residual region instead of enlarging the prompt.
When passing multiple --ref images, label each reference by index and role
inside the prompt, then state the relationships explicitly:
Image 1: base scene and composition.
Image 2: subject identity reference.
Image 3: style reference.
Place the subject from Image 2 into Image 1. Apply only Image 3's palette and
brushwork. Preserve Image 1's framing, background, perspective, and lighting.
There is no --parallel flag. For multiple candidates from the same prompt,
prefer one server-side batch request:
ima2 gen "four poster candidates" -n 4 -d ./out --quality high
ima2 multimode "four different poster directions" --max-images 4
For truly different prompts, independent CLI jobs can run concurrently against the same server. Capture request IDs with JSON output, then monitor or cancel:
ima2 gen "variation 1" --quality high --json
ima2 gen "variation 2" --quality high --json
ima2 ps --json
ima2 cancel <requestId>
Treat capabilities.limits.maxParallel as advisory client-side queue guidance only.
It is not a guaranteed server-side semaphore.
Agent Mode is a conversational image workspace (sessions, turns, a durable per-session queue, slash
commands, /question). It is served at /api/agent/* and lives in the web UI — there is no
ima2 agent CLI command. From the CLI, drive generation with ima2 gen, ima2 edit,
ima2 multimode, and ima2 node generate instead.
Use JSON when another agent needs to reason about active work:
ima2 inflight ls --json
ima2 inflight ls --kind multimode --terminal --json
Expect job fields such as requestId, kind, phase, startedAt, prompt,
model, and sessionId. Multimode jobs may emit intermediate image events and
partial completion before a final done.
Build a structured image prompt from a message or transcript:
ima2 prompt build --message "make this product prompt clearer" --json
ima2 prompt build --messages @conversation.json --json
Preview a local markdown/text prompt source before committing:
ima2 prompt import preview ./prompts.md --json
Import a JSON export body:
ima2 prompt import json ./prompts-export.json --folder __root__
Import a raw image into history:
ima2 history import ./local-image.png
Inspect the running server defaults, including defaults.cli.image and
defaults.cli.video in JSON:
ima2 defaults --json
Inspect local effective defaults without contacting a server:
ima2 defaults --local --json
Discover live model IDs and lane status before choosing a CLI target:
ima2 models
ima2 models --kind image --lane oauth --json
ima2 models --kind video --json
ima2 models --json has the stable shape
{"ok":true,"kinds":{"image":[],"video":[]}}. It requires the server; an
unreachable server returns SERVER_UNREACHABLE and exits 3.
Persist the server-side model defaults shared by GPT OAuth and API provider paths:
The built-in OAuth image default is gpt-5.6-luna; Grok image and video code
defaults are grok-imagine-image-quality and grok-imagine-video respectively.
Use grok-imagine-video-1.5 explicitly when its quality or 1080p capabilities
are needed.
ima2 defaults set model gpt-5.5
Persist the fail-closed CLI image and video targets separately:
ima2 defaults set image oauth/gpt-5.6-luna
ima2 defaults set video grok/grok-imagine-video
ima2 defaults reset image
ima2 defaults reset video
Setting a CLI target validates the live catalog. Unknown models and lanes are rejected, and locked/disconnected/key-missing lanes cannot become defaults.
Persist the default reasoning policy:
ima2 defaults set reasoning high
Restart a running server after changing persisted defaults:
ima2 serve
Request flags such as --model and --reasoning-effort are per-call overrides.
They do not change persistent defaults.
Use ima2 capabilities --json as the source of truth for:
Use only models from:
valid.imageModels.supported
Do not pick models from:
valid.imageModels.unsupported
Discover writable configuration keys:
ima2 config keys --json
.env values.ima2 capabilities --json before guessing model names.ima2 skill path when an agent needs the installed Markdown skill path.ima2 skill <name> refs to discover reference modules for front/uiux skills.ima2 skill <name> ref <refname> to load a specific reference module on demand.ima2 skill install --dir <path> to install skills to the agent's skill directory.ima2 inflight ls --json or ima2 ps --json to inspect active jobs.Generate AI videos through a configured Grok or MCP lane. Grok OAuth requires a SuperGrok subscription; MCP lanes require their own connected subscription.
ima2 models --kind video
ima2 defaults set video grok/grok-imagine-video
ima2 video "a cat playing piano" # text-to-video, uses saved default
ima2 video "animate this" --model grok/grok-imagine-video --ref photo.png
ima2 video "cinematic" --model grok/grok-imagine-video --ref a.png --ref b.png
Targets use --model <lane>/<model>; a bare ID is accepted only when it is
unique across lanes. Generate-mode video also accepts an explicit
--provider <grok|grok-api|runway|higgsfield>. --provider auto is removed.
Runway and Higgsfield are MCP lanes. They submit POST /api/mcp/generate (202)
and the CLI waits on SSE until completion. MCP generation supports -n 1 only,
and --ref values must be generated gallery filenames, not arbitrary local
paths. Core-only planning/session flags are rejected with FLAG_NOT_SUPPORTED.
ima2 gen and ima2 video also accept --character <element-id|name> on MCP
lanes: the element must be a character element in the assets workspace with
a provider binding for the selected lane (Runway: stateless refs + optional
@tag). Fail-closed envelopes: CHARACTER_ELEMENT_NOT_FOUND,
CHARACTER_ELEMENT_AMBIGUOUS, CHARACTER_BINDING_MISSING,
CAPABILITY_MISMATCH (core lane or model without image_references),
BINDING_NOT_READY, and server-side CHARACTER_ELEMENT_CONFLICT /
CHARACTER_REFS_EXCEED_PROVIDER_CAP.
Runway is available when connected; Higgsfield remains locked until its paid
lane is enabled. Inspect current state with ima2 models --kind video.
ima2 upscale <generated-file> upscales through the MCP media-action pipeline:
images take --scale-factor 2|4|8|16 (above 2 requires --flavor sublime),
--flavor, --sharpen, --smart-grain, --ultra-detail; videos take no
parameters. Multishot generation is POST /api/mcp/multishot (CLI surface
planned). Video edit is the 2-step edit-video-preview → edit-video-submit
media action; stage-1 returns a synchronous keyframe preview.
| Refs | Mode | Max Duration |
|---|---|---|
| 0 | text-to-video | 15s |
| 1 | image-to-video | 15s |
| 2-7 | reference-to-video | 10s |
grok-imagine-video-1.5 supports image-to-video and supports 1080p for prompt-only text-to-video and single image/frame image-to-video. Prompt-only 1.5 text-to-video is implemented as an internal white-canvas image-to-video anchor because upstream 1.5 rejects raw T2V. The old grok-imagine-video-1.5-preview string is accepted as a compatibility alias. 1.5 does not support reference_images Ref2V, V2V edit, or extension. For 2+ references, use grok-imagine-video and keep duration at 10s or less. ima2 may auto-retry a rejected 1.5 Ref2V request with the base model; read effectiveModel and modelFallback from the final result before naming or reporting the output.
| Flag | Values | Default |
|---|---|---|
--duration | 1–15 (seconds) | 5 |
--resolution | 480p, 720p, 1080p (1.5 T2V canvas shim or I2V) | 480p |
--aspect-ratio | auto, 1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3 | auto |
--model | <lane>/<model>; Grok base or 1.5 (preview alias accepted) | grok-imagine-video after selecting the Grok lane |
--topic | any string | (none) |
--session | session ID | (none) |
-o, --out | output file path | saved under configured generated dir |
--json | (flag) | false |
--topic is legacy/best-effort series context. Prefer branch-local artifact
continuity with ima2 video continue, Classic "Continue here", gallery video
drag, or Node parent-video generation. Those flows use the previous generated
video's last frame plus its stored revisedPrompt lineage.
Continuation is last-frame image-to-video, not true video-to-video. Only a single still frame carries over — motion, trajectory and camera movement are not preserved. Expect the next clip to start from that pose rather than to inherit momentum. Frames are pulled server-side with ffmpeg when the source is a generated file, falling back to in-browser canvas capture when ffmpeg is unavailable.
ima2 video "episode 1: morning routine" --topic "daily-vlog"
ima2 video "episode 2: commute" --model grok/grok-imagine-video --topic "daily-vlog"
Prompts are NOT sent directly to the video model. A Grok planner rewrites your prompt with web search context for better results. The revisedPrompt in the response shows what was actually sent. Default planner model is grok-4.3 (configurable in settings UI).
Override the planner model per-request:
ima2 video "prompt" --model grok/grok-imagine-video --planner-model gpt-5.5
ima2 video "prompt" --model grok/grok-imagine-video --planner-model gpt-5.4
| Surface | Files | Responsibility |
|---|---|---|
| Image search/planner | lib/grokImageAdapter.ts | Web-search context and final image prompt for Grok image generation/editing. |
| Video planner | lib/grokVideoAdapter.ts, lib/grokVideoPlannerPrompt.ts | Final video prompt for T2V/I2V/Ref2V, duration pacing, and continuity lineage when present. |
| Video analyzer | routes/videoExtended.ts | First/last-frame analysis prompt for recreating or continuing an existing generated video. |
| Agent/runtime prompt use | lib/agentRuntime.ts, card/template planner modules | Higher-level orchestration surfaces that may create image/video prompt inputs but do not replace the video planner contract. |
For video, the Grok 4.3 planner must produce one focused English prompt with:
core subject, expected action/motion, camera/composition, environment/style,
dialogue/audio intent, ending frame/continuity handoff, and constraints. If
videoContinuity exists, the lineage is authoritative context: continue from
the latest clip's final frame and final audio/dialogue state without restarting
the scene. The planner also applies duration pacing: use the selected seconds as
the full clip runtime, expand even short requests into a production-level
sequence, and make the clip feel complete through composition, blocking, camera
movement, motion rhythm, sound/dialogue timing, and an ending hold.
Blank video prompts are blocked. Weak natural-language prompts are allowed, but agents should always write an active prompt that includes:
The planner is model-aware: it adjusts for grok-imagine-video (simpler, bolder
composition at 480p) vs grok-imagine-video-1.5 (finer detail, 1080p textures).
For 1.5 text-to-video, the server uses a white-canvas shim internally; the
planner automatically writes a fresh-scene prompt without referencing a source
image.
Anti-slop: the planner rejects generic prestige phrases ("AAA trailer", "senior VFX artist", "shot on RED"), filler lighting ("volumetric", "neon glow"), and unmotivated dark/moody defaults. Write what the camera actually sees.
ima2 grok login # authenticate (manual-paste flow)
ima2 grok status # verify connection
ima2 serve # server must be running
SSE streaming events: planning → submitted → progress (0-100%) → done.
The submitted and done payloads include requestedModel, effectiveModel, and modelFallback so agents can report when a requested 1.5-preview Ref2V job actually ran on grok-imagine-video. CLI --json prints video.requestedModel, video.effectiveModel, and video.modelFallback; use path/filename for local chaining.
ima2 capabilities --json | jq '.valid.videoModels'
Generate a high-quality still image first, then animate it. This produces better results than text-to-video alone because the video model has a concrete visual anchor.
Critical rule for i2v: Compose ALL characters and the environment together in ONE image. Do NOT use individual portrait refs for i2v — the video model needs a single composed scene to animate from.
Keyframe image provider rule (MANDATORY):
provider: oauth) with quality: high, maximum resolution matching the target video aspect ratio. For 16:9 video use 1792x1024. For 1:1 use 1024x1024. For 9:16 use 1024x1792.provider: grok, model grok-imagine-image-quality). Only aspect ratio must match — resolution does not matter because i2v accepts any resolution source image and internally rescales.ref2v vs i2v decision:
| Scenario | Use | Why |
|---|---|---|
| Need 2+ character identity lock from separate refs | ref2v (grok-imagine-video, max 7 refs, max 10s) | Refs lock character appearance |
| Single composed scene with all elements | i2v (1.5-preview or base, 1 ref) | Better motion quality from composed start |
| Continue from previous video | video continue (last frame as i2v ref) | Lineage metadata preserved |
# Multi-character scene: compose BOTH characters in one image first
# Primary: GPT Image 2 at high quality, max resolution, aspect ratio matching 16:9 video
ima2 gen "cinematic wide shot of Bruce Lee in yellow tracksuit facing Elon Musk in dark gi, underground fight arena, dramatic lighting, 16:9" --quality high --size 1792x1024 -o scene.png
# Fallback if GPT fails: Grok quality model, match aspect ratio only
# ima2 gen "same prompt" --provider grok --model grok-imagine-image-quality --size 1824x1024 -o scene.png
# Then animate from the composed scene
ima2 video "Bruce throws a rapid jeet kune do combination" --ref scene.png --duration 10 --resolution 720p --aspect-ratio 16:9
Create a sequence of connected clips using --topic for narrative continuity. Each generation receives context from previous clips in the same topic.
# Scene 1: Establishing shot
ima2 video "wide establishing shot of a busy Tokyo street at night, neon signs" \
--topic "tokyo-night" --duration 5
# Scene 2: Medium shot (planner sees Scene 1's revised prompt)
ima2 video "medium shot following a person walking through the crowd" \
--topic "tokyo-night" --duration 5
# Scene 3: Close-up (planner sees Scenes 1+2)
ima2 video "close-up of rain drops on a neon sign reflection" \
--topic "tokyo-night" --duration 5
The planner receives previous prompts from the same topic as continuity context. This is best-effort prompt guidance, not a guarantee that subjects, palette, or style will remain identical. For branch-local continuation, use ima2 video continue instead.
The highest-quality video production workflow. Since Grok i2v accepts only one image input, pack the entire action sequence into a single 3×3 (9-panel) storyboard grid image. The i2v model reads the panels as a visual script and animates the progression.
Full workflow:
keyframe image (GPT high)
→ GPT i2i with reference → 9-panel storyboard grid
→ Grok i2v (reads panels, animates sequence)
→ extract last frame
→ GPT i2i with last frame → next 9-panel storyboard
→ Grok i2v
→ repeat
Step 1 — Opening keyframe (GPT Image 2, quality: high, max resolution matching target aspect ratio):
ima2 gen "cinematic wide shot of two fighters in a dojo, dramatic lighting" \
--quality high --size 1792x1024 --storyboard
Fallback: Grok grok-imagine-image-quality, match aspect ratio only — resolution does not matter because i2v internally rescales.
Step 2 — 9-panel storyboard grid (GPT Image 2 with keyframe as reference):
# Use the keyframe as reference, prompt describes 9 sequential panels
ima2 gen "Using this scene as reference, create a 3x3 storyboard grid (9 panels, thin black borders) showing a 15-second action sequence. Panel 1 (0s): ... Panel 2 (2s): ... Panel 9 (15s): ... Maintain identical character designs across all panels." \
--ref keyframe.png --quality high --size 1024x1024
9-panel storyboard rules:
Step 3 — Animate storyboard via i2v:
ima2 video "This is a 9-panel storyboard. Animate the full sequence as one continuous 15-second clip following panels left-to-right, top-to-bottom. Panel 1: ... Panel 9: ... Sound: [describe music, SFX, dialogue]. Camera: [describe movement per beat]." \
--ref storyboard.png --duration 15 --resolution 720p --model grok-imagine-video-1.5
i2v prompt rules for storyboard input:
Step 4 — Extract last frame and repeat:
# Extract last frame via ffmpeg
ffmpeg -sseof -0.1 -i clip.mp4 -frames:v 1 -q:v 2 -update 1 lastframe.jpg -y
# Generate next storyboard using last frame as reference
ima2 gen "Using this fight scene last frame as reference, create a 3x3 storyboard grid..." \
--ref lastframe.jpg --quality high --size 1024x1024
# Animate next storyboard
ima2 video "This is a 9-panel storyboard..." --ref storyboard2.png --duration 15
Fallback: continueFromVideo — If a storyboard image triggers content moderation (common with intense action/fight scenes), fall back to video continue with a detailed text prompt instead:
ima2 video continue "detailed action description with sound and camera direction" \
--video "$PREV_CLIP" --duration 15
Clip duration is flexible — use 15s for action-dense sequences with many beats, 10s for transitions, 5s for quick cuts. The 9-panel storyboard works best with 15s clips (each panel ≈ 1.5-2s of screen time).
Music and sound are MANDATORY in i2v prompts — describe the score (orchestral, percussion, taiko drums), sound effects (impacts, whooshes, crashes), dialogue lines, and audio transitions. "No music" or undefined audio produces flat, lifeless output.
To continue from an existing video's last frame:
# Get the last generated video filename
LAST=$(ima2 ls -n 1 --json | jq -r '.items[0].filename')
# True extension keeps the original clip and appends new motion
ima2 video extend "the camera slowly pulls back revealing the full scene" --video "$LAST" --duration 6
# Branch-local sequel keeps revisedPrompt lineage and starts from the last frame
ima2 video continue "from the last frame, the camera slowly pulls back, no music, footsteps echo, end on a still wide shot" --video "$LAST"
Or in the UI: use "Continue here" on a video, drag a video from gallery/history
to the prompt composer, or create a child from a video node. These flows attach
the previous video's last frame and carry a branch-local videoContinuity
lineage stack. The stack stores up to 4 revised prompts using
keep-start-plus-latest-3: start clip is preserved, and the newest three clips
stay in context.
ima2 video extend is xAI native extension: it returns original+extension as a
combined artifact. ima2 video continue is ima2 branch continuation: it creates
a new clip from the generated video's last frame and persists lineage metadata.
Generate a product showcase video from a product image:
# Step 1: Generate or provide product image
ima2 gen "clean product photo of wireless earbuds on white background" -o product.png
# Step 2: Create dynamic product video
ima2 video "sleek product reveal with rotating camera, premium feel, studio lighting" \
--ref product.png --duration 10 --resolution 720p --aspect-ratio 16:9
For maintaining visual style across multiple videos (e.g., social media series):
# First video establishes the style
ima2 video "minimalist animation of a coffee cup, flat design, pastel colors" \
--topic "coffee-series" --duration 5
# Subsequent videos inherit style via planner context
ima2 video "same style, now showing latte art being poured" \
--topic "coffee-series" --duration 5
ima2 video "same style, steam rising from the cup" \
--topic "coffee-series" --duration 5
#!/bin/bash
PROMPTS=("sunrise over ocean" "waves crashing" "seagulls flying" "sunset colors")
TOPIC="ocean-day"
for prompt in "${PROMPTS[@]}"; do
ima2 video "$prompt" --topic "$TOPIC" --duration 5 --json >> results.jsonl
sleep 2 # rate limiting
done
grok-imagine-video effective modelgrok-imagine-video-1.5 prompt-only text-to-video via the white-canvas I2V shim, and for image-to-video with a single image/frame sourceEdit an existing video with a text prompt. This uses xAI's real video edit endpoint and saves the result as a generated video artifact.
# Get the local video file from a previous generation
VIDEO_FILE=$(ima2 video "ocean waves" --json | jq -r '.path')
# Edit: change style
ima2 video edit "Make the water glow neon blue, bioluminescent" --video "$VIDEO_FILE"
# Edit: add object
ima2 video edit "Add a sailboat in the distance" --video "$VIDEO_FILE"
# Edit: change mood
ima2 video edit "Make it stormy with dark clouds" --video "$VIDEO_FILE"
Constraints: grok-imagine-video only, input mp4 <=8.7s. Use -o/--out if you also need a local copy outside the generated directory.
Extend a video from its last frame using xAI's video extension endpoint. The output combines the source video and extension, but continuity quality is provider-dependent.
Constraints: grok-imagine-video only, extension duration 2-10s. 1.5-preview is not supported for extension.
# Generate initial clip
VIDEO_FILE=$(ima2 video "a bird takes flight from a branch" --duration 5 --json | jq -r '.path')
# Extend: add 5 more seconds
ima2 video extend "the bird soars higher into the clouds" --video "$VIDEO_FILE" --duration 5
# Chain extensions for longer videos
EXTENDED=$(ima2 video extend "camera follows the bird" --video "$VIDEO_FILE" --duration 5 --json | jq -r '.filename')
ima2 video extend "bird lands on a distant tree" --video "$EXTENDED" --duration 5
Extract frames from generated videos for use as references or analysis.
# Extract last frame
ima2 video frame 1780226256355_50252101.mp4 --last -o lastframe.png
# Extract frame at specific timestamp
ima2 video frame 1780226256355_50252101.mp4 --position 2.5 -o frame_2s.png
# Use extracted frame as reference for new generation
ima2 video "continue this scene" --ref lastframe.png
Analyze first and last video frames with Grok 4.3 image understanding to get a structured recreation prompt. This infers motion from frames; it is not full temporal video understanding.
# Analyze a generated filename
ima2 video analyze 1780226256355_50252101.mp4
# Output: structured prompt with shot type, inferred camera movement, lighting, color, motion, mood
# Use the analysis to recreate with variations
ANALYSIS=$(ima2 video analyze 1780226256355_50252101.mp4 --json | jq -r '.analysis')
ima2 video "$ANALYSIS but in anime style" --ref reference.png
The API does not expose a separate audio on/off or audio-track control. Treat audio as prompt-compiled: describe dialogue, music, no-music, room tone, or sound-effects-only behavior in the video prompt. Output is provider-dependent, but the prompt must be explicit when audio matters.
# Explicit sound direction
ima2 video "ocean waves crashing on rocks with seagull calls and distant thunder"
# Music direction
ima2 video "timelapse of city at night, lo-fi hip hop background music"
# Dialogue
ima2 video "person speaking to camera: Hello world, welcome to my channel"
# No music / room tone
ima2 video "quiet forest scene, no background music, only subtle wind and leaves rustling"
# Sound effects only
ima2 video "no music, only footsteps, cloth movement, rain hits, and one radio click"
For continuity clips, always define the final audio state: whether dialogue finishes before the cut, music resolves or continues, or a sound effect carries into the next clip.
Use this structure for serious video generation, Ref2V, extension prompts, and multi-shot continuity. A static visual description is not enough. Write like a director calling a shot, not filling out a form.
Opening frame: composition, depth layers, spatial staging, material/texture.
Motivated movement: what changes and why — reveal, follow, discover, tension.
Camera intent: the specific move that serves this scene (macro push-in, orbit,
lateral slider, rack focus, locked overhead, handheld, crane).
Visual turning point: a shift in focus, scale, light, or subject state.
Dialogue: speaker (by visual appearance, not name), exact line in original
language, timing — or "no dialogue".
Sound: music style with swell/cut/resolve behavior, or "no background music,
room tone only", or specific SFX (footsteps, rain, machine hum, impact).
Settling final frame: stable pose, camera angle, background, lighting, held
audio state — self-explanatory for continuation.
Negative constraints: no visible subtitles/text unless requested, preserve
identity/style.
When creating a sequence, write both motions explicitly: "A motion" for the first clip and "B motion" for the continuation. For last-frame Ref2V, use ref 1 as identity/style and ref 2 as current state/last frame.
Shot discipline (cross-vendor official guidance):
Example — product reveal (10s, 1.5, 1080p):
A single continuous macro shot begins inches above a matte black desk surface,
tight on the brushed aluminum edge of wireless earbuds catching a narrow softbox
reflection. The camera glides laterally as focus racks from the charging case
texture to the earbud stem, revealing the full product silhouette against soft
warm-gray negative space. A gentle ambient hum, no music. The camera settles
into a medium close-up with the product centered, soft rim light from behind,
holding steady on the final composition.
Guide the video toward a desired final scene using reference images:
# Start frame + end frame concept
ima2 video "smooth transition from day to night" \
--ref sunrise.png --ref nightsky.png
The planner treats reference images as subject/style/composition guidance. This is best-effort guidance, not a guaranteed final-frame constraint.
Guide character identity across multiple videos using reference photos:
# Provide face references for consistency
ima2 video "person walking through a park, smiling" \
--ref face_front.png --ref face_side.png --ref face_smile.png
# Same character in different scenes
ima2 video "same person now sitting at a cafe" \
--ref face_front.png --ref face_side.png --topic "character-series"
Turn a product image into a dynamic showcase video:
# Step 1: Generate or provide product image
ima2 gen "clean product photo of wireless earbuds on white background" -o product.png
# Step 2: Create product video
ima2 video "sleek product reveal, rotating camera, premium studio lighting" \
--ref product.png --duration 10 --aspect-ratio 16:9
# Step 3: Extend with lifestyle shot
PRODUCT_VID=$(ima2 video "product reveal" --ref product.png --json | jq -r '.path')
ima2 video extend "person puts on the earbuds and smiles" --video "$PRODUCT_VID" --duration 5
Agents: run ima2 tools list --json for the live view; this section is the bundled-snapshot projection.
ima2 (5)| tool | executable via | description |
|---|---|---|
ima2.generate_image | agent runtime | Generate one or more images. Supports fanout: provide one prompt per variant. |
ima2.generate_video | agent runtime | Generate a single video with Grok Imagine. If the session has a last image, it is used as the image-to-video source automatically; prompt-on |
ima2.get_generation_errors | agent runtime | Read-only lookup of the session's recent generation failures (failed queue jobs and error turns). Use when the user asks why a generation fa |
ima2.get_image_context | agent runtime | Load the session image context manifest (previous images, current image, locks). Runs automatically before image generation. |
ima2.web_search | agent runtime | Search the web for factual visual references before generating. Only available when web search is enabled for the session. |
mcp.higgsfield (73)| tool | executable via | description |
|---|---|---|
mcp.higgsfield.animation_actions | — | Read-only catalog of the 3D rig animation library (678 actions: locomotion, gestures, dancing, combat, daily actions). Search by name or bro |
mcp.higgsfield.balance | — | Get the user's available credits and current subscription plan. For transaction history, call transactions instead. |
mcp.higgsfield.cancel_trial_auto_renewal | — | Cancel the auto-renewal of the Higgsfield MCP 3-day free Plus trial. Call this when the user asks to cancel the trial, cancel auto-renewal, |
mcp.higgsfield.confirm_billing_purchase | — | INTERNAL — invoked ONLY by the plans widget on an explicit user Confirm click. Do NOT call this tool yourself; it charges the user's real sa |
mcp.higgsfield.confirm_trial_cancel | — | INTERNAL — invoked ONLY by the cancel-trial confirmation widget on an explicit user click of 'Cancel auto-renewal'. Do NOT call this tool yo |
mcp.higgsfield.create_voice | — | Open the Create Voice Apps UI. Call this immediately when the user asks to create a voice, call the Create Voice tool, or needs a local brow |
mcp.higgsfield.create_voice_from_confirmed_audio | — | Backend-only creation of a cloned voice from an already confirmed audio upload. Do not call this tool until audio_media_id and name are alre |
mcp.higgsfield.create_website | — | Start a new full-stack website. Creates the website and a git repo: a React 19 + TanStack Start app, server-rendered, in ONE Cloudflare Work |
mcp.higgsfield.deploy_game | — | Deploy a built browser game from an uploaded zip archive and get a shareable play URL. Deploying also lists the game in the Higgsfield marke |
mcp.higgsfield.deploy_website | — | Build and deploy the website via CI, then return its live URL. Every deploy ships the live site at the website's public URL (there is no sep |
mcp.higgsfield.dubbing | — | Dub a video into another language: translate the spoken audio, synthesize it in the target language, and lip-sync the result back onto the v |
mcp.higgsfield.explainer_video | — | Assemble an explainer / narrated video from its per-block clips and voice takes: stitches two or more existing video clips into one MP4, in |
mcp.higgsfield.generate_3d | — | Generate a 3D GLB mesh. Use models_explore(type:'3d') to pick a model and see its medias[].roles and parameters. Apps UI local file: c |
mcp.higgsfield.generate_audio | — | Generate speech/voice audio (text-to-speech). DEFAULT model: seed_audio (Seed Audio 1.0 by ByteDance) — use it unless the user explicitly as |
mcp.higgsfield.generate_image | POST /api/mcp/generate | Generate an image. Apps UI local file media: call media_upload_widget; do not ask for Claude chat attachments because remote tools cannot |
mcp.higgsfield.generate_video | POST /api/mcp/generate | Generate a video. Apps UI local file: call media_upload_widget; do not ask for Claude chat attachments; remote tools cannot read them. Web |
mcp.higgsfield.get_explainer_presets | — | Show the explainer video style presets (CMS-managed catalog). Returns preset ids, names, and preview media. When the user picks one, resolve |
mcp.higgsfield.get_game_creation_bundle_file | — | Read a safe text file or directory from the game-generation resource folder. Use this after get_game_creation_instructions when the instruct |
mcp.higgsfield.get_game_creation_instructions | — | REQUIRED before creating or editing any browser game. Reads the game-generation SKILL.md resource and returns the current list of files avai |
mcp.higgsfield.get_website_creation_bundle_file | — | Read a safe text file or directory from the website-builder-flow resource folder. Use this after get_website_creation_instructions when the |
mcp.higgsfield.get_website_creation_instructions | — | REQUIRED before creating or editing any website with the website tools (create_website / website_repo_access / deploy_website / website_db / |
mcp.higgsfield.get_workflow_bundle_file | — | Read a safe text file or directory from a workflow's resource folder. Use this after get_workflow_instructions when the SKILL.md requires a |
mcp.higgsfield.get_workflow_instructions | — | Discover and load multi-step content-generation workflows (each a bundled SKILL.md that orchestrates the generate_* tools). Call with NO arg |
mcp.higgsfield.job_display | — | Show a single generation result in the UI widget by job ID. Pass exactly one job ID — to display multiple generations, call this tool once p |
mcp.higgsfield.job_status | — | Check the status and results of an async job. Returns instantly. For non-terminal jobs the response includes poll_after_seconds — wait that |
mcp.higgsfield.list_voices | — | List available voices for speech and voice tools. Returns built-in preset voices plus the user's own custom voices. Each voice has a voice_i |
mcp.higgsfield.list_websites | — | List the websites you own — each with its id, name, slug, and live URL. Use this to find the id of a website you created earlier so you can |
mcp.higgsfield.list_workspaces | — | List every workspace the user can access (their private workspace plus any shared/team workspaces). The is_selected field marks which work |
mcp.higgsfield.media_confirm | — | Confirm file uploads after using the upload_url method. Call this after the curl uploads succeed. Supports confirming multiple uploads at on |
mcp.higgsfield.media_import_url | — | Import an HTTPS image, video, or audio URL into Higgsfield storage and return a confirmed media_id. Use this before generate_image/generate_ |
mcp.higgsfield.media_upload | — | Upload media for use in generation, or general files (documents, archives, code) for sharing. Returns presigned URLs for clients that can up |
mcp.higgsfield.media_upload_widget | — | Open the Higgsfield upload widget for a user-provided local image, video, or audio file. Call this immediately when the user says they have |
mcp.higgsfield.models_explore | — | Find generation models. Use recommend with goal + input context; use get for model constraints. |
mcp.higgsfield.motion_control | — | Animate an existing character image with the motion and camera movement from a reference video using Kling 3.0 Motion Control. Use this when |
mcp.higgsfield.outpaint_image | — | Expand or uncrop an existing image by outpainting beyond the original frame while preserving the source content. Use this when the user asks |
mcp.higgsfield.participate_in_contest | — | Enter the website in the current Higgsfield app contest, together with the social-media links promoting it. A website not yet PUBLISHED to t |
mcp.higgsfield.personal_clipper_create | — | Turn YouTube videos into ready-to-share clips. This is a long-running job and can take up to 30+ minutes. Before starting, ask the user how |
mcp.higgsfield.personal_clipper_jobs | — | Show recent clipping jobs. |
mcp.higgsfield.personal_clipper_status | — | Check clip creation progress. |
mcp.higgsfield.presets_show | — | Show available Higgsfield presets for image-to-video generation. Returns preset ids, names, previews, and descriptions. |
mcp.higgsfield.publish_game | — | Publish a deployed game to the Higgsfield marketplace. This does not deploy anything: the game must already be live — use deploy_game first, |
mcp.higgsfield.publish_website | — | Publish the website: lists the website's CURRENT LIVE production deploy on the Higgsfield community feed ('show in feed'), where other users |
mcp.higgsfield.reframe | — | Expand or reframe an existing video to a new aspect ratio while preserving the source content. Use this when the user asks to make a video v |
mcp.higgsfield.remove_background | — | Remove or cut out the background from an existing image or video. Use this when the user asks for background removal, a transparent backgrou |
mcp.higgsfield.rename_website | — | Rename the website's SUBDOMAIN (the slug in its public URL). The site is re-deployed under the new subdomain and the OLD subdomain STOPS WOR |
mcp.higgsfield.resolve_explainer_preset | — | Resolve an explainer video style preset (from get_explainer_presets) into a style reference media_id: the backend imports the preset's style |
mcp.higgsfield.reveal_generation | — | Confirm the user has rights to the content of an ip_detected generation and flip its status to completed. Backend accepts only seedance- |
mcp.higgsfield.select_workspace | — | Set or clear the active workspace — the one all subsequent MCP operations bill against and read from (generations, balance, transactions, up |
mcp.higgsfield.shorts_studio_create | — | Start a Shorts Studio short: restyle one uploaded source video (4s–120s) into a set of AI-generated short-form clips using a style preset. P |
mcp.higgsfield.shorts_studio_create_preset | — | Create a user-owned Shorts Studio style preset from reference media (videos + images). This just stores a STYLE — no generation, no credits. |
mcp.higgsfield.shorts_studio_list_presets | — | Browse Shorts Studio style presets — the visual STYLE a short is restyled toward. Use this when the user wants to make a short and needs to |
mcp.higgsfield.shorts_studio_list_sessions | — | List the caller's past Shorts Studio sessions (newest first) to find a session_id to poll with shorts_studio_status. |
mcp.higgsfield.shorts_studio_status | — | Poll one Shorts Studio session. Returns {id, status, job_ids}. status='completed' means every clip job is terminal (not necessarily successf |
mcp.higgsfield.show_characters | — | Soul Characters widget — reusable trained identity models. Actions: list (browse), train (needs name + 5-20 ref images, ~10 min, non-b |
mcp.higgsfield.show_generations | — | Browse past completed non-Marketing Studio generations and render them directly in the widget. Returns generations with {id, type, status, m |
mcp.higgsfield.show_marketing_studio | — | When replying to the user, do not say ms_image — refer to it as "DTC Ads". |
mcp.higgsfield.show_marketing_studio_generations | — | Browse past completed Marketing Studio generations only. Returns Marketing Studio video and ad/image generations with {id, type, status, mod |
mcp.higgsfield.show_medias | — | List your uploaded media files by type. Returns media IDs, URLs, and creation timestamps. Pass media IDs as value in the medias array of gen |
mcp.higgsfield.show_plans_and_credits | — | Open the single combined pricing widget for everything billing-related. The widget has two tabs the user can switch between: *Upgrade Plan |
mcp.higgsfield.show_reference_elements | — | Elements widget — reusable characters / environments / props per workspace. Actions: |
mcp.higgsfield.sync_agents | — | Sync Agents — imports the user's user-authored Skills and a personality dump from the current host LLM into Higgsfield. One trigger, one upl |
mcp.higgsfield.transactions | — | List the user's credit transactions (spend/refund/grant/deduct), newest first. Paginated: if next_cursor is not null, pass it as cursor to g |
mcp.higgsfield.upscale_image | POST /api/mcp/media-action | Upscale and enhance an existing image. Use this when the user asks to upscale, enhance, or increase the resolution of an image to 2K/4K. Thi |
mcp.higgsfield.upscale_video | POST /api/mcp/media-action | Upscale and enhance an existing video. Use this when the user asks to upscale, enhance, sharpen, denoise, restore, or convert a video to hig |
mcp.higgsfield.video_analysis_create | — | Start a scene-by-scene analysis of a video. Provide EXACTLY ONE of: (a) video_input_id — UUID of a video the user has uploaded via media_upl |
mcp.higgsfield.video_analysis_jobs | — | List the user's video analyses in the current workspace, newest first. Paginate by passing the previous response's cursor. |
mcp.higgsfield.video_analysis_status | — | Get the status and result of a video analysis. Poll this after video_analysis_create until status='completed' (scenes populated) or 'failed' |
mcp.higgsfield.virality_predictor | — | Virality Predictor predicts a video's virality potential, engagement, attention, audience response, retention risk, hook strength, and creat |
mcp.higgsfield.voice_change | — | Replace the spoken voice in a video with a different voice while keeping the original timing and visuals, then re-merge the new audio onto t |
mcp.higgsfield.website_db | — | Inspect the website's database (D1 / SQLite), READ-ONLY. The website has ONE database — the live site's real data. Pick an operation: 'table |
mcp.higgsfield.website_repo_access | — | Get direct git access to a website's repo to edit it — THE way to get the website's code. Returns the repo URL, branch, slug, and a scoped t |
mcp.higgsfield.website_secrets | — | Manage a website's SECRETS (environment variables: API keys, tokens). Set them HERE instead of hardcoding them in source. One tool, three op |
mcp.higgsfield.website_status | — | Get the website's deploy status — the live URL and the status of the last deploy. Use to check a deploy that returned 'pending', or to fetch |
mcp.runway (14)| tool | executable via | description |
|---|---|---|
mcp.runway.complete_upload | — | Finalize a file upload and get the asset URL. |
mcp.runway.edit_video | POST /api/mcp/media-action | Edit an existing video with Runway's Aleph 2.0 in-context video editor: it changes ONLY what you ask for and preserves everything else — sub |
mcp.runway.feedback | — | Call this when you (the AI agent) get stuck using Runway tools. |
mcp.runway.generate_image | POST /api/mcp/generate | Generate OR edit an image using a Runway-hosted image model. This is the only image tool — there is no separate "edit_image" tool. Pass the |
mcp.runway.generate_multishot_video | — | Generate a multi-shot video — 3 to 5 connected scenes from a single story or per-shot prompts. Powered by Kling 3.0 (standard at 720p, pro a |
mcp.runway.generate_product_marketing_video | — | Generate a polished creative product ad video from a product URL or product image plus a campaign idea. |
mcp.runway.generate_video | POST /api/mcp/generate | Generate OR edit a video using a Runway-hosted video model. Pass the source video as referenceVideo to edit/restyle it. |
mcp.runway.get_task | — | Gets details for a Runway task by ID — used to check status and retrieve the result of a generation/edit task once it completes. Generation |
mcp.runway.init_upload | — | Initialize a file upload to Runway. Returns temporary upload URLs for direct upload. |
mcp.runway.list_recent | — | Lists recent uploaded and generated assets for the authenticated workspace. Returns asset IDs, media types, task IDs when available, and reu |
mcp.runway.list_workspaces | — | Lists every Runway workspace the authenticated user belongs to, with role and a compact disabled-model summary per workspace. The MCP connec |
mcp.runway.upscale_image | POST /api/mcp/media-action | Upscale an existing image to a higher resolution (2x, 4x, 8x, or 16x) using Runway's AI image upscaler. Use this to sharpen, denoise, and in |
mcp.runway.upscale_video | POST /api/mcp/media-action | Upscale an existing video to a higher resolution (up to 4K) using Runway's AI video upscaler. Use this to sharpen, clean up, and increase th |
mcp.runway.whoami | — | Returns the authenticated Runway user profile, the workspace this MCP connection is pinned to (chosen at sign-in), and the list of image/vid |
评论 (0)
暂无评论,成为第一个评论者吧!