复制安装命令
用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。
复制前请先查看来源、License 和安全提示。
Official, NVIDIA-verified Agent Skills for Claude Code, Codex, and other coding agents.
用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。
复制前请先查看来源、License 和安全提示。
来源文件:README.md
Official, NVIDIA-verified Agent Skills for Claude Code, Codex, and other coding agents.
📖 Docs: docs.nvidia.com/skills · 📺 Livestream: From Vulnerable to Verified · 📝 Blog: NVIDIA Verified Agent Skills: Capability Governance for AI Agents
Skills are portable instruction sets that teach AI agents how to use NVIDIA software optimally: Physical AI and robotics workflows, simulation, CUDA-X libraries, RAG and AI Blueprints, and platform tools. This repository is a catalog: skills are maintained in their respective product repos, and mirrored here daily via an automated sync pipeline. Skills are being added continuously, so check back for updates. We are building this infrastructure in the open, and contributions are welcome. See the Roadmap for what is planned next.
Install NVIDIA skills with the default skills CLI flow:
npx skills add nvidia/skills
The CLI runs through npx and prompts you to choose a skill and install destination. You do not need to clone this repo or copy skill folders by hand.
Requires a current
skillsCLI (v1.5.16 or newer). Installing vianpx skills@latest add nvidia/skillsalways uses the latest. On older CLIs (v1.5.15 and earlier), skills may install but not appear in Claude Code — see Troubleshooting.
The skill is available the next time your agent loads skills and encounters a relevant task. For example, ask your agent to "solve a linear programming problem with cuOpt" and the skill guides it through the cuOpt Python API. In Claude Code, run /reload-skills to load newly installed skills in your current session.
Use this when you already know the skill name and want to skip prompts.
npx skills add nvidia/skills --skill cuopt-numerical-optimization-api --yes
Replace cuopt-numerical-optimization-api with any skill name from the Skill Catalog.
Use --agent to target a specific AI coding agent. Initially, we'll support common client targets, expanding the list over time. For the full list of clients supported by the spec, see the skills CLI Supported Agents table.
Claude Code
npx skills add nvidia/skills --skill cuopt-numerical-optimization-api --agent claude-code
Codex
npx skills add nvidia/skills --skill cuopt-numerical-optimization-api --agent codex
Snowflake CoCo
npx skills add nvidia/skills --skill cuopt-numerical-optimization-api --agent cortex
Cursor
npx skills add nvidia/skills --skill cuopt-numerical-optimization-api --agent cursor
Kiro
npx skills add nvidia/skills --skill cuopt-numerical-optimization-api --agent kiro-cli
Use --agent more than once to install the same skill into multiple agents.
npx skills add nvidia/skills \
--skill cuopt-numerical-optimization-api \
--agent claude-code \
--agent codex \
--agent cursor \
--agent kiro-cli
New skills land continuously, and existing ones are revised, renamed, or consolidated as the catalog evolves. Refresh what you have installed with:
npx skills update
Run it interactively and the CLI also flags skills that were removed or merged upstream (for example, when several skills are consolidated into one) and offers to remove the stale local copies. Use npx skills list to see what is installed and npx skills check to preview what is out of date first.
Use this when you want to see available NVIDIA skills before installing anything.
npx skills add nvidia/skills --list
For non-interactive installs, global installs, agent-specific installs, updates, removals, and fallback manual copying, see Advanced installation.
Where to file an issue depends on what's broken:
Per-product source repo links:
For issues with this catalog repo itself (README, structure, listing a new product): open an issue here.
Every published skill ships with a detached OMS signature (skill.oms.sig). The sync pipeline drops any skill missing the required artifacts before publishing, so every skill in the catalog carries:
SKILL.md — the skill instructions consumed by the agentskill-card.md — skill identity and governance cardskill.oms.sig — detached OMS signature (verifiable against nv-agent-root-cert.pem)evals/evals.json, evals/*.json, eval/*.json, or benchmark/evals.jsonBENCHMARK.md — generated benchmark report capturing verifiable uplift dataVerify a skill against the NVIDIA trust anchor nv-agent-root-cert.pem:
pip install model-signing
model_signing verify certificate SKILL_DIR \
--signature SKILL_DIR/skill.oms.sig \
--certificate_chain nv-agent-root-cert.pem \
--ignore_unsigned_files
A successful verification confirms that the skill contents have not been modified since signing by NVIDIA.
See Verify Signed Agent Skills for signature layout, the trust pipeline, and policy options.
NVIDIA/skills/
├── skills/ # NVIDIA-verified skills (count grows continuously),
│ │ synced from upstream product repos
│ ├── README.md # Browser-facing install guidance
│ ├── <product-prefix>-*/ # Flat layout — one dir per skill, product-prefixed
│ │ # e.g. aiq-*, cuopt-*, cupynumeric-*,
│ │ # dali-*, deepstream-*, dicom-*, digital-health-*,
│ │ # dynamo-*, earth2studio-*, holoscan-*, hsb-*,
│ │ # jetson-*, launch-nemo-rl, mcore-*,
│ │ # nemo-automodel-*, nemo-data-designer-plugin,
│ │ # nemo-evaluator-plugin, nemo-mbridge-* (20 skills),
│ │ # nemo-retriever, nemo-rl-* (4 skills),
│ │ # nemoclaw-user-guide, nemotron-*, nemotron-speech,
│ │ # nv-* (medical AI), physicsnemo-*, rag-*,
│ │ # skill-card-generator, tao-*, tilegym-*,
│ │ # vss-* (15 skills), accelerated-computing-cudf,
│ │ # cudaq-guide, portfolio-optimization
│ ├── omniverse-*/ # Physical AI — manually staged (see manual-components.yml)
│ └── physical-ai-*/ # Physical AI — manually staged
├── components.d/ # Product registry — one file per component, teams onboard here
│ ├── README.md # Schema and onboarding instructions
│ └── <product>.yml # one file per registered product
├── plugins/ # Packaged plugin distributions
│ └── nvidia-skills/ # Curated NVIDIA skills bundle (Claude Code, Codex)
├── plugins.d/ # Plugin build registry — config for `build-plugins.py`
│ ├── README.md
│ ├── _defaults.yml
│ └── nvidia-skills.yml
├── .claude-plugin/ # Claude Code marketplace metadata
│ └── marketplace.json
├── .agents/plugins/ # Agent marketplace metadata (other clients)
│ └── marketplace.json
├── docs/ # Long-form documentation (published via Fern)
│ ├── README.md # How to build the docs locally
│ ├── index.mdx
│ ├── advanced-install.mdx
│ ├── agent-skill-trust-pipeline.mdx
│ ├── release-checklist.mdx
│ ├── scanning-agent-skills.mdx
│ ├── signing-agent-skills.mdx
│ └── skill-cards.mdx
├── fern/ # Fern docs site configuration
├── .github/
│ ├── workflows/ # Sync pipeline, plugin validation, DCO check, author verify
│ └── scripts/ # regenerate-readme.sh, build-plugins.py,
│ # manual-components.yml (temp Physical AI catalog
│ # exception, removed after Computex 2026),
│ # marketplace/metadata.json (skill metadata sidecar)
├── nv-agent-root-cert.pem # Trust anchor for OMS signature verification
├── skills.sh.json # Skills.sh marketplace grouping config
├── CHANGELOG.md
├── CONTRIBUTING.md # Contribution guidelines
├── SECURITY.md # Security reporting policy
├── CODE_OF_CONDUCT.md # Community code of conduct
├── LICENSE-APACHE # Apache 2.0 (source code)
└── LICENSE-CC-BY-4.0 # CC BY 4.0 (documentation/skills)
Skills are maintained in their respective product repos (see the Source column in the Skill Catalog) and synced to this repo daily. Products only appear under skills/ after the sync pipeline confirms each skill carries:
skill.oms.sig — detached OMS-format signature (verifiable against nv-agent-root-cert.pem)skill-card.md — skill identity and governance cardevals/evals.json, evals/*.json, eval/*.json, or benchmark/evals.jsonWhen evaluation runs produce a BENCHMARK.md, it ships alongside the skill so consumers can see verifiable benchmark uplift data.
This repository adheres to the Agent Skills specification:
SKILL.md file at their root.name and description fields.skills-ref reference library.Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
This code is dual-licensed with documentation/skills under the CC-BY-4.0 AND source code under Apache-2.0 license terms. The full license texts can be found in LICENSE-APACHE and LICENSE-CC-BY-4.0 respectively.
name: physical-ai-event-video-generation
description: Run the PAIDF Orchestration Event Video Generation DAG on Kubernetes - image-to-video anomaly generation, auto-labeling, and anomaly dataset generation. Select for requests about event video generation, anomaly video generation, image-to-video synthesis, Cosmos3 image2video, anomaly dataset creation, safety/surveillance SDG, or generating person-falling, person-climbing, person-running, fighting, smoking/vaping, fire/smoke, or shoplifting video clips from a seed image. Runs environment setup first when controller readiness is unknown. Not for person-crop clothing/attribute augmentation (that is image-attribute-augmentation-workflow) and not for video style transfer.
version: "1.0.0"
license: CC-BY-4.0 AND Apache-2.0
metadata:
owner: NVIDIA
service: physical-ai-data-factory
version: 1.0.0
reviewed: '2026-09-02'
author: NVIDIA
tags:
- physical-ai
- paidf-orchestration
- event-video-generation
- cosmosRun the Event Video Generation DAG end to end: seed-image input preparation, Cosmos3 image-to-video anomaly augmentation, auto-labeling (detection and tracking, captioning, anomaly visual QA, person-attribute visual QA, person attribute search), anomaly dataset generation, and result retrieval.
The workflow builds one DAG per compute platform from
airflow/dags/workflows/event_video_generation_dag/:
| Platform | DAG ID | Manifest |
|---|---|---|
| Kubernetes | event_video_generation_dag_k8s | event_video_generation_k8s_manifest.yaml |
Kubernetes is the only platform whose manifest is checked in, so
event_video_generation_dag_k8s is the only DAG this repository registers. A DAG is
registered only if its manifest exists; a missing manifest means the DAG is absent from Airflow
rather than broken. List the DAGs Airflow actually loaded before triggering, and never name a DAG
ID that is not in that list.
There is a single end-to-end pipeline — there are no generation-only or labeling-only DAG variants. If a user asks for video generation without auto-labeling, tell them the checked-in DAG does not offer that flow rather than inventing a DAG ID.
If the user wants to enter their own payload directly in the Airflow UI rather than have you
construct and trigger one, your job is limited to getting them to the UI: confirm controller
readiness, ensure make port-forward is running (see
airflow-direct-api.md), and report the
reachable URL. Do not render a payload, run preflight, or trigger a run yourself in this case —
the user is doing that from the UI. Resume monitoring (step 6 below) once they tell you a run has
been triggered; you can find it via the Airflow API without needing the payload they used.
Before building any payload, collect all of the following from the user. Do not fall back to repository defaults, CI payloads, or any hardcoded endpoint URL or bucket path.
| Required | Field | What to ask |
|---|---|---|
| Always | input_path | S3 (or HTTP/HTTPS) URL to a single seed image or a directory of seed images |
| Always | output_directory | Writable S3 URL where results should be written |
| Always | service mode | external (user provides endpoint URLs) or internal (DAG deploys services in-cluster) |
| External mode | cosmos.vlm_service_url | Full HTTPS URL for the VLM inference endpoint |
| External mode | cosmos.llm_service_url | Full HTTPS URL for the LLM inference endpoint |
| External mode | cosmos.image2video_service_url | Full HTTPS URL for the Cosmos3 image-to-video inference endpoint |
| Optional | max_images | Number of images to process from a directory (default: 10; 0 or negative = all) |
| Optional | cosmos.num_augmentation | Anomaly videos generated per image (default: 1) |
| Optional | cosmos.variable_distribution | anomaly_type / env_type sampling distribution (see payload-contract.md) |
If the user does not provide a required value, ask for it explicitly before proceeding. Do not invent or reuse values from previous runs or checked-in files.
Always run the following readiness checks before triggering a run. The checks are short-circuiting — stop at the first failure and route to the environment-setup skill immediately.
Before any check, establish the cluster connection. The cluster is reached only through credentials the user supplies — they are never part of the repository. Check whether the cluster credential file path is already exported in the shell environment; if not, ask the user for the absolute path before running any cluster command. Never assume a path or fall back to any on-disk default — see setup-and-preflight.md for the full procedure.
The controller (Airflow) and DAG compute tasks run on the same cluster unless a different remote cluster connection was configured. GPU capacity is checked on this cluster.
Controller pods — check that the Airflow controller pods (not DAG task pods) are Running.
DAG task pods in Pending or Failed state are normal and must not be mistaken for controller
failures:
kubectl get pods -n sdg-workflow -l "release=sdg-workflow-controller"
All pods matching the release=sdg-workflow-controller label must be Running. If the
namespace is absent, this is a first-install condition — route to the environment-setup skill,
do not diagnose further.
Airflow API — reachable only if check 1 passes. First establish AIRFLOW_URL from the
Kubernetes ClusterIP (always routable from the host, no port-forward required):
AIRFLOW_URL="http://$(kubectl get svc -n sdg-workflow \
sdg-workflow-controller-api-server \
-o jsonpath='{.spec.clusterIP}'):8080"
Then confirm the target DAG is loaded and is_paused: False. See
airflow-direct-api.md for the full auth + check sequence.
If the API is unreachable, route to the environment-setup skill.
Pools — only if check 2 passes. Required pools with open slots: k8s_gpu_1,
default_pool, and the image2video pool for the chosen mode
(external_image2video_service_pool for external, internal_image2video_service_pool for
internal).
Compute-cluster GPUs — check the cluster (using the cluster connection established above):
kubectl get nodes \
-o custom-columns='NAME:.metadata.name,GPU_ALLOC:.status.allocatable.nvidia\.com/gpu'
# Also check pods already consuming GPUs — capacity ≠ availability on a shared cluster
kubectl get pods -n sdg-workflow \
--field-selector=status.phase=Running -o wide
The compute cluster is shared — other users' runs may be active. Report GPUs as free-versus-total, not just allocatable. Two independent GPU sources, only one of which is mode-dependent:
detection_and_tracking, captioning, and visual_qa all run on the k8s_gpu_task profile
(1 GPU each) — this cost applies in both external and internal mode, since these do in-pod
inference rather than calling an endpoint. event_and_person_attribute_search and
augmentation run on CPU profiles and cost nothing. External mode therefore needs a minimum
of three GPUs, not zero.external_services: false): one GPU per VLM/LLM replica,
plus two GPUs per image2video replica (gpu_count: 2, host_ipc: true) — four GPUs for
one replica of each service.Stale failed pods — before triggering, check for accumulated failed pods in the compute namespace and report them. They are retained by design and do not affect run correctness, but they consume namespace quota and clutter log searches:
kubectl get pods -n sdg-workflow \
--field-selector=status.phase=Failed \
-o custom-columns='NAME:.metadata.name,AGE:.metadata.creationTimestamp,DAG:.metadata.labels.dag_id'
Clean up only pods whose dag_id label matches a run you own, after confirming with the user.
Document each check result explicitly.
If any check fails: invoke the environment-setup skill automatically — do not wait for the user to say "set up" or ask them to name the skill.
If the user's request implies first-time or explicit deployment ("deploy", "install", "set up", "reinstall", "redeploy", "full setup"): invoke the environment-setup skill even if all checks pass, and confirm the planned commands first.
If all checks pass and the user only wants to run the workflow: proceed directly to payload and trigger.
scripts/upload_images.py: validate/upload a local seed image or flat directory of seed images.scripts/payload.py: render or validate a standalone
EventVideoGenerationDagPayloadConfig-compatible JSON.scripts/summarize_results.py: summarize a downloaded anomaly_dataset/dataset.json dataset.Run commands from this skill directory. Credentials must be inherited from the shell that launched the agent; never ask the user to paste secret values into the prompt.
Determine the input source.
For local data, validate before upload:
python scripts/upload_images.py --path /path/to/seed-images --validate-only
Then upload:
python scripts/upload_images.py \
--path /path/to/seed-images --destination-path event-video-generation/my-run
For an existing storage URL, use it unchanged after confirming it names either a single
image or a flat directory of images (.jpg, .jpeg, .png, .bmp, .gif, .tiff,
.webp). Unlike person-crop workflows, there is no subdirectory convention — every matching
file directly under input_path is one input image. When input_path names a directory,
the DAG sorts matching files and takes the first max_images of them.
Select service mode.
external requires explicit VLM, LLM, and image2video endpoint URLs.internal lets the DAG's service lifecycle deploy all three services in-cluster.k8s_gpu_1, but image2video claims
two GPUs per replica (gpu_count: 2, host_ipc: true) — four GPUs for the internally
deployed endpoints alone. That's on top of, not instead of, the three GPUs
detection_and_tracking/captioning/visual_qa always claim from k8s_gpu_task regardless
of service mode (see the readiness-check GPU breakdown above): external mode needs a
minimum of three allocatable GPUs, internal mode a minimum of seven — not zero and four.Read payload-contract.md, then render a payload from the values collected above. Do not copy checked-in dev or CI payloads — they contain deployment- specific endpoint URLs and bucket paths that must not be inherited by user runs.
External:
python scripts/payload.py render \
--input-path s3://bucket/input/seed-images/ \
--output-directory s3://bucket/output/event-video-generation/ \
--service-mode external \
--vlm-url https://vlm.example/v1 \
--llm-url https://llm.example/v1 \
--image2video-url https://image2video.example/v1 \
--max-images 10 --num-augmentation 3 \
--variable-distribution assets/variable-distribution.json \
--output /tmp/evg-payload.json
Internal:
python scripts/payload.py render \
--input-path s3://bucket/input/seed-images/ \
--output-directory s3://bucket/output/event-video-generation/ \
--service-mode internal \
--max-images 10 --num-augmentation 3 \
--output /tmp/evg-payload.json
Show the user the rendered payload (or its validated contents) and get explicit confirmation before proceeding. Only continue to preflight and triggering if they confirm; if they want changes, re-render and re-confirm.
Preflight the DAG through the Airflow API. Check that the DAG is loaded, required pools have slots, and controller pods are healthy — see airflow-direct-api.md#preflight-direct-path. Confirm presence only; never print credential values.
Submit exactly one DAG run. Pass the payload from step 3 as conf.payload — see
airflow-direct-api.md#trigger-a-run for the
full request shape. Record and return the dag_run_id, input path, output directory, and
service mode.
Immediately after triggering — without waiting to be asked — monitor the run until it reaches
a terminal state (success or failed). Poll the Airflow API every 60–120 seconds:
# Poll run state
RESPONSE=$(curl -s -H "Authorization: Bearer $TOKEN" \
"$AIRFLOW_URL/api/v2/dags/$DAG_ID/dagRuns/$RUN_ID")
RESPONSE="$RESPONSE" python3 -c "import json, os; print(json.loads(os.environ['RESPONSE'])['state'])"
For a per-task breakdown when state is running or failed, see
airflow-direct-api.md.
Stop polling as soon as the run state is success or failed. Use the polling loop that
fits your runtime — a shell while loop, a background process, or a tool-native scheduler.
Do not block the user waiting for each poll; report state changes as they occur.
Tell the user they can also watch progress live in the Airflow UI. make port-forward runs in
the foreground and never exits, so start it as a background job — and prefer that the user runs
it in their own terminal, since an agent-owned forward dies with the session. Resolve the host's
real address rather than reporting a placeholder or localhost, which is meaningless from
another machine:
HOST_IP=$(hostname -I | awk '{print $1}')
echo "Airflow UI: http://$HOST_IP:8080"
Default credentials are admin/admin, defined in deploy/values.yaml under
airflow.createUserJob.defaultUser (not webserver.defaultUser). Update them before
production use.
For a full per-task breakdown see airflow-direct-api.md.
To stop an in-progress run: open the Airflow UI, find the active DagRun, locate the running task, and mark it Failed (task menu → Mark Failed). This triggers the DAG's shutdown path, cleaning up Deployments, Services, and GPU pods. Do not delete the DagRun or the DAG — that bypasses cleanup and leaves stale cluster resources.
After the run reaches success or failed, ask the user:
"Would you like to download and analyze the results?"
Do not download automatically — wait for confirmation.
If the user confirms, use whatever AWS credentials are already available in the shell
environment (standard AWS_ACCESS_KEY_ID / AWS_SECRET_ACCESS_KEY / AWS_DEFAULT_REGION,
an AWS profile, or instance role). Never ask the user to paste credentials into the prompt.
Run artifacts live under <output_directory>/<run_id>/, where <output_directory> is the
payload value and <run_id> is the dag_run_id from step 5. The final dataset is in
anomaly_dataset/:
aws s3 sync "<output_directory>/<run_id>/anomaly_dataset/" /tmp/evg-results/
python scripts/summarize_results.py --results-dir /tmp/evg-results
To inspect intermediate generated videos instead, sync <output_directory>/<run_id>/cosmos/
and read metadata.json from each <video_key>/<augmentation_index>/ folder.
Read outputs.md before interpreting files.
image-attribute-augmentation-workflow.
评论 (0)
暂无评论,成为第一个评论者吧!