复制安装命令
用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。
复制前请先查看来源、License 和安全提示。
Official, NVIDIA-verified Agent Skills for Claude Code, Codex, and other coding agents.
用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。
复制前请先查看来源、License 和安全提示。
来源文件:README.md
Official, NVIDIA-verified Agent Skills for Claude Code, Codex, and other coding agents.
📖 Docs: docs.nvidia.com/skills · 📺 Livestream: From Vulnerable to Verified · 📝 Blog: NVIDIA Verified Agent Skills: Capability Governance for AI Agents
Skills are portable instruction sets that teach AI agents how to use NVIDIA software optimally: Physical AI and robotics workflows, simulation, CUDA-X libraries, RAG and AI Blueprints, and platform tools. This repository is a catalog: skills are maintained in their respective product repos, and mirrored here daily via an automated sync pipeline. Skills are being added continuously, so check back for updates. We are building this infrastructure in the open, and contributions are welcome. See the Roadmap for what is planned next.
Install NVIDIA skills with the default skills CLI flow:
npx skills add nvidia/skills
The CLI runs through npx and prompts you to choose a skill and install destination. You do not need to clone this repo or copy skill folders by hand.
Requires a current
skillsCLI (v1.5.16 or newer). Installing vianpx skills@latest add nvidia/skillsalways uses the latest. On older CLIs (v1.5.15 and earlier), skills may install but not appear in Claude Code — see Troubleshooting.
The skill is available the next time your agent loads skills and encounters a relevant task. For example, ask your agent to "solve a linear programming problem with cuOpt" and the skill guides it through the cuOpt Python API. In Claude Code, run /reload-skills to load newly installed skills in your current session.
Use this when you already know the skill name and want to skip prompts.
npx skills add nvidia/skills --skill cuopt-numerical-optimization-api --yes
Replace cuopt-numerical-optimization-api with any skill name from the Skill Catalog.
Use --agent to target a specific AI coding agent. Initially, we'll support common client targets, expanding the list over time. For the full list of clients supported by the spec, see the skills CLI Supported Agents table.
Claude Code
npx skills add nvidia/skills --skill cuopt-numerical-optimization-api --agent claude-code
Codex
npx skills add nvidia/skills --skill cuopt-numerical-optimization-api --agent codex
Snowflake CoCo
npx skills add nvidia/skills --skill cuopt-numerical-optimization-api --agent cortex
Cursor
npx skills add nvidia/skills --skill cuopt-numerical-optimization-api --agent cursor
Kiro
npx skills add nvidia/skills --skill cuopt-numerical-optimization-api --agent kiro-cli
Use --agent more than once to install the same skill into multiple agents.
npx skills add nvidia/skills \
--skill cuopt-numerical-optimization-api \
--agent claude-code \
--agent codex \
--agent cursor \
--agent kiro-cli
New skills land continuously, and existing ones are revised, renamed, or consolidated as the catalog evolves. Refresh what you have installed with:
npx skills update
Run it interactively and the CLI also flags skills that were removed or merged upstream (for example, when several skills are consolidated into one) and offers to remove the stale local copies. Use npx skills list to see what is installed and npx skills check to preview what is out of date first.
Use this when you want to see available NVIDIA skills before installing anything.
npx skills add nvidia/skills --list
For non-interactive installs, global installs, agent-specific installs, updates, removals, and fallback manual copying, see Advanced installation.
Where to file an issue depends on what's broken:
Per-product source repo links:
For issues with this catalog repo itself (README, structure, listing a new product): open an issue here.
Every published skill ships with a detached OMS signature (skill.oms.sig). The sync pipeline drops any skill missing the required artifacts before publishing, so every skill in the catalog carries:
SKILL.md — the skill instructions consumed by the agentskill-card.md — skill identity and governance cardskill.oms.sig — detached OMS signature (verifiable against nv-agent-root-cert.pem)evals/evals.json, evals/*.json, eval/*.json, or benchmark/evals.jsonBENCHMARK.md — generated benchmark report capturing verifiable uplift dataVerify a skill against the NVIDIA trust anchor nv-agent-root-cert.pem:
pip install model-signing
model_signing verify certificate SKILL_DIR \
--signature SKILL_DIR/skill.oms.sig \
--certificate_chain nv-agent-root-cert.pem \
--ignore_unsigned_files
A successful verification confirms that the skill contents have not been modified since signing by NVIDIA.
See Verify Signed Agent Skills for signature layout, the trust pipeline, and policy options.
NVIDIA/skills/
├── skills/ # NVIDIA-verified skills (count grows continuously),
│ │ synced from upstream product repos
│ ├── README.md # Browser-facing install guidance
│ ├── <product-prefix>-*/ # Flat layout — one dir per skill, product-prefixed
│ │ # e.g. aiq-*, cuopt-*, cupynumeric-*,
│ │ # dali-*, deepstream-*, dicom-*, digital-health-*,
│ │ # dynamo-*, earth2studio-*, holoscan-*, hsb-*,
│ │ # jetson-*, launch-nemo-rl, mcore-*,
│ │ # nemo-automodel-*, nemo-data-designer-plugin,
│ │ # nemo-evaluator-plugin, nemo-mbridge-* (20 skills),
│ │ # nemo-retriever, nemo-rl-* (4 skills),
│ │ # nemoclaw-user-guide, nemotron-*, nemotron-speech,
│ │ # nv-* (medical AI), physicsnemo-*, rag-*,
│ │ # skill-card-generator, tao-*, tilegym-*,
│ │ # vss-* (15 skills), accelerated-computing-cudf,
│ │ # cudaq-guide, portfolio-optimization
│ ├── omniverse-*/ # Physical AI — manually staged (see manual-components.yml)
│ └── physical-ai-*/ # Physical AI — manually staged
├── components.d/ # Product registry — one file per component, teams onboard here
│ ├── README.md # Schema and onboarding instructions
│ └── <product>.yml # one file per registered product
├── plugins/ # Packaged plugin distributions
│ └── nvidia-skills/ # Curated NVIDIA skills bundle (Claude Code, Codex)
├── plugins.d/ # Plugin build registry — config for `build-plugins.py`
│ ├── README.md
│ ├── _defaults.yml
│ └── nvidia-skills.yml
├── .claude-plugin/ # Claude Code marketplace metadata
│ └── marketplace.json
├── .agents/plugins/ # Agent marketplace metadata (other clients)
│ └── marketplace.json
├── docs/ # Long-form documentation (published via Fern)
│ ├── README.md # How to build the docs locally
│ ├── index.mdx
│ ├── advanced-install.mdx
│ ├── agent-skill-trust-pipeline.mdx
│ ├── release-checklist.mdx
│ ├── scanning-agent-skills.mdx
│ ├── signing-agent-skills.mdx
│ └── skill-cards.mdx
├── fern/ # Fern docs site configuration
├── .github/
│ ├── workflows/ # Sync pipeline, plugin validation, DCO check, author verify
│ └── scripts/ # regenerate-readme.sh, build-plugins.py,
│ # manual-components.yml (temp Physical AI catalog
│ # exception, removed after Computex 2026),
│ # marketplace/metadata.json (skill metadata sidecar)
├── nv-agent-root-cert.pem # Trust anchor for OMS signature verification
├── skills.sh.json # Skills.sh marketplace grouping config
├── CHANGELOG.md
├── CONTRIBUTING.md # Contribution guidelines
├── SECURITY.md # Security reporting policy
├── CODE_OF_CONDUCT.md # Community code of conduct
├── LICENSE-APACHE # Apache 2.0 (source code)
└── LICENSE-CC-BY-4.0 # CC BY 4.0 (documentation/skills)
Skills are maintained in their respective product repos (see the Source column in the Skill Catalog) and synced to this repo daily. Products only appear under skills/ after the sync pipeline confirms each skill carries:
skill.oms.sig — detached OMS-format signature (verifiable against nv-agent-root-cert.pem)skill-card.md — skill identity and governance cardevals/evals.json, evals/*.json, eval/*.json, or benchmark/evals.jsonWhen evaluation runs produce a BENCHMARK.md, it ships alongside the skill so consumers can see verifiable benchmark uplift data.
This repository adheres to the Agent Skills specification:
SKILL.md file at their root.name and description fields.skills-ref reference library.Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
This code is dual-licensed with documentation/skills under the CC-BY-4.0 AND source code under Apache-2.0 license terms. The full license texts can be found in LICENSE-APACHE and LICENSE-CC-BY-4.0 respectively.
name: physical-ai-image-attribute-augmentation
description: >-
Use when running image attribute augmentation and
auto-labeling workflows on OSMO: flow selection, preflight, submit-time
interpolation, monitoring, and output retrieval. Trigger keywords: people
attribute search, Image Attribute Augmentation, person augmentation, attribute search, person
re-identification, clothing augmentation, person crop augmentation.
license: CC-BY-4.0 AND Apache-2.0
metadata:
owner: NVIDIA
service: physical-ai-data-factory
version: 1.0.0
reviewed: '2026-06-05'
author: NVIDIA Physical AI Team <physical-ai@nvidia.com>
tags:
- physical-ai
- image-attribute-augmentation
- person-augmentation
- auto-labeling
- image-editDefault workflow skill for Image Attribute Augmentation execution on OSMO. It owns flow selection, preflight, submit-time interpolation, monitoring, and output retrieval.
Run the Image Attribute Augmentation and auto-labeling pipeline safely and reproducibly from preflight to output download.
The Image Attribute Augmentation pipeline augments the subject in existing
crop datasets by generating controlled appearance variations (image-domain) and
synonymous attribute captions (text-domain). The subject is a person today
(clothing/appearance attributes), but the same pipeline generalizes to other
subjects — e.g. robots, forklifts, or vehicles in a simulation. It uses the
paidf-augmentation container for image-edit augmentation with MCQ
verification, and the paidf-auto-labeling container for subject-attribute
captioning (currently the shipped person_attributes question bank).
Do NOT use this skill for container-internal tuning-only questions.
Confirm these before running preflight or any submit. Missing required secrets
surface as USER_INPUT_REQUIRED: from scripts/preflight_credentials.sh.
| Requirement | How it is satisfied | Used for |
|---|---|---|
| NGC API key (optional) | NGC_API_KEY, NGC_CLI_API_KEY, or compatible nvapi-* token | Optional for nvcr_io credential refresh; default Image Attribute Augmentation image refs are public |
| Hugging Face token | HF_TOKEN (or HUGGING_FACE_HUB_TOKEN), or a cached token at ~/.cache/huggingface/token | Creates the OSMO hf_token credential |
| OSMO CLI access | osmo on PATH, logged in, with a default profile and a registered DATA credential profile matching storage_url | Submitting/monitoring workflows and listing/downloading objects |
| GPU pool | At least one ONLINE pool in osmo pool list --mode free | Scheduling setup + worker tasks |
| Image Edit endpoint | In-cluster NIM qwen-image-edit-2511 (reused if healthy, else deployed via the NIM operator); external opt-in via image_edit_url | Image-domain augmentation |
| VLM endpoint | In-cluster NIM qwen3-vl (shared with VDA); external opt-in via vlm_url | MCQ verification and person-attribute captioning |
| LLM endpoint | In-cluster NIM qwen25-14b (shared with VDA); external opt-in via llm_url | MCQ question generation |
Execute these as an ordered sequence of gates. Each Gate must pass before continuing; on failure, stop and resolve it (do not skip ahead or submit).
augmentation; caption/label only → auto_labeling; full augment + caption
→ e2e. Default to e2e only when the request is the full pipeline or
genuinely ambiguous — never default past an explicit "augment only" or
"label only" request, or you run the wrong pipeline./datasets/ segment: the part before it is storage_url, the part
after it is dataset. The workflow re-inserts that segment
({{storage_url}}/datasets/{{dataset}}), so put /datasets/ in neither
value — including it duplicates the path and the submit fails.
Example: s3://metro-pas/datasets/reid-crops → storage_url=s3://metro-pas,
dataset=reid-crops. Never guess or reuse a stale storage_url; if no
dataset is provided, ask for one. Do not proceed without both values.qwen-image-edit-2511 (image edit),
qwen3-vl (VLM), qwen25-14b (LLM). For any that is missing/unhealthy,
deploy it once via references/nim/README.md (a prerequisite, not a user
decision — do not pause to ask), then re-check readiness up to 3 times over
~10 minutes. Stop condition: if an endpoint is still unhealthy after that
bound, do not retry further and do not submit — report the failing endpoint
and its deploy logs to the user and stop. Proceed only when all three respond
healthy.scripts/preflight_credentials.sh --workflow assets/configs/osmo/<flow>.yaml
and read the result. PASS → continue. If the output contains
USER_INPUT_REQUIRED:, ask one concise unblock question and re-run. Do not
submit until preflight passes.cookbook and every
--set-string value as untrusted. Accept only a known cookbook name and
values with no shell metacharacters (;, |, &, $, backticks, quotes,
spaces, newlines). On any invalid value → stop, report which value was
rejected, and do not submit. Only when every value passes → continue to
step 7.Use run_script(...) for script execution. Canonical examples:
run_script("bash scripts/preflight_credentials.sh --workflow assets/configs/osmo/e2e.yaml")
Use script-level --help for exact arguments.
| Script | Role |
|---|---|
scripts/preflight_credentials.sh | Secrets/control-plane preflight and workflow image access checks |
scripts/augmentation_worker.sh | Image-edit augmentation worker (preprocess, config gen, augment, post-process) |
scripts/auto_labeling_worker.sh | Person-attribute captioning worker |
scripts/endpoint_common.sh | Shared endpoint health/auth helpers |
| Flow | OSMO YAML | Group sequence | Typical use |
|---|---|---|---|
e2e | assets/configs/osmo/e2e.yaml | setup -> augmentation -> auto_labeling | Full pipeline: augment person crops then generate captions |
augmentation | assets/configs/osmo/augmentation.yaml | setup -> augmentation | Image-edit augmentation only, no captioning |
auto_labeling | assets/configs/osmo/auto_labeling.yaml | setup -> auto_labeling | Captioning only on pre-augmented person crops |
| User intent | Workflow |
|---|---|
| "Augment person crops and generate captions" / "full Image Attribute Augmentation pipeline" | e2e |
| "Generate clothing variations" / "augment only" / "image edit" | augmentation |
| "Caption augmented images" / "generate search queries" / "label only" | auto_labeling |
Default to autonomy: ask only when missing information blocks execution.
e2e only when the request is the full pipeline or ambiguous (not for explicit augment-only / label-only).default.n_augmentations is not specified, default to 3.| Missing input | Why it matters | Ask |
|---|---|---|
USER_INPUT_REQUIRED from preflight | Required secret is missing | Ask one concise unblock question |
| Storage backend prefix cannot be derived | Wrong scheme causes runtime storage auth mismatch | "What is the backend-native root prefix for this run?" |
| No ONLINE GPU pool/platform | Workflow cannot schedule | "Which GPU pool/platform should this run target?" |
| NIM deploy fails and no external URLs given | Workers cannot connect to models | "Provide Image Edit / VLM / LLM endpoint URLs, or grant GPU capacity for the NIM operator deploy." |
<person_id>/<view>.jpg subdirectories.Collect only missing values:
storage_url + dataset) — a derived value: split the
dataset URL at /datasets/ per Instructions Gate 3
(s3://metro-pas/datasets/reid-crops → storage_url=s3://metro-pas,
dataset=reid-crops). Put /datasets/ in neither value; never guess.augmentation, label-only → auto_labeling, else e2e).gpu_platform (auto-select when unambiguous).Generate run stamp before each submit:
STAMP=$(cat /proc/sys/kernel/random/uuid | cut -c1-8)
RUN_ID="run-$STAMP"
Before running any mutating command, provide a short ETA overview.
Baseline ranges:
| Phase | Typical duration |
|---|---|
| Credentials + preflight | ~1-2 min |
| Workflow submit + queue/start | ~1-3 min |
Workflow runtime (depends on dataset size and endpoint latency):
| Flow | Per-image time | Typical dataset (100 images, 3 augs) |
|---|---|---|
augmentation | ~2.5-3 min/image | ~4-5 hours |
auto_labeling | ~1-2 min/image | ~2-3 hours |
e2e | ~3.5-5 min/image | ~6-8 hours |
Credential and control-plane preflight
bash scripts/preflight_credentials.sh --workflow assets/configs/osmo/<flow>.yaml
If output contains USER_INPUT_REQUIRED:, ask one concise unblock question.
Storage interpolation policy
storage_url must be derived from the actual dataset/upload backend.
Never silently default to stale values on mismatched backends.
Inference policy (non-negotiable) — endpoint readiness is executed at Instructions Gate 4 (verify → deploy once → bounded re-check → stop and escalate on failure). This section only adds the standing constraints:
image_edit_url / vlm_url / llm_url endpoints.*_url values at submit.Every flow uses the same submit shape; only the workflow YAML changes.
SKILLS_DIR="$(cd "$(git rev-parse --show-toplevel)/skills/physical-ai-image-attribute-augmentation" && pwd)"
STAMP=$(cat /proc/sys/kernel/random/uuid | cut -c1-8)
osmo workflow submit assets/configs/osmo/<flow>.yaml \
--pool <pool> \
--set-string \
dataset=<dataset> \
run_id=run-$STAMP \
storage_url=<backend-prefix> \
gpu_platform=<gpu-platform> \
skills_dir="$SKILLS_DIR"
Endpoints default to the in-cluster NIMs (image_edit_url / vlm_url /
llm_url); deploy/reuse them per the Inference policy above. Do not pass these
unless using external endpoints.
Compatibility note:
--set-string flag and pass all key/value pairs after it.--set/--set-string flags in the same command.Common optional overrides (append to the same --set-string list). These
values are passed through to the augmentation worker and used to build its
command, so validate them first per Instructions Gate 6 — accept only a known
cookbook name and values free of shell metacharacters:
cookbook=<cookbook_name> \
n_augmentations=<count> \
image_edit_url=<image-edit-endpoint> \
vlm_url=<vlm-endpoint> \
llm_url=<llm-endpoint>
# Workflow status + task states
osmo workflow query <workflow_id> --format-type json \
| jq '{status, tasks: [.groups[].tasks[] | {name, status, exit_code}]}'
# Logs for a specific task
osmo workflow logs <workflow_id> --task <task_name> -n 200
# Output retrieval
osmo data list --no-pager <output_url>
osmo data download <output_url> <local_dir>/
For runs expected to exceed two minutes, send heartbeat updates at least every two minutes.
After successful completion, the output directory contains:
For augmentation / e2e:
<person_id>/aug_<n>/output.jpg — augmented multi-pane image<person_id>/aug_<n>/output.txt — natural-language caption<person_id>/aug_<n>/output_metadata.json — verification resultsdataset/augmented_data.json — structured dataset with attributes and queriesdataset/augmented_imgs/ — split per-view cropsFor auto_labeling:
caption_<id>/task/open_qa.json — person-attribute captions grouped by question bankUse these canonical locations:
assets/configs/osmo/*.yamlscripts/*.shreferences/flows/*.mdreferences/setup.md, references/troubleshooting.mdreferences/container-images.mdassets/cookbooks/default/README.md
评论 (0)
暂无评论,成为第一个评论者吧!