复制安装命令
用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。
复制前请先查看来源、License 和安全提示。
Official, NVIDIA-verified Agent Skills for Claude Code, Codex, and other coding agents.
用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。
复制前请先查看来源、License 和安全提示。
来源文件:README.md
Official, NVIDIA-verified Agent Skills for Claude Code, Codex, and other coding agents.
📖 Docs: docs.nvidia.com/skills · 📺 Livestream: From Vulnerable to Verified · 📝 Blog: NVIDIA Verified Agent Skills: Capability Governance for AI Agents
Skills are portable instruction sets that teach AI agents how to use NVIDIA software optimally: Physical AI and robotics workflows, simulation, CUDA-X libraries, RAG and AI Blueprints, and platform tools. This repository is a catalog: skills are maintained in their respective product repos, and mirrored here daily via an automated sync pipeline. Skills are being added continuously, so check back for updates. We are building this infrastructure in the open, and contributions are welcome. See the Roadmap for what is planned next.
Install NVIDIA skills with the default skills CLI flow:
npx skills add nvidia/skills
The CLI runs through npx and prompts you to choose a skill and install destination. You do not need to clone this repo or copy skill folders by hand.
Requires a current
skillsCLI (v1.5.16 or newer). Installing vianpx skills@latest add nvidia/skillsalways uses the latest. On older CLIs (v1.5.15 and earlier), skills may install but not appear in Claude Code — see Troubleshooting.
The skill is available the next time your agent loads skills and encounters a relevant task. For example, ask your agent to "solve a linear programming problem with cuOpt" and the skill guides it through the cuOpt Python API. In Claude Code, run /reload-skills to load newly installed skills in your current session.
Use this when you already know the skill name and want to skip prompts.
npx skills add nvidia/skills --skill cuopt-numerical-optimization-api --yes
Replace cuopt-numerical-optimization-api with any skill name from the Skill Catalog.
Use --agent to target a specific AI coding agent. Initially, we'll support common client targets, expanding the list over time. For the full list of clients supported by the spec, see the skills CLI Supported Agents table.
Claude Code
npx skills add nvidia/skills --skill cuopt-numerical-optimization-api --agent claude-code
Codex
npx skills add nvidia/skills --skill cuopt-numerical-optimization-api --agent codex
Snowflake CoCo
npx skills add nvidia/skills --skill cuopt-numerical-optimization-api --agent cortex
Cursor
npx skills add nvidia/skills --skill cuopt-numerical-optimization-api --agent cursor
Kiro
npx skills add nvidia/skills --skill cuopt-numerical-optimization-api --agent kiro-cli
Use --agent more than once to install the same skill into multiple agents.
npx skills add nvidia/skills \
--skill cuopt-numerical-optimization-api \
--agent claude-code \
--agent codex \
--agent cursor \
--agent kiro-cli
New skills land continuously, and existing ones are revised, renamed, or consolidated as the catalog evolves. Refresh what you have installed with:
npx skills update
Run it interactively and the CLI also flags skills that were removed or merged upstream (for example, when several skills are consolidated into one) and offers to remove the stale local copies. Use npx skills list to see what is installed and npx skills check to preview what is out of date first.
Use this when you want to see available NVIDIA skills before installing anything.
npx skills add nvidia/skills --list
For non-interactive installs, global installs, agent-specific installs, updates, removals, and fallback manual copying, see Advanced installation.
Where to file an issue depends on what's broken:
Per-product source repo links:
For issues with this catalog repo itself (README, structure, listing a new product): open an issue here.
Every published skill ships with a detached OMS signature (skill.oms.sig). The sync pipeline drops any skill missing the required artifacts before publishing, so every skill in the catalog carries:
SKILL.md — the skill instructions consumed by the agentskill-card.md — skill identity and governance cardskill.oms.sig — detached OMS signature (verifiable against nv-agent-root-cert.pem)evals/evals.json, evals/*.json, eval/*.json, or benchmark/evals.jsonBENCHMARK.md — generated benchmark report capturing verifiable uplift dataVerify a skill against the NVIDIA trust anchor nv-agent-root-cert.pem:
pip install model-signing
model_signing verify certificate SKILL_DIR \
--signature SKILL_DIR/skill.oms.sig \
--certificate_chain nv-agent-root-cert.pem \
--ignore_unsigned_files
A successful verification confirms that the skill contents have not been modified since signing by NVIDIA.
See Verify Signed Agent Skills for signature layout, the trust pipeline, and policy options.
NVIDIA/skills/
├── skills/ # NVIDIA-verified skills (count grows continuously),
│ │ synced from upstream product repos
│ ├── README.md # Browser-facing install guidance
│ ├── <product-prefix>-*/ # Flat layout — one dir per skill, product-prefixed
│ │ # e.g. aiq-*, cuopt-*, cupynumeric-*,
│ │ # dali-*, deepstream-*, dicom-*, digital-health-*,
│ │ # dynamo-*, earth2studio-*, holoscan-*, hsb-*,
│ │ # jetson-*, launch-nemo-rl, mcore-*,
│ │ # nemo-automodel-*, nemo-data-designer-plugin,
│ │ # nemo-evaluator-plugin, nemo-mbridge-* (20 skills),
│ │ # nemo-retriever, nemo-rl-* (4 skills),
│ │ # nemoclaw-user-guide, nemotron-*, nemotron-speech,
│ │ # nv-* (medical AI), physicsnemo-*, rag-*,
│ │ # skill-card-generator, tao-*, tilegym-*,
│ │ # vss-* (15 skills), accelerated-computing-cudf,
│ │ # cudaq-guide, portfolio-optimization
│ ├── omniverse-*/ # Physical AI — manually staged (see manual-components.yml)
│ └── physical-ai-*/ # Physical AI — manually staged
├── components.d/ # Product registry — one file per component, teams onboard here
│ ├── README.md # Schema and onboarding instructions
│ └── <product>.yml # one file per registered product
├── plugins/ # Packaged plugin distributions
│ └── nvidia-skills/ # Curated NVIDIA skills bundle (Claude Code, Codex)
├── plugins.d/ # Plugin build registry — config for `build-plugins.py`
│ ├── README.md
│ ├── _defaults.yml
│ └── nvidia-skills.yml
├── .claude-plugin/ # Claude Code marketplace metadata
│ └── marketplace.json
├── .agents/plugins/ # Agent marketplace metadata (other clients)
│ └── marketplace.json
├── docs/ # Long-form documentation (published via Fern)
│ ├── README.md # How to build the docs locally
│ ├── index.mdx
│ ├── advanced-install.mdx
│ ├── agent-skill-trust-pipeline.mdx
│ ├── release-checklist.mdx
│ ├── scanning-agent-skills.mdx
│ ├── signing-agent-skills.mdx
│ └── skill-cards.mdx
├── fern/ # Fern docs site configuration
├── .github/
│ ├── workflows/ # Sync pipeline, plugin validation, DCO check, author verify
│ └── scripts/ # regenerate-readme.sh, build-plugins.py,
│ # manual-components.yml (temp Physical AI catalog
│ # exception, removed after Computex 2026),
│ # marketplace/metadata.json (skill metadata sidecar)
├── nv-agent-root-cert.pem # Trust anchor for OMS signature verification
├── skills.sh.json # Skills.sh marketplace grouping config
├── CHANGELOG.md
├── CONTRIBUTING.md # Contribution guidelines
├── SECURITY.md # Security reporting policy
├── CODE_OF_CONDUCT.md # Community code of conduct
├── LICENSE-APACHE # Apache 2.0 (source code)
└── LICENSE-CC-BY-4.0 # CC BY 4.0 (documentation/skills)
Skills are maintained in their respective product repos (see the Source column in the Skill Catalog) and synced to this repo daily. Products only appear under skills/ after the sync pipeline confirms each skill carries:
skill.oms.sig — detached OMS-format signature (verifiable against nv-agent-root-cert.pem)skill-card.md — skill identity and governance cardevals/evals.json, evals/*.json, eval/*.json, or benchmark/evals.jsonWhen evaluation runs produce a BENCHMARK.md, it ships alongside the skill so consumers can see verifiable benchmark uplift data.
This repository adheres to the Agent Skills specification:
SKILL.md file at their root.name and description fields.skills-ref reference library.Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
This code is dual-licensed with documentation/skills under the CC-BY-4.0 AND source code under Apache-2.0 license terms. The full license texts can be found in LICENSE-APACHE and LICENSE-CC-BY-4.0 respectively.
name: tao-launch-workflow
description: >-
The mandatory pre-launch gate and four-verb execution contract for every TAO
workflow or action. Invoke BEFORE launching anything side-effecting — AutoML,
train, evaluate, inference, export, TensorRT engine generation, or
DEFT/application workflows — on any execution platform. Covers platform
selection, credentials, image confirmation, dataset intake, preflight, the
launch review, job records, monitoring, and failure/retry classification.
Trigger phrases include "train this model", "run AutoML", "launch on
SLURM/docker/k8s/brev/virtualenv", "evaluate my checkpoint", "start a TAO
job".
license: Apache-2.0
compatibility: Requires the packaged TAO skill bank helper scripts.
metadata:
author: NVIDIA Corporation
version: "0.1.1"
allowed-tools: Read Bash
tags:
- tao
- workflow
- launchStandalone install? If this session was not initialized by the TAO skill bank plugin, run the
tao-setupskill first (host preflight, credentials, cross-skill discovery).
Use this skill before launching any TAO workflow or model action.
Run the platform helper, ask for platform and monitoring preferences, then run the selected platform detail helper before asking for credentials.
This gate is model-agnostic. Apply it to every TAO model, data action, and application workflow before launching side-effecting work.
Do not create runner scripts, launch scripts, compatibility shims, workspace folders, state files, logs, or dependency-install side effects until the launch preflight passes.
Preflight passes only after all of these are true:
image=<override>.If any item is missing, ask for the missing input and stop before generating artifacts. This applies to AutoML, normal train/eval/infer/export/TRT, and DEFT/application workflows.
When preflight work clears a blocker, keep track of the original user request. After the fix, rerun the relevant preflight and continue toward that request; do not stop at "blocker fixed" unless the user explicitly asked only for the repair.
Once the launch gate passes and the producing model/data skill has authored the
spec-bundle (schema: tao-artifacts), execution is exactly four verbs. Every
platform skill implements them over its native CLI — the bank ships five
(tao-run-on-docker, -slurm, -kubernetes, -brev, -virtualenv), and any
externally installed platform skill joins the same contract (§ External
platform skills); nothing else is platform-specific.
$BANK = ${TAO_SKILL_BANK_PATH}.
tao-data-io only on a
frame mismatch (remote URIs, cross-host paths, PTM fetches, tier-C result
uploads). Then lint the assembled command with redact_secrets.py lint and
open the record and launch, in that order:
JOB_ID=$("$BANK/scripts/tao_job_record.py" open --platform <p> --image <img> \
--network-arch <arch> --action <action> --storage-tier <A|B|C> --results-root <root>)
# <native launch, naming the backend object after $JOB_ID>
"$BANK/scripts/tao_job_record.py" mark "$JOB_ID" --state RUNNING --backend-ref <ref>
PENDING RUNNING COMPLETE ERROR CANCELED UNKNOWN; the native sub-state
(ImagePullBackOff, PENDING-resources, slurm COMPLETING) rides in the
transition message. Never read "what's running" from records — poll the backend.mark <id> --state CANCELED.Record-then-launch is the ordering invariant. open mints the id and binds
results_dir before any launch, and the id it returns is the only handle the
launch can use — a submit that skipped the gate or the open has no id, so it
cannot launch. This is what keeps a run recoverable across a context break:
results_dir is recorded before the backend object (which K8s TTL or docker
--rm may later delete) ever exists.
When the producing spec-bundle declares execution, preserve it as model-owned
action semantics across every application that reuses that model skill. The
selected platform consumes the lifecycle; an application must not copy its
commands into a private launcher. Platform-independent pre/post commands,
runtime attestations, helper dependencies, distributed intent, and completion
evidence belong in the producer's spec-bundle. Scheduler syntax, mounts,
secrets, timeouts, ranks, and child-exit preservation remain platform-owned.
No registry, no interface file: a platform skill declares the contract by
documenting the four verbs, and you verify by reading before first use. A
skill with only native primitives may be used by inferring the mapping
(bank invariants still bind; the mapping goes in the launch review; persist
what worked). Rules and the no-equivalent hard floor:
references/external-platforms.md.
When status reaches ERROR, read the log tail and classify before any
retry — infrastructure faults are retriable (new record, --retry-of, up
to 10), program faults never are. Full criteria, the two judgment calls
(device-side asserts, downstream tracebacks), and the post-turn poller
rules: references/failure-analysis-retry.md.
After the user confirms what they want to do, ask which execution platform
should run it. Discover the choices from the platform skills installed in this
session — you already see them by name and description (tao-run-on-docker,
-slurm, -kubernetes, -brev, -virtualenv, plus any externally installed one such as the
official brev-cli skill). There is no central platform registry to read. If
your runtime surfaces only the core router skills (e.g. Codex), list the bank's
platform skills by reading skills/platform/tao-run-on-*/SKILL.md frontmatter
(name + one-line description) under ${TAO_SKILL_BANK_PATH}.
Then ask:
Use long_running_enabled=true and status_interval_minutes=5 when the user
accepts the defaults.
When monitoring is enabled, do not send a final summary just because several
polls have elapsed or the job is still PENDING. Keep the turn attached and
emit status every status_interval_minutes until a terminal state or explicit
user stop/detach request. If the runtime environment cannot keep the chat turn
open, say that clearly and leave a durable watcher/log path; do not imply that
chat updates will continue after the turn ends.
Final-answer rule: a final response ends chat-side monitoring. While
long_running_enabled=true and any launched job is non-terminal, status
messages must be sent as in-progress updates and the agent must continue
polling. Only send a final response when the workflow reaches terminal state,
the user explicitly asks to detach/stop monitoring, or the runtime genuinely
cannot keep the turn open; in that last case, say it is a runtime limitation
and provide the exact durable status command/log path.
When intake inputs are missing, ask with the exact prompt shape in
references/intake-prompts.md (one consolidated ask, concrete examples,
no invented defaults).
After model ownership resolution, inspect the selected model's
references/skill_info.yaml. If it declares backend_contracts, resolve the
implementation before selecting an image or authoring a spec. An explicit
backend wins when it supports the model/action; otherwise apply the packaged
backend_selection policy and show its rationale. The selected backend
metadata in skill_info.yaml owns its image. The referenced backend contract
owns the entrypoint, configuration schema, data mappings, topology, checkpoint
format, output layout, and status behavior. Never use a legacy top-level image
fallback for a multi-backend frontend, and never treat one backend as a version
of another.
Pass action, backend, and workload hints to the model resolver. When metadata
declares a backend planner, use it. The shared Cosmos frontend, for example,
uses scripts/cosmos_workflow.py plan to generate backend-native TOML and a
launch sequence.
Before creating specs, runner scripts, workspaces, logs, state files, or submitting a job, resolve the image for the selected model/action:
${TAO_SKILL_BANK_PATH:-~/tao-skill-bank}/scripts/resolve_tao_image.py \
--skill-bank ${TAO_SKILL_BANK_PATH:-~/tao-skill-bank} \
--model <network> --action <action> --backend <auto-or-explicit> \
--workload <workload-hint> --format text
If the helper is unavailable, read skills/models/<network>/config.json
directly. Resolve image fields in this order:
backend_contracts.<selected-backend>.container_image, when presentactions.<action>.container_imageactions.<action>.imagecontainer_imageimageShow the exact image and ask:
Container image for <network>/<action>:
default=<resolved image>
Use this image, or provide image=<override>?
If the user accepts, pass the resolved image as the job image. If the user
overrides, require a non-empty image reference and pass that value instead.
Do not silently launch on the default image. This confirmation applies to
training, AutoML recommendations, evaluation, inference, export, TensorRT
engine generation, and application workflows that submit TAO containers.
After the user chooses a platform, get the credential list for only that
platform from the chosen skill itself — its ## Credentials section and, if
present, references/skill_info.yaml (required_credentials, credential_groups,
optional_credentials). The launch preflight (check_tao_launch_preflight.py)
reads that same per-skill skill_info.yaml to enforce the credential gate; a
credential-free platform (e.g. Docker) may ship only prose, in which case rely on
its Preflight section.
Ask only for credentials that platform actually needs, plus model-specific
credentials from the selected model skill. Do not ask for Brev credentials on
SLURM, Kubernetes, or Docker. Do not ask for SLURM credentials on Brev,
Kubernetes, or Docker. Ask S3 credentials only when the selected
platform and the dataset/result URIs require s3:// access.
Credentials may already be present in the process environment or in a
user-approved secret env file such as ~/.tao/secrets.env or
~/.config/tao/.env; source such files only when needed and never print,
grep, cat, paste, or log their contents. Verify only variable presence.
For initial launch intake, ask for required credentials and required credential
groups only. Treat the helper's optional credentials/settings section as
reference material; do not request those values unless their only_when
condition applies, the selected workflow cannot proceed without them, or the
user asks to customize that setting.
When the helper output includes a "Required credential groups" section, satisfy one credential from each group before proceeding. Explain each requested value using the helper's description and "How to get it" text.
For SLURM, user-facing prompts should ask for SSH_KEY_PATH first. Mention
SSH_AUTH_SOCK only if the user says they already use an SSH agent.
If a required CLI/library is missing, say exactly what is missing and why it is needed, then ask before installing. Examples:
aws.After user approval and installation, rerun the same preflight. Do not create runner files or launch jobs between the failed check and the rerun.
Accept dataset inputs in either mode:
custom.train_dataset.annotation_path=<root>/annotations.json and
custom.train_dataset.media_path=<root>.custom.train_dataset.annotation_path=<TRAIN_ANNOTATION_PATH>
and custom.train_dataset.media_path=<TRAIN_MEDIA_PATH>.Ask for dataset examples that match the selected platform:
s3://bucket/path/train and
s3://bucket/path/eval unless the platform profile mounts shared storage./data/tao/<model>/train, or direct spec paths visible inside the planned
container mount.DOCKER_HOST, not paths on the local agent machine.Do not assume "dataset root" is the only acceptable input. When direct spec paths are supplied, validate the exact spec paths rather than appending default filenames.
Run the selected platform's preflight checks before any launch artifact is
created — prefer the packaged helper scripts/check_tao_launch_preflight.py
(--platform <p> --container-image <img> --path <label>=<path> ...). It verifies
credentials, client tools, platform/cluster/object-store access, dataset paths
from the compute frame, GPU/runtime health, and image-architecture fit; treat any
failure as blocking. Never use --skip-platform-access for a real launch.
See references/platform-preflight.md for the full per-platform detail (SLURM
SSH/key setup + resource defaults, docker/remote-docker GPU + bind-mount checks,
Brev/Kubernetes API + object-store checks, annotation content-field checks, and
data staging).
Before any side-effecting launch, show a concise review:
For AutoML, also show the algorithm, metric/direction, recommendation budget,
search parameters, ranges, and generated/default recommendation details as
described in skills/applications/tao-run-automl/SKILL.md. Ask for confirmation after
this review. If the user supplied a time limit, flag any plan that exceeds it
and offer concrete reductions before launch.
Never end a successful launch review with only “nothing was launched.” End
with one direct action prompt, for example: Ready to materialize the sealed plan and submit the job. Reply "launch", "go ahead", or "yes" to proceed.
The next unambiguous affirmative chat message authorizes materialization,
job-record creation, submission, and the previously reviewed monitoring mode;
execute immediately without another intake or confirmation round.
When the model contract declares a structured status path or metric extractor, poll it alongside the native backend. Scheduler/container completion is not a successful training result by itself: require the model's terminal structured success record, collect concrete checkpoint events, and return final train loss plus every epoch validation-complete loss. Do not promote validation heartbeat/batch metrics or a train-loss line to epoch validation loss. If the process fails before its native logger exists, invoke the packaged status finalizer or report the real process exit failure; use raw log parsing only as a fallback.
评论 (0)
暂无评论,成为第一个评论者吧!