复制安装命令
用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。
复制前请先查看来源、License 和安全提示。
Official, NVIDIA-verified Agent Skills for Claude Code, Codex, and other coding agents.
用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。
复制前请先查看来源、License 和安全提示。
来源文件:README.md
Official, NVIDIA-verified Agent Skills for Claude Code, Codex, and other coding agents.
📖 Docs: docs.nvidia.com/skills · 📺 Livestream: From Vulnerable to Verified · 📝 Blog: NVIDIA Verified Agent Skills: Capability Governance for AI Agents
Skills are portable instruction sets that teach AI agents how to use NVIDIA software optimally: Physical AI and robotics workflows, simulation, CUDA-X libraries, RAG and AI Blueprints, and platform tools. This repository is a catalog: skills are maintained in their respective product repos, and mirrored here daily via an automated sync pipeline. Skills are being added continuously, so check back for updates. We are building this infrastructure in the open, and contributions are welcome. See the Roadmap for what is planned next.
Install NVIDIA skills with the default skills CLI flow:
npx skills add nvidia/skills
The CLI runs through npx and prompts you to choose a skill and install destination. You do not need to clone this repo or copy skill folders by hand.
Requires a current
skillsCLI (v1.5.16 or newer). Installing vianpx skills@latest add nvidia/skillsalways uses the latest. On older CLIs (v1.5.15 and earlier), skills may install but not appear in Claude Code — see Troubleshooting.
The skill is available the next time your agent loads skills and encounters a relevant task. For example, ask your agent to "solve a linear programming problem with cuOpt" and the skill guides it through the cuOpt Python API. In Claude Code, run /reload-skills to load newly installed skills in your current session.
Use this when you already know the skill name and want to skip prompts.
npx skills add nvidia/skills --skill cuopt-numerical-optimization-api --yes
Replace cuopt-numerical-optimization-api with any skill name from the Skill Catalog.
Use --agent to target a specific AI coding agent. Initially, we'll support common client targets, expanding the list over time. For the full list of clients supported by the spec, see the skills CLI Supported Agents table.
Claude Code
npx skills add nvidia/skills --skill cuopt-numerical-optimization-api --agent claude-code
Codex
npx skills add nvidia/skills --skill cuopt-numerical-optimization-api --agent codex
Snowflake CoCo
npx skills add nvidia/skills --skill cuopt-numerical-optimization-api --agent cortex
Cursor
npx skills add nvidia/skills --skill cuopt-numerical-optimization-api --agent cursor
Kiro
npx skills add nvidia/skills --skill cuopt-numerical-optimization-api --agent kiro-cli
Use --agent more than once to install the same skill into multiple agents.
npx skills add nvidia/skills \
--skill cuopt-numerical-optimization-api \
--agent claude-code \
--agent codex \
--agent cursor \
--agent kiro-cli
New skills land continuously, and existing ones are revised, renamed, or consolidated as the catalog evolves. Refresh what you have installed with:
npx skills update
Run it interactively and the CLI also flags skills that were removed or merged upstream (for example, when several skills are consolidated into one) and offers to remove the stale local copies. Use npx skills list to see what is installed and npx skills check to preview what is out of date first.
Use this when you want to see available NVIDIA skills before installing anything.
npx skills add nvidia/skills --list
For non-interactive installs, global installs, agent-specific installs, updates, removals, and fallback manual copying, see Advanced installation.
Where to file an issue depends on what's broken:
Per-product source repo links:
For issues with this catalog repo itself (README, structure, listing a new product): open an issue here.
Every published skill ships with a detached OMS signature (skill.oms.sig). The sync pipeline drops any skill missing the required artifacts before publishing, so every skill in the catalog carries:
SKILL.md — the skill instructions consumed by the agentskill-card.md — skill identity and governance cardskill.oms.sig — detached OMS signature (verifiable against nv-agent-root-cert.pem)evals/evals.json, evals/*.json, eval/*.json, or benchmark/evals.jsonBENCHMARK.md — generated benchmark report capturing verifiable uplift dataVerify a skill against the NVIDIA trust anchor nv-agent-root-cert.pem:
pip install model-signing
model_signing verify certificate SKILL_DIR \
--signature SKILL_DIR/skill.oms.sig \
--certificate_chain nv-agent-root-cert.pem \
--ignore_unsigned_files
A successful verification confirms that the skill contents have not been modified since signing by NVIDIA.
See Verify Signed Agent Skills for signature layout, the trust pipeline, and policy options.
NVIDIA/skills/
├── skills/ # NVIDIA-verified skills (count grows continuously),
│ │ synced from upstream product repos
│ ├── README.md # Browser-facing install guidance
│ ├── <product-prefix>-*/ # Flat layout — one dir per skill, product-prefixed
│ │ # e.g. aiq-*, cuopt-*, cupynumeric-*,
│ │ # dali-*, deepstream-*, dicom-*, digital-health-*,
│ │ # dynamo-*, earth2studio-*, holoscan-*, hsb-*,
│ │ # jetson-*, launch-nemo-rl, mcore-*,
│ │ # nemo-automodel-*, nemo-data-designer-plugin,
│ │ # nemo-evaluator-plugin, nemo-mbridge-* (20 skills),
│ │ # nemo-retriever, nemo-rl-* (4 skills),
│ │ # nemoclaw-user-guide, nemotron-*, nemotron-speech,
│ │ # nv-* (medical AI), physicsnemo-*, rag-*,
│ │ # skill-card-generator, tao-*, tilegym-*,
│ │ # vss-* (15 skills), accelerated-computing-cudf,
│ │ # cudaq-guide, portfolio-optimization
│ ├── omniverse-*/ # Physical AI — manually staged (see manual-components.yml)
│ └── physical-ai-*/ # Physical AI — manually staged
├── components.d/ # Product registry — one file per component, teams onboard here
│ ├── README.md # Schema and onboarding instructions
│ └── <product>.yml # one file per registered product
├── plugins/ # Packaged plugin distributions
│ └── nvidia-skills/ # Curated NVIDIA skills bundle (Claude Code, Codex)
├── plugins.d/ # Plugin build registry — config for `build-plugins.py`
│ ├── README.md
│ ├── _defaults.yml
│ └── nvidia-skills.yml
├── .claude-plugin/ # Claude Code marketplace metadata
│ └── marketplace.json
├── .agents/plugins/ # Agent marketplace metadata (other clients)
│ └── marketplace.json
├── docs/ # Long-form documentation (published via Fern)
│ ├── README.md # How to build the docs locally
│ ├── index.mdx
│ ├── advanced-install.mdx
│ ├── agent-skill-trust-pipeline.mdx
│ ├── release-checklist.mdx
│ ├── scanning-agent-skills.mdx
│ ├── signing-agent-skills.mdx
│ └── skill-cards.mdx
├── fern/ # Fern docs site configuration
├── .github/
│ ├── workflows/ # Sync pipeline, plugin validation, DCO check, author verify
│ └── scripts/ # regenerate-readme.sh, build-plugins.py,
│ # manual-components.yml (temp Physical AI catalog
│ # exception, removed after Computex 2026),
│ # marketplace/metadata.json (skill metadata sidecar)
├── nv-agent-root-cert.pem # Trust anchor for OMS signature verification
├── skills.sh.json # Skills.sh marketplace grouping config
├── CHANGELOG.md
├── CONTRIBUTING.md # Contribution guidelines
├── SECURITY.md # Security reporting policy
├── CODE_OF_CONDUCT.md # Community code of conduct
├── LICENSE-APACHE # Apache 2.0 (source code)
└── LICENSE-CC-BY-4.0 # CC BY 4.0 (documentation/skills)
Skills are maintained in their respective product repos (see the Source column in the Skill Catalog) and synced to this repo daily. Products only appear under skills/ after the sync pipeline confirms each skill carries:
skill.oms.sig — detached OMS-format signature (verifiable against nv-agent-root-cert.pem)skill-card.md — skill identity and governance cardevals/evals.json, evals/*.json, eval/*.json, or benchmark/evals.jsonWhen evaluation runs produce a BENCHMARK.md, it ships alongside the skill so consumers can see verifiable benchmark uplift data.
This repository adheres to the Agent Skills specification:
SKILL.md file at their root.name and description fields.skills-ref reference library.Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
This code is dual-licensed with documentation/skills under the CC-BY-4.0 AND source code under Apache-2.0 license terms. The full license texts can be found in LICENSE-APACHE and LICENSE-CC-BY-4.0 respectively.
license: Apache-2.0
name: doca-gpunetio-ib-write-bw
description: >
Use this skill when the user is building, running, or interpreting
the doca/tools/gpunetio_ib_write_bw client+server benchmark — a CUDA
kernel on the client posts RDMA WRITE work requests through the
doca-gpunetio device-side surface to measure sustained GPU-driven
WRITE bandwidth on a GPU+IB-device pair. Trigger even when the user
does not explicitly mention "doca-gpunetio-ib-write-bw" or
"GPUNetIO" — typical implicit phrasings include "measure WRITE BW
when the GPU posts the WRs", "BW swings between runs on the same
flags", "is the NIC saturated or am I CPU-bound on the CUDA
kernel", "meson compile fails for the GPUNetIO bw tool",
"nvidia_peermem isn't picking up my GPU buffer", or "GPU-initiated
WRITE throughput vs CPU-initiated perftest". Refuse and route
elsewhere for general doca-gpunetio library work, DOCA install, the
GPU-initiated WRITE latency analog, the CPU-initiated upstream
perftest, or application-level end-to-end throughput — those belong
to other skills.
metadata:
kind: tool
compatibility: >
Requires DOCA SDK on Linux with a BlueField DPU or ConnectX NIC,
NVIDIA GPU, CUDA toolkit and nvcc, loaded `nvidia_peermem`, and an
InfiniBand RNIC paired with the GPU. Uses `pkg-config` for
doca-gpunetio, doca-rdma, and doca-common, plus the installed
gpunetio_ib_write_bw sources. Run only on a trusted, non-shared IB
fabric during the benchmark window.Where to start: This is a tool skill for the GPUNetIO-
flavored ib_write_bw benchmark shipped under
doca/tools/gpunetio_ib_write_bw/ (a client + server pair,
built from source against the installed DOCA via meson).
It measures sustained RDMA WRITE bandwidth when the WRs are
posted from a CUDA kernel through the doca-gpunetio
device-side surface, with the GPU on the data path. Open
TASKS.md and start at
## configure for the GPU-NIC
pairing precondition and the build pattern; jump to
## run for the smoke-before-bulk flow.
Open CAPABILITIES.md when the question
is what this tool actually measures, how the result
decomposes (GPU occupancy vs NIC issue rate vs link
saturation), or how the result reads against the GPI
sister tool and the upstream CPU-initiated perftest
ib_write_bw. If DOCA is not installed yet, route to
doca-setup first; if the
user is still deciding between the GPI and GPUNetIO
programming surfaces, the picture in
../../libs/doca-gpunetio/CAPABILITIES.md#capabilities-and-modes
and
../../libs/doca-gpi/CAPABILITIES.md#capabilities-and-modes
is the first stop.
The CLASSES of doca-gpunetio-ib-write-bw questions this
skill is built to answer, each with one worked example. The
class is the load-bearing piece; the worked example is one
instance.
CAPABILITIES.md ## Capabilities and modes
TASKS.md ## configure +
TASKS.md ## run. The same shape
answers "measure GPUNetIO-driven WRITE BW between a
host GPU and a BlueField DPU".CAPABILITIES.md ## Observability
TASKS.md ## test.perftest ib_write_bw?" — worked example:
"my team has a CPU-initiated WRITE BW number on this
same NIC; should I expect the GPUNetIO number to match
or be different?". Answered by the "GPU-initiated
path adds (or removes) overhead vs the CPU-initiated
path" rule in
CAPABILITIES.md ## Capabilities and modes.CAPABILITIES.md ## Capabilities and modes
TASKS.md ## use.CAPABILITIES.md ## Error taxonomy
layer 5 + the steady-state guidance in
TASKS.md ## test.gpunetio_ib_write_bw even link?".
Answered by the version overlay in
CAPABILITIES.md ## Version compatibility
which cross-links the canonical detection chain in
doca-version.This skill serves external developers and performance engineers who need a reproducible measurement of sustained RDMA WRITE bandwidth when the WRs are posted from a CUDA kernel through doca-gpunetio, on the user's actual install and GPU-NIC pair. Concretely:
perftest-style path before
committing an application design to one of them.It is not for users debugging the doca-gpunetio
library itself (route to
../../libs/doca-gpunetio/SKILL.md),
and not a substitute for the perftest upstream
ib_write_bw (which measures CPU-initiated WRITE BW).
The doca-gpunetio-ib-write-bw tool is shipped as C plus
a CUDA .cu translation unit under
doca/tools/gpunetio_ib_write_bw/, split into a client/
subtree and a server/ subtree. The verified surface (per
client/{main.c,common.h,common.c,kernel.cu,perftest.c} and
server/{main.c,common.h,common.c,perftest.c}): host-side
build via meson against the installed DOCA pkg-config
modules (doca-gpunetio, doca-rdma, doca-common); the
device-side build via nvcc against the DOCA GPU NetIO
device-side header set; the OOB descriptor exchange via a
TCP socket between client and server. There is no Python /
Rust / Go binding — the tool is a pair of CLI binaries.
The skill's job is to keep the operator-side workflow
language-neutral; the device-side CUDA surface is not
wrappable in another language.
Load this skill when the user is — or the agent needs to —
build and run the gpunetio_ib_write_bw client + server on
real hosts with DOCA installed plus a CUDA Toolkit matched
to the DOCA install, and a GPU + IB device pair on the
host's PCIe topology. Concretely:
doca-gpi
library — doca/tools/ ships no GPI benchmark binary) or
the classic CPU-initiated perftest path.Do not load this skill for general DOCA orientation,
library API work, or installation. For those, use
doca-public-knowledge-map,
../../libs/doca-gpunetio/SKILL.md,
or doca-setup. Do not load
it for application-level end-to-end throughput either —
this benchmark measures the WR-submission path through
GPUNetIO, not the user's full pipeline.
This is a thin loader. Substantive material lives in two companion files:
CAPABILITIES.md — what the tool measures (the
sustained-WRITE-BW primitive driven by a client-side
CUDA kernel through doca-gpunetio), the
runtime-surface selection rule (GPUNetIO vs GPI vs
CPU-initiated), the GPU-NIC pairing precondition, the
throughput-decomposition guide (GPU compute occupancy
vs NIC issue rate vs link saturation), the version
overlay (DOCA .pc PLUS CUDA Toolkit), the layered
error taxonomy (config-syntax / build-time / GPU-NIC-
pairing / GPUNetIO-lifecycle / RDMA-connection /
measurement-soundness / version / cross-cutting), the
observability surface (stdout report, DOCA log levels,
OOB-socket exchange), and the safety overlay (the
"GPU-side handle is a credential" rule from
doca-gpunetio; the cross-cutting hardware-safety
meta-policy).TASKS.md — step-by-step workflows for the in-scope
task verbs: install (preconditions — DOCA install,
CUDA Toolkit, GPU + NIC pair, OOB connectivity),
configure (build-tree under
doca/tools/gpunetio_ib_write_bw/ and the meson
build wrapping the shipped DOCA), build (the
meson setup + meson compile pattern from the
public DOCA build documentation), modify (do not
patch the shipped tool source; modify the invocation
and the surrounding environment instead), run (smoke-
before-bulk; client + server bring-up order; reading
the per-iteration report), test (the eval loop —
steady-state, NUMA placement, NIC saturation cross-
check), debug (walk the error taxonomy layer by
layer), use (how a BW result feeds a class-of-
workload decision), plus a Deferred task verbs
block routing out-of-scope questions.The skill assumes a host where DOCA is already installed,
a CUDA Toolkit matched to the install is present, and the
operator has whatever privileges the public install profile
expects for binding a doca_dev, a doca_gpu, and an OOB
TCP socket.
This skill is agent guidance, not a samples or scripts bundle. To keep the boundary clean, it deliberately does not contain — and pull requests should not add:
--help and main.c ARGP
registration establish. The flag surface is small
(device name, GPU PCIe address, GID index, server IP on
the client side); the agent re-reads the binary's
--help on the installed version before quoting flag
strings. Throughput numbers are device-, firmware-,
version-, and topology-specific.client/{main.c,kernel.cu,perftest.c,common.{c,h}}
and server/{main.c,perftest.c,common.{c,h}} files are
the verified worked example; the agent's job is to
route the user there and prescribe minimum-diff
modification per the universal modify-a-sample workflow
in
doca-programming-guide.CAPABILITIES.md ## Observability;
if the user wants to script against it, the right
answer is "read the live source, write the parser
against your installed binary".samples/, bindings/, or reference/ subtree.
This is a thin loader for a shipped tool tree;
substantive material lives in the source tree and in
the GPUNetIO library docs.SKILL.md first to confirm the user's
question is in scope (the user actually wants to
measure sustained kernel-initiated WRITE BW through
GPUNetIO, not learn GPUNetIO as a library or do a
CPU-initiated measurement).perftest, the throughput-decomposition guide, the
version overlay, the error taxonomy, the observability
surface, and the safety overlay, see
CAPABILITIES.md.install, configure,
build, modify, run, test, debug, use — see
TASKS.md.../../libs/doca-gpunetio/SKILL.md —
the library this tool wraps. The per-GPU doca_gpu
context, the GPU-visible doca_gpu_eth_* and RDMA-side
handles, the CUDA-side persistent-kernel pattern, the
dual capability-discovery rule (DOCA cap-query AND
cudaGetDeviceProperties), and the env preconditions
(nvidia_peermem loaded, CUDA buffers registered with
DOCA) live there.../../libs/doca-rdma/SKILL.md —
the underlying RDMA library. The RDMA queue this tool
binds is created and connected via doca-rdma; the
queue lifecycle, transport type (RC vs UC vs UD),
permission matrix, and connection method are owned
there.../../libs/doca-verbs/SKILL.md —
the raw-verbs escape hatch beneath doca-rdma /
doca-gpunetio. This tool stays on the higher-level
surfaces; doca-verbs is the right place only if the
user needs a specific WR flag / QP attribute the
GPUNetIO + RDMA surfaces do not expose.../doca-gpunetio-ib-write-lat/SKILL.md —
the latency analog of this tool. Same physical
operation; same runtime framework; different metric
class (BW vs latency). The two together carry the
full GPUNetIO-side throughput / latency picture.doca-gpi — the GPI
programming surface (CUDA-kernel-initiated RDMA), the
alternative runtime framework for the same physical
operation. doca/tools/ ships no GPI ib_write_lat /
ib_write_bw benchmark binary, so the GPI comparison is
against the library surface, not a sibling tool. The
selection rule in
CAPABILITIES.md ## Capabilities and modes
is the decision aid.doca-version — the
canonical version-detection chain, four-way match rule,
NGC container semantics, and headers-win-over-docs
rule. The ## Version compatibility section in this
skill is a thin overlay; the body lives there.doca-setup — env
preparation, install verification, GPU + CUDA Toolkit
pairing, nvidia_peermem load, hugepages, NUMA, and
the I have no install yet path with the public NGC
DOCA container.doca-public-knowledge-map —
routing to the public DOCA documentation set (DOCA GPU
NetIO, DOCA RDMA pages on docs.nvidia.com) and the
docs.nvidia.com/cuda/ pointer for the CUDA Toolkit.doca-debug — the
cross-cutting debug ladder. The tool surfaces its own
error taxonomy; when the cause is below DOCA, the
taxonomy hands off here.doca-hardware-safety —
the bundle-wide hardware-safety meta-policy. The
## Safety policy overlay cross-links it.
评论 (0)
暂无评论,成为第一个评论者吧!