SkillAtlasSkill 详情

doca-dpa-hl-tracer

Official, NVIDIA-verified Agent Skills for Claude Code, Codex, and other coding agents.

审核状态:已审核Quality 72Security 90

复制安装命令

用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。

复制前请先查看来源、License 和安全提示。

项目 README

来源文件:README.md

抓取于 2026年8月9日

NVIDIA Agent Skills

Official, NVIDIA-verified Agent Skills for Claude Code, Codex, and other coding agents.

NVIDIA Agent Skills Spec License

📖 Docs: docs.nvidia.com/skills  ·  📺 Livestream: From Vulnerable to Verified  ·  📝 Blog: NVIDIA Verified Agent Skills: Capability Governance for AI Agents


Skills are portable instruction sets that teach AI agents how to use NVIDIA software optimally: Physical AI and robotics workflows, simulation, CUDA-X libraries, RAG and AI Blueprints, and platform tools. This repository is a catalog: skills are maintained in their respective product repos, and mirrored here daily via an automated sync pipeline. Skills are being added continuously, so check back for updates. We are building this infrastructure in the open, and contributions are welcome. See the Roadmap for what is planned next.


Quickstart

Install NVIDIA skills with the default skills CLI flow:

npx skills add nvidia/skills

The CLI runs through npx and prompts you to choose a skill and install destination. You do not need to clone this repo or copy skill folders by hand.

Requires a current skills CLI (v1.5.16 or newer). Installing via npx skills@latest add nvidia/skills always uses the latest. On older CLIs (v1.5.15 and earlier), skills may install but not appear in Claude Code — see Troubleshooting.

The skill is available the next time your agent loads skills and encounters a relevant task. For example, ask your agent to "solve a linear programming problem with cuOpt" and the skill guides it through the cuOpt Python API. In Claude Code, run /reload-skills to load newly installed skills in your current session.

Install One Skill Without Prompts

Use this when you already know the skill name and want to skip prompts.

npx skills add nvidia/skills --skill cuopt-numerical-optimization-api --yes

Replace cuopt-numerical-optimization-api with any skill name from the Skill Catalog.

Install for a Specific Agent

Use --agent to target a specific AI coding agent. Initially, we'll support common client targets, expanding the list over time. For the full list of clients supported by the spec, see the skills CLI Supported Agents table.

Claude Code

npx skills add nvidia/skills --skill cuopt-numerical-optimization-api --agent claude-code

Codex

npx skills add nvidia/skills --skill cuopt-numerical-optimization-api --agent codex

Snowflake CoCo

npx skills add nvidia/skills --skill cuopt-numerical-optimization-api --agent cortex

Cursor

npx skills add nvidia/skills --skill cuopt-numerical-optimization-api --agent cursor

Kiro

npx skills add nvidia/skills --skill cuopt-numerical-optimization-api --agent kiro-cli

Use --agent more than once to install the same skill into multiple agents.

npx skills add nvidia/skills \
  --skill cuopt-numerical-optimization-api \
  --agent claude-code \
  --agent codex \
  --agent cursor \
  --agent kiro-cli

Keep Skills Up to Date

New skills land continuously, and existing ones are revised, renamed, or consolidated as the catalog evolves. Refresh what you have installed with:

npx skills update

Run it interactively and the CLI also flags skills that were removed or merged upstream (for example, when several skills are consolidated into one) and offers to remove the stale local copies. Use npx skills list to see what is installed and npx skills check to preview what is out of date first.

Browse the Catalog

Use this when you want to see available NVIDIA skills before installing anything.

npx skills add nvidia/skills --list

For non-interactive installs, global installs, agent-specific installs, updates, removals, and fallback manual copying, see Advanced installation.


Skill Catalog

ProductDescriptionSkills
AIQNVIDIA AI-Q Blueprint - deploy local AI-Q services and run shallow or deep research workflows as agent skills.aiq-research, aiq-deploy
CUDA-QCUDA Quantum — onboarding guide for installation, test programs, GPU simulation, QPU hardware, and quantum applications.cudaq-guide
cuDFOfficial NVIDIA-authored guidance for NVIDIA cuDF GPU DataFrames, pandas acceleration, dask-cuDF, ETL, joins, groupby, CSV/Parquet I/O, nullable semantics, and multi-GPU DataFrame workloads.accelerated-computing-cudf
cuOptGPU-accelerated optimization — vehicle routing, linear programming, quadratic programming, installation, server deployment, and developer tools.cuopt-install, cuopt-multi-objective-exploration, cuopt-numerical-optimization-api, cuopt-numerical-optimization-formulation, cuopt-routing-api-python, cuopt-server-api-python
cuPyNumericNumPy and SciPy on multi-node multi-GPU systems — skills to help with installing cuPyNumeric, migrating existing NumPy code, and doing parallel I/Ocupynumeric-hdf5, cupynumeric-install, cupynumeric-migration-readiness, cupynumeric-parallel-data-load
DALIGPU-accelerated data loading and processing with NVIDIA DALI.dali-dynamic-mode
Data DesignerBuild declarative synthetic dataset generation pipelines with NeMo Data Designer.data-designer
DeepStreamAgentic skills for guided DeepStream development.amc-run-sample-calibration, amc-run-video-calibration, amc-setup-calibration-stack, deepstream-dev, deepstream-generate-pipeline, deepstream-import-vision-model, deepstream-profile-pipeline, deepstream-sop
Digital HealthAgent skills for the clinical ASR evaluation flywheel — term curation, synthetic clinical-speech benchmark generation, KER (Keyword Error Rate) scoring, and fine-tune guidance.digital-health-clinical-asr-setup, digital-health-clinical-asr-build, digital-health-clinical-asr-eval, digital-health-clinical-asr-finetune
DOCATeach AI agents to use the NVIDIA DOCA SDK on BlueField DPUs and ConnectX NICs — setup, libraries, services, tools, deployment, and debugging.doca-bare-metal-deployment, doca-bf3-deployment, doca-bf4-deployment, doca-collectx-deployment, doca-container-deployment, doca-debug, doca-hardware-safety, doca-programming-guide, doca-public-knowledge-map, doca-setup, doca-structured-tools-contract, doca-upgrade, doca-version, doca-aes-gcm, doca-argp, doca-comch, doca-common, doca-compress, doca-devemu, doca-dma, doca-dpa, doca-dpdk-bridge, doca-erasure-coding, doca-eth, doca-flow, doca-flow-dpa-provider, doca-gpi, doca-gpunetio, doca-mgmt, doca-pcc, doca-pcc-ztr-rttcc-algo, doca-rdma, doca-rdmi, doca-rmax, doca-sha, doca-sta, doca-telemetry, doca-telemetry-exporter, doca-urom, doca-verbs, doca-argus, doca-dms, doca-firefly, doca-urom-svc, doca-bench, doca-bench-extension, doca-caps, doca-comm-channel-admin, doca-dpa-hl-tracer, doca-flow-dpa-perf, doca-flow-grpc-server, doca-flow-perf, doca-flow-tune, doca-gpunetio-ib-write-bw, doca-gpunetio-ib-write-lat, doca-pcc-counters, doca-sha-offload-engine, doca-socket-relay, doca-spcx-cc, doca-telemetry-utils
DynamoNVIDIA Dynamo deployment bring-up on Kubernetes — pick and deploy recipes, start router modes, validate disagg NIXL/UCX/NCCL interconnect, and triage day-2 failures.dynamo-interconnect-check, dynamo-recipe-runner, dynamo-router-starter, dynamo-troubleshoot
Earth2StudioOpen-source deep-learning framework for exploring, building and deploying AI weather/climate workflows.earth2studio-create-datasource, earth2studio-create-diagnostic, earth2studio-create-prognostic, earth2studio-data-fetch, earth2studio-deterministic-forecast, earth2studio-discover, earth2studio-install
HoloHubBuild, run, debug, benchmark, and develop HoloHub applications and Holoscan Modules with validated lifecycle workflows.holohub-app-lifecycle, holohub-debug-build-run, holohub-module-lifecycle
Holoscan SDKInstall and set up the Holoscan SDK on any platform (container, Debian, Python, Conda, or source).holoscan-install-debian, holoscan-install-source, holoscan-install-wheel, holoscan-install-conda, holoscan-install-container, holoscan-setup
Holoscan Sensor BridgeAgent-ready skills for Holoscan Sensor Bridge devkit workflows, including demo environment bring-up, FPGA flashing for Lattice and VB1940 hardware, example application execution, QA test-plan automation, and support for configuring and using the Holoscan Sensor Bridge FPGA intellectual property (IP) core.hsb-setup, hsb-flash, hsb-app, hsb-test, hsb-ip-def, hsb-ip-packetizer, hsb-ip-create-top
Isaac for Healthcare WorkflowsAgent-ready skills for Isaac for Healthcare agentic and catheter-navigation workflows, covering task authoring, data pipelines, policy training and validation, CT-derived digital twins, DRR rendering, and interactive catheter simulation.i4h-workflow, i4h-workflow-setup, i4h-workflow-create, i4h-workflow-scene-edit, i4h-workflow-dataset-teleop, i4h-workflow-dataset-replay, i4h-workflow-dataset-mimic, i4h-workflow-dataset-annotate, i4h-workflow-dataset-convert, i4h-workflow-finetune, i4h-workflow-validate, i4h-workflow-e2e, i4h-lerobot-viz, i4h-catheter-navigation, i4h-catheter-navigation-setup, i4h-catheter-navigation-digital-twin, i4h-catheter-navigation-render-drr, i4h-catheter-navigation-viewport, i4h-catheter-navigation-smoke, i4h-catheter-navigation-e2e
Jetson BSPAgentic skills for setting up and customizing an NVIDIA Jetson Linux Board Support Package (BSP) — pick a target, prepare image and sources, customize IO (camera, PCIe, USB, pinmux, clocks, and more), then promote, flash, and validate.jetson-build-source, jetson-customize-camera, jetson-customize-clocks, jetson-customize-fan, jetson-customize-mgbe, jetson-customize-nvpmodel, jetson-customize-pcie, jetson-customize-pinmux, jetson-customize-uphy, jetson-customize-usb, jetson-derive-carrier, jetson-download-bsp, jetson-flash-image, jetson-generate-kb, jetson-init-image, jetson-init-source, jetson-init-target, jetson-link-docs, jetson-optimize-memory, jetson-print-bsp-info, jetson-promote-image, jetson-quick-start, jetson-set-target, jetson-validate-image
Jetson DeviceDevice-side agent skills for working with a live NVIDIA Jetson after boot — diagnostics, memory auditing, headless setup, inference memory tuning, LLM serving and benchmarking, packaging guidance, and speculative decoding.jetson-diagnostic, jetson-headless-mode, jetson-inference-mem-tune, jetson-llm-benchmark, jetson-llm-serve, jetson-memory-audit, jetson-package, jetson-print-device-info, jetson-speculative-decoding
Medical AI SkillsAgent-ready medical AI skills built on MONAI for DICOM handling, NVIDIA-hosted medical imaging model workflows, segmentation, synthesis, and evidence-oriented evaluation.dicom-metadata-extract, dicom-series-preflight, dicom-series-to-volume, nv-generate-ct-rflow, nv-generate-mr, nv-generate-mr-brain, nv-generate-mr-brain-finetune, nv-generate-vae-finetune, nv-reason-cxr, nv-segment-ct, nv-segment-ct-finetune, nv-segment-ctmr
Megatron-CoreLarge-scale distributed training — model parallelism, pipeline parallelism, and mixed precision.mcore-create-issue, mcore-linting-and-formatting, mcore-run-on-slurm, mcore-split-pr, mcore-testing
NeMo AutoModelNeMo AutoModel - PyTorch-native distributed training for LLMs/VLMs with Hugging Face support, recipes, launchers, and validation workflows.nemo-automodel-distributed-training, nemo-automodel-launcher-config, nemo-automodel-model-onboarding, nemo-automodel-recipe-development
NeMo MBridgeNeMo MBridge - PyTorch-native bridge between Hugging Face and Megatron-Core for checkpoint conversion, training recipes, and NVIDIA GPU performance workflows.nemo-mbridge-mlm-bridge-training, nemo-mbridge-multi-node-slurm, nemo-mbridge-perf-activation-recompute, nemo-mbridge-perf-cpu-offloading, nemo-mbridge-perf-cuda-graphs, nemo-mbridge-perf-expert-parallel-overlap, nemo-mbridge-perf-hierarchical-context-parallel, nemo-mbridge-perf-megatron-fsdp, nemo-mbridge-perf-memory-tuning, nemo-mbridge-perf-moe-comm-overlap, nemo-mbridge-perf-moe-dispatcher-selection, nemo-mbridge-perf-moe-hardware-configs, nemo-mbridge-perf-moe-long-context, nemo-mbridge-perf-moe-optimization-workflow, nemo-mbridge-perf-moe-vlm-training, nemo-mbridge-perf-parallelism-strategies, nemo-mbridge-perf-sequence-packing, nemo-mbridge-perf-tp-dp-comm-overlap, nemo-mbridge-recipe-recommender, nemo-mbridge-resiliency
NeMo PlatformNeMo Platform brings NVIDIA NeMo libraries together under one CLI, Python SDK, and web UInemo-evaluator-plugin, nemo-data-designer-plugin
NeMo RelaySkills to help get started and use NeMo Relay - a runtime for instrumenting and controlling AI agents across harnesses, applications, and frameworks.nemo-relay-install, nemo-relay-get-started, nemo-relay-instrument-calls, nemo-relay-instrument-context-isolation, nemo-relay-instrument-typed-wrappers, nemo-relay-plugin-adaptive-tuning, nemo-relay-plugin-build, nemo-relay-plugin-observability, nemo-relay-migrate-from-flow, nemo-relay-debug-runtime-integration
NeMo RetrieverNeMo Retriever - deploy NeMo Retriever Library locally, extract information from corpus of data, and answer questions against the corpus.nemo-retriever
NeMo-RLRLHF training on Ray — GRPO, DPO, and SFT for LLMs and VLMs with FSDP2 and Megatron-Core.launch-nemo-rl, nemo-rl-auto-research, nemo-rl-brev-etiquette, nemo-rl-docs, nemo-rl-session-memory
NemoClawSecure agent sandboxing — run OpenClaw inside NVIDIA OpenShell with managed inference, policy management, remote deployment, sandbox monitoring.nemoclaw-user-guide
NemotronAuthor end-to-end model development, customization, evaluation, and deployment pipelines using the NVIDIA AI stack.nemotron-customize, nemotron-retrieval-recipes, nemotron-policy-generator
Nemotron SpeechDeploy and operate NVIDIA Nemotron Speech (Riva) NIMs — ASR, TTS, and NMT, cloud-hosted via build.nvidia.com or self-hosted on your own GPU.nemotron-speech, nemotron-asr-finetune
Physical AIPhysical AI skills for simulation, synthetic data generation, training, validation and deployment and more.omniverse-cad-to-simready, omniverse-realtime-viewer, omniverse-usd-performance-tuning, physical-ai-infrastructure-setup-and-resilient-scaling, physical-ai-neural-reconstruction, physical-ai-defect-image-generation, physical-ai-video-data-augmentation, physical-ai-people-attribute-search
PhysicsNeMoNVIDIA PhysicsNeMo - Open-source deep-learning framework for building, training, and fine-tuning deep learning models using state-of-the-art Physics-ML methods.physicsnemo-discover, physicsnemo-shard-tensor
Portfolio OptimizationGPU-accelerated Mean-CVaR portfolio optimization with NVIDIA cuOpt — CVaR optimization, efficient frontier, scenario generation, backtesting, and rebalancing.portfolio-optimization
RAG BlueprintRAG pipeline — deploy, configure, troubleshoot, and manage retrieval augmented generation with Docker Compose or Helm.rag-blueprint, rag-eval, rag-perf
Skill Card GeneratorReads an agent skill's source files and produces a skill card plus a review table. Use when a skill directory exists and a governance card needs to be generated or updated.skill-card-generator
TAO ToolkitNVIDIA TAO Toolkit - fine-tune and optimize 100+ pretrained vision AI models with your own data using low-code microservices, then export production-ready models for edge or cloud deployment.tao-analyze-changenet-rca, tao-finetune-huggingface-model, tao-port-huggingface-model, tao-run-automl, tao-run-automl-deft-pipeline, tao-run-deft-aoi, tao-run-inference-service, tao-train-single-step, paidf-anomalygen, tao-analyze-gaps-visual-changenet, tao-analyze-gaps-vlm-bcq, tao-convert-dataset-format, tao-generate-image-grounding, tao-generate-referring-expressions, tao-generate-video-reasoning-annotations, tao-mine-aoi-images, tao-route-visual-changenet-samples, tao-validate-dataset-format, tao-finetune-clip, tao-finetune-cosmos-embed, tao-finetune-cosmos-reason, tao-train-action-recognition, tao-train-bevfusion, tao-train-centerpose, tao-train-deformable-detr, tao-train-depth-anything-v2, tao-train-dino, tao-train-fast-foundation-stereo, tao-train-foundation-stereo, tao-train-grounding-dino, tao-train-image-classification, tao-train-mask-auto-encoder, tao-train-mask-auto-label, tao-train-mask-grounding-dino, tao-train-mask2former, tao-train-metric-learning-recognition, tao-train-nvdinov2, tao-train-nvpanoptix3d, tao-train-ocdnet, tao-train-ocrnet, tao-train-oneformer, tao-train-optical-inspection, tao-train-pointpillars, tao-train-pose-classification, tao-train-reid, tao-train-rtdetr, tao-train-segformer, tao-train-sparse4d, tao-train-visual-changenet, tao-run-on-brev, tao-run-on-docker, tao-run-on-kubernetes, tao-run-on-local-docker, tao-run-on-slurm, tao-run-platform, tao-setup-nvidia-gpu-host, tao-launch-workflow, tao-list-capabilities
TileGymTile-based GPU programming — adding new kernels, cross-framework conversion, and performance optimization.tilegym-adding-cutile-kernel, tilegym-converting-cutile-to-julia, tilegym-converting-cutile-to-triton, tilegym-cutile-autotuning, tilegym-cutile-python, tilegym-improve-cutile-kernel-perf, tilegym-monkey-patch-kernels-to-transformers
Video Search and SummarizationVSS Blueprint — deploy profiles, search and summarize video, generate analysis reports, manage alerts and incidents, query VIOS sensors, and use the RTVI VLM microservice.vss-ask-video, vss-deploy-dense-captioning, vss-deploy-detection-tracking-2d, vss-deploy-detection-tracking-3d, vss-deploy-profile, vss-deploy-video-embedding, vss-generate-video-calibration, vss-generate-video-report, vss-manage-alerts, vss-manage-video-io-storage, vss-query-analytics, vss-search-archive, vss-setup-behavior-analytics, vss-setup-video-analytics-api, vss-summarize-video

Getting Help & Contributing

Where to file an issue depends on what's broken:

  • Skill content issues (a specific skill has a bug, missing functionality, or incorrect content) — file in the source repo for that product, using the per-product table below.
  • Catalog issues (catalog README errors, sync workflow problems, distribution channels, signing/verification flow, docs in this repo) — file here using the catalog issue templates: Bug Report, Feature Request, or Documentation Request or Correction.
  • Questions or general discussion — use Discussions. The issue tracker is reserved for bug reports, feature proposals with a design, and documentation issues.
  • Security vulnerabilities — follow the disclosure process in SECURITY.md; do not open a public issue.

Per-product source repo links:

ProductIssuesDiscussionsContributingSecurity
AIQIssuesDiscussionsContributingSecurity
CUDA-QIssuesDiscussionsContributingSecurity
cuDFIssuesDiscussionsContributingSecurity
cuOptIssuesDiscussionsContributingSecurity
cuPyNumericIssues—Contributing—
DALIIssues—Contributing—
Data DesignerIssuesDiscussionsContributingSecurity
DeepStreamIssues—ContributingSecurity
Digital HealthIssues—ContributingSecurity
DOCAIssues—ContributingSecurity
DynamoIssuesDiscussionsContributingSecurity
Earth2StudioIssuesDiscussionsContributing—
HoloHubIssues—ContributingSecurity
Holoscan SDKIssues—ContributingSecurity
Holoscan Sensor BridgeIssues—Contributing—
Isaac for Healthcare WorkflowsIssues—ContributingSecurity
Jetson BSPIssues—ContributingSecurity
Jetson DeviceIssues—ContributingSecurity
Medical AI SkillsIssues—ContributingSecurity
Megatron-CoreIssuesDiscussionsContributing—
NeMo AutoModelIssuesDiscussionsContributingSecurity
NeMo MBridgeIssuesDiscussionsContributingSecurity
NeMo PlatformIssuesDiscussionsContributingSecurity
NeMo RelayIssuesDiscussionsContributingSecurity
NeMo RetrieverIssuesDiscussionsContributingSecurity
NeMo-RLIssuesDiscussionsContributingSecurity
NemoClawIssuesDiscussionsContributingSecurity
NemotronIssuesDiscussionsContributingSecurity
Nemotron SpeechIssues—ContributingSecurity
Physical AIIssues—ContributingSecurity
PhysicsNeMoIssuesDiscussionsContributingSecurity
Portfolio OptimizationIssuesDiscussionsContributingSecurity
RAG BlueprintIssuesDiscussionsContributingSecurity
Skill Card GeneratorIssues—ContributingSecurity
TAO ToolkitIssuesDiscussionsContributingSecurity
TileGymIssues—ContributingSecurity
Video Search and SummarizationIssuesDiscussionsContributingSecurity

For issues with this catalog repo itself (README, structure, listing a new product): open an issue here.


Verifying Skills

Every published skill ships with a detached OMS signature (skill.oms.sig). The sync pipeline drops any skill missing the required artifacts before publishing, so every skill in the catalog carries:

  • SKILL.md — the skill instructions consumed by the agent
  • skill-card.md — skill identity and governance card
  • skill.oms.sig — detached OMS signature (verifiable against nv-agent-root-cert.pem)
  • A Tier-3 evaluation dataset — accepted at evals/evals.json, evals/*.json, eval/*.json, or benchmark/evals.json
  • BENCHMARK.md — generated benchmark report capturing verifiable uplift data

Verify a skill against the NVIDIA trust anchor nv-agent-root-cert.pem:

pip install model-signing
model_signing verify certificate SKILL_DIR \
  --signature SKILL_DIR/skill.oms.sig \
  --certificate_chain nv-agent-root-cert.pem \
  --ignore_unsigned_files

A successful verification confirms that the skill contents have not been modified since signing by NVIDIA.

See Verify Signed Agent Skills for signature layout, the trust pipeline, and policy options.


Roadmap

  • ✅ Public skills catalog with NVIDIA-verified skills across multiple products
  • ✅ Automated sync pipeline with skills mirrored from product repos daily
  • ✅ Security scanning for all published skills covering instruction safety and supply-chain integrity
  • ✅ Skills signing so every published skill carries a verifiable NVIDIA signature
  • ✅ Skills universal evaluation criteria and task-specific criteria
  • ✅ Skill Card with machine-readable metadata for identity, provenance, quality, and behavioral boundaries
  • ✅ Sync-time compliance gates — signature drift detection and missing-artifact enforcement
  • ✅ Syndication to external marketplaces — Skills.sh, Codex plugin, Claude Code plugin, ClawHub, Hermes Hub
  • 🔲 Syndication to additional MCP hubs and partner channels

Repository Structure

NVIDIA/skills/
├── skills/                      # NVIDIA-verified skills (count grows continuously),
│   │                              synced from upstream product repos
│   ├── README.md                 # Browser-facing install guidance
│   ├── <product-prefix>-*/       # Flat layout — one dir per skill, product-prefixed
│   │                               # e.g. aiq-*, cuopt-*, cupynumeric-*,
│   │                               # dali-*, deepstream-*, dicom-*, digital-health-*,
│   │                               # dynamo-*, earth2studio-*, holoscan-*, hsb-*,
│   │                               # jetson-*, launch-nemo-rl, mcore-*,
│   │                               # nemo-automodel-*, nemo-data-designer-plugin,
│   │                               # nemo-evaluator-plugin, nemo-mbridge-* (20 skills),
│   │                               # nemo-retriever, nemo-rl-* (4 skills),
│   │                               # nemoclaw-user-guide, nemotron-*, nemotron-speech,
│   │                               # nv-* (medical AI), physicsnemo-*, rag-*,
│   │                               # skill-card-generator, tao-*, tilegym-*,
│   │                               # vss-* (15 skills), accelerated-computing-cudf,
│   │                               # cudaq-guide, portfolio-optimization
│   ├── omniverse-*/              # Physical AI — manually staged (see manual-components.yml)
│   └── physical-ai-*/            # Physical AI — manually staged
├── components.d/                # Product registry — one file per component, teams onboard here
│   ├── README.md                 # Schema and onboarding instructions
│   └── <product>.yml             # one file per registered product
├── plugins/                     # Packaged plugin distributions
│   └── nvidia-skills/            # Curated NVIDIA skills bundle (Claude Code, Codex)
├── plugins.d/                   # Plugin build registry — config for `build-plugins.py`
│   ├── README.md
│   ├── _defaults.yml
│   └── nvidia-skills.yml
├── .claude-plugin/              # Claude Code marketplace metadata
│   └── marketplace.json
├── .agents/plugins/             # Agent marketplace metadata (other clients)
│   └── marketplace.json
├── docs/                        # Long-form documentation (published via Fern)
│   ├── README.md                 # How to build the docs locally
│   ├── index.mdx
│   ├── advanced-install.mdx
│   ├── agent-skill-trust-pipeline.mdx
│   ├── release-checklist.mdx
│   ├── scanning-agent-skills.mdx
│   ├── signing-agent-skills.mdx
│   └── skill-cards.mdx
├── fern/                        # Fern docs site configuration
├── .github/
│   ├── workflows/                # Sync pipeline, plugin validation, DCO check, author verify
│   └── scripts/                  # regenerate-readme.sh, build-plugins.py,
│                                 # manual-components.yml (temp Physical AI catalog
│                                 # exception, removed after Computex 2026),
│                                 # marketplace/metadata.json (skill metadata sidecar)
├── nv-agent-root-cert.pem       # Trust anchor for OMS signature verification
├── skills.sh.json               # Skills.sh marketplace grouping config
├── CHANGELOG.md
├── CONTRIBUTING.md              # Contribution guidelines
├── SECURITY.md                  # Security reporting policy
├── CODE_OF_CONDUCT.md           # Community code of conduct
├── LICENSE-APACHE               # Apache 2.0 (source code)
└── LICENSE-CC-BY-4.0            # CC BY 4.0 (documentation/skills)

Skills are maintained in their respective product repos (see the Source column in the Skill Catalog) and synced to this repo daily. Products only appear under skills/ after the sync pipeline confirms each skill carries:

  • skill.oms.sig — detached OMS-format signature (verifiable against nv-agent-root-cert.pem)
  • skill-card.md — skill identity and governance card
  • A Tier-3 evaluation dataset — accepted at evals/evals.json, evals/*.json, eval/*.json, or benchmark/evals.json

When evaluation runs produce a BENCHMARK.md, it ships alongside the skill so consumers can see verifiable benchmark uplift data.


Standards & Compatibility

This repository adheres to the Agent Skills specification:

  • Skills are portable directories with a SKILL.md file at their root.
  • Metadata uses YAML frontmatter with required name and description fields.
  • Skills follow a progressive disclosure model — lightweight metadata loads at startup, full instructions load on activation.
  • Validate your skill using the skills-ref reference library.

License

Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.

This code is dual-licensed with documentation/skills under the CC-BY-4.0 AND source code under Apache-2.0 license terms. The full license texts can be found in LICENSE-APACHE and LICENSE-CC-BY-4.0 respectively.

其他

低风险

  • 来源需自行核对维护者身份。
  • 未检测到明显脚本安装指令。
  • 可能需要外部 token、网络权限或第三方服务。
  • 未检测到高风险命令。
  • 扫描发现:0 条。

Codex — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/NVIDIA/skills.git
  3. 将 "skills/doca-dpa-hl-tracer" 文件夹复制到 Codex 的 skills 目录中。
  4. 重启 Codex 让新的 skill 生效。

Codex — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Codex 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Codex 让新的 skill 生效。

Claude Code — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/NVIDIA/skills.git
  3. 将 "skills/doca-dpa-hl-tracer" 文件夹复制到 Claude Code 的 skills 目录中。
  4. 重启 Claude Code 让新的 skill 生效。

Claude Code — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Claude Code 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Claude Code 让新的 skill 生效。

Cursor — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/NVIDIA/skills.git
  3. 将 "skills/doca-dpa-hl-tracer" 文件夹复制到 Cursor 的 skills 目录中。
  4. 重启 Cursor 让新的 skill 生效。

Cursor — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Cursor 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Cursor 让新的 skill 生效。

GitHub Copilot — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/NVIDIA/skills.git
  3. 将 "skills/doca-dpa-hl-tracer" 文件夹复制到 GitHub Copilot 的 skills 目录中。
  4. 重启 GitHub Copilot 让新的 skill 生效。

GitHub Copilot — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 GitHub Copilot 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 GitHub Copilot 让新的 skill 生效。

Windsurf — Git Clone 安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 克隆仓库:git clone https://github.com/NVIDIA/skills.git
  3. 将 "skills/doca-dpa-hl-tracer" 文件夹复制到 Windsurf 的 skills 目录中。
  4. 重启 Windsurf 让新的 skill 生效。

Windsurf — 手动复制安装

  1. 安装前请先查看来源仓库和风险报告。
  2. 从源仓库下载 SKILL.md 及相关文件。
  3. 在 Windsurf 的 skills 目录中创建新文件夹。
  4. 将所有 skill 文件复制到新文件夹中。
  5. 重启 Windsurf 让新的 skill 生效。
查看 SKILL.md 原文
license: Apache-2.0
name: doca-dpa-hl-tracer
description: >
  Use this skill when the user runs doca_dpa_hl_tracer to
  capture/decode DPA-side traces at the programming-events
  layer (kernel entry/exit, sync points, comm primitive
  calls, RDMA WR submission, completion drain) — picking
  TRACE vs CRIT, tuning the JSON config (file-size limits
  + file_size_limit_policy, thread priorities/cores),
  decoding against the matching DPA-side ELF, or
  diagnosing empty/noisy captures. Trigger even when the
  user does not explicitly mention "DOCA DPA tracer" or
  "high-level tracer" — typical implicit phrasings include
  "DPA kernel returns wrong result but host completions
  look clean", "kernel-entry to first-comm latency is
  huge", "RDMA WR to drain gap on the DPA", "trace file
  truncated mid-run", "TRACE doubled my DPA latency", or
  "tracer wrote a file but parser shows zero events".
  Refuse and route elsewhere for writing DPA kernels,
  DPA-Comms/DPA-Verbs programming, raw per-cycle DPA
  profiling, host-side doca-dpa debugging, or production
  DPA telemetry — those belong to other skills.
metadata:
  kind: tool
compatibility: >
  Requires DOCA SDK installed at /opt/mellanox/doca on
  Linux (Ubuntu 22.04/24.04 or RHEL/SLES) with a BlueField
  device whose DPA processor is exposed to the host, plus
  the DOCA DPA Tools optional component (binary at
  /opt/mellanox/doca/tools/doca_dpa_hl_tracer). Requires a
  DPACC-built DPA-side ELF and a live doca-dpa-launched
  workload for events to fire.

DOCA DPA High-Level Tracer

Where to start: This is a tool skill for invoking doca_dpa_hl_tracer — the documented host-side CLI that captures DPA-side execution traces in higher-level terms (DPA programming events: kernel entry / exit, sync points, comm primitive calls, RDMA WR submission, completions) rather than raw cycle counts. Open TASKS.md and start at ## configure for the mode-vs-overhead decision and the JSON config layout, then ## run for the capture → decode → render pipeline. Open CAPABILITIES.md when the question is what does this tool actually trace, which DPA programming events does it expose, what is the trace-overhead vs fidelity tradeoff, or how does it slot into a DPA debug loop alongside doca-dpa and doca-debug. If DPA is not the right surface for the user's question (e.g. the bug is host-side, the bug is in the DPACC-produced image, the user wants raw cycle counts), the path-selection rule in CAPABILITIES.md ## Capabilities and modes routes the agent before any capture is attempted.

Example questions this skill answers well

The CLASSES of doca_dpa_hl_tracer questions this skill is built to answer, each with one worked example. The class is the load-bearing piece; the worked example is one instance.

  • "My DPA kernel is doing the wrong thing — where do I look?" — worked example: "my host-side doca_dpa_kernel_launch_update_* completes, but the kernel's reported result is wrong; no host-side DOCA_ERROR_*". Answered by the when DPA-side high-level tracing is the right surface gate in CAPABILITIES.md ## Capabilities and modes
  • "My DPA kernel is slow at a granularity that doesn't show up in cycle profiles — how do I see kernel-entry to first-comm-call latency?" — worked example: "my DPA kernel runs but the time between launch and the first RDMA WR submission is bigger than I expected". Answered by the event-taxonomy table in CAPABILITIES.md ## Capabilities and modes
    • the iterative loop in TASKS.md ## test which treats trace overhead, mode (TRACE vs CRIT), and capture window as axes to tune.
  • "How do I capture a trace without burying the DPA in observation overhead?" — worked example: "TRACE mode is producing too much data and my measured DPA latency went up by 2x compared to without the tracer". Answered by the mode-vs-overhead tradeoff in CAPABILITIES.md ## Capabilities and modes
    • the CRIT-first guidance in TASKS.md ## configure (start with critical-events-only; widen to TRACE only when the bug demands per-event detail).
  • "My trace file got truncated mid-run — how should I configure the file-size limits?" — worked example: "binary trace file hit 5 GB and the capture stopped". Answered by the log_file_max_size_in_bytes / bin_file_max_size_in_bytes / file_size_limit_policy triple in CAPABILITIES.md ## Capabilities and modes
  • "Is the tracer on my install, and is it paired with the matching doca-dpa library and DPACC compiler version?" — worked example: "is the tracer ABI on my install compatible with the DPA image my DPACC just produced?". Answered by the version-overlay in CAPABILITIES.md ## Version compatibility, which redirects to the canonical doca-version chain and adds the tracer ↔ doca-dpa library ↔ DPACC compiler match rule.
  • "The capture file looks empty / decode failed — is the install broken, no events fired, or am I tracing the wrong thing?" — worked example: "doca_dpa_hl_tracer ran, wrote a file, but the parser shows zero events". Answered by the layered error taxonomy in CAPABILITIES.md ## Error taxonomy (install / device-binding / DPA-image-instrumented / capture-window / decode-vs-elf / overhead-saturated / version / cross-cutting) + the layered walk in TASKS.md ## debug.

Audience

This skill serves external developers, platform operators, and AI agents who have already brought up a DPA-side workload through doca-dpa and now need higher-level visibility into what the DPA kernel is actually doing on the wire — DPA programming events ordering, sync gaps, comm-call latencies, RDMA-WR / completion timing — without dropping all the way down to raw cycle counters. Concretely:

  • A DPA developer who can launch their kernel cleanly from the host side but whose kernel's result is wrong or whose DPA-side performance is below expectation, and who needs a DPA-side ground truth before triaging.
  • A platform operator running a DPA-using workload (RDMA offload from accelerator, custom CC algorithm via doca-pcc) and needs to localize a regression to the DPA side without instrumenting the application.
  • An AI agent producing a DPA-side trace report as evidence for the host-side doca-dpa TASKS.md ## debug ladder when the host side reports clean completions but the DPA-side behaviour is wrong.

It is not for users debugging the tracer binary itself, not a substitute for the live public DOCA DPA Tools guide, not the right place for users learning how to write a DPA kernel (that audience belongs in doca-dpa plus the public DOCA DPA / DPACC / DPA-Comms / DPA-Verbs guides), and not the right place for raw per-instruction cycle profiling (different surface, different tool — route via doca-public-knowledge-map ## DOCA tools).

The tracer is shipped as a CLI binary under /opt/mellanox/doca/tools/, not a library you link against. The skill uses the same kind: tool three-file shape as the rest of the bundle so the agent's task-verb contract is uniform across libraries, services, and tools.

Language scope

doca_dpa_hl_tracer is a C++ host-side CLI. Its inputs are its JSON config file, the DPA-side ELF (the doca_dpa_app-class image produced by DPACC), and a running DPA-side workload that the host-side doca-dpa lifecycle already started. Its outputs are a binary trace file (bin_file) and a human-readable log file (log_file). The skill keeps the workflow guidance language-neutral — the DPA-side workload it traces can be C compiled by DPACC or any other DPA translation unit DPACC accepts — and routes per-language questions to the public DPA / DPACC guides via doca-public-knowledge-map.

When to load this skill

Load this skill when the user is — or the agent needs to — invoke doca_dpa_hl_tracer on a real host with DOCA installed against a BlueField with a DPA processor visible to the host, and the host-side doca-dpa lifecycle has already brought a DPA workload up at least once. Concretely:

  • Capturing a DPA-side trace to localize a DPA kernel's wrong-result or wrong-ordering behaviour when the host-side doca-dpa lifecycle reports clean completions.
  • Capturing a DPA-side trace to localize a DPA-side performance gap (kernel-entry to first-comm latency, RDMA-WR-issue to completion gap, sync-point dwell time) at a granularity above raw cycle counts.
  • Choosing between TRACE and CRIT capture modes based on the bug-vs-overhead tradeoff and the available capture window.
  • Tuning the JSON config (thread priorities, core affinities, file size limits, file-size-limit policy) so the capture itself does not perturb the workload more than the bug it is investigating.
  • Decoding a captured bin_file against the matching DPA-side ELF to render the human-readable event stream.
  • Capturing a side-effect-bounded trace as prerequisite evidence for a host-side doca-dpa TASKS.md ## debug ladder step.

Do not load this skill for general DOCA orientation, DPA-side programming model questions, raw cycle profiling, or DOCA / DPACC install. For those, route to doca-public-knowledge-map, doca-dpa, or doca-setup.

What this skill provides

This is a thin loader. Substantive material lives in two companion files:

  • CAPABILITIES.md — what doca_dpa_hl_tracer captures: the DPA programming event taxonomy (kernel entry / exit, sync points, comm primitive calls, RDMA WR submission and completion drain), the two documented capture modes (TRACE for full per-event, CRIT for critical-events only), the trace-overhead-vs-fidelity tradeoff, the config-file shape (receiver / binary-writer / file-writer / printer threads with priority + core affinity, file size limits, file_size_limit_policy), the capture-window + workload-must-be-running invariant, the ELF-must-match-image rule for decode, the version-availability overlay (tracer ↔ doca-dpa library ↔ DPACC compiler), the layered error taxonomy (install / device-binding / image-instrumented / capture-window / decode / overhead-saturated / version / cross-cutting), the observability surface (binary trace file + log file + tool's own stderr), and the safety policy (capture is bounded; tracing is not a production observability surface).
  • TASKS.md — step-by-step workflows for the in-scope task verbs: install (route to host-side DOCA install + DPA prerequisites), configure (mode + JSON config layout + capture window), build (route to install — the binary is shipped, the DPA-side application is user-built by DPACC), modify (refuse — do not patch the binary; modify the JSON config and the invocation instead), run (the capture flow with --mode, --config-file, --output-file), test (iterative loop tuning mode, window, and overhead), debug (walk the error taxonomy), use (consume the decoded trace in a doca-dpa debug session), plus a Deferred task verbs block.

The skill assumes a host where DOCA is already installed at the standard location, a BlueField with a DPA processor is present and visible to the host, the DPACC compiler is installed at a version matched to the host-side DOCA, the DPA-side application image (the ELF the tracer decodes against) is on disk and matches what the doca-dpa lifecycle loaded, and the operator has the privileges the public DOCA DPA Tools guide requires.

What this skill deliberately does not ship

This skill is agent guidance, not a samples or scripts bundle. To keep the boundary clean, it deliberately does not contain — and pull requests should not add:

  • Specific flag strings, event names, or mode tokens beyond what the public DOCA DPA Tools page and --help document. The DPA programming events surface evolves release to release; --help on the installed binary is the authoritative inventory.
  • Pre-baked example traces or expected event timings. Trace output is workload-, DPA-image-, BlueField-, and firmware-specific; a captured example pinned to one setup misleads operators elsewhere.
  • Wrappers, parsers, or rendering scripts in any language that consume the binary trace format. The format is documented; users who want to script against it should read the live guide and write the parser against their installed version.
  • A specific tuning recommendation derived from a single trace. A DPA-side perf decision (move a sync, batch a comm call, change a launch argument) is a workload question and the skill prescribes how to capture and read traces — it refuses to translate a captured gap into a kernel-rewrite recommendation without the user's own analysis.
  • A samples/ or reference/ subtree. This is a thin loader for a shipped CLI; substantive material lives on the public page, in --help, and in doca-dpa.

Loading order

  1. Read this SKILL.md first to confirm the user's question is in scope (DPA-side high-level tracing, not DPA-side programming and not raw cycle profiling).
  2. For the event taxonomy, capture modes, overhead tradeoff, JSON config layout, version overlay, error taxonomy, observability, and safety policy, see CAPABILITIES.md.
  3. For the documented invocations and the capture → decode → render workflow — install, configure, build, modify, run, test, debug, use — see TASKS.md.

Related skills

  • doca-dpa — the host-side DPA control library whose loaded application image the tracer captures. Pair them in every DPA debug session: doca-dpa brings the workload up; the tracer captures what the workload does at the DPA programming event layer. Conflating the library with the tracer is the most common DPA-debug first-touch error.
  • doca-debug — the cross-cutting debug ladder. The tracer slots in at the runtime layer as the DPA-side ground truth before any DPA-side perf or correctness conclusion is made.
  • doca-public-knowledge-map — routing to the public DOCA DPA Tools page on docs.nvidia.com and the rest of the public DOCA documentation set.
  • doca-version — canonical DOCA version-handling rules. The ## Version compatibility section in CAPABILITIES.md is a concise overlay that redirects here for the body and adds the tracer ↔ doca-dpa library ↔ DPACC compiler matching rule.
  • doca-setup — env preparation, install verification, DPACC compiler install / verification, BlueField mode (the DPA processor must be exposed before any tracing is meaningful), and the I have no install yet path with the public NGC DOCA container.
  • doca-structured-tools-contract — the bundle's detect → prefer → fall back → report contract for structured helper tools. The command appendix in TASKS.md honors this contract.
  • doca-programming-guide — general DOCA programming patterns shared by every library / tool surface, including the cross-library DOCA_ERROR_* taxonomy this tool's host-side error layer overlays on top of when host-side doca-dpa calls fail in tandem.

The DPA-side companion libraries doca-dpa-comms (comm primitives the DPA kernel itself calls) and doca-dpa-verbs (RDMA verbs the DPA kernel itself calls) are different artifacts that the tracer's DPA programming events surface visibly names; for the DPA-side programming model itself, route through doca-public-knowledge-map to the public DOCA DPA-Comms and DPA-Verbs guides and to the shipped /opt/mellanox/doca/samples/doca_dpa/ samples. This tool traces their use; it does not redefine them.

发现问题?提交给管理员复核

评分:

评论 (0)

暂无评论,成为第一个评论者吧!