Doca Bench logo

Doca Bench

OrganizationPopular
NVIDIA
doca-bench

Run `doca_bench` (DOCA 2.7.0 or newer) to measure throughput, bulk latency, precision latency, or maximum bandwidth for RDMA, Compress, AES-GCM, SHA, DMA, EC, Ethernet, Comch, or GPUNetIO on a host or BlueField Arm. Use it to discover enabled benchmark libraries, capture a reproducible command/version/device/environment baseline, compare stable runs against a declared tolerance, or diagnose configuration, device-binding, workload-precondition, and measurement failures. Trigger for requests such as measuring BlueField compression speed, NIC RDMA throughput, crypto latency, or a pre-upgrade baseline. Do not use for application end-to-end timing, custom benchmark code, DOCA installation, or binary patches.

Overview

PublisherNVIDIA
Repositoryskills
Skill namedoca-bench
Stars
3.3K
Forks
397
Bundled files
7
LicenseApache-2.0
Links
  • Markdown instructions

    A SKILL.md file the model loads on demand, so it only costs tokens when a request actually matches.

  • Works with any LLM

    AI skills are plain Markdown, not provider-specific code, so this works with GPT, Claude, Gemini, Grok, or a local model.

  • 7 bundled files

    Scripts, templates, and references the model can read while it works. Files are read-only and never executed.

  • Open source

    Published by NVIDIA on GitHub. Read the source before you install it.

Installation

Install the Doca Bench AI skill in TypingMind to use it with any LLM, or drop it into another agent that reads SKILL.md.

1

Install in TypingMind

TypingMind installs a skill straight from its GitHub folder — it reads SKILL.md, bundles the resource files, and stores the result locally.

  1. Open the app and go to Plugins → Skills.
  2. Choose "Install from GitHub".
  3. Paste the skill folder URL below and confirm.
  4. Enable the skill in any chat where you want it available.
Plugins → Skills → Add skill → From GitHub URL, then paste the folder URL and press Continue.
2

Install in another agent

Any agent that reads the Agent Skills format can use this skill — copy the folder into that agent's skills directory.

Claude Code — .claude/skills
git clone --depth 1 https://github.com/NVIDIA/skills.git /tmp/skills
mkdir -p .claude/skills
cp -r /tmp/skills/skills/doca-bench .claude/skills/doca-bench
Restart Claude Code after copying so it picks up the new skill.

Use it in TypingMind

Enable Doca Bench in any TypingMind chat and the model takes it from there. Its name and description sit in the system prompt, and the moment a request matches, the model loads the full instructions itself — you never invoke it by hand, and it costs no tokens until it is actually used.

The model loads Doca Bench on its own as soon as a request matches it.

Works with any AI model

AI skills are plain Markdown instructions rather than provider-specific code, so Doca Bench is not tied to the model it was written for. Install it once in TypingMind and use it with GPT-5, Claude, Gemini, Grok, DeepSeek, Mistral, Llama, or a local model you run yourself — all on your own API keys.

  • Loaded only when it is needed

    The system prompt carries just the name and description. The instructions are fetched on the first matching request, so an idle skill costs nothing.

  • Switch models mid-chat

    Because the skill is instructions rather than code, changing model does not break it — the next model reads the same SKILL.md.

Skill instructions

This is the SKILL.md content the model loads. Read it before installing — a skill is instructions your model will follow.

DOCA Bench (doca_bench)

Where to start: This is a tool skill for invoking doca_bench, the cross-library micro-benchmark harness. Open TASKS.md and start at ## configure for the three-axis decision (target library × workload shape × measurement axis), then ## run for the smoke-before-bulk flow. Open CAPABILITIES.md when the question is what doca_bench can measure, which DOCA libraries it can drive, or how to interpret throughput / latency / op-rate output without fooling yourself on warm-up or steady-state. If DOCA is not installed yet, route to doca-setup first; if the install version is < 2.7.0, doca_bench is not shipped on this host.

Example questions this skill answers well

The CLASSES of doca_bench questions this skill is built to answer, each with one worked example. The class is the load-bearing piece; the worked example is one instance.

  • "What does this DOCA library actually deliver on this device?" — worked example: "throughput of DOCA Compress on my BlueField-3". Answered by the three-axis configuration in CAPABILITIES.md ## Capabilities and modes
    • the smoke-before-bulk flow in TASKS.md ## run. The same shape answers "send-side throughput of DOCA RDMA"doca_bench is cross-library, not single-library.
  • "Which DOCA libraries can doca_bench actually drive on this install?" — worked example: "is doca_sha enumerable on a granular-build install". Answered by the built-in query system surfaced in CAPABILITIES.md ## Capabilities and modes
    • TASKS.md ## configure step 2 (probe-before-bench). Empty enumeration = library not installed, not bench failure.
  • "Is this number reliable, or did I miss the warm-up?" — worked example: "why does my first-second number differ from my steady-state number". Answered by the measurement-soundness overlay in CAPABILITIES.md ## Error taxonomy layer 5 + TASKS.md ## test (the eval-loop overlay treats warm-up / steady-state / outliers as re-iteration triggers, not one-shot facts).
  • "Bench reports zero throughput / hangs at start / disagrees with the public docs." — worked example: "doca_bench shows zero ops for AES-GCM but doca_caps says the device supports it". Answered by the layered error taxonomy in CAPABILITIES.md ## Error taxonomy (config-syntax → device-binding → library-precondition → workload-precondition → measurement-soundness → version → cross-cutting) + TASKS.md ## debug.
  • "How do I capture a baseline I can later regression-test against?" — worked example: "snapshot decompress throughput on this BlueField + DOCA version before a firmware update". Answered by the CSV output + version-overlay rule in TASKS.md ## test (capture command line + version + device + as-deployed environment alongside the numbers; quoting numbers without the four-tuple is the cross-version regression-hunt failure mode).
  • "doca_bench returns nothing for library X — what does that mean?" — worked example: "empty output for DOCA SHA". Answered by the empty-output interpretation rules in TASKS.md ## debug + CAPABILITIES.md ## Error taxonomy. Re-route through doca-caps for the coarse per-device per-library capability ground truth, then back into bench once the capability is confirmed present.

Audience

This skill serves external operators, developers, and AI agents who need a reproducible, vendor-supported way to measure DOCA library performance on the user's actual install and device. Concretely:

  • An external developer choosing between DOCA libraries (e.g. COMPRESS vs SHA vs DMA throughput) before committing an application design.
  • A platform operator validating a tuning change (NUMA pinning, driver upgrade, firmware burn) by re-running a captured doca_bench baseline against the new state.
  • An SRE / performance engineer producing a "this is what the device delivers today" artifact that downstream consumers (capacity planning, regression bisection) can cite.
  • An AI agent answering "what throughput / latency should I expect from DOCA library X on device Y?" honestly — with a measured number, the command line that produced it, and the version + device + environment that scopes it — instead of guessing from datasheet headlines.

It is not for users debugging the doca_bench source code, and not a substitute for the live public DOCA Bench guide on docs.nvidia.com.

doca_bench is shipped as a tool (a single CLI binary plus a companion app for the remote half of remote-memory / RDMA / Eth scenarios), not a library you link against. The skill uses the same kind: tool three-file shape as the rest of the bundle so the agent's task-verb contract (configure / build / modify / run / test / debug) is uniform across libraries, services, and tools — even when individual verbs collapse to a routing stub for a shipped binary.

When to load this skill

Load this skill when the user is — or the agent needs to — invoke doca_bench on a real host with DOCA ≥ 2.7.0 installed (or inside the public NGC DOCA container with the equivalent version) to measure performance of a DOCA library. Concretely:

  • Picking which DOCA library to benchmark for a candidate workload (RDMA vs COMPRESS vs DMA, etc.).
  • Picking which measurement axis to ask for (throughput vs bulk latency vs precision latency vs max-bandwidth) — the four modes defined in tools/bench/doca_bench/configuration.hpp are not interchangeable.
  • Probing the install's granular-build state so the agent can honestly report "this library is not exposed on this install" instead of inventing a workload.
  • Capturing a documented baseline (command line + version + device
    • as-deployed environment + numbers) for later regression hunts.
  • Requiring the workload owner to predeclare acceptable variance and obtaining two consecutive runs within that tolerance before reporting a stable result; otherwise escalating the variance.
  • Diagnosing why a bench run reported zero / unstable / unexpected results (the error-taxonomy walk in TASKS.md ## debug).

Do not load this skill for general DOCA orientation, library API work, or installation. For those, use doca-public-knowledge-map, the matching libs/<library> skill, or doca-setup. Do not load it for application-level end-to-end benchmarking either — doca_bench measures the DOCA library surface, not the user's application above it.

What this skill provides

This is a thin loader. Substantive material lives in two companion files:

  • CAPABILITIES.md — what doca_bench can measure (the cross-library scope, the three-axis configuration model, the documented operating modes, the warm-up / pipeline / multi-core concepts that constrain measurement soundness), the version overlay (doca-bench-specific facts on top of the canonical doca-version rules), the layered error taxonomy (config-syntax / device-binding / library-precondition / workload-precondition / measurement-soundness / version / cross-cutting), the observability surface (screen + CSV output, real-time stats, query system), and the safety posture (the public guide's "not for production" warning, the host vs BlueField execution rule, the companion-app attack surface).
  • TASKS.md — step-by-step workflows for the in-scope task verbs: configure (the three-axis decision + the probe-before-bench step), build (route to install — the binary is shipped, the companion app is shipped), modify (refuse — do not patch the bench binary; modify the bench invocation instead), run (the smoke-before-bulk flow), test (the eval loop — warm-up, steady-state, outliers, cross-version), debug (walk the error taxonomy layer by layer), plus a Deferred task verbs block routing out-of-scope questions and a Command appendix of doca_bench-specific invocation classes.

The skill assumes a host where DOCA ≥ 2.7.0 is already installed (or the public NGC DOCA container is running at an equivalent version) and the operator has whatever permissions the public guide requires for doca_bench to bind devices and allocate resources on their platform.

What this skill deliberately does not ship

This skill is agent guidance, not a samples or scripts bundle. To keep the boundary clean, it deliberately does not contain — and pull requests should not add:

  • Specific flag strings or scenario / metric / attribute names beyond what the public DOCA Bench guide documents. The flag surface evolves and is install-specific; the documented invocations + --help on the installed version are the authoritative answer. Inventing a flag is the most common hallucination failure for this skill.
  • Pre-baked example output or expected throughput numbers. Bench output is device-, version-, firmware-, NUMA-, and tuning-specific. A captured number pinned to one platform and one DOCA version misleads operators on a different platform / version.
  • Wrappers, parsers, or scripts in any language that consume doca_bench CSV or stdout. The output formats are documented; if a user wants to script against them, the right answer is "read the live guide, write the parser against your installed version".
  • A samples/ or reference/ subtree. This is a thin loader for a documented CLI; substantive material lives on the public page and in --help.

Loading order

  1. Read this SKILL.md first to confirm the user's question is in scope (the user actually wants to invoke doca_bench for measurement, not learn about a DOCA library in general).
  2. For what doca_bench measures, the three-axis model, the version overlay, the error taxonomy, observability surface, and safety posture, see CAPABILITIES.md.
  3. For the documented invocations and the smoke-before-bulk workflow — configure, build, modify, run, test, debug — see TASKS.md.

Related skills

  • doca-public-knowledge-map — routing to the public DOCA Bench page on docs.nvidia.com and the rest of the public DOCA documentation set.
  • doca-version — the canonical version-detection chain, four-way match rule, NGC container semantics, and headers-win-over-docs rule. The ## Version compatibility section in this skill is a thin overlay on top of doca-version; the body lives there.
  • doca-structured-tools-contract — the bundle-wide contract for structured-output helper tools. Bench-runner / bench-snapshot executables that satisfy the detect-prefer-fallback-report loop are deferred to PR2; the contract is consumed here in advance so the ## Command appendix in TASKS.md is infra-aware from PR1.
  • doca-setup — env preparation, install verification, hugepages, NUMA awareness, and the I have no install yet path with the public NGC DOCA container.
  • doca-debug — the cross-cutting debug ladder. Bench surfaces its own error taxonomy in CAPABILITIES.md ## Error taxonomy; when the cause turns out to be below DOCA (driver, firmware, NUMA), the bench taxonomy hands off to doca-debug.
  • doca-caps — the sibling DOCA tool for the coarse per-device per-library capability snapshot. Bench probes capability at finer grain via its own query system; doca_caps is the cheaper first step to confirm the device is even visible to DOCA.
  • The matching libs/<library> skill — e.g. doca-comch, doca-compress — for the workload-side preconditions, capability-query rules, and error-taxonomy overlays of the library under test. Bench drives the library; the library skill explains what "healthy" means for it.

Bundled files

The model reads these on demand while the skill is loaded. They are exposed as readable files and are never executed.

Frequently asked questions

What does the Doca Bench AI skill do?

Run `doca_bench` (DOCA 2.7.0 or newer) to measure throughput, bulk latency, precision latency, or maximum bandwidth for RDMA, Compress, AES-GCM, SHA, DMA, EC, Ethernet, Comch, or GPUNetIO on a host or BlueField Arm. Use it to discover enabled benchmark libraries, capture a reproducible command/version/device/environment baseline, compare stable runs against a declared tolerance, or diagnose configuration, device-binding, workload-precondition, and measurement failures. Trigger for requests such as measuring BlueField compression speed, NIC RDMA throughput, crypto latency, or a pre-upgrade b...

Why use Doca Bench on TypingMind?

Because you install it once and use it with any model. Doca Bench is plain Markdown rather than provider-specific code, so the same skill runs on GPT-5, Claude, Gemini, Grok, or a local model — and you can switch model mid-chat without it breaking. TypingMind runs on your own API keys, so you pay providers directly instead of a per-seat subscription, and your skills and chats stay in your own storage.

How do I install Doca Bench in TypingMind?

Open Plugins → Skills → Install from GitHub in TypingMind and paste https://github.com/NVIDIA/skills/tree/main/skills/doca-bench. TypingMind reads its SKILL.md and bundles its files and installs it as a skill you can enable per chat.

Which AI models can use Doca Bench?

Any model you connect in TypingMind. AI skills are plain Markdown instructions rather than provider-specific code, so GPT, Claude, Gemini, Grok, and local models can all load this skill when a request matches it.

How many AI models can I use with Doca Bench?

As many as you like. As long as a model supports skills, you can use Doca Bench with it — GPT, Claude, Gemini, Grok, DeepSeek, Mistral, Llama and more — all on TypingMind with your own API keys.

Is the Doca Bench AI skill free?

Yes. It is published on GitHub by NVIDIA under the Apache-2.0 license. You only pay your own AI provider for the tokens you use.

What are AI skills?

An AI skill is a reusable instruction bundle that teaches an AI model how to do one specific task. It follows the open Agent Skills format: a SKILL.md file with a name and description, plus any scripts, templates or reference files the model may need. The model reads the instructions only when your request matches the skill, so an installed skill costs nothing until it is used.

How are AI skills different from plugins or MCP servers?

A plugin or MCP server gives a model new tools to call — code that runs somewhere and returns a result. An AI skill gives the model knowledge and process instead: how to approach a task, which steps to follow, what good output looks like. Skills are plain Markdown, so they need no server, no API key and no runtime, and they work with any model.

View all

Set up your own AI workspace now

Get notified about new features and future giveaways by subscribing to our newsletter 👇