Run Benchmark logo

Run Benchmark

OrganizationPopular
AgibotTech
run-benchmark

Launch a geniesim_benchmark task locally (typically inside the GUI Docker container) against a user-provided inference server, using the `geniesim benchmark run` CLI verb. Trigger: When the user asks to "run geniesim", "本地跑仿真", "启动仿真任务", "run a benchmark", "launch <some>_<config>.yaml", or wants to execute a benchmark task config (anything under `geniesim_benchmark/config/*.yaml`) against a remote inference host (ip:port).

Overview

PublisherAgibotTech
Repositorygenie_sim
Skill namerun-benchmark
Stars
1.4K
Forks
119
Bundled files
Instructions only
LicenseMPL-2.0
Links
  • Markdown instructions

    A SKILL.md file the model loads on demand, so it only costs tokens when a request actually matches.

  • Works with any LLM

    AI skills are plain Markdown, not provider-specific code, so this works with GPT, Claude, Gemini, Grok, or a local model.

  • Self-contained

    Everything the model needs lives in the instructions — no extra files to sync.

  • Open source

    Published by AgibotTech on GitHub. Read the source before you install it.

Installation

Install the Run Benchmark AI skill in TypingMind to use it with any LLM, or drop it into another agent that reads SKILL.md.

1

Install in TypingMind

TypingMind installs a skill straight from its GitHub folder — it reads SKILL.md, bundles the resource files, and stores the result locally.

  1. Open the app and go to Plugins → Skills.
  2. Choose "Install from GitHub".
  3. Paste the skill folder URL below and confirm.
  4. Enable the skill in any chat where you want it available.
Plugins → Skills → Add skill → From GitHub URL, then paste the folder URL and press Continue.
2

Install in another agent

Any agent that reads the Agent Skills format can use this skill — copy the folder into that agent's skills directory.

Claude Code — .claude/skills
git clone --depth 1 https://github.com/AgibotTech/genie_sim.git /tmp/genie_sim
mkdir -p .claude/skills
cp -r /tmp/genie_sim/source/geniesim_benchmark/skills/run-benchmark .claude/skills/run-benchmark
Restart Claude Code after copying so it picks up the new skill.

Use it in TypingMind

Enable Run Benchmark in any TypingMind chat and the model takes it from there. Its name and description sit in the system prompt, and the moment a request matches, the model loads the full instructions itself — you never invoke it by hand, and it costs no tokens until it is actually used.

The model loads Run Benchmark on its own as soon as a request matches it.

Works with any AI model

AI skills are plain Markdown instructions rather than provider-specific code, so Run Benchmark is not tied to the model it was written for. Install it once in TypingMind and use it with GPT-5, Claude, Gemini, Grok, DeepSeek, Mistral, Llama, or a local model you run yourself — all on your own API keys.

  • Loaded only when it is needed

    The system prompt carries just the name and description. The instructions are fetched on the first matching request, so an idle skill costs nothing.

  • Switch models mid-chat

    Because the skill is instructions rather than code, changing model does not break it — the next model reads the same SKILL.md.

Skill instructions

This is the SKILL.md content the model loads. Read it before installing — a skill is instructions your model will follow.

When to Use

  • User wants to run a geniesim_benchmark task on their workstation (not the Challenge platform).
  • User has an inference server already running somewhere reachable and provides its ip:port.
  • User references any task config under source/geniesim_benchmark/src/geniesim_benchmark/config/.

Do not use for:

  • Submitting jobs to the Challenge platform → challenge-submit-job.
  • Verifying an inference server is healthy → use the check-inference skill.
  • Adding a new benchmark task → add-benchmark-task.

Critical Patterns

  1. Always collect three required inputs first:
    • Task config (basename, full path, or substring — the CLI resolves all three).
    • Inference IP.
    • Inference port.
  2. The runtime needs omni_python / Isaac Sim on the host. Inside the Genie Sim Docker image (geniesim docker upgeniesim docker into), that's already the case. Outside the container the user needs Isaac Sim installed system-wide.
  3. Working directory: anywhere under the repo root works — the CLI walks up to find scripts/ and uses find_spec to locate the benchmark package.
  4. Confirm before launching. The task spawns a full simulator and typically holds a GPU; ask before kicking it off if there's any ambiguity.

Workflow

Step 1 — Collect inputs

If the user hasn't named a config, list candidates:

bash
geniesim benchmark categories      # show category counts
geniesim benchmark robots          # show robot counts
geniesim benchmark list --robot=<R> --category=<C>

Then ask via AskUserQuestion:

  • Task config: free-text (use the basename — e.g. g2op_if_pick_block_color).
  • Inference host as ip:port.

Step 2 — Probe inference (optional but recommended)

Before sinking minutes into Isaac Sim startup, sanity-check the server (uses the bundled corobot payload):

bash
geniesim benchmark check-inference --infer-host=<IP>:<PORT>

See the check-inference skill to override the payload.

Step 3 — Run the task

Inside the GUI container (geniesim docker into):

bash
geniesim benchmark run <CONFIG> --infer-host=<IP>:<PORT>

Example:

bash
geniesim benchmark run g2op_if_pick_block_color --infer-host=<IP>:<PORT>

Step 4 — Pass-through overrides (when asked)

geniesim benchmark run forwards any unknown --key=value to the benchmark's ParameterServer. Common ones:

FlagMeaning
--app.headless=trueNo GUI (required on remote / batch hosts)
--benchmark.num_episode=NOverride episode count
--benchmark.seed=NRNG / instance-sampling seed
--benchmark.record=truePersist episode logs to output_dir
--benchmark.policy_class=…Use a different policy class

Full schema: source/geniesim_benchmark/src/geniesim_benchmark/config/params.py.

Commands (copy-paste summary for the user)

bash
# Host — start the container (GUI by default; add --headless on remote/batch hosts)
cd /path/to/main
geniesim docker up

# Host — drop into a shell inside the running container
geniesim docker into
# inside container:
geniesim status                                       # verify the stack is healthy
geniesim benchmark check-inference --infer-host=<IP>:<PORT>
geniesim benchmark run <CONFIG> --infer-host=<IP>:<PORT>

Notes

  • If geniesim isn't on $PATH (the launcher wasn't installed), substitute python3 -m geniesim_cli benchmark … — same args, same behaviour.
  • <CONFIG> accepts the bare basename (g2op_if_pick_block_color), a full path, or a unique substring.
  • For batch evaluations, prefer geniesim benchmark batch --category=… --robot=… over a shell loop — it forwards extras consistently and prints a per-config pass/fail summary.
  • The new CLI replaces the older ad-hoc omni_python app/app.py --config … invocation. The new form normalizes interpreter selection, host shorthand, and config resolution.

Frequently asked questions

What does the Run Benchmark AI skill do?

Launch a geniesim_benchmark task locally (typically inside the GUI Docker container) against a user-provided inference server, using the `geniesim benchmark run` CLI verb. Trigger: When the user asks to "run geniesim", "本地跑仿真", "启动仿真任务", "run a benchmark", "launch <some>_<config>.yaml", or wants to execute a benchmark task config (anything under `geniesim_benchmark/config/*.yaml`) against a remote inference host (ip:port).

Why use Run Benchmark on TypingMind?

Because you install it once and use it with any model. Run Benchmark is plain Markdown rather than provider-specific code, so the same skill runs on GPT-5, Claude, Gemini, Grok, or a local model — and you can switch model mid-chat without it breaking. TypingMind runs on your own API keys, so you pay providers directly instead of a per-seat subscription, and your skills and chats stay in your own storage.

How do I install Run Benchmark in TypingMind?

Open Plugins → Skills → Install from GitHub in TypingMind and paste https://github.com/AgibotTech/genie_sim/tree/main/source/geniesim_benchmark/skills/run-benchmark. TypingMind reads its SKILL.md and installs it as a skill you can enable per chat.

Which AI models can use Run Benchmark?

Any model you connect in TypingMind. AI skills are plain Markdown instructions rather than provider-specific code, so GPT, Claude, Gemini, Grok, and local models can all load this skill when a request matches it.

How many AI models can I use with Run Benchmark?

As many as you like. As long as a model supports skills, you can use Run Benchmark with it — GPT, Claude, Gemini, Grok, DeepSeek, Mistral, Llama and more — all on TypingMind with your own API keys.

Is the Run Benchmark AI skill free?

Yes. It is published on GitHub by AgibotTech under the MPL-2.0 license. You only pay your own AI provider for the tokens you use.

What are AI skills?

An AI skill is a reusable instruction bundle that teaches an AI model how to do one specific task. It follows the open Agent Skills format: a SKILL.md file with a name and description, plus any scripts, templates or reference files the model may need. The model reads the instructions only when your request matches the skill, so an installed skill costs nothing until it is used.

How are AI skills different from plugins or MCP servers?

A plugin or MCP server gives a model new tools to call — code that runs somewhere and returns a result. An AI skill gives the model knowledge and process instead: how to approach a task, which steps to follow, what good output looks like. Skills are plain Markdown, so they need no server, no API key and no runtime, and they work with any model.

View all

Set up your own AI workspace now

Get notified about new features and future giveaways by subscribing to our newsletter 👇