Claw Score logo

Claw Score

OrganizationPopular
openclaw
claw-score

Audit or refresh OpenClaw maturity scorecard docs from root taxonomy, maturity scores, and QA evidence artifacts without using maintainer discrawl data or committed inventory reports.

Overview

Publisheropenclaw
Repositoryopenclaw
Skill nameclaw-score
Stars
391K
Forks
82.2K
Bundled files
51
Links
  • Markdown instructions

    A SKILL.md file the model loads on demand, so it only costs tokens when a request actually matches.

  • Works with any LLM

    AI skills are plain Markdown, not provider-specific code, so this works with GPT, Claude, Gemini, Grok, or a local model.

  • 51 bundled files

    Scripts, templates, and references the model can read while it works. Files are read-only and never executed.

  • Open source

    Published by openclaw on GitHub. Read the source before you install it.

Installation

Install the Claw Score AI skill in TypingMind to use it with any LLM, or drop it into another agent that reads SKILL.md.

1

Install in TypingMind

TypingMind installs a skill straight from its GitHub folder — it reads SKILL.md, bundles the resource files, and stores the result locally.

  1. Open the app and go to Plugins → Skills.
  2. Choose "Install from GitHub".
  3. Paste the skill folder URL below and confirm.
    https://github.com/openclaw/openclaw/tree/main/.agents/skills/claw-score
  4. Enable the skill in any chat where you want it available.
Plugins → Skills → Add skill → From GitHub URL, then paste the folder URL and press Continue.
2

Install in another agent

Any agent that reads the Agent Skills format can use this skill — copy the folder into that agent's skills directory.

Claude Code — .claude/skills
git clone --depth 1 https://github.com/openclaw/openclaw.git /tmp/openclaw
mkdir -p .claude/skills
cp -r /tmp/openclaw/.agents/skills/claw-score .claude/skills/claw-score
Restart Claude Code after copying so it picks up the new skill.

Use it in TypingMind

Enable Claw Score in any TypingMind chat and the model takes it from there. Its name and description sit in the system prompt, and the moment a request matches, the model loads the full instructions itself — you never invoke it by hand, and it costs no tokens until it is actually used.

The model loads Claw Score on its own as soon as a request matches it.

Works with any AI model

AI skills are plain Markdown instructions rather than provider-specific code, so Claw Score is not tied to the model it was written for. Install it once in TypingMind and use it with GPT-5, Claude, Gemini, Grok, DeepSeek, Mistral, Llama, or a local model you run yourself — all on your own API keys.

  • Loaded only when it is needed

    The system prompt carries just the name and description. The instructions are fetched on the first matching request, so an idle skill costs nothing.

  • Switch models mid-chat

    Because the skill is instructions rather than code, changing model does not break it — the next model reads the same SKILL.md.

Skill instructions

This is the SKILL.md content the model loads. Read it before installing — a skill is instructions your model will follow.

claw-score

Use this skill when working on the OpenClaw maturity scorecard in this repo. This is the openclaw-local version of the maintainer claw-score workflow: it keeps the taxonomy and scorecard concepts, but excludes discrawl and the old committed inventory/ report tree.

Authority

This skill owns the operational workflow for:

  • taxonomy.yaml
  • qa/maturity-scores.yaml
  • docs/concepts/qa-e2e-automation.md
  • qa/scenarios/index.yaml

Keep person-specific, maintainer-private, Discord archive, and discrawl facts out of this repo. If a score needs private evidence, use the redacted qa-evidence.json artifact shape generated by OpenClaw QA workflows.

Source Model

  • taxonomy.yaml is the hand-edited source of truth for surfaces, levels, QA profiles, categories, feature coverage IDs, docs refs, LTS overrides, and completeness-instruction paths.
  • Each feature has exactly one coverageIds entry. Keep that evidence ID unique to the feature; broader many-to-many evidence mapping is not part of the current taxonomy schema.
  • Coverage IDs use dotted namespace.behavior form, with lowercase alphanumeric/dash segments. Profile, surface, and category IDs may remain dashed or dotted.
  • Keep categories and feature names unique, product-shaped, and broader than raw coverage IDs. Do not promote generic IDs into standalone feature names.
  • Avoid duplicate coverage-ID bundles under different feature names in one category.
  • qa/maturity-scores.yaml is the committed aggregate source for Quality, Completeness, and LTS review state.
  • extensions/qa-lab/src/scorecard-taxonomy.ts exports readValidatedQaMaturityScoreSources; use it to validate score output.
  • Generated public docs are docs/maturity/scorecard.md and docs/maturity/taxonomy.md; both come from pnpm maturity:render. Do not hand-edit generated Markdown to change score results.
  • qa-evidence.json artifacts provide per-run QA scorecard evidence. Release profile artifacts are the source of truth for Coverage. They can enrich generated artifact docs, but they are not committed as inventory.

Commands

Run from the openclaw repo root.

Validate taxonomy YAML structure and the maturity score schema after source edits:

bash
node --import tsx --input-type=module <<'NODE'
import fs from "node:fs";
import YAML from "yaml";
import { readValidatedQaMaturityScoreSources } from "./extensions/qa-lab/src/scorecard-taxonomy.ts";

for (const file of ["taxonomy.yaml", "qa/scenarios/index.yaml"]) {
  YAML.parse(fs.readFileSync(file, "utf8"));
}
readValidatedQaMaturityScoreSources();
NODE

Check docs when touching docs prose:

bash
pnpm check:docs

Run focused QA/profile checks when changing coverage IDs or profile membership:

bash
pnpm openclaw qa coverage --json

Full Generation Runs

For a direct full scorecard run that publishes the generated-doc pull request, use floating main resolution by default:

bash
gh workflow run maturity-scorecard.yml \
  --repo openclaw/openclaw \
  --ref main \
  -f ref=main \
  -f expected_sha='' \
  -f publish_pull_request=true \
  -f allow_failures=true

Do not resolve main locally and pass that commit as both ref and expected_sha for an ordinary manual generation run. OpenClaw's main moves quickly, so the caller-selected commit can become stale before validation. The workflow then correctly rejects publication when the pull request base contains newer maturity inputs, and QA never starts.

With ref=main and a blank expected_sha, the workflow's floating_default_branch path fetches and freezes the current remote default branch inside validation before handing an immutable revision to downstream jobs. Use an explicit SHA only when the requested evidence must remain bound to that exact revision, such as a release-candidate workflow call or an artifact-only historical reproduction. If that exact-revision run also requests publication and main has changed relevant inputs, expect validation to fail and dispatch again from floating main instead.

Scoring Workflow

When asked to score or refresh a surface:

  1. Read the surface in taxonomy.yaml.
  2. Read the surface completeness rubric under .agents/skills/claw-score/references/completeness/.
  3. Gather public repo evidence from docs, source, tests, and QA scenario metadata.
  4. Prefer existing release profile qa-evidence.json artifacts for executed proof.
  5. Update qa/maturity-scores.yaml only for Quality, Completeness, and LTS review state backed by public or redacted artifact evidence.
  6. Run the schema validation command from this skill.
  7. Run pnpm check:docs if docs prose changed, and focused QA coverage checks if coverage IDs or profile membership changed.

For subjective score changes, make the smallest defensible edit and leave the evidence path in the PR or task summary. Keep manual prose in current docs and keep score data in qa/maturity-scores.yaml.

Default Completeness Process

Completeness is scored against the intended operator-visible workflow for each category, not against test breadth or implementation quality. The completeness reference files under references/completeness/ define the category scope and any surface-specific variation from this default process.

By default, Completeness measures how fully OpenClaw exposes the intended surface capability set to the user, operator, author, or maintainer persona for that surface. Score whether each category delivers the full expected workflow, including setup, normal use, status or inspection, recovery, and important platform, provider, channel, security, or lifecycle variants where they apply.

Treat Surface-Specific Scoring Questions and Surface-Specific Guidance as higher-priority instructions for that surface. The surface instructions may flesh out, narrow, or intentionally conflict with the default ideas here; when they do, follow the surface instructions and make the score rationale reflect that surface-specific instruction. If a reference file does not include surface-specific questions or guidance, apply this default process to the surface's Category Scope.

For each category, ask:

  • Can the intended user or operator complete the category workflow end to end?
  • Are the taxonomy features present as supported capabilities rather than isolated implementation fragments?
  • Are the important lifecycle stages represented: setup, normal operation, status/inspection, recovery, and upgrade or removal where relevant?
  • Are the important environment, provider, platform, channel, or security branches present for this surface?
  • Do the known gaps leave major user-visible capability branches missing?

Default guidance:

  • Favor higher Completeness when the category supports the full operator-visible workflow described by taxonomy and category evidence.
  • Lower Completeness when only the happy path exists, when important variants are undocumented or unimplemented, or when recovery/status paths are missing.
  • Do not lower Completeness because tests are thin; that is Coverage.
  • Do not lower Completeness because implementation quality is fragile; that is Quality.

Default Completeness bands:

  • Clawesome (95-100): complete across expected workflows, variants, and recovery branches, with only minor polish gaps.
  • Stable (80-95): the expected workflow set is broadly present, with only bounded missing branches.
  • Beta (70-80): the main workflow exists, but meaningful branches or recovery paths are still absent.
  • Alpha (50-70): only a partial capability set is present; users can complete some core tasks but not the full expected workflow.
  • Experimental (0-50): the category exposes only fragments of the intended capability.

Decision Context

Record an optional decision beside score and label for surface and category Quality/Completeness, or beside supported for category LTS. In taxonomy.yaml, use optional level_decision beside the canonical surface level.

Each record contains value, rationale, reviewer, evidence_refs, and revalidate_when. Use an integer from 0–100 for Quality/Completeness, a boolean for LTS, and a declared taxonomy level ID for level_decision. Supply nonempty text fields and at least one evidence reference. Name the actual reviewer and the condition that should trigger another review.

Leave unavailable history absent: it is unknown, not an invitation to invent reviewers, rationale, or evidence. A record does not overwrite the current score, support flag, or canonical level. If its value differs, retain both; generated docs show a non-gating mismatch, including under strict input validation.

Do not attach decisions to Coverage, computed rollups, surface LTS summaries, or the copied level in score aggregates. Decision context does not change coverage identity, score calculations, support commitments, or release gates.

Score Semantics

  • Coverage: deterministic release validation coverage derived from the release profile qa-evidence.json.scorecard feature fulfillment data.
  • Quality: reliability, maintainability, operator safety, and regression confidence for the category.
  • Completeness: how much of the intended operator-visible workflow exists for the category. Use the default completeness process plus any surface-specific variation before changing this score.
  • LTS: derived from Quality, release-evidence Coverage, and human_lts_override; do not hand-edit generated Markdown to change LTS status.

Bands:

  • Clawesome: 95-100
  • Stable: 80-95
  • Beta: 70-80
  • Alpha: 50-70
  • Experimental: 0-50

Artifacts

Do not add the maintainer repo's docs/kevinslin/maturity-scorecard/inventory/ tree to openclaw. Evidence-enriched scorecard outputs belong in short-lived artifacts, not committed generated docs, unless this repo adds an explicit renderer/check workflow first.

Bundled files

The model reads these on demand while the skill is loaded. They are exposed as readable files and are never executed.

Frequently asked questions

What does the Claw Score AI skill do?

Audit or refresh OpenClaw maturity scorecard docs from root taxonomy, maturity scores, and QA evidence artifacts without using maintainer discrawl data or committed inventory reports.

Why use Claw Score on TypingMind?

Because you install it once and use it with any model. Claw Score is plain Markdown rather than provider-specific code, so the same skill runs on GPT-5, Claude, Gemini, Grok, or a local model — and you can switch model mid-chat without it breaking. TypingMind runs on your own API keys, so you pay providers directly instead of a per-seat subscription, and your skills and chats stay in your own storage.

How do I install Claw Score in TypingMind?

Open Plugins → Skills → Install from GitHub in TypingMind and paste https://github.com/openclaw/openclaw/tree/main/.agents/skills/claw-score. TypingMind reads its SKILL.md and bundles its files and installs it as a skill you can enable per chat.

Which AI models can use Claw Score?

Any model you connect in TypingMind. AI skills are plain Markdown instructions rather than provider-specific code, so GPT, Claude, Gemini, Grok, and local models can all load this skill when a request matches it.

How many AI models can I use with Claw Score?

As many as you like. As long as a model supports skills, you can use Claw Score with it — GPT, Claude, Gemini, Grok, DeepSeek, Mistral, Llama and more — all on TypingMind with your own API keys.

Is the Claw Score AI skill free?

It is published on GitHub by openclaw. Check the repository for licensing terms. You only pay your own AI provider for the tokens you use.

What are AI skills?

An AI skill is a reusable instruction bundle that teaches an AI model how to do one specific task. It follows the open Agent Skills format: a SKILL.md file with a name and description, plus any scripts, templates or reference files the model may need. The model reads the instructions only when your request matches the skill, so an installed skill costs nothing until it is used.

How are AI skills different from plugins or MCP servers?

A plugin or MCP server gives a model new tools to call — code that runs somewhere and returns a result. An AI skill gives the model knowledge and process instead: how to approach a task, which steps to follow, what good output looks like. Skills are plain Markdown, so they need no server, no API key and no runtime, and they work with any model.

View all

Set up your own AI workspace now

Get notified about new features and future giveaways by subscribing to our newsletter 👇