Ops Inspector logo

Ops Inspector

OrganizationPopular
TencentCloudBase
ops-inspector

AIOps-style CloudBase inspection skill (v3). Use when users need health checks, log diagnosis, alarm interpretation (CPU alert normal?, peak QPS), metrics via queryEnv(action=metrics), or fault playbooks for 429 / function 404 / ACCESS_TOKEN_INVALID / zero invocations. Triggers on 巡检, 诊断, 告警, 峰值 QPS, 限频, 调用量为 0, troubleshooting.

Overview

PublisherTencentCloudBase
RepositoryCloudBase-AI-Toolkit
Skill nameops-inspector
Stars
1.1K
Forks
141
Bundled files
2
LicenseMIT
Links
  • Markdown instructions

    A SKILL.md file the model loads on demand, so it only costs tokens when a request actually matches.

  • Works with any LLM

    AI skills are plain Markdown, not provider-specific code, so this works with GPT, Claude, Gemini, Grok, or a local model.

  • 2 bundled files

    Scripts, templates, and references the model can read while it works. Files are read-only and never executed.

  • Open source

    Published by TencentCloudBase on GitHub. Read the source before you install it.

Installation

Install the Ops Inspector AI skill in TypingMind to use it with any LLM, or drop it into another agent that reads SKILL.md.

1

Install in TypingMind

TypingMind installs a skill straight from its GitHub folder — it reads SKILL.md, bundles the resource files, and stores the result locally.

  1. Open the app and go to Plugins → Skills.
  2. Choose "Install from GitHub".
  3. Paste the skill folder URL below and confirm.
  4. Enable the skill in any chat where you want it available.
Plugins → Skills → Add skill → From GitHub URL, then paste the folder URL and press Continue.
2

Install in another agent

Any agent that reads the Agent Skills format can use this skill — copy the folder into that agent's skills directory.

Claude Code — .claude/skills
git clone --depth 1 https://github.com/TencentCloudBase/CloudBase-AI-Toolkit.git /tmp/CloudBase-AI-Toolkit
mkdir -p .claude/skills
cp -r /tmp/CloudBase-AI-Toolkit/config/source/skills/ops-inspector .claude/skills/ops-inspector
Restart Claude Code after copying so it picks up the new skill.

Use it in TypingMind

Enable Ops Inspector in any TypingMind chat and the model takes it from there. Its name and description sit in the system prompt, and the moment a request matches, the model loads the full instructions itself — you never invoke it by hand, and it costs no tokens until it is actually used.

The model loads Ops Inspector on its own as soon as a request matches it.

Works with any AI model

AI skills are plain Markdown instructions rather than provider-specific code, so Ops Inspector is not tied to the model it was written for. Install it once in TypingMind and use it with GPT-5, Claude, Gemini, Grok, DeepSeek, Mistral, Llama, or a local model you run yourself — all on your own API keys.

  • Loaded only when it is needed

    The system prompt carries just the name and description. The instructions are fetched on the first matching request, so an idle skill costs nothing.

  • Switch models mid-chat

    Because the skill is instructions rather than code, changing model does not break it — the next model reads the same SKILL.md.

Skill instructions

This is the SKILL.md content the model loads. Read it before installing — a skill is instructions your model will follow.

Sibling skills (local only)

Sibling CloudBase skills ship beside this skill. Use local relative paths such as ../auth-tool-cloudbase/SKILL.md.

If a referenced sibling skill file is missing from this environment, ask the user to install the full CloudBase plugin (or the missing skill). Do not HTTP-fetch remote skill or protocol markdown into the agent context.

Activation Contract

Use this first when

  • The user wants to check the health or status of CloudBase resources (cloud functions, CloudRun, databases, storage, etc.).
  • The user reports errors, failures, or abnormal behavior and wants a quick diagnosis.
  • The user asks for an "inspection", "health check", "巡检", "诊断", or "troubleshooting" of their CloudBase environment.
  • The user wants to review recent error logs across services.
  • The user asks 告警解读 questions: whether a CPU 告警 is normal, what 峰值 QPS was, or whether throttle/error metrics look healthy.
  • The symptom matches a v3 fault playbook: 429 / 限频, 云函数 404, ACCESS_TOKEN_INVALID, or 调用量为 0.

Read before writing code if

  • The inspection reveals code-level issues in cloud functions or CloudRun services — then read the relevant implementation skill before suggesting fixes.
  • The user wants to fix a problem found during inspection rather than just diagnose it.

Then also read

  • Alarm interpretation baselines -> references/alarm-interpretation.md
  • Fault playbooks (429 / 404 / token / zero calls) -> references/fault-playbooks.md
  • Cloud function issues -> ../cloud-functions/SKILL.md
  • CloudRun issues -> ../cloudrun-development/SKILL.md
  • Database issues -> ../postgresql-development-cloudbase/SKILL.md for CloudBase PG / PostgreSQL, ../relational-database-mcp-cloudbase/SKILL.md for MySQL, or ../cloudbase-document-database-web-sdk/SKILL.md for NoSQL
  • Auth readiness (token failures) -> ../auth-tool-cloudbase/SKILL.md
  • Platform overview -> ../cloudbase-platform/SKILL.md

Do NOT use for

  • Deploying new resources or writing application code. This skill is read-only and diagnostic.
  • Replacing proper monitoring/alerting infrastructure. It provides point-in-time inspection, not continuous monitoring.
  • Directly fixing problems — it diagnoses and recommends; actual fixes should use the appropriate implementation skill.
  • Fetching metrics by guessing cloud API Actions. Never use callCloudApi for monitor curves — always use queryEnv(action="metrics").

Common mistakes / gotchas

  • Running a full inspection without first confirming the environment is bound (auth tool must show logged-in and env-bound state).
  • Ignoring CLS log service status — if CLS is not enabled, queryLogs will fail; always check first with queryLogs(action="checkLogService").
  • Searching logs without a time range — this can return excessive or irrelevant results. Always scope searches to a relevant time window.
  • Treating a single error log as the root cause without correlating across resources. A function error may stem from a database or config issue.
  • Answering "峰值 QPS" / "CPU 告警是否正常" from screenshots or memory instead of queryEnv(action="metrics").
  • Calling callCloudApi with invented GetMonitorData / DescribeCurveData parameters — the metrics branch already wraps Manager SDK.

Minimal checklist

  • Environment is bound and accessible (queryEnv(action="info"))
  • Metrics pulled with queryEnv(action="metrics") when the question involves QPS / CPU / throttle / invocation volume
  • CLS log service is enabled (queryLogs(action="checkLogService")) when log diagnosis is needed
  • Matching fault playbook selected when symptoms match 429 / function 404 / ACCESS_TOKEN_INVALID / 调用量为 0
  • Time range is specified for any log or metrics searches
  • Findings are summarized with severity levels, 告警解读, and actionable recommendations

How to use this skill (for a coding agent)

Ops Inspector v3 additions

v3 adds two mandatory capabilities on top of log/resource inspection:

  1. 告警解读 — pull metrics, compare to baselines in references/alarm-interpretation.md, answer CPU-alert / peak-QPS style questions in plain language.
  2. 故障剧本 — when symptoms match, follow references/fault-playbooks.md instead of ad-hoc tool fishing.

Inspection Modes

ModeWhen to useScope
Full inspectionUser asks for a general health check / 巡检 / 全面检查All resource types + core metrics
Targeted inspectionUser reports a specific error or asks about a specific resourceOne resource type or playbook
Alarm interpretationUser asks CPU 告警是否正常 / 峰值 QPS / 是否限流Metrics-first, then logs
Fault playbook429 / function 404 / ACCESS_TOKEN_INVALID / 调用量为 0Playbook steps only

Full Inspection Workflow

Follow these steps in order for a comprehensive environment health check:

Step 1 — Environment Check

queryEnv(action="info")

Confirm the environment is accessible. Record the envId for console link generation.

Step 2 — Metrics snapshot (v3)

queryEnv(action="metrics", envId="<EnvId>", metricName="GatewayTraceEnvQPS")
queryEnv(action="metrics", envId="<EnvId>", metricName="FunctionInvocation")
queryEnv(action="metrics", envId="<EnvId>", metricName="MysqlCpuUsageRate")

Use returned Summary.max / avg / allZero / peakTimestamp. Add FunctionError, FunctionThrottle, or CloudRun Tke* metrics when those resources exist. Read references/alarm-interpretation.md before concluding.

Step 3 — Log Service Status

queryLogs(action="checkLogService")

If CLS is not enabled, note this as a warning — log-based diagnosis will be unavailable. Recommend enabling CLS in the console: https://tcb.cloud.tencent.com/dev?envId=${envId}#/devops/log

Step 4 — Cloud Functions Inspection

queryFunctions(action="listFunctions")

For each function, check:

  • Status: Is the function in an active/deployed state?
  • Recent errors: queryFunctions(action="listFunctionLogs", functionName="<name>", startTime="<recent>")
  • Common issues:
    • Timeout errors (execution exceeded limit)
    • Memory limit exceeded
    • Runtime errors (unhandled exceptions)
    • Cold start frequency
    • Zero invocations while traffic is expected → Playbook 4

Step 5 — CloudRun Services Inspection

queryCloudRun(action="list")

For each service, check:

  • Status: Is the service running?
  • Detail: queryCloudRun(action="detail", detailServerName="<name>")
  • Metrics: queryEnv(action="metrics", metricName="TkeQPSService", resourceID="<serviceName>") (resourceID required)
  • Common issues:
    • Service not running (scaled to zero or crashed)
    • Image pull failures
    • OOMKilled events
    • Health check failures

Step 6 — Error Log Aggregation (if CLS is enabled)

queryLogs(action="searchLogs", queryString="ERROR", service="tcb", startTime="<24h-ago>", limit=50)
queryLogs(action="searchLogs", queryString="ERROR", service="tcbr", startTime="<24h-ago>", limit=50)

Look for patterns:

  • Repeated error messages (same error many times)
  • Cascading failures (errors in multiple services around the same time)
  • Timeout / 429 / 404 / ACCESS_TOKEN_INVALID patterns → jump to the matching playbook

Step 7 — Summary Report

Generate a structured report:

markdown
# CloudBase Resource Inspection Report

**Environment**: ${envId}
**Inspection Time**: ${timestamp}

## Overall Health: ✅ Healthy / ⚠️ Warnings Found / ❌ Issues Found

## 告警解读
| 问题 | 指标 | 窗口峰值 | 基线 | 结论 |
|------|------|----------|------|------|
| 峰值 QPS | GatewayTraceEnvQPS | ... | package default 500 unless known | ... |
| CPU 告警是否正常 | MysqlCpuUsageRate | ... | warn≥80 / crit≥90 | ... |

### Cloud Functions
| Function | Status | Recent Errors | Invocations | Severity |
|----------|--------|---------------|-------------|----------|
| ... | ... | ... | ... | ... |

### CloudRun Services
| Service | Status | Issues | Severity |
|---------|--------|--------|----------|
| ... | ... | ... | ... |

### Error Log Summary
- Total errors in last 24h: N
- Top error patterns: ...

## Recommendations
1. ...
2. ...

## Console Links
- Cloud Functions: https://tcb.cloud.tencent.com/dev?envId=${envId}#/scf
- CloudRun: https://tcb.cloud.tencent.com/dev?envId=${envId}#/platform-run
- Logs: https://tcb.cloud.tencent.com/dev?envId=${envId}#/devops/log
- Monitor: https://tcb.cloud.tencent.com/dev?envId=${envId}#/devops

Targeted Inspection Workflow

When the user specifies a resource type or a specific resource:

  1. Cloud function errors: queryFunctions(action="listFunctionLogs", functionName="<name>") then queryLogs(action="searchLogs", queryString="* AND functionName:<name> AND level:ERROR", ...)
  2. CloudRun errors: queryCloudRun(action="detail", detailServerName="<name>") then queryLogs(action="searchLogs", queryString="ERROR", service="tcbr", ...)
    • If logs show DB / Redis connection failures (ECONNREFUSED, timeout, "could not connect"): check whether VpcConf is set and matches the database VPC. See cloudrun-development/references/vpc-and-database.md.
  3. Database issues: Check queryPgDatabase(action="context"|"metadata"|"objects") for CloudBase PG, queryMysqlDatabase for MySQL, or readNoSqlDatabaseStructure for NoSQL depending on type; for CPU/disk alerts also pull MysqlCpuUsageRate / MysqlStorageUsage metrics
  4. General error search: queryLogs(action="searchLogs", queryString="<error-keyword>", ...)
  5. Alarm / QPS questions: follow references/alarm-interpretation.md
  6. 429 / function 404 / ACCESS_TOKEN_INVALID / 调用量为 0: follow references/fault-playbooks.md

AIOps Methodology

This skill follows AIOps principles for intelligent inspection:

  1. Data Collection: Gather metrics (queryEnv metrics), logs, and resource states via MCP tools — never via ad-hoc callCloudApi
  2. Pattern Recognition: Identify recurring errors, anomaly patterns, and correlations across services
  3. Baseline Comparison: Compare metric Summary values to skill baselines (告警解读)
  4. Root Cause Hypothesis: Based on error patterns + metrics, suggest likely root causes
  5. Actionable Recommendations: Provide specific, prioritized remediation steps with links to relevant skills and console pages

Severity Levels

LevelIconMeaning
CriticalService is down or data is at risk; requires immediate action
Warning⚠️Errors detected but service is still partially functional; investigate soon
Infoℹ️No errors found; informational status only
HealthyResource is operating normally

Preferred Tool Map

OperationMCP Tool Call
Check environmentqueryEnv(action="info")
Query metrics (QPS/CPU/invocations)queryEnv(action="metrics", envId, metricName="...")
Check CLS statusqueryLogs(action="checkLogService")
List cloud functionsqueryFunctions(action="listFunctions")
Get function detailqueryFunctions(action="getFunctionDetail", functionName="<name>")
Get function logsqueryFunctions(action="listFunctionLogs", functionName="<name>", startTime="<time>", endTime="<time>")
Get function log detailqueryFunctions(action="getFunctionLogDetail", requestId="<id>")
List CloudRun servicesqueryCloudRun(action="list")
Get CloudRun detailqueryCloudRun(action="detail", detailServerName="<name>")
Search CLS logsqueryLogs(action="searchLogs", queryString="<query>", service="tcb|tcbr", startTime="<time>", endTime="<time>")
Check NoSQL structurereadNoSqlDatabaseStructure(action="listCollections")
Check PostgreSQL contextqueryPgDatabase(action="context")
Check PostgreSQL metadataqueryPgDatabase(action="metadata", limit=20)
Check MySQL statusqueryMysqlDatabase(action="getContext")
Auth provider readinessqueryAppAuth / auth-tool skill (for ACCESS_TOKEN_INVALID)

Common CLS Query Patterns

ScenarioqueryString
All errorsERROR
Function timeouttimeout OR 超时
Function OOMOOM OR out of memory OR 内存超限
CloudRun crashcrash OR OOMKilled OR Error
Specific function errorsfunctionName:<name> AND level:ERROR
5xx HTTP errorsstatusCode:>499
429 / throttle429 OR throttle OR 限流 OR FREQUENCY
Function 404404 OR FUNCTION_NOT_FOUND
Token invalidACCESS_TOKEN_INVALID OR token invalid
Cold start issuescoldStart OR 冷启动

Time Range Guidance

  • Quick check: Last 1 hour (startTime = 1 hour ago)
  • Standard inspection: Last 24 hours
  • Trend analysis: Last 7 days
  • Specific incident: Narrow to the reported time window

Always use ISO-like YYYY-MM-DD HH:mm:ss for metrics startTime/endTime, e.g., "2026-08-17 00:00:00".

Related Skills

  • cloud-functions — Cloud function development, deployment, and debugging
  • cloudrun-development — CloudRun backend deployment and management
  • cloudbase-platform — General platform knowledge and console navigation
  • postgresql-development-cloudbase — CloudBase PostgreSQL / PG diagnostics and schema/RLS checks
  • relational-database-mcp-cloudbase — MySQL database management and diagnostics
  • auth-tool-cloudbase — Auth provider readiness for token failures

Bundled files

The model reads these on demand while the skill is loaded. They are exposed as readable files and are never executed.

Frequently asked questions

What does the Ops Inspector AI skill do?

AIOps-style CloudBase inspection skill (v3). Use when users need health checks, log diagnosis, alarm interpretation (CPU alert normal?, peak QPS), metrics via queryEnv(action=metrics), or fault playbooks for 429 / function 404 / ACCESS_TOKEN_INVALID / zero invocations. Triggers on 巡检, 诊断, 告警, 峰值 QPS, 限频, 调用量为 0, troubleshooting.

Why use Ops Inspector on TypingMind?

Because you install it once and use it with any model. Ops Inspector is plain Markdown rather than provider-specific code, so the same skill runs on GPT-5, Claude, Gemini, Grok, or a local model — and you can switch model mid-chat without it breaking. TypingMind runs on your own API keys, so you pay providers directly instead of a per-seat subscription, and your skills and chats stay in your own storage.

How do I install Ops Inspector in TypingMind?

Open Plugins → Skills → Install from GitHub in TypingMind and paste https://github.com/TencentCloudBase/CloudBase-AI-Toolkit/tree/main/config/source/skills/ops-inspector. TypingMind reads its SKILL.md and bundles its files and installs it as a skill you can enable per chat.

Which AI models can use Ops Inspector?

Any model you connect in TypingMind. AI skills are plain Markdown instructions rather than provider-specific code, so GPT, Claude, Gemini, Grok, and local models can all load this skill when a request matches it.

How many AI models can I use with Ops Inspector?

As many as you like. As long as a model supports skills, you can use Ops Inspector with it — GPT, Claude, Gemini, Grok, DeepSeek, Mistral, Llama and more — all on TypingMind with your own API keys.

Is the Ops Inspector AI skill free?

Yes. It is published on GitHub by TencentCloudBase under the MIT license. You only pay your own AI provider for the tokens you use.

What are AI skills?

An AI skill is a reusable instruction bundle that teaches an AI model how to do one specific task. It follows the open Agent Skills format: a SKILL.md file with a name and description, plus any scripts, templates or reference files the model may need. The model reads the instructions only when your request matches the skill, so an installed skill costs nothing until it is used.

How are AI skills different from plugins or MCP servers?

A plugin or MCP server gives a model new tools to call — code that runs somewhere and returns a result. An AI skill gives the model knowledge and process instead: how to approach a task, which steps to follow, what good output looks like. Skills are plain Markdown, so they need no server, no API key and no runtime, and they work with any model.

View all

Set up your own AI workspace now

Get notified about new features and future giveaways by subscribing to our newsletter 👇