Add Dictation logo

Add Dictation

OrganizationPopular
cursor
add-dictation

Use when the user runs /add-dictation or wants speech turned into text with Grok speech-to-text: a mic button that dictates into the composer, live captions, or transcribing recorded audio (files, uploads, URLs) with word timestamps, diarization, subtitles, meeting notes. STT, transcribe, transcription. For a voice agent that talks back use /add-voice.

Overview

Publishercursor
Repositoryplugins
Skill nameadd-dictation
Stars
8K
Forks
728
Bundled files
Instructions only
Links
  • Markdown instructions

    A SKILL.md file the model loads on demand, so it only costs tokens when a request actually matches.

  • Works with any LLM

    AI skills are plain Markdown, not provider-specific code, so this works with GPT, Claude, Gemini, Grok, or a local model.

  • Self-contained

    Everything the model needs lives in the instructions — no extra files to sync.

  • Open source

    Published by cursor on GitHub. Read the source before you install it.

Installation

Install the Add Dictation AI skill in TypingMind to use it with any LLM, or drop it into another agent that reads SKILL.md.

1

Install in TypingMind

TypingMind installs a skill straight from its GitHub folder — it reads SKILL.md, bundles the resource files, and stores the result locally.

  1. Open the app and go to Plugins → Skills.
  2. Choose "Install from GitHub".
  3. Paste the skill folder URL below and confirm.
  4. Enable the skill in any chat where you want it available.
Plugins → Skills → Add skill → From GitHub URL, then paste the folder URL and press Continue.
2

Install in another agent

Any agent that reads the Agent Skills format can use this skill — copy the folder into that agent's skills directory.

Claude Code — .claude/skills
git clone --depth 1 https://github.com/cursor/plugins.git /tmp/plugins
mkdir -p .claude/skills
cp -r /tmp/plugins/grok-voice/skills/add-dictation .claude/skills/add-dictation
Restart Claude Code after copying so it picks up the new skill.

Use it in TypingMind

Enable Add Dictation in any TypingMind chat and the model takes it from there. Its name and description sit in the system prompt, and the moment a request matches, the model loads the full instructions itself — you never invoke it by hand, and it costs no tokens until it is actually used.

The model loads Add Dictation on its own as soon as a request matches it.

Works with any AI model

AI skills are plain Markdown instructions rather than provider-specific code, so Add Dictation is not tied to the model it was written for. Install it once in TypingMind and use it with GPT-5, Claude, Gemini, Grok, DeepSeek, Mistral, Llama, or a local model you run yourself — all on your own API keys.

  • Loaded only when it is needed

    The system prompt carries just the name and description. The instructions are fetched on the first matching request, so an idle skill costs nothing.

  • Switch models mid-chat

    Because the skill is instructions rather than code, changing model does not break it — the next model reads the same SKILL.md.

Skill instructions

This is the SKILL.md content the model loads. Read it before installing — a skill is instructions your model will follow.

Add Dictation

Add Grok Speech to Text to an existing app: a mic button that dictates into the composer, live captions, or transcripts of recorded audio. Run on /add-dictation, typed Dictate, or clear “transcribe” intent. Cursor has no mic; wire the app, not the IDE.

Docs

Pick the path

NeedPath
Tap, speak, tap, text appears. Uploaded files. URLs.Batch POST https://api.x.ai/v1/stt (default)
Text appears while speaking: captions, long dictation, push-to-talkStreaming wss://api.x.ai/v1/stt through a backend relay

Batch is the default for a composer mic button: one request, no socket, the key never leaves the server. Go streaming only when the UX needs interim text.

Auth

  • Bearer XAI_API_KEY, server side only. The STT docs document no ephemeral-token flow, and browsers cannot set WebSocket headers, so browser streaming goes through your backend relay. Do not invent a token flow.
  • Never put the key in a client bundle. Do not paste keys in chat.

Steps

  1. Map the app

    • Composer or input component, where the text should land (insert at cursor vs replace), server framework, package manager.
    • The microphone icon belongs to dictation. If /add-voice is installed, its waveform primary button stays as is; add the mic as a secondary ghost button beside it.
    • Existing mic capture? If /add-voice ran, its PCM capture can feed streaming STT; pass its rate as sample_rate. 16 kHz is the model’s native rate; other supported rates (8000, 16000, 22050, 24000, 44100, 48000) are resampled server side.
  2. Batch path (default)

    • Client: MediaRecorderBlobPOST to your own route. The endpoint auto-detects containers (WAV, MP3, OGG, Opus, FLAC, AAC, MP4, M4A, MKV, WebM), so send whatever MediaRecorder produces.
    • Server: forward as multipart/form-data. Option fields first, file last; fields after file may be ignored. file or url, max 500 MB.
ts
// server (any runtime with fetch + FormData)
export async function transcribe(blob: Blob, filename: string) {
  const form = new FormData();
  form.append("format", "true");     // written-form numbers/currency; requires language
  form.append("language", "en");
  // form.append("keyterm", "Acme"); // repeat per term, ≤100 terms × 50 chars
  form.append("file", blob, filename); // last
  const res = await fetch("https://api.x.ai/v1/stt", {
    method: "POST",
    headers: { Authorization: `Bearer ${process.env.XAI_API_KEY}` },
    body: form,
  });
  if (!res.ok) throw new Error(`STT ${res.status}`); // 400 bad input, 413 >500 MB, 429 back off, 502 url fetch failed, 503 retry
  return (await res.json()) as {
    text: string; language: string; duration: number;
    words?: { text: string; start: number; end: number; speaker?: number }[];
    channels?: { index: number; text: string; words: unknown[] }[];
  };
}
ts
// client
const mime = MediaRecorder.isTypeSupported("audio/webm;codecs=opus") ? "audio/webm;codecs=opus" : "audio/mp4";
const rec = new MediaRecorder(stream, { mimeType: mime });
const parts: BlobPart[] = [];
rec.ondataavailable = (e) => parts.push(e.data);
rec.onstop = async () => {
  const fd = new FormData();
  fd.append("file", new Blob(parts, { type: mime }), "dictation");
  const { text } = await (await fetch("/api/dictation", { method: "POST", body: fd })).json();
  insertAtCursor(text);
};
rec.start(); // second tap: rec.stop()
  1. Streaming path
    • Relay: server holds the key, upgrades the browser socket, forwards binary frames and client control messages up, JSON events down. Build the query string server side.
ts
import { WebSocketServer, WebSocket } from "ws";

new WebSocketServer({ port: 8788 }).on("connection", (client) => {
  const q = new URLSearchParams({ sample_rate: "16000", encoding: "pcm", interim_results: "true", language: "en" });
  const up = new WebSocket(`wss://api.x.ai/v1/stt?${q}`, { headers: { Authorization: `Bearer ${process.env.XAI_API_KEY}` } });
  up.on("message", (d) => client.send(d.toString()));                       // transcript.* and error events
  client.on("message", (d, isBinary) => up.readyState === WebSocket.OPEN && up.send(d, { binary: isBinary })); // audio + finalize/audio.done
  const end = () => { client.close(); up.close(); };
  up.on("close", end); up.on("error", end); client.on("close", end);
});
  • Browser capture: PCM16 little-endian, mono, 16 kHz, 100 ms frames = 3,200 bytes, raw binary, no base64. Wait for transcript.created before sending. MediaRecorder output is a container, not raw frames; do not stream it.
ts
const ws = new WebSocket(relayUrl); ws.binaryType = "arraybuffer";
const ctx = new AudioContext({ sampleRate: 16000 }); // if ctx.sampleRate !== 16000, downsample in the worklet
await ctx.audioWorklet.addModule("/pcm16-worklet.js"); // Float32 → Int16LE, posts one 3,200-byte frame per 100 ms
const node = new AudioWorkletNode(ctx, "pcm16");
ctx.createMediaStreamSource(stream).connect(node);
let ready = false;
node.port.onmessage = (e) => ready && ws.readyState === WebSocket.OPEN && ws.send(e.data);

let committed = "", locked = "", live = "";
ws.addEventListener("message", (e) => {
  const ev = JSON.parse(e.data);
  if (ev.type === "transcript.created") ready = true;
  else if (ev.type === "transcript.partial") {
    if (ev.speech_final) { committed += ev.text + " "; locked = ""; live = ""; } // complete stitched utterance
    else if (ev.is_final) { locked += ev.text + " "; live = ""; }                // chunk final: text will not change
    else live = ev.text;                                                          // interim: may change
    render(committed + locked + live);
  } else if (ev.type === "transcript.done") ws.close();                          // after audio.done
  else if (ev.type === "error") showError(ev.message);                           // most errors close the socket
});
// stop: ws.send(JSON.stringify({ type: "audio.done" }))
// push-to-talk release: ws.send(JSON.stringify({ type: "Finalize" })) then keep streaming (docs show both `finalize` and `Finalize`; the examples use `Finalize`)
  1. Options (query params for streaming, form fields for batch)
WantSet
Text while speakinginterim_results=true
“one hundred dollars” → $100streaming: language=en; batch: format=true + language=en
Product names, jargonkeyterm= repeated
Not cut off mid-sentence while dictating numberssmart_turn=0.7&smart_turn_timeout=3000
Faster or slower end of utteranceendpointing= ms, default 400
Who said what (meetings)diarize=truewords[].speaker
Agent and customer on separate channelsmultichannel=true&channels=2 (PCM only, not Opus)
Keep “um”, “uh”filler_words=true (removed by default)
Low bandwidth or mobileencoding=opus, exactly one raw Opus packet per frame, omit sample_rate
Raw audio to batch`audio_format=pcm
Quiet or telephony audiolower vad_threshold (streaming default 0.08, batch 0.5)
  1. Python twin (only if the server is Python)
python
import os, requests
r = requests.post(
    "https://api.x.ai/v1/stt",
    headers={"Authorization": f"Bearer {os.environ['XAI_API_KEY']}"},
    data=[("format", "true"), ("language", "en")],
    files={"file": ("dictation.webm", blob, "audio/webm")},  # requests sends data fields before files
)
r.raise_for_status(); text = r.json()["text"]
# streaming: websockets.connect(url, additional_headers={"Authorization": f"Bearer {key}"}); await ws.send(pcm_bytes)
  1. Smoke
    • Batch: curl -X POST https://api.x.ai/v1/stt -H "Authorization: Bearer $XAI_API_KEY" -F language=en -F file=@short.wav → 200 with text. Same call with -F format=true and no language → 400.
    • Streaming: dictate two sentences with a pause between them. Expect interim text, then a final; no duplicated or vanished words at the utterance boundary (if words vanish, the stitched speech_final text did not include the chunk finals: append instead of replacing locked). audio.donetranscript.done, socket closes.
    • Search the client bundle for XAI_API_KEY; it must not be there.
    • Debug from logs with /debug-voice; swap its hook points to transcript.* events.

Out of scope

  • Speech that talks back (/add-voice), speaking text (/add-read-aloud)
  • Inventing an STT token flow, endpoints, or event names not in the docs

Frequently asked questions

What does the Add Dictation AI skill do?

Use when the user runs /add-dictation or wants speech turned into text with Grok speech-to-text: a mic button that dictates into the composer, live captions, or transcribing recorded audio (files, uploads, URLs) with word timestamps, diarization, subtitles, meeting notes. STT, transcribe, transcription. For a voice agent that talks back use /add-voice.

Why use Add Dictation on TypingMind?

Because you install it once and use it with any model. Add Dictation is plain Markdown rather than provider-specific code, so the same skill runs on GPT-5, Claude, Gemini, Grok, or a local model — and you can switch model mid-chat without it breaking. TypingMind runs on your own API keys, so you pay providers directly instead of a per-seat subscription, and your skills and chats stay in your own storage.

How do I install Add Dictation in TypingMind?

Open Plugins → Skills → Install from GitHub in TypingMind and paste https://github.com/cursor/plugins/tree/main/grok-voice/skills/add-dictation. TypingMind reads its SKILL.md and installs it as a skill you can enable per chat.

Which AI models can use Add Dictation?

Any model you connect in TypingMind. AI skills are plain Markdown instructions rather than provider-specific code, so GPT, Claude, Gemini, Grok, and local models can all load this skill when a request matches it.

How many AI models can I use with Add Dictation?

As many as you like. As long as a model supports skills, you can use Add Dictation with it — GPT, Claude, Gemini, Grok, DeepSeek, Mistral, Llama and more — all on TypingMind with your own API keys.

Is the Add Dictation AI skill free?

It is published on GitHub by cursor. Check the repository for licensing terms. You only pay your own AI provider for the tokens you use.

What are AI skills?

An AI skill is a reusable instruction bundle that teaches an AI model how to do one specific task. It follows the open Agent Skills format: a SKILL.md file with a name and description, plus any scripts, templates or reference files the model may need. The model reads the instructions only when your request matches the skill, so an installed skill costs nothing until it is used.

How are AI skills different from plugins or MCP servers?

A plugin or MCP server gives a model new tools to call — code that runs somewhere and returns a result. An AI skill gives the model knowledge and process instead: how to approach a task, which steps to follow, what good output looks like. Skills are plain Markdown, so they need no server, no API key and no runtime, and they work with any model.

View all

Set up your own AI workspace now

Get notified about new features and future giveaways by subscribing to our newsletter 👇