Media AI Skills
708 open-source Media AI skills that teach any AI model a new workflow.
Search and filter AI skills
AI skills directory results
Render Multiworld
Assemble a silent, music-led 3-world product-tour ad — trim and hard-cut-concat the per-world WIDE-arrival + top-down-macro clips, composite the HTML/Playwright brand end card ("FIND YOUR DAILY." +…
Worldlabs
Generate photorealistic 3D worlds and environments with the World Labs Marble API — Gaussian Splat scenes from text prompts or reference images. Use when the user says "generate a 3D world", "create…
Video To Landing Page
Turn any video into a cinematic scroll-driven landing page — Apple-style hero where scrolling progresses the visible frame through the video. Use when the user provides a video file and asks for "a…
Ai Video Generator
Generate short-form videos with AI — script writing, text-to-speech narration, stock footage selection, subtitle generation, and video assembly. Use when: creating TikTok/YouTube Shorts/Reels content,…
Translator
Zoom AI Services Translator for synchronous text translation and asynchronous batch file translation. Use for plain-text translation, one-target-language jobs, S3 text archives, Build-platform JWT…
Openai Whisper Api
Transcribe audio via OpenAI Audio Transcriptions API (Whisper).
Muapi Youtube Shorts
Auto-generate viral 9:16 YouTube Shorts (or TikTok / Reels clips) from a long-form video. Thin platform-aware wrapper around the AI Clipping skill — picks sensible defaults for short-form social…
Ra Audio To Subtitles
Generate production subtitle artifacts from the final narration audio or final merged video using Volcengine Doubao ASR word timestamps. Use for local IndexTTS2 videos, Xiaohei page videos,…
Oma Video
Create short, explainer, or recorded-demo videos through the OMA video CLI. Use for scripts, narration, assets, composition, and video delivery.
Deepgram Voice
Select and tune a Deepgram TTS voice - curated voice list, full Aura voice catalog via API key, and tuning parameters
Render Myth Vs Fact
Assemble a myth-vs-fact kinetic-typography explainer video ad (≈29.5s, 9:16) from N myth/fact pairs + hook / turn / punch copy + palette + a brand end-card PNG + a VO track — a hook, 3 red-strike MYTH…
Elevenlabs Tts
This skill converts text to high-quality audio files using ElevenLabs API. Use this skill when users request text-to-speech generation, audio narration, or voice synthesis with customizable voice…
Video Processing
Process video files with ffmpeg automation. Use when: compressing videos for upload; extracting audio from video; resizing for social formats; clipping segments; merging multiple videos; generating…
Cometchat Flutter V5 Sdk
Add voice & video calling to any Flutter app FROM SCRATCH with the headless CometChat Calls SDK v5 (`cometchat_calls_sdk`, pub.dev) — no UI Kit. init→login→generateCallToken→joinSession, which hands…
Openai Whisper
Local speech-to-text with the Whisper CLI (no API key).
Omh External Connector Readiness
[omh] External connector readiness - assess whether a named plugin, connector, API, data provider, or multimodal route is safe, affordable, fresh, and observable; use executor-runtime-readiness for…
Libei Macro Hedge
Use when evaluating markets through Li Bei's macro hedge lens: regime identification, trend + contrarian timing, position management by conviction, and risk control for drawdown.
Voice Activity Detection (VAD)
Detect speech segments in audio using VAD tools like Silero VAD, SpeechBrain VAD, or WebRTC VAD. Use when preprocessing audio for speaker diarization, filtering silence, or segmenting audio into…
Oma Voice
Generate speech or transcribe audio locally with Voicebox. Use for narration, voice assets, dictation, and meeting transcription.
Render Narrated Ugc Wardrobe Stitch
Assemble a narrated-UGC "stitch reply" ad from a config — a single spoken VO carries a verbatim testimonial while ~30 per-cut i2v clips (one creator across ~5 wardrobes in ~3 worlds, plus product…
Automate This
Analyze a screen recording of a manual process and produce targeted, working automation scripts. Extracts frames and audio narration from video files, reconstructs the step-by-step workflow, and…
Short Drama Review Normalizer
Internal deterministic consent gate for meta-short-drama. Normalizes draft review and post-revision confirmation replies, and fails closed before external image/video generation.
Bio Chipseq Peak Calling
ChIP-seq peak calling using MACS3 (or MACS2). Call narrow peaks for transcription factors or broad peaks for histone modifications. Supports input control, fragment size modeling, and various output…
Ra Local Talking Head Cut
Produce a polished local talking-head or narrated screen-recording rough cut without a cloud editor. Use when Codex must clean Chinese or mixed Chinese-English speech, correct product terminology…
Render Offer Ad
Render a punchy ~12s vertical (9:16) music-only direct-response OFFER ad as a 4-beat kinetic-typography film — HEADLINE slam → real PRODUCT drop → CLAIM/proof → CTA pill — from one config of copy…
Fathom
Fetch meetings, transcripts, summaries, and action items from Fathom API. Use when user asks to get Fathom recordings, sync meeting transcripts, or fetch recent calls.
Whisper Transcription
Transcribe audio and video files to text using OpenAI Whisper. Use when: converting podcasts to blog posts; creating video subtitles; extracting quotes from interviews; repurposing video content to…
Ai News
Fetch and summarize recent AI news from curated RSS feeds (Hugging Face, VentureBeat, The Verge, OpenAI, Anthropic, DeepMind, etc.) and YouTube channels (Yannic Kilcher, Two Minute Papers, AI…
Cometchat Flutter V6 Components
The Flutter v6 UI Kit widget catalog — which drop-in widgets exist, which barrel each comes from, and how to compose or swap them (conversations, messages, users, groups, threads, calling, bubbles).…
Youtube Devrel
When the user wants to create developer YouTube content, technical screencasts, or video tutorials. Trigger phrases include "YouTube," "developer video," "screencast," "video tutorial," "live coding,"…
Competition Stego Media
Internal downstream skill for ctf-sandbox-orchestrator. CTF-sandbox workflow for image, audio, video, document, and container steganography. Use when the user asks to inspect metadata, alpha or…
Deepapi
Use DeepAPI for all web search, deep research, and web scraping (websites, LinkedIn, GitHub, X/Twitter, YouTube, Instagram) instead of built-in search, research, fetch, or browser tools. Prefer…
Ffmpeg Video Editing
Video editing with ffmpeg including cutting, trimming, concatenating segments, and re-encoding. Use when working with video files (.mp4, .mkv, .avi) for: removing segments, joining clips, extracting…
Ra Video Download
Download source video or audio from Douyin, YouTube, Bilibili, Twitter/X, Xiaohongshu, and other yt-dlp-supported URLs into the content-creation workspace. Use when the user says 下载视频, 下载音频, 保存这个链接,…
Huashu Speech Coach
演讲与分享教练。基于Patrick Winston(MIT AI教授)的How to Speak方法论,帮助准备线下培训、技术分享、B站教程视频等演讲场景。当用户提到"演讲"、"分享"、"培训"、"讲课"、"PPT演讲"、"开场"、"结尾"、"如何讲"、"演讲结构"时使用此技能。
Render Photo Grid Card
Render a 'photo-grid promo card' video from a config — a real-DOM card with a FIXED header/footer and a CONTINUOUSLY SCROLLING 2-row grid of MIXED tiles (video clips, product/lifestyle stills,…
Bankr Communities
Bankr Space ↔ bankr.bot/agents two-way sync (BANKR-PROJECT-SYNC.md Paths B+C). Original tweets from GET /agent-profiles/:id/tweets shown on Spaces. Holder votes: yes/no or multiple-choice polls…
Yuv Design System
Yuval Avidani's YUV.AI brand and design system. Apply ONLY when YUV.AI-branded output is requested — presentations, decks, keynotes, portfolio, brand site, profile, speaker bio, brand assets, or any…
Youtube Downloader
Download and process YouTube content for research. Use when: downloading competitor videos for analysis; extracting audio for podcasts; getting transcripts for content repurposing; archiving webinars;…
Check Command Injection
Analyzes PHP code for command injection vulnerabilities. Detects shell_exec, exec, system, passthru with user input, missing escapeshellarg/escapeshellcmd.
Srt From Script
Build an SRT subtitle file from a 3-shot short-drama script (ai-video-script OUTPUT FORMAT). Reads each SHOT_N block's DURATION_S + VOICEOVER, emits cumulative-timestamped SRT cues. Pure…
Maui Safe Area
.NET MAUI safe area and edge-to-edge layout guidance for .NET 10+. Covers the new SafeAreaEdges property, SafeAreaRegions enum, per-edge control, keyboard avoidance, Blazor Hybrid CSS safe areas,…
Content Engine
为X、LinkedIn、TikTok、YouTube、新闻通讯和跨平台重新利用的多平台活动创建平台原生内容系统。适用于当用户需要社交媒体帖子、帖子串、脚本、内容日历,或一个源资产在多个平台上清晰适配时。
Filler Word Processing
Process filler word annotations to generate video edit lists. Use when working with timestamp annotations for removing speech disfluencies (um, uh, like, you know) from audio/video content.
Ra Video Production Director
End-to-end video production orchestration for the content-creation workspace. Use when the user asks to make, recreate, package, render, QC, or archive a video; when the user says 制作待制作队列, 按交接稿制作,…
Elevenlabs Voice
Select and tune an ElevenLabs TTS voice - curated voice list, custom/cloned voices via API key, and tuning parameters
Render Podcast Skit
Assemble a two-host fake-podcast skit ad from a config — per-line lipsync clips hard-concatenated in script order, scaled/padded to 1080×1920, WHITE bottom-center captions (up to 5 words per cue,…
Yuv Reel Covers
Generate unified, on-brand Instagram Reel covers for Yuval (YUV.AI Neon Phoenix system) — the signature look is a giant Hebrew headline BEHIND the subject cutout + a punch line IN FRONT (depth…
Infinitetalk
自媒体创作者与内容创作者在制作数字人播报或视频配音时,只需输入单张人像与音频,即可自动生成唇形、表情、动作完美同步的无限时长说话视频。一键打造高质量虚拟主播内容,告别繁琐拍摄,让音视频创作更高效!
Fireflies Transcript
Fetch raw Fireflies.ai meeting transcripts. Use only when the user explicitly invokes /fireflies-transcript; for YouTube, use youtube-transcript.
Whisper Transcription
Transcribe audio/video to text with word-level timestamps using OpenAI Whisper. Use when you need speech-to-text with accurate timing information for each word.
