Media AI Skills
708 open-source Media AI skills that teach any AI model a new workflow.
Search and filter AI skills
AI skills directory results
Huashu Video Check
基于MrBeast策略检查视频标题、封面和开头钩子。当用户提到"视频标题"、"封面图"、"点击率"、"CTR"、"观看时长"时使用。
Render Proof Points Overlay
Build the deterministic PIL pill overlays for a 'perfect-score + proof-points' UGC video ad (white 10/10 score header with a medal + orange sub with a finger-down + 3-4 green-check proof pills) and…
Short Drama
制作多集 AI 微短剧:建立剧集圣经和角色参考,完成分集剧本、逐镜 I2V、对白审计、配音字幕 BGM 与成片,保持跨镜跨集一致性。 当用户说“AI/横屏/竖屏/微短剧、拍短剧、分集剧本、连续剧情视频、做几集短剧”时使用。 单条非剧情视频用 auto-short-video;只写单条脚本用 video-script;只生成一个视频片段用 ai-video-gen。
Yuv Video Director
Yuval's all-in-one AI video pipeline. Turns an idea/script into a finished, on-brand MP4 by orchestrating HyperFrames (HTML→deterministic video render), Lottie (branded motion graphics), ManimCE (math…
Video Sdk/Linux
Zoom Video SDK for Linux - C++ headless bots, raw audio/video capture/injection, Qt/GTK integration, Docker support
Remotion
Creates programmatic video with React via Remotion — compositions, sequences, and motion graphics animated with useCurrentFrame() and rendered to MP4. USE WHEN video, animation, motion graphics, video…
Remotion Best Practices
Best practices for Remotion - Video creation in React
Subtitle Burner
Burn an SRT subtitle file into an MP4 via ffmpeg's subtitles filter (libass). Single-pass re-encode of video; audio copied as-is. Uses a verified managed Noto Sans CJK font when available. Used by…
Infinitetalk
自媒体创作者与内容创作者在制作数字人播报或视频重配音时,只需单张人像图片与音频,即可一键生成唇形、头部运动及身体姿态精准同步的无限时长说话视频。轻松搞定虚拟主播内容,大幅节省真人拍摄与后期剪辑时间!
Speech To Text
Transcribe video to timestamped text using Whisper tiny model (pre-installed).
Ra Video Wash Pipeline
End-to-end Chinese video washing pipeline. Use when the user provides a Bilibili, Douyin, Xiaohongshu, YouTube, web video URL, or local video file and asks for 视频洗稿, 视频二创, 洗稿并制作视频, 链接视频改写, 提取逐字稿后洗稿,…
Huashu Video Outline
快速生成2-3个视频大纲方案,含标题、封面建议和结构设计。当用户提到"视频大纲"、"视频结构"、"脚本大纲"、"视频选题"时使用。
Fish Audio
Generate expressive audio clips using Fish Audio S2 TTS with bracket emotion tags. Record voice memos, narration, audio messages, or any spoken content.
Render Search Grid
Render a 'search-grid' (Pinterest search-moodboard) video from a config — a real-DOM page with four continuous beats (masonry search grid + typing hook with counter-drift columns → 3 cards slide in…
Skill Account Diagnosis
账号诊断/起号体检:读取已完善的画像 Profile + 近期内容数据,诊断垂直度、定位清晰度、限流降权信号、流量池阶段,给出病因→证据→处方式的起号意见与发布建议。当用户说账号诊断/起号体检/为什么没流量/是不是被限流了/账号定位诊断/起号建议/怎么起号时使用
Zoom Video Sdk Macos
Zoom Video SDK for macOS native desktop apps. Use when building custom macOS video sessions with native UI control, tokenized join, and desktop-oriented media/device workflows.
Remotion
Remotion renderer for json-render that turns JSON timeline specs into videos. Use when working with @json-render/remotion, building video compositions from JSON, creating video catalogs, or rendering…
Sonos
Sonos control: search, queue, playlists, rooms/groups, volume, YouTube.
Sag
ElevenLabs text-to-speech with mac-style say UX.
Multicoin Thesis Vc
Use when evaluating crypto investments through a Multicoin-style thesis lens: high-conviction sector bets, performance chains, token value capture, and crypto-native market structure.
Ai Tool Picker
Figure out which AI tool actually fits the task in front of you — chatbot, coding assistant, image model, agent, or none — instead of forcing one tool onto everything. Use when asked which AI tool…
Render Song Mv
Assemble a song-driven music-video ad from a config — a generated sung track carries the whole narration across N tableaux (one keyframe -> one i2v clip per lyric beat) with NO separate voiceover,…
Zoom Video Sdk React Native
Zoom Video SDK for React Native. Use when building custom mobile video session experiences with @zoom/react-native-videosdk, event listeners, helper-based APIs, and backend JWT token flows.
Ship Learn Next
Transform learning content (like YouTube transcripts, articles, tutorials) into actionable implementation plans using the Ship-Learn-Next framework. Use when user wants to turn advice, lessons, or…
Render Split Screen Creator
Assemble a split-screen creator ad from a config — a two-zone vertical composite where a supplied AI-creator lip-sync take fills the BOTTOM ~48% while real 16:9 product/demo clips run uncropped in the…
Speech Skill
Convert text to spoken audio and transcribe audio to text, across multiple speech providers (OpenAI, ElevenLabs, Deepgram, Groq, Sarvam AI).
Audio Descriptions
Use when reviewing rendered HTML, interactive components, or design-system patterns related to Provide audio descriptions for video. Check native semantics first, then inspect keyboard behavior, focus…
Video Merger
Concatenate a directory of numbered MP4 segments (1_*.mp4, 2_*.mp4, ...) into one MP4 with optional fade transitions, unified resolution/fps/codec. Pure ffmpeg wrapper, no LLM. Trigger when a workflow…
Ad Ready
Generate advertising images automatically from a product URL + brand profile. ✅ USE WHEN: - User provides a product URL (e-commerce link) - Want automated product scraping + image generation - Have a…
Render Stopmotion Hand Swatch Cycle
Assemble a stop-motion hand-swatch-cycle product-demo ad from a config — a sequence of still PLATES (one hand swiping a single-barrel cosmetic across a cream skin-patch, the barrel + swatch changing…
Gpt Image 2
Generate and edit images using OpenAI's GPT Image 2 API. Interactive skill that guides users through image creation with style presets, cost-aware draft/final workflow, thinking mode, carousels, and…
Video Sdk Web
Build and debug browser-based Zoom Video SDK for Web integrations using @zoom/videosdk. Use for custom video sessions, joining or leaving sessions, JWT auth, audio/video and video-player rendering,…
Video Still Animator
Turn a single still image (PNG/JPG) into a short MP4 with a slow Ken-Burns zoom and a silent audio track. Pure ffmpeg wrapper. Designed as the on_failure substitute for AI video-gen steps that get…
Media Processor
影片与视频编辑及自媒体创作者在处理多媒体素材时,当需批量转换音视频或图像格式、调整分辨率及压缩文件时必用。一键调用底层工具自动完成转换与压缩,高效输出符合预期的媒体文件,大幅节省手动处理时间。
Sherpa Onnx Tts
Local text-to-speech via sherpa-onnx (offline, no cloud)
Ra 实操策划
实操长片策划工作流(B站/YouTube 新赛道,实操+真人出镜)。产出可直接照着录的策划稿:测试题组(含可粘贴 prompt)、 结构时间轴(全身出镜段 vs PiP 段)、出镜口播稿、录屏操作清单、悬念与彩蛋设计。 Use when the user says 做实操长片, 实操视频, 实测视频, 对比实测, xx做成实操, 出个实测策划, 这个选题拍实操, or ra-选题 routes a…
Common Accessibility
Enforce WCAG 2.2 AA compliance with semantic HTML, ARIA roles, keyboard navigation, and color contrast standards for web UIs. Use when building interactive components, adding form labels, fixing focus…
Prd V09 Launch Channels Orb
Allocate launch channels using the Owned / Rented / Borrowed (ORB) framework during PRD v0.9 Go-to-Market. Triggers on requests to pick launch channels, distribute the offer, build a channel…
Ai Marketing Videos
Create AI marketing videos for ads, promos, product launches, and brand content. Models: Veo, Seedance, Wan, FLUX for visuals, Kokoro for voiceover. Types: product demos, testimonials, explainers,…
Skill Evaluation
Evaluate any agent skill against a merged framework — Anthropic's Claude Code best practices plus Matt Pocock's writing-great-skills methodology — across 4 axes (Trigger, Structure, Steering,…
Video Sdk/Windows
Zoom Video SDK for Windows - C++ integration for video sessions, raw audio/video capture, screen sharing, recording, and real-time communication
Muapi Fashion Try On
Virtually try on different outfits by combining a person's photo and a clothing item, then optionally generate a professional fashion model video.
Digital Health Clinical Asr Eval
Stage 3 of Clinical ASR Flywheel. Score a NeMo manifest, produce the five-section KER leaderboard (by-ipa_source diagnostic). Not for ASR auth (/riva-asr).
Conversational Ux
Design voice and conversational interfaces — dialog flows, error recovery, and persona. Use when the interface speaks and listens rather than being tapped. For graphical input collection, use…
Adb Claw
Your eyes, hands, and ears on Android. See the screen (screenshot + indexed UI tree), interact (tap, swipe, scroll, type, clear-field), navigate via deep links (bypass CJK text input limits), wait for…
Paradigm Crypto Research
Use when evaluating crypto protocols through a Paradigm-style research lens: mechanism design, protocol fundamentals, L2/MEV/incentives, and research-driven crypto investing.
Ra 洗稿
视频脚本洗稿和二创脚本工作流。Use when the user provides an existing transcript, oral draft, viral script, competitor video text, rough topic notes, or wants 洗稿, 二创改写, 爆款视频脚本重写, 逐字稿改成短视频脚本, rewrite with ra-人话, and…
Render Vignette
Assemble a short-form 'vignette' ad from clean product cutouts composited over a kinetic background video — birefnet cutout, then a cold-open text card + product carousel + annotated specimen-sheet…
Remote Desktop Testing Linux
Test any GUI app or change on a Daytona Linux (Ubuntu xfce4 + noVNC) remote desktop sandbox. Use to launch a GUI program, sync a local project, take a screenshot, record a video, or share a clickable…
Songsee
Generate spectrograms and feature-panel visualizations from audio with the songsee CLI.
Placeholder Token Networks
Use when evaluating crypto networks through a Placeholder-style lens: open networks, token investment frameworks, network participation, and long-duration crypto network ownership.
