Gen logo

Gen

Organization
vm0-ai
gen

Use Okou generation pipelines for images, video, talking-avatar videos, voice, presentations, websites, reports, and designs.

Overview

Publishervm0-ai
Repositoryvm0-skills
Skill namegen
Stars
76
Forks
18
Bundled files
Instructions only
Links
  • Markdown instructions

    A SKILL.md file the model loads on demand, so it only costs tokens when a request actually matches.

  • Works with any LLM

    AI skills are plain Markdown, not provider-specific code, so this works with GPT, Claude, Gemini, Grok, or a local model.

  • Self-contained

    Everything the model needs lives in the instructions — no extra files to sync.

  • Open source

    Published by vm0-ai on GitHub. Read the source before you install it.

Installation

Install the Gen AI skill in TypingMind to use it with any LLM, or drop it into another agent that reads SKILL.md.

1

Install in TypingMind

TypingMind installs a skill straight from its GitHub folder — it reads SKILL.md, bundles the resource files, and stores the result locally.

  1. Open the app and go to Plugins → Skills.
  2. Choose "Install from GitHub".
  3. Paste the skill folder URL below and confirm.
  4. Enable the skill in any chat where you want it available.
Plugins → Skills → Add skill → From GitHub URL, then paste the folder URL and press Continue.
2

Install in another agent

Any agent that reads the Agent Skills format can use this skill — copy the folder into that agent's skills directory.

Claude Code — .claude/skills
git clone --depth 1 https://github.com/vm0-ai/vm0-skills.git /tmp/vm0-skills
mkdir -p .claude/skills
cp -r /tmp/vm0-skills/gen .claude/skills/gen
Restart Claude Code after copying so it picks up the new skill.

Use it in TypingMind

Enable Gen in any TypingMind chat and the model takes it from there. Its name and description sit in the system prompt, and the moment a request matches, the model loads the full instructions itself — you never invoke it by hand, and it costs no tokens until it is actually used.

The model loads Gen on its own as soon as a request matches it.

Works with any AI model

AI skills are plain Markdown instructions rather than provider-specific code, so Gen is not tied to the model it was written for. Install it once in TypingMind and use it with GPT-5, Claude, Gemini, Grok, DeepSeek, Mistral, Llama, or a local model you run yourself — all on your own API keys.

  • Loaded only when it is needed

    The system prompt carries just the name and description. The instructions are fetched on the first matching request, so an idle skill costs nothing.

  • Switch models mid-chat

    Because the skill is instructions rather than code, changing model does not break it — the next model reads the same SKILL.md.

Skill instructions

This is the SKILL.md content the model loads. Read it before installing — a skill is instructions your model will follow.

Gen

Use this skill to generate user-facing artifacts through the Okou CLI:

bash
okou generate -h

okou generate is the source of truth. Always inspect the current command help when exact flags, models, styles, or providers matter.

Core Commands

  • okou generate image - billed image file generation; supports built-in models, image editing/reference inputs, image style registry selection, and connector guidance.
  • okou generate video - billed video file generation; supports built-in video models, first/last frames, reference media, audio controls, and connector guidance.
  • okou generate avatar-video - billed JoggAI talking-avatar video generation; supports built-in public avatar and voice discovery, script or audio input, and JoggAI connector guidance.
  • okou generate voice - billed speech audio generation; supports built-in voices and connector guidance.
  • okou generate presentation - returns an Open Design resource-selection packet for an HTML presentation that the agent authors and hosts.
  • okou generate website - returns website authoring instructions / an Open Design packet that the agent uses to build and host a static site.
  • okou generate report, docs-design, poster, dashboard-design, mobile-app-design - return Open Design resource-selection packets for static HTML artifacts.
  • okou generate text, code, document, audio - list connector-backed options and print connector skill-invocation guidance; these do not have built-in platform pipelines unless the CLI help says otherwise.

Run okou generate <type> with no generation input to list available providers for that artifact type. Add --all when unavailable or not-yet-authorized connectors are relevant.

Generation Workflow

  1. Identify the artifact type from the user's request.

    • Use image for raster images, edits, references, thumbnails, icons, illustrations, and visual assets.
    • Use video for generated motion, animated frames, product clips, or reference-driven video.
    • Use avatar-video for a talking avatar driven by a narration script or public audio URL.
    • Use voice for speech audio from text.
    • Use Open Design artifact commands for websites, decks, reports, posters, docs, dashboards, and mobile UI prototypes.
    • Use connector-listing commands for text, code, document, or non-speech audio generation.
  2. Discover current capability before committing.

    • Run okou generate <type> -h for flags and built-in model support.
    • Run okou generate <type> to see available providers.
    • If the user asked for a connector/provider by name, run okou generate <type> --provider <name> to get invocation guidance, then follow that provider skill.
  3. Determine whether style discovery is needed.

    • For styled images, inspect the current style registry from okou generate image -h.
    • Do not hardcode style names or style descriptions in this skill. Treat the registry printed by the CLI as the live source.
    • Choose a registered style when the user's wording clearly matches a trigger, named style, or visual direction in the registry.
    • To obtain style-specific prompt guidance, run okou generate image --style <id> --prompt "<brief>" --compile, follow the returned packet, then generate with --compiled-prompt "<final prompt>".
    • Use --raw-prompt "<final prompt>" when the user explicitly wants no registry style, wants photorealism/model-native output, supplies a fully specified prompt, or no registry style is a good match. For video keyframes, preserve the selected video's visual direction; do not add an unrelated image style.
    • When no style is obvious but style materially affects the result, ask the user to choose among a few options summarized from the live registry or ask whether to proceed without a style.
  4. Decide provider and model.

    • Prefer --provider built-in when the user wants a direct artifact, does not name a connector, and the built-in pipeline supports the request.
    • Use connector guidance when the user names a provider, needs a provider-specific capability, or the requested artifact type is connector-only.
    • Ask the user when the choice changes cost, latency, fidelity, licensing, account usage, or final format in a way that is not implied by the request.
    • Otherwise choose a sensible default from the CLI help and proceed.
  5. Build the prompt.

    • Preserve the user's core intent, constraints, audience, brand, source materials, aspect ratio, duration, size, format, and delivery target.
    • Add operational details only when they improve generation reliability: composition, visual hierarchy, must-include/must-avoid elements, target medium, and reference handling.
    • For avatar video, discover public avatar and voice IDs through the CLI before generation. Never invent either ID, and use exactly one of script or audio URL input.
    • For style-guided image generation, let the selected registry style drive stylistic details through the compilation packet.
    • For prompt text that is long or quote-sensitive, use a file and a safely quoted argument, or stdin when the selected prompt mode supports it.
  6. Execute and wait for completion.

    • For video, follow Video preview below before submitting a video job, including template and connector routes.
    • Run the selected okou generate <type> command.
    • For commands that return an Open Design resource-selection packet, follow the packet: author the artifact, verify it locally if needed, and host static outputs with okou host.
    • For commands that return /f/ file URLs, keep the URL and metadata for the user.
    • If generation fails because of missing credits, run okou doctor credit.
    • If connector auth fails, run okou doctor check-connector using the environment name or URL from the provider guidance.
  7. Deliver the result.

    • Give the user the generated URL or hosted artifact URL.
    • Mention important parameters used: provider, model, selected style or raw prompt mode, size/aspect ratio, duration, voice, or site slug.
    • If the output is temporary or provider-hosted with expiration, download or host a durable copy when appropriate.

Video preview

  1. Before generating a video, prepare a few keyframes matching the user's subject, style, and aspect ratio. Reuse suitable supplied or already approved images.
  2. Show the actual images through accessible links, briefly describe the intended motion, and ask the user to confirm. End the turn and wait; do not start a video job before confirmation. Revise the preview if requested.
  3. After confirmation, generate the video. Use the approved images as first/last frames or references when supported by the selected model; otherwise follow the approved visual direction in the prompt. Reuse existing confirmation for an unchanged preview.

When using BytePlus/Seedance, choose one supported input mode: first/last frames (--first-frame-image-url, --last-frame-image-url) or reference media (--image-url, --video-url, --audio-url). Never combine these groups in one request. If the user's requirements need both modes, explain the tradeoff before choosing; do not silently drop supplied inputs. Correct conflicting inputs before retrying.

Keep the preview message short: the images, a brief motion description, and one confirmation question.

Asking vs. Choosing

For video, obtain the preview confirmation above before proceeding.

Ask the user before generation when:

  • The prompt is too underspecified to produce a useful artifact.
  • Multiple provider/style/model choices are plausible and materially different.
  • The command will spend meaningful credits and the request did not imply that spend.
  • The user requested brand, person, legal, medical, financial, or other high-stakes accuracy that needs source material.
  • Required inputs are missing, such as source images, brand assets, narration text, dimensions, or audience.

Proceed without asking when:

  • The user gave enough context for a reasonable first version.
  • The built-in default clearly fits the request.
  • The user asked for speed or explicitly delegated choices.
  • The missing details can be safely inferred and refined after the first artifact.

Common Patterns

List current providers:

bash
okou generate image
okou generate video
okou generate voice

Discover public JoggAI avatars and voices, then generate through the built-in pipeline:

bash
okou generate avatar-video --provider built-in --list-avatars
okou generate avatar-video --provider built-in --list-voices
okou generate avatar-video --provider built-in --avatar-id "<avatar-id>" --voice-id "<voice-id>" --script "<script>"

Get JoggAI connector skill guidance for BYOK operations:

bash
okou generate avatar-video --provider joggai

Inspect current image styles and flags:

bash
okou generate image -h

Compile a styled image prompt after selecting a live registry style, then use the packet's final prompt:

bash
okou generate image --style "<style-id>" --prompt "<brief>" --compile
okou generate image --provider built-in --compiled-prompt "<final prompt>"

Generate an unstyled/model-native image:

bash
okou generate image --provider built-in --raw-prompt "<prompt>"

Use connector guidance instead of built-in generation:

bash
okou generate video --provider "<connector-name>"

Generate a static Open Design artifact:

bash
okou generate website --prompt "<brief>"

Then follow the returned packet to build and host the artifact.

Frequently asked questions

What does the Gen AI skill do?

Use Okou generation pipelines for images, video, talking-avatar videos, voice, presentations, websites, reports, and designs.

Why use Gen on TypingMind?

Because you install it once and use it with any model. Gen is plain Markdown rather than provider-specific code, so the same skill runs on GPT-5, Claude, Gemini, Grok, or a local model — and you can switch model mid-chat without it breaking. TypingMind runs on your own API keys, so you pay providers directly instead of a per-seat subscription, and your skills and chats stay in your own storage.

How do I install Gen in TypingMind?

Open Plugins → Skills → Install from GitHub in TypingMind and paste https://github.com/vm0-ai/vm0-skills/tree/main/gen. TypingMind reads its SKILL.md and installs it as a skill you can enable per chat.

Which AI models can use Gen?

Any model you connect in TypingMind. AI skills are plain Markdown instructions rather than provider-specific code, so GPT, Claude, Gemini, Grok, and local models can all load this skill when a request matches it.

How many AI models can I use with Gen?

As many as you like. As long as a model supports skills, you can use Gen with it — GPT, Claude, Gemini, Grok, DeepSeek, Mistral, Llama and more — all on TypingMind with your own API keys.

Is the Gen AI skill free?

It is published on GitHub by vm0-ai. Check the repository for licensing terms. You only pay your own AI provider for the tokens you use.

What are AI skills?

An AI skill is a reusable instruction bundle that teaches an AI model how to do one specific task. It follows the open Agent Skills format: a SKILL.md file with a name and description, plus any scripts, templates or reference files the model may need. The model reads the instructions only when your request matches the skill, so an installed skill costs nothing until it is used.

How are AI skills different from plugins or MCP servers?

A plugin or MCP server gives a model new tools to call — code that runs somewhere and returns a result. An AI skill gives the model knowledge and process instead: how to approach a task, which steps to follow, what good output looks like. Skills are plain Markdown, so they need no server, no API key and no runtime, and they work with any model.

View all

Set up your own AI workspace now

Get notified about new features and future giveaways by subscribing to our newsletter 👇