Gemini Image Generator logo

Gemini Image Generator

Community
shinpr

MCP server for AI image generation and editing with automatic prompt optimization and quality presets. Supports Nano Banana (Gemini), OpenAI GPT Image, and BytePlus Seedream.

Publishershinpr
Repositorymcp-image
LanguageTypeScript
Forks
37
Stars
166
Available tools
1
Transport typestdio
Categories
LicenseMIT
Links
  • Connect tools to AI workflows

    Gemini Image Generator exposes MCP capabilities that can be used by compatible AI clients and agents.

  • 1 available tools

    Browse the callable actions below, including names and descriptions when provided by the server.

  • Ready-to-copy setup

    Use the installation snippets to configure this server in your preferred MCP client.

  • Open source signals

    166 stars and 37 forks from the linked repository.

MCP Image Generator 🍌

Generate and edit images from Codex, Cursor, Claude Code, or any MCP client. mcp-image adds visual direction to your request before sending it to Gemini, OpenAI, or BytePlus Seedream.

npm version npm downloads License: MIT

Tell it what image to create or what to change in an existing image, and what it is for. The result is saved to disk and returned to your assistant.

What It Does

Before generating an image, mcp-image rewrites short requests into more specific prompts. It keeps what you asked for and fills in details such as composition, lighting, and camera angle. The more detail you provide, the less it changes.

You ask:

"A photo of a roast chicken dinner for a recipe site. It should look like it was actually cooked, and it should be partway through being carved so you can tell how juicy it is."

mcp-image sends to the image model:

"... a beautifully roasted whole chicken, golden-brown and glistening, resting on a rustic wooden cutting board. One leg is partially carved, revealing tender, succulent white meat and rich, glistening juices pooling around the carving knife ... shallow depth of field focused on the carved chicken."

Roast chicken, generated with prompt enhancement

Generated with Gemini using the default fast quality preset.

What carried through:

  • for a recipe site: one clear subject, with everything else kept subordinate
  • actually cooked: uneven browning and juices across the board
  • partway through being carved: the cut face and slices beside it
  • how juicy it is: close framing and shallow depth of field around the cut

Baseline from the same request, with prompt enhancement disabled.

Set SKIP_PROMPT_ENHANCEMENT=true to send the original prompt to the image model unchanged.

Quick Start

You need Node.js 22 or later, an MCP-compatible client, and an API key for one image provider.

1. Get an API key

All three providers generate and edit images. Gemini is the default and requires the least configuration.

ProviderImage sizeOutput formatSetup
Gemini (default)1K, 2K, 4KAutomaticGet a key, then set GEMINI_API_KEY
OpenAI1K, 2K, 4KPNG or JPEGGet a key, then set IMAGE_PROVIDER=openai and OPENAI_API_KEY
BytePlus Seedream1K, 2KPNG or JPEGGet an AP region key, then set IMAGE_PROVIDER=seedream and ARK_API_KEY

Google Search grounding is available with Gemini only. OpenAI may require organization verification before it can generate images.

The examples below use Gemini. Replace the provider settings if you prefer OpenAI or Seedream.

2. Configure your MCP client

Codex

Add this to ~/.codex/config.toml:

toml
[mcp_servers.mcp-image]
command = "npx"
args = ["-y", "mcp-image"]

[mcp_servers.mcp-image.env]
GEMINI_API_KEY = "your_gemini_api_key_here"
IMAGE_OUTPUT_DIR = "/absolute/path/to/images"

Cursor

Add this to ~/.cursor/mcp.json for all projects, or .cursor/mcp.json in a project:

json
{
  "mcpServers": {
    "mcp-image": {
      "command": "npx",
      "args": ["-y", "mcp-image"],
      "env": {
        "GEMINI_API_KEY": "your_gemini_api_key_here",
        "IMAGE_OUTPUT_DIR": "/absolute/path/to/images"
      }
    }
  }
}

Claude Code

Run this in your project directory:

bash
claude mcp add mcp-image --env GEMINI_API_KEY=your-api-key --env IMAGE_OUTPUT_DIR=/absolute/path/to/images -- npx -y mcp-image

Add --scope user after mcp-image to make it available in every project.

Never commit API keys to version control. Use an absolute IMAGE_OUTPUT_DIR in MCP configuration because the server's working directory depends on the client. If omitted, images are written to ./output relative to that working directory.

3. Generate an image

Restart your MCP client after changing its configuration, then ask your AI assistant:

text
Generate a product photo of a ceramic coffee mug on a wooden desk.

The generated file is saved in the configured output directory and returned to the assistant as an MCP resource.

bash
pnpm install
pnpm run build

Configure the MCP client to run the local build instead of npx -y mcp-image:

bash
node /absolute/path/to/mcp-image/dist/index.js

More Examples

Edit an existing image

Give the assistant an absolute path to the source image:

text
Edit /path/to/image.jpg so the person is facing right.

Control the result

  • Generate a high-quality product photo of a smartphone with clear text on the screen.
  • Generate a cinematic desert landscape in a 21:9 aspect ratio.
  • Keep the knight's appearance consistent with the previous image.

See the tool reference for the options your assistant can pass explicitly.

Configuration

Changing the provider changes both prompt enhancement and image generation. The way you ask for an image stays the same.

Quality

IMAGE_QUALITY accepts fast (default), balanced, or quality. Set it in the MCP server environment:

bash
IMAGE_QUALITY=balanced

Use fast to try ideas quickly, balanced for everyday use, and quality for complex scenes or images where small details matter. Higher settings can take longer and cost more; results vary by provider.

All three providers support these presets for generation and editing. You can override the default with the quality option on each request.

Environment variables

VariableDefaultDescription
IMAGE_PROVIDERgeminiDefault provider: gemini, openai, or seedream
GEMINI_API_KEY-API key for Gemini
OPENAI_API_KEY-API key for OpenAI
ARK_API_KEY-ModelArk AP API key for Seedream
IMAGE_OUTPUT_DIR./outputDirectory where generated images are saved; use an absolute path in MCP configuration
IMAGE_QUALITYfastDefault quality preset: fast, balanced, or quality
SKIP_PROMPT_ENHANCEMENTfalseSet to true to send prompts through unchanged

You can configure keys for more than one provider and switch per request. A request-level provider option takes precedence over IMAGE_PROVIDER.

Tool Reference

Your MCP client calls this tool for you. Open the reference when you need to check an option or provider limitation.

ParameterTypeRequiredDescription
promptstringYesImage description or editing instruction
qualitystringNofast, balanced, or quality; overrides IMAGE_QUALITY
providerstringNogemini, openai, or seedream; overrides IMAGE_PROVIDER
inputImagePathstringNoAbsolute path to an input image for editing
fileNamestringNoOutput filename; .png, .jpg, or .jpeg selects the format for OpenAI and Seedream
aspectRatiostringNo1:1 (default), 2:3, 3:2, 3:4, 4:3, 4:5, 5:4, 9:16, 16:9, 21:9, 1:4, 1:8, 4:1, or 8:1
imageSizestringNo1K, 2K, or 4K; availability depends on the provider
blendImagesbooleanNoAdd blending guidance when combining visual elements
maintainCharacterConsistencybooleanNoKeep a character's appearance consistent across images
useWorldKnowledgebooleanNoAdd context for historical figures, landmarks, and factual scenes
useGoogleSearchbooleanNoGemini only. Use Google Search grounding for current information
purposestringNoIntended use, such as cookbook cover or social media post

Troubleshooting

API key not found

Check that the key for the selected provider is present in the MCP server's environment:

  • Gemini: GEMINI_API_KEY
  • OpenAI: OPENAI_API_KEY
  • Seedream: ARK_API_KEY

Restart the MCP client after changing its configuration.

Input image file not found

Use an absolute path and make sure the MCP server can read the file. Input images can be PNG, JPEG, or WebP and must be no larger than 10 MB. Seedream editing accepts PNG and JPEG only.

Provider rejects a request

Check the requested size in the provider table. useGoogleSearch works with Gemini only, and Seedream does not support 4K. For OpenAI permission errors, check your organization settings. For quota or rate-limit errors, check the selected provider account.

Image Generation Prompt Skill

This repository also includes an Agent Skill for assistants that already have access to an image generation tool. It teaches the prompt-writing approach used by mcp-image and works independently of this server.

Install it with:

bash
npx mcp-image skills install --path <skills-directory>

For example, use ~/.codex/skills, ~/.cursor/skills, or ~/.claude/skills as the destination.

License

MIT License. See LICENSE for details.


Need help? Open an issue or check Troubleshooting.

Installation

TypingMind
Prerequisites:

Node.js 18+

{
  "mcpServers": {
    "mcp-image": {
      "command": "npx",
      "args": [
        "-y",
        "mcp-image"
      ],
      "env": {
        "GEMINI_API_KEY": "your_gemini_api_key_here",
        "IMAGE_OUTPUT_DIR": "/absolute/path/to/images"
      }
    }
  }
}

Available Tools

  • generate_image

    Generate image with specified prompt and optional parameters

Use Gemini Image Generator MCP with multiple AI models

TypingMind connects MCP tools at the workspace level, so once Gemini Image Generator is connected, you can use it with different AI models in TypingMind instead of setting it up separately for each model. This MCP runs locally through the TypingMind MCP connector on your device.

Setup guide to use the local connector

Use this when the MCP server needs access to local files, apps, or private resources on your computer.

1

Open the MCP settings

In TypingMind, go to Settings, Advanced Settings, then Model Context Protocol and choose Setup Connector.

  1. Open TypingMind in your browser.
  2. Click the Settings icon.
  3. Go to Advanced Settings.
  4. Open the Model Context Protocol section.
  5. Click Setup Connector and choose This Device.
TypingMind MCP connector setup screen with This Device selected
2

Run the connector command

Choose This Device, copy the command from TypingMind, and run it in Terminal. Keep the process running while you use MCP.

  1. Copy the setup command shown by TypingMind.
  2. Open Terminal on macOS or Windows Terminal on Windows.
  3. Paste and run the command.
  4. Approve the package install if Terminal asks you to proceed.
  5. Keep the Terminal window running while using MCP tools.
3

Add Gemini Image Generator as a server

When the connector status is Ready, click Edit Servers and paste the MCP server configuration.

  1. Wait until the connector status shows Ready.
  2. Click Edit Servers.
  3. Paste the Gemini Image Generator MCP server configuration.
  4. Save the server list.
  5. Refresh if you want to confirm the connector is still ready.
TypingMind MCP settings showing active server and Edit Servers button
{
  "mcpServers": {
    "gemini-image-generator": {
      "command": "npx",
      "args": [
        "-y",
        "mcp-image"
      ]
    }
  }
}
4

Use it across models

Save the server list, open Plugins, enable the Gemini Image Generator MCP tools, then select any supported AI model in TypingMind and use the tools in chat or assign them to an AI agent.

  1. Open the Plugins page in TypingMind.
  2. Enable the Gemini Image Generator MCP tools.
  3. Start a chat and choose the AI model you want to use.
  4. Use the MCP tools in chat or assign them to an AI agent.
  5. Switch to another AI model whenever needed without reconnecting MCP.
TypingMind chat using enabled MCP tools with a selected AI model
Can you use Gemini Image Generator to help me with this task?
Gemini Image Generator
Sure. I read it.
Here is what I found using Gemini Image Generator.

Frequently asked questions

What is the Gemini Image Generator MCP server used for?

Gemini Image Generator is an MCP server that lets compatible AI clients connect to external tools and context. In TypingMind, you can add this MCP server once and make its tools available in your AI workspace.

Can I use Gemini Image Generator MCP with multiple AI models in TypingMind?

Yes. TypingMind connects MCP tools at the workspace level, so you can use Gemini Image Generator with different AI models such as Claude, ChatGPT, Gemini, or other models you have configured in TypingMind without setting up the MCP server separately for each model.

Why use Gemini Image Generator MCP with TypingMind?

TypingMind is one of the best frontends for LLM chat because it brings multiple AI models, prompts, plugins, AI agents, API keys, and MCP tools into one workspace. With Gemini Image Generator connected, you can use its MCP tools across your preferred models while keeping your chat workflow organized in TypingMind.

How do I connect Gemini Image Generator MCP to TypingMind?

Gemini Image Generator runs through the TypingMind local MCP connector. This is best when the MCP server needs access to local files, desktop apps, command-line tools, or private resources on your computer.

What tools does Gemini Image Generator MCP provide in TypingMind?

Gemini Image Generator exposes 1 MCP tools that can be enabled from the TypingMind Plugins page and used in chat or assigned to AI agents.

Do I need to share my API keys with TypingMind to use Gemini Image Generator MCP?

No. TypingMind is local-first and lets you keep your model providers, API keys, prompts, and MCP configuration under your control. If Gemini Image Generator requires authentication, add the required headers, OAuth settings, or local configuration for that MCP server when you create the connection.

Related MCP Servers

View all

Set up your own AI workspace now

Get notified about new features and future giveaways by subscribing to our newsletter 👇