Skip to main content
Updated July 26, 2026
These extensions put real media generation behind an agent. Official vendor servers create and edit images, video, avatars, and speech; cross-modal platforms open a whole model catalog through one account; a couple of skills need no account at all. Pick by modality - the platform entries appear under every modality they serve.

Remotion logo
Programmatic video in React
300Kestimated installs

What it is

Remotion’s official best-practices skill teaches the agent to build videos as React code - compositions, timing, rendering, and media handling in the Remotion framework - instead of leaving those framework rules to the model’s guesses.

When to use

Video you want versioned and reproducible as code: data-driven clips, templated social video, motion graphics tied to your product. For prompt-to-video generation, use a hosted route like Runway instead.

What you need

A working React/Node project with Remotion’s render toolchain; cloud rendering and some commercial uses carry separate licensing costs.

HyperFrames logo
Plan-render-review video pipeline
240Kestimated installs

What it is

HeyGen’s open-source agent video framework: a CLI, skills, and plugins that take a video from planning through generation, rendering, visual inspection, and revision instead of one-shot prompting.

When to use

Video projects that mix code, media assets, narration, and browser rendering - anywhere you want the agent to inspect and revise its own output before you see it.

What you need

Node 22+ and FFmpeg locally. Optional TTS, transcription, image, or cloud providers need their own credentials and can cost money.

Higgsfield logo
Seven-skill media generation bundle
100Kestimated installs

What it is

Higgsfield’s official cross-agent bundle of seven creation skills - spanning image, video, audio, reusable characters, product photography, website, and game-asset generation - built around its CLI and hosted platform.

When to use

Specialized creative workflows such as consistent characters, product shots, and explainer videos, from one publisher-maintained install rather than separate tools per job.

What you need

A Higgsfield account with CLI authentication; generation consumes plan credits or metered usage.

ElevenLabs logo
Speech, voices, music, sound effects
100Kestimated installs

What it is

ElevenLabs’ official local MCP server, with official companion skills, covering speech synthesis, transcription, voice cloning and changing, sound effects, and music generation - one publisher-maintained audio stack instead of separate wrappers per operation.

When to use

When one integration should cover most audio jobs - narration, custom voices, effects, or generated music - billed against a single ElevenLabs account. It is also a legitimate route to AI music, which most vendors do not offer.

What you need

An ElevenLabs account and API key; the server installs as a local Python package and generation consumes plan credits or metered usage.

Algorithmic Art logo

Algorithmic Art

Original generative art as code
95Kestimated installs

What it is

Anthropic’s official skill for creating original generative artwork: p5.js-oriented code, deterministic seeds, and an interactive viewer workflow. Every piece stays editable code rather than an opaque image file.

When to use

Code-driven visual experiments - posters, backgrounds, art studies - with no account or API key involved. It brings artistic process, not a hosted image model; for photo-style generation use a vendor route.

Slack GIF Creator logo
Looping GIFs that fit Slack limits
60Kestimated installs

What it is

Anthropic’s official skill for designing compact looping animations that satisfy Slack’s format constraints, with validation utilities and GIF optimization built in.

When to use

Quick expressive GIFs for chat - reactions, celebrations, tiny explainers. It is not a live Slack integration and not a video editor; it makes small files that actually upload and play well.

ComfyUI logo
Node-based generation workflows
55Kestimated installs

What it is

Comfy Org’s official Comfy Cloud MCP and skills run node-based generation workflows - repeatable graphs over a broad model and node ecosystem. Despite the name, it is a generation-workflow tool, not a user-interface design tool.

When to use

When you want workflow-level control - the same graph rerun and refined - rather than a single prompt-to-image endpoint. Its verified evidence covers image and video generation, and supports audio generation as well.

What you need

A Comfy Cloud account with OAuth; generation consumes plan credits or metered usage.

Install

MCP
Replicate logo
Thousands of models via one MCP
35Kestimated installs

What it is

Replicate’s official integration - a hosted MCP server plus eight official skills - lets the agent search thousands of hosted models, inspect their schemas, run predictions, and fetch results across image, video, and audio generation.

When to use

When you want model choice instead of one vendor: compare and run whatever the catalog offers, with no local GPU setup. As a cross-modal platform it covers all three modalities on this page from one account.

What you need

A Replicate account and API token. Every model run is pay-per-use, so the agent can spend real money - review costs and the terms of the models it picks.

Install

MCP
Adobe for Creativity logo
Adobe’s creative apps from Claude
30Kestimated installs

What it is

Adobe’s official Claude plugin bundling 50+ tools across Photoshop, Lightroom, Illustrator, Firefly, Premiere, Express, InDesign, and Stock, so one creative task can move across several Adobe products without wiring each one up.

When to use

Edit-heavy image work - retouching, asset creation, stock, resizing - plus social and video variants of the same asset. The bundle includes video editing capability through Premiere alongside its image tools.

What you need

Currently a Claude-only plugin. Limited signed-out use works; higher limits and full workflows may require paid Adobe access.

Install

Plugin
MiniMax logo
Speech, image, video, music in one
25Kestimated installs

What it is

MiniMax’s official MCP server exposes its speech synthesis, voice cloning, image generation, video generation, and music APIs from one package - one of the few verified publisher routes that genuinely covers all three modalities on this page.

When to use

When a single vendor account should back several modalities at once - including generated music, which few vendor routes offer.

What you need

A local Python package plus a MiniMax API key that matches the regional API host; available models and regional availability differ by capability, and generation consumes metered usage.

Install

MCP
OpenAI Image Generation logo
Images via OpenAI’s native tooling
25Kestimated installs

What it is

OpenAI’s official image-generation skill: repeatable instructions for new images, edits, transparent-background work, and output verification with OpenAI image tooling. It ships first-party with Codex and installs as a portable skill elsewhere.

When to use

When the agent already has OpenAI image tooling available and you want a packaged workflow rather than a separate service connection.

What you need

The skill itself is free, but the host must expose OpenAI’s image tool or API, which has its own access and usage requirements.

Install

Skill
fal logo
1,000+ hosted models, pay per run
25Kestimated installs

What it is

fal’s official hosted MCP connects the agent to more than 1,000 hosted generative-media models through nine focused tools: discover a model, inspect its price and schema, upload inputs, then run and monitor jobs.

When to use

Model breadth with cost visibility - pick the right model per job across image, video, and audio without committing to one vendor. As a cross-modal platform it covers all three modalities on this page.

What you need

A fal account and API key; usage is pay per model run. Claude Code works with the bearer-key endpoint, but Claude Desktop and claude.ai custom connectors currently cannot connect because the endpoint does not yet support OAuth.

HeyGen logo
Avatar videos and dubbing
25Kestimated installs

What it is

HeyGen’s official three-skill package turns a photo or brief into reusable avatars, avatar-led videos, and translated or dubbed video, executing through the HeyGen CLI or hosted MCP.

When to use

Scripted, localized, avatar-led video - especially keeping one avatar identity consistent across many videos. For HeyGen’s broader code-driven video framework, see HyperFrames.

What you need

A HeyGen account; the skills use the CLI with an API key, or fall back to the hosted MCP with OAuth when no key is set. Generation consumes plan credits.

Recraft logo
Production graphics and vector work
25Kestimated installs

What it is

Recraft’s official hosted MCP exposes image generation plus design-oriented editing - vectorization, upscaling, background work, and custom brand styles.

When to use

Production-ready graphic assets: brand-consistent raster and vector output where the deliverable matters more than raw model breadth.

What you need

A Recraft account via OAuth; the hosted route consumes subscription credits and is separate from Recraft’s local API-unit route.

Install

MCP
FLUX logo

FLUX

Direct FLUX.2 generation and editing
25Kestimated installs

What it is

Black Forest Labs’ official hosted MCP brings FLUX.2 image generation, editing, variations, and browsing into the agent directly from the model’s publisher rather than through an aggregator.

When to use

When FLUX quality or its editing controls are the specific reason for the choice. The same model family is also available through Replicate and fal if you prefer a multi-model platform.

What you need

A Black Forest Labs account; generation consumes publisher credits. The route launched only weeks before this page was researched, so its adoption numbers are still early.

Install

MCP
Runway logo
Runway image and video generation
25Kestimated installs

What it is

Runway’s official hosted MCP generates images and video with Runway’s own and selected partner models through one OAuth route - a broad creative studio billed against your existing Runway plan.

When to use

Hosted image and video generation on a Runway plan you already have, without managing separate model integrations.

What you need

A Runway account; generation consumes plan credits. Partner models such as Kling and GPT-Image are capabilities of this one route - you reach them through Runway, not as separate integrations.

Install

MCP
DaVinci Resolve MCP logo

DaVinci Resolve MCP

Agent control of Resolve editing
20Kestimated installs

What it is

A community MCP server that gives the agent broad local control of DaVinci Resolve through Blackmagic’s official Scripting API - project, media, timeline, color, Fusion, Fairlight, rendering, and analysis workflows.

When to use

Automating real editing, grading, media organization, and render work inside a professional editor you already use, rather than calling a hosted generation service.

What you need

This is a community server, not an official Blackmagic extension. It requires the paid DaVinci Resolve Studio edition, a local Python/Node setup, and the scripting API enabled - and it has write access to your projects, media, and renders, so use it only where that level of local control is acceptable.

Install

MCP
OpenAI Transcription logo
Transcripts with diarization guidance
15Kestimated installs

What it is

OpenAI’s official transcription skill: a repeatable route for transcribing audio with OpenAI tools, including diarization guidance and transcript output handling.

When to use

Turning recordings into usable transcripts inside an agent workflow. For meeting products that produce their own transcripts, look at the Communication category instead.

What you need

The host must expose the required OpenAI tooling, which has its own access and usage requirements; audio sent for transcription is processed by the OpenAI service.

Install

Skill
AssemblyAI logo
Building transcription features
15Kestimated installs

What it is

AssemblyAI’s official skill gives the agent current guidance for building transcription, streaming speech, and voice-agent features with AssemblyAI’s SDKs and APIs, preventing stale-SDK and wrong-model mistakes.

When to use

Building speech features into your own product. It is developer guidance, not a turnkey transcribe-this-file tool - for that, use OpenAI Transcription or a vendor MCP.

What you need

An AssemblyAI API key for actual transcription work; processing consumes metered API usage.

Install

Skill
Deepgram logo
Speech-to-text and TTS via one CLI
15Kestimated installs

What it is

Deepgram’s official CLI includes an MCP server exposing audio transcription, speech synthesis, text analysis, model discovery, and account usage checks through one publisher-maintained route.

When to use

Speech input and output backed by Deepgram’s APIs - transcribe audio in, synthesize speech out - from a single package.

What you need

The dg CLI installed and authenticated locally with a Deepgram API key; processing consumes metered usage. The route is official but new - adoption of this specific package is early.

Install

MCP
Picsart logo
Picsart creative API workflows
10Kestimated installs

What it is

Picsart’s official agent package exposes its image-generation and editing workflows through portable skills, with a Codex plugin as the native OpenAI route.

When to use

Teams already using Picsart’s creative APIs who want the agent wired to Picsart tooling rather than a generic model endpoint.

What you need

Picsart API credentials where the workflows call the service; generation consumes plan credits or metered usage.

Cartesia logo
Low-latency voice and TTS
8Kestimated installs

What it is

Cartesia’s official MCP server and companion skills expose speech generation, voice management, and related Cartesia API operations to the agent.

When to use

Focused low-latency voice and text-to-speech work. For a broader audio suite including music and sound effects, ElevenLabs covers more ground.

What you need

A Cartesia account and API key; generation consumes plan credits or metered usage.

Install

MCP
Leonardo.Ai logo
Leonardo’s model catalog by MCP
5Kestimated installs

What it is

Leonardo.Ai’s official hosted MCP lets the agent create images through Leonardo’s generation platform and its model catalog - including catalog access to models such as Ideogram.

When to use

Teams already on Leonardo’s platform and production API who want the same account and model catalog behind the agent.

What you need

A Leonardo.Ai account; generation consumes publisher credits. Leonardo’s docs do not state the authentication method Claude Desktop and claude.ai custom connectors require, so treat desktop and web Claude compatibility as unverified.

Install

MCP

Choose by use case

Want one integration instead of a separate tool per model?

Replicate - thousands of hosted models across image, video, and audio, with official skills

fal - 1,000+ models with price and schema inspection before each run

MiniMax - one vendor’s own speech, image, video, and music behind a single key

ComfyUI - repeatable node-based workflows rather than single endpoints

All four bill for usage - review model costs before letting the agent run them.

Making still images or graphics?

Recraft for production-ready graphics, vector work, and brand styles

FLUX for direct publisher access to the FLUX.2 family (also on Replicate and fal)

OpenAI Image Generation when the agent already has OpenAI image tooling

Adobe for Creativity for edit-heavy retouching and assets across the Adobe suite

Leonardo.Ai and Picsart put an existing vendor account behind the agent

No legitimate Midjourney route exists - FLUX and the multi-model platforms are the honest paths.

Producing video?

Runway for hosted generation on a plan you already pay for

HeyGen for avatar-led, translated, and dubbed video with one consistent identity

Higgsfield for characters, product shots, and explainer workflows

HyperFrames for a plan-generate-render-review pipeline the agent drives

Remotion when the video is versioned React code

DaVinci Resolve MCP to automate edits inside Resolve Studio rather than generate clips

Generating voice, speech, or music?

ElevenLabs - one official stack for TTS, voice cloning, sound effects, and music

Cartesia - focused low-latency voice and text-to-speech

MiniMax - speech and music alongside its image and video, from one key

No official Suno or Udio route exists - ElevenLabs and MiniMax are the legitimate music routes.

Turning recordings into text?

OpenAI Transcription - a packaged transcribe-a-file workflow with diarization guidance

Deepgram - speech in and speech out through one CLI-backed MCP

AssemblyAI - guidance for building transcription into your own product

Meeting products that make their own transcripts live in Communication.

Need visuals without any paid service?

Algorithmic Art - original generative artwork as editable code, no API key

Slack GIF Creator - compact looping GIFs for chat, no account needed

Design & UI

Interface design, presentations, and canvas tools live there.

Writing

Prose drafting, style, and translation.

Communication

Meeting-transcription products live there.