Skip to main content
Updated July 26, 2026
These extensions put real media generation behind an agent. Official vendor servers create and edit images, video, avatars, and speech; cross-modal platforms open a whole model catalog through one account; a couple of skills need no account at all. Pick by modality - the platform entries appear under every modality they serve.

Remotion logo
Programmatic video in React
300Kestimated installs
What it isRemotion’s official best-practices skill teaches the agent to build videos as React code - compositions, timing, rendering, and media handling in the Remotion framework - instead of leaving those framework rules to the model’s guesses.
When to useVideo you want versioned and reproducible as code: data-driven clips, templated social video, motion graphics tied to your product. For prompt-to-video generation, use a hosted route like Runway instead.
What you needA working React/Node project with Remotion’s render toolchain; cloud rendering and some commercial uses carry separate licensing costs.
HyperFrames logo
Plan-render-review video pipeline
240Kestimated installs
What it isHeyGen’s open-source agent video framework: a CLI, skills, and plugins that take a video from planning through generation, rendering, visual inspection, and revision instead of one-shot prompting.
When to useVideo projects that mix code, media assets, narration, and browser rendering - anywhere you want the agent to inspect and revise its own output before you see it.
What you needNode 22+ and FFmpeg locally. Optional TTS, transcription, image, or cloud providers need their own credentials and can cost money.
Higgsfield logo
Seven-skill media generation bundle
100Kestimated installs
What it isHiggsfield’s official cross-agent bundle of seven creation skills - spanning image, video, audio, reusable characters, product photography, website, and game-asset generation - built around its CLI and hosted platform.
When to useSpecialized creative workflows such as consistent characters, product shots, and explainer videos, from one publisher-maintained install rather than separate tools per job.
What you needA Higgsfield account with CLI authentication; generation consumes plan credits or metered usage.
ElevenLabs logo
Speech, voices, music, sound effects
100Kestimated installs
What it isElevenLabs’ official local MCP server, with official companion skills, covering speech synthesis, transcription, voice cloning and changing, sound effects, and music generation - one publisher-maintained audio stack instead of separate wrappers per operation.
When to useWhen one integration should cover most audio jobs - narration, custom voices, effects, or generated music - billed against a single ElevenLabs account. It is also a legitimate route to AI music, which most vendors do not offer.
What you needAn ElevenLabs account and API key; the server installs as a local Python package and generation consumes plan credits or metered usage.
InstallMCPSkill
Algorithmic Art logo

Algorithmic Art

Original generative art as code
95Kestimated installs
What it isAnthropic’s official skill for creating original generative artwork: p5.js-oriented code, deterministic seeds, and an interactive viewer workflow. Every piece stays editable code rather than an opaque image file.
When to useCode-driven visual experiments - posters, backgrounds, art studies - with no account or API key involved. It brings artistic process, not a hosted image model; for photo-style generation use a vendor route.
Slack GIF Creator logo
Looping GIFs that fit Slack limits
60Kestimated installs
What it isAnthropic’s official skill for designing compact looping animations that satisfy Slack’s format constraints, with validation utilities and GIF optimization built in.
When to useQuick expressive GIFs for chat - reactions, celebrations, tiny explainers. It is not a live Slack integration and not a video editor; it makes small files that actually upload and play well.
ComfyUI logo
Node-based generation workflows
55Kestimated installs
What it isComfy Org’s official Comfy Cloud MCP and skills run node-based generation workflows - repeatable graphs over a broad model and node ecosystem. Despite the name, it is a generation-workflow tool, not a user-interface design tool.
When to useWhen you want workflow-level control - the same graph rerun and refined - rather than a single prompt-to-image endpoint. Its verified evidence covers image and video generation, and supports audio generation as well.
What you needA Comfy Cloud account with OAuth; generation consumes plan credits or metered usage.
InstallMCP
Replicate logo
Thousands of models via one MCP
35Kestimated installs
What it isReplicate’s official integration - a hosted MCP server plus eight official skills - lets the agent search thousands of hosted models, inspect their schemas, run predictions, and fetch results across image, video, and audio generation.
When to useWhen you want model choice instead of one vendor: compare and run whatever the catalog offers, with no local GPU setup. As a cross-modal platform it covers all three modalities on this page from one account.
What you needA Replicate account and API token. Every model run is pay-per-use, so the agent can spend real money - review costs and the terms of the models it picks.
InstallMCP
Adobe for Creativity logo
Adobe’s creative apps from Claude
30Kestimated installs
What it isAdobe’s official Claude plugin bundling 50+ tools across Photoshop, Lightroom, Illustrator, Firefly, Premiere, Express, InDesign, and Stock, so one creative task can move across several Adobe products without wiring each one up.
When to useEdit-heavy image work - retouching, asset creation, stock, resizing - plus social and video variants of the same asset. The bundle includes video editing capability through Premiere alongside its image tools.
What you needCurrently a Claude-only plugin. Limited signed-out use works; higher limits and full workflows may require paid Adobe access.
InstallPlugin
MiniMax logo
Speech, image, video, music in one
25Kestimated installs
What it isMiniMax’s official MCP server exposes its speech synthesis, voice cloning, image generation, video generation, and music APIs from one package - one of the few verified publisher routes that genuinely covers all three modalities on this page.
When to useWhen a single vendor account should back several modalities at once - including generated music, which few vendor routes offer.
What you needA local Python package plus a MiniMax API key that matches the regional API host; available models and regional availability differ by capability, and generation consumes metered usage.
InstallMCP
OpenAI Image Generation logo
Images via OpenAI’s native tooling
25Kestimated installs
What it isOpenAI’s official image-generation skill: repeatable instructions for new images, edits, transparent-background work, and output verification with OpenAI image tooling. It ships first-party with Codex and installs as a portable skill elsewhere.
When to useWhen the agent already has OpenAI image tooling available and you want a packaged workflow rather than a separate service connection.
What you needThe skill itself is free, but the host must expose OpenAI’s image tool or API, which has its own access and usage requirements.
InstallSkill
fal logo
1,000+ hosted models, pay per run
25Kestimated installs
What it isfal’s official hosted MCP connects the agent to more than 1,000 hosted generative-media models through nine focused tools: discover a model, inspect its price and schema, upload inputs, then run and monitor jobs.
When to useModel breadth with cost visibility - pick the right model per job across image, video, and audio without committing to one vendor. As a cross-modal platform it covers all three modalities on this page.
What you needA fal account and API key; usage is pay per model run. Claude Code works with the bearer-key endpoint, but Claude Desktop and claude.ai custom connectors currently cannot connect because the endpoint does not yet support OAuth.
InstallPluginMCP
HeyGen logo
Avatar videos and dubbing
25Kestimated installs
What it isHeyGen’s official three-skill package turns a photo or brief into reusable avatars, avatar-led videos, and translated or dubbed video, executing through the HeyGen CLI or hosted MCP.
When to useScripted, localized, avatar-led video - especially keeping one avatar identity consistent across many videos. For HeyGen’s broader code-driven video framework, see HyperFrames.
What you needA HeyGen account; the skills use the CLI with an API key, or fall back to the hosted MCP with OAuth when no key is set. Generation consumes plan credits.
Recraft logo
Production graphics and vector work
25Kestimated installs
What it isRecraft’s official hosted MCP exposes image generation plus design-oriented editing - vectorization, upscaling, background work, and custom brand styles.
When to useProduction-ready graphic assets: brand-consistent raster and vector output where the deliverable matters more than raw model breadth.
What you needA Recraft account via OAuth; the hosted route consumes subscription credits and is separate from Recraft’s local API-unit route.
InstallMCP
FLUX logo

FLUX

Direct FLUX.2 generation and editing
25Kestimated installs
What it isBlack Forest Labs’ official hosted MCP brings FLUX.2 image generation, editing, variations, and browsing into the agent directly from the model’s publisher rather than through an aggregator.
When to useWhen FLUX quality or its editing controls are the specific reason for the choice. The same model family is also available through Replicate and fal if you prefer a multi-model platform.
What you needA Black Forest Labs account; generation consumes publisher credits. The route launched only weeks before this page was researched, so its adoption numbers are still early.
InstallMCP
Runway logo
Runway image and video generation
25Kestimated installs
What it isRunway’s official hosted MCP generates images and video with Runway’s own and selected partner models through one OAuth route - a broad creative studio billed against your existing Runway plan.
When to useHosted image and video generation on a Runway plan you already have, without managing separate model integrations.
What you needA Runway account; generation consumes plan credits. Partner models such as Kling and GPT-Image are capabilities of this one route - you reach them through Runway, not as separate integrations.
InstallMCP
DaVinci Resolve MCP logo

DaVinci Resolve MCP

Agent control of Resolve editing
20Kestimated installs
What it isA community MCP server that gives the agent broad local control of DaVinci Resolve through Blackmagic’s official Scripting API - project, media, timeline, color, Fusion, Fairlight, rendering, and analysis workflows.
When to useAutomating real editing, grading, media organization, and render work inside a professional editor you already use, rather than calling a hosted generation service.
What you needThis is a community server, not an official Blackmagic extension. It requires the paid DaVinci Resolve Studio edition, a local Python/Node setup, and the scripting API enabled - and it has write access to your projects, media, and renders, so use it only where that level of local control is acceptable.
InstallMCP
OpenAI Transcription logo
Transcripts with diarization guidance
15Kestimated installs
What it isOpenAI’s official transcription skill: a repeatable route for transcribing audio with OpenAI tools, including diarization guidance and transcript output handling.
When to useTurning recordings into usable transcripts inside an agent workflow. For meeting products that produce their own transcripts, look at the Communication category instead.
What you needThe host must expose the required OpenAI tooling, which has its own access and usage requirements; audio sent for transcription is processed by the OpenAI service.
InstallSkill
AssemblyAI logo
Building transcription features
15Kestimated installs
What it isAssemblyAI’s official skill gives the agent current guidance for building transcription, streaming speech, and voice-agent features with AssemblyAI’s SDKs and APIs, preventing stale-SDK and wrong-model mistakes.
When to useBuilding speech features into your own product. It is developer guidance, not a turnkey transcribe-this-file tool - for that, use OpenAI Transcription or a vendor MCP.
What you needAn AssemblyAI API key for actual transcription work; processing consumes metered API usage.
InstallSkill
Deepgram logo
Speech-to-text and TTS via one CLI
15Kestimated installs
What it isDeepgram’s official CLI includes an MCP server exposing audio transcription, speech synthesis, text analysis, model discovery, and account usage checks through one publisher-maintained route.
When to useSpeech input and output backed by Deepgram’s APIs - transcribe audio in, synthesize speech out - from a single package.
What you needThe dg CLI installed and authenticated locally with a Deepgram API key; processing consumes metered usage. The route is official but new - adoption of this specific package is early.
InstallMCP
Picsart logo
Picsart creative API workflows
10Kestimated installs
What it isPicsart’s official agent package exposes its image-generation and editing workflows through portable skills, with a Codex plugin as the native OpenAI route.
When to useTeams already using Picsart’s creative APIs who want the agent wired to Picsart tooling rather than a generic model endpoint.
What you needPicsart API credentials where the workflows call the service; generation consumes plan credits or metered usage.
Cartesia logo
Low-latency voice and TTS
8Kestimated installs
What it isCartesia’s official MCP server and companion skills expose speech generation, voice management, and related Cartesia API operations to the agent.
When to useFocused low-latency voice and text-to-speech work. For a broader audio suite including music and sound effects, ElevenLabs covers more ground.
What you needA Cartesia account and API key; generation consumes plan credits or metered usage.
InstallMCP
Leonardo.Ai logo
Leonardo’s model catalog by MCP
5Kestimated installs
What it isLeonardo.Ai’s official hosted MCP lets the agent create images through Leonardo’s generation platform and its model catalog - including catalog access to models such as Ideogram.
When to useTeams already on Leonardo’s platform and production API who want the same account and model catalog behind the agent.
What you needA Leonardo.Ai account; generation consumes publisher credits. Leonardo’s docs do not state the authentication method Claude Desktop and claude.ai custom connectors require, so treat desktop and web Claude compatibility as unverified.
InstallMCP

Choose by use case

Want one integration instead of a separate tool per model?

Replicate - thousands of hosted models across image, video, and audio, with official skills

fal - 1,000+ models with price and schema inspection before each run

MiniMax - one vendor’s own speech, image, video, and music behind a single key

ComfyUI - repeatable node-based workflows rather than single endpoints

All four bill for usage - review model costs before letting the agent run them.

Making still images or graphics?

Recraft for production-ready graphics, vector work, and brand styles

FLUX for direct publisher access to the FLUX.2 family (also on Replicate and fal)

OpenAI Image Generation when the agent already has OpenAI image tooling

Adobe for Creativity for edit-heavy retouching and assets across the Adobe suite

Leonardo.Ai and Picsart put an existing vendor account behind the agent

No legitimate Midjourney route exists - FLUX and the multi-model platforms are the honest paths.

Producing video?

Runway for hosted generation on a plan you already pay for

HeyGen for avatar-led, translated, and dubbed video with one consistent identity

Higgsfield for characters, product shots, and explainer workflows

HyperFrames for a plan-generate-render-review pipeline the agent drives

Remotion when the video is versioned React code

DaVinci Resolve MCP to automate edits inside Resolve Studio rather than generate clips

Generating voice, speech, or music?

ElevenLabs - one official stack for TTS, voice cloning, sound effects, and music

Cartesia - focused low-latency voice and text-to-speech

MiniMax - speech and music alongside its image and video, from one key

No official Suno or Udio route exists - ElevenLabs and MiniMax are the legitimate music routes.

Turning recordings into text?

OpenAI Transcription - a packaged transcribe-a-file workflow with diarization guidance

Deepgram - speech in and speech out through one CLI-backed MCP

AssemblyAI - guidance for building transcription into your own product

Meeting products that make their own transcripts live in Communication.

Need visuals without any paid service?

Algorithmic Art - original generative artwork as editable code, no API key

Slack GIF Creator - compact looping GIFs for chat, no account needed

Design & UI

Interface design, presentations, and canvas tools live there.

Writing

Prose drafting, style, and translation.

Communication

Meeting-transcription products live there.