Best Image, Video, and Audio Plugins, Skills, and MCP Servers in 2026
Research-backed image, video, and voice plugins, skills, and MCP servers from official vendors and cross-modal platforms.
Updated July 26, 2026
These extensions put real media generation behind an agent. Official vendor servers create and edit images, video, avatars, and speech; cross-modal platforms open a whole model catalog through one account; a couple of skills need no account at all. Pick by modality - the platform entries appear under every modality they serve.
Remotion’s official best-practices skill teaches the agent to build videos as React code - compositions, timing, rendering, and media handling in the Remotion framework - instead of leaving those framework rules to the model’s guesses.
Video you want versioned and reproducible as code: data-driven clips, templated social video, motion graphics tied to your product. For prompt-to-video generation, use a hosted route like Runway instead.
HeyGen’s open-source agent video framework: a CLI, skills, and plugins that take a video from planning through generation, rendering, visual inspection, and revision instead of one-shot prompting.
Video projects that mix code, media assets, narration, and browser rendering - anywhere you want the agent to inspect and revise its own output before you see it.
Higgsfield’s official cross-agent bundle of seven creation skills - spanning image, video, audio, reusable characters, product photography, website, and game-asset generation - built around its CLI and hosted platform.
Specialized creative workflows such as consistent characters, product shots, and explainer videos, from one publisher-maintained install rather than separate tools per job.
ElevenLabs’ official local MCP server, with official companion skills, covering speech synthesis, transcription, voice cloning and changing, sound effects, and music generation - one publisher-maintained audio stack instead of separate wrappers per operation.
When one integration should cover most audio jobs - narration, custom voices, effects, or generated music - billed against a single ElevenLabs account. It is also a legitimate route to AI music, which most vendors do not offer.
Anthropic’s official skill for creating original generative artwork: p5.js-oriented code, deterministic seeds, and an interactive viewer workflow. Every piece stays editable code rather than an opaque image file.
Code-driven visual experiments - posters, backgrounds, art studies - with no account or API key involved. It brings artistic process, not a hosted image model; for photo-style generation use a vendor route.
Anthropic’s official skill for designing compact looping animations that satisfy Slack’s format constraints, with validation utilities and GIF optimization built in.
Quick expressive GIFs for chat - reactions, celebrations, tiny explainers. It is not a live Slack integration and not a video editor; it makes small files that actually upload and play well.
Comfy Org’s official Comfy Cloud MCP and skills run node-based generation workflows - repeatable graphs over a broad model and node ecosystem. Despite the name, it is a generation-workflow tool, not a user-interface design tool.
When you want workflow-level control - the same graph rerun and refined - rather than a single prompt-to-image endpoint. Its verified evidence covers image and video generation, and supports audio generation as well.
Replicate’s official integration - a hosted MCP server plus eight official skills - lets the agent search thousands of hosted models, inspect their schemas, run predictions, and fetch results across image, video, and audio generation.
When you want model choice instead of one vendor: compare and run whatever the catalog offers, with no local GPU setup. As a cross-modal platform it covers all three modalities on this page from one account.
A Replicate account and API token. Every model run is pay-per-use, so the agent can spend real money - review costs and the terms of the models it picks.
Adobe’s official Claude plugin bundling 50+ tools across Photoshop, Lightroom, Illustrator, Firefly, Premiere, Express, InDesign, and Stock, so one creative task can move across several Adobe products without wiring each one up.
Edit-heavy image work - retouching, asset creation, stock, resizing - plus social and video variants of the same asset. The bundle includes video editing capability through Premiere alongside its image tools.
MiniMax’s official MCP server exposes its speech synthesis, voice cloning, image generation, video generation, and music APIs from one package - one of the few verified publisher routes that genuinely covers all three modalities on this page.
A local Python package plus a MiniMax API key that matches the regional API host; available models and regional availability differ by capability, and generation consumes metered usage.
OpenAI’s official image-generation skill: repeatable instructions for new images, edits, transparent-background work, and output verification with OpenAI image tooling. It ships first-party with Codex and installs as a portable skill elsewhere.
fal’s official hosted MCP connects the agent to more than 1,000 hosted generative-media models through nine focused tools: discover a model, inspect its price and schema, upload inputs, then run and monitor jobs.
Model breadth with cost visibility - pick the right model per job across image, video, and audio without committing to one vendor. As a cross-modal platform it covers all three modalities on this page.
A fal account and API key; usage is pay per model run. Claude Code works with the bearer-key endpoint, but Claude Desktop and claude.ai custom connectors currently cannot connect because the endpoint does not yet support OAuth.
HeyGen’s official three-skill package turns a photo or brief into reusable avatars, avatar-led videos, and translated or dubbed video, executing through the HeyGen CLI or hosted MCP.
Scripted, localized, avatar-led video - especially keeping one avatar identity consistent across many videos. For HeyGen’s broader code-driven video framework, see HyperFrames.
A HeyGen account; the skills use the CLI with an API key, or fall back to the hosted MCP with OAuth when no key is set. Generation consumes plan credits.
Black Forest Labs’ official hosted MCP brings FLUX.2 image generation, editing, variations, and browsing into the agent directly from the model’s publisher rather than through an aggregator.
When FLUX quality or its editing controls are the specific reason for the choice. The same model family is also available through Replicate and fal if you prefer a multi-model platform.
A Black Forest Labs account; generation consumes publisher credits. The route launched only weeks before this page was researched, so its adoption numbers are still early.
Runway’s official hosted MCP generates images and video with Runway’s own and selected partner models through one OAuth route - a broad creative studio billed against your existing Runway plan.
A Runway account; generation consumes plan credits. Partner models such as Kling and GPT-Image are capabilities of this one route - you reach them through Runway, not as separate integrations.
A community MCP server that gives the agent broad local control of DaVinci Resolve through Blackmagic’s official Scripting API - project, media, timeline, color, Fusion, Fairlight, rendering, and analysis workflows.
Automating real editing, grading, media organization, and render work inside a professional editor you already use, rather than calling a hosted generation service.
This is a community server, not an official Blackmagic extension. It requires the paid DaVinci Resolve Studio edition, a local Python/Node setup, and the scripting API enabled - and it has write access to your projects, media, and renders, so use it only where that level of local control is acceptable.
OpenAI’s official transcription skill: a repeatable route for transcribing audio with OpenAI tools, including diarization guidance and transcript output handling.
Turning recordings into usable transcripts inside an agent workflow. For meeting products that produce their own transcripts, look at the Communication category instead.
The host must expose the required OpenAI tooling, which has its own access and usage requirements; audio sent for transcription is processed by the OpenAI service.
AssemblyAI’s official skill gives the agent current guidance for building transcription, streaming speech, and voice-agent features with AssemblyAI’s SDKs and APIs, preventing stale-SDK and wrong-model mistakes.
Building speech features into your own product. It is developer guidance, not a turnkey transcribe-this-file tool - for that, use OpenAI Transcription or a vendor MCP.
Deepgram’s official CLI includes an MCP server exposing audio transcription, speech synthesis, text analysis, model discovery, and account usage checks through one publisher-maintained route.
The dg CLI installed and authenticated locally with a Deepgram API key; processing consumes metered usage. The route is official but new - adoption of this specific package is early.
Picsart’s official agent package exposes its image-generation and editing workflows through portable skills, with a Codex plugin as the native OpenAI route.
Leonardo.Ai’s official hosted MCP lets the agent create images through Leonardo’s generation platform and its model catalog - including catalog access to models such as Ideogram.
A Leonardo.Ai account; generation consumes publisher credits. Leonardo’s docs do not state the authentication method Claude Desktop and claude.ai custom connectors require, so treat desktop and web Claude compatibility as unverified.