Video, long video, images, 3D models & rigged characters, game environments with Unity export, conversational avatars, music, speech, LLMs, RAG and user memory — the exact engine ContentX uses for its own commercial productions. Self-hosted GPUs answer first, frontier cloud models fill in automatically.
Pay-as-you-go · no seats, no minimums · hard budget caps per key · cancel anytime
Self-hosted first means the cheap tier is actually cheap — every number below is the all-in price you pay.
Concrete output, not abstract rates. Pick any one of these:
$25 buys
~4 minutes of AI video
enough for a full commercial, a music-video cut or 15+ social clips
$25 buys
~830 images
storyboards, thumbnails, product shots and sprite sheets
$25 buys
~600 3D meshes
or 30 fully rigged, animated game characters
$25 buys
~100M chat tokens
roughly 50,000 grounded answers in your app's knowledge base
Pay-as-you-go with no seats, no minimums and no platform fee. Every price already includes smart routing, automatic fallbacks, request tracing and per-key budget caps — the number in the table is the number on your invoice.
Our own GPU fleet serves FLUX, Wan 2.2, LTX-2, SkyReels-V2, FramePack, Trellis 2, Hunyuan3D, HY-WorldPlay, MuseTalk, FasterLivePortrait, ACE-Step, Whisper and the Qwen chat tier — near-zero marginal cost, prompts never leave our infrastructure.
Veo, Sora, Seedance, Kling, Claude, GPT, Gemini, Kimi and 400+ open models sit behind the same key. Every alias has a fallback chain, so a provider outage re-routes inside the same API call.
Every row is one line of code away: send the alias as model on the OpenAI-compatible endpoint https://litellm.intelli-verse-x.ai/v1.
Veo 3.1, Sora 2, Seedance 2.0, Kling 3.0, Wan 2.2 and LTX-2 — routed by duration, budget and reference needs. Native-audio options for trailers, UGC ads and cinematics.
| Alias | Served by | Modality | Context | Price |
|---|---|---|---|---|
veo-3.1Cinematic quality with synced dialogue & SFX baked in. | Google Veo 3.1 (native audio) Fallbacks: veo-3.1-fast → Seedance → Kling | text-to-video + image-to-video, 720p–4K | ≤8s/clip | $0.8 /sec of video |
veo-3.1-fastOur default shorts engine — 4x cheaper than Veo standard. veo-3.1-lite (from $0.10/s) also routed. | Google Veo 3.1 Fast Fallbacks: Seedance fast → Kling → Wan | text-to-video + image-to-video, 720p–4K | ≤8s/clip | $0.2 /sec of video |
sora-2sora-2-pro ($0.60/s, 1080p) also routed. | OpenAI Sora 2 Fallbacks: Veo 3.1 → Seedance | text-to-video with audio, up to ~12s | ≤12s/clip | $0.2 /sec of video |
seedance-2Up to 9 reference images for character-consistent shots — the workhorse for our cinematics. Fast tier at $0.48/s. | ByteDance Seedance 2.0 Fallbacks: seedance-2-fast → Kling → Wan | text/image/reference-to-video, up to 15s | ≤15s/clip | $0.607 /sec (720p + audio) |
kling-3.0Motion brush + element consistency (1–4 refs). $0.22/s with audio. | Kuaishou Kling 3.0 Fallbacks: Wan → Hailuo | text/image-to-video with camera control | ≤10s/clip | $0.168 /sec (audio off) |
wan-2.2Self-hosted with scale-to-zero — the cheapest clip tier on the platform. | Wan2.2 TI2V-5B · in-cluster GPUs Fallbacks: Wan 2.6 cloud → Hailuo → Hunyuan | text/image-to-video, 720p | ≤8s/clip | $0.1 /sec of video |
ltx-2Self-hosted, open-weights — priced vs the hosted LTX-2 rate. | Lightricks LTX-2 19B · in-cluster GPUs Fallbacks: Wan 2.2 → cloud clip tier | text/image-to-video with native audio, up to 4K | ≤10s/clip | $0.12 /sec (1080p) |
Self-hosted FramePack and SkyReels-V2 generate continuous 60–120 second shots that hosted APIs can't — and our pipelines stitch them into narrated 10–60 minute videos.
| Alias | Served by | Modality | Context | Price |
|---|---|---|---|---|
framepackAnti-drift long takes — no hosted API offers this length; priced vs the cheapest hosted per-second clip rate. | FramePack (HunyuanVideo) · in-cluster GPUs Fallbacks: FramePack F1 → cloud FramePack (10–30s) | image-to-video, continuous 60s+ shots | ≤60s+/shot | $0.1 /sec of video |
skyreels-v2Diffusion-forcing for infinite-length generation. Our pipelines stitch these into narrated 10–60 minute videos with TTS, music and captions. | SkyReels-V2-DF-14B 720p · in-cluster GPUs Fallbacks: FramePack → clip tier + stitching | text/image-to-video, up to 120s | ≤120s/shot | $0.1 /sec of video |
Self-hosted FLUX plus Gemini Nano Banana 2 and gpt-image-1.5 — sprites, textures, storyboards, thumbnails and store creatives.
| Alias | Served by | Modality | Context | Price |
|---|---|---|---|---|
nano-banana-2Primary storyboard & sprite engine in our own pipelines. 2K/4K output supported. | Gemini 3.1 Flash Image · Google AI Fallbacks: FLUX self-hosted → nano-banana-pro | image generation + editing (text-in-image, multi-turn) | — | $0.134 /1K-res image |
flux-devSelf-hosted with scale-to-zero — priced vs the cheapest hosted FLUX.2 [klein] rate. Sprites, textures and marketing stills. | FLUX.1 dev/schnell · in-cluster ComfyUI GPUs Fallbacks: Gemini Nano Banana 2 → PiAPI FLUX | image generation (txt2img, img2img, LoRA) | — | $0.03 /image (1MP) |
gpt-image-1.5 | OpenAI gpt-image-1.5 Fallbacks: gpt-image-1 also routed | image generation + editing | — | $16 / $64 /1M image tokens |
Self-hosted Trellis 2 and Hunyuan3D 2.5 Turbo for meshes, UniRig auto-rigging and HY-Motion text-to-animation — game-ready FBX characters, props and sprite sheets from a prompt or concept art.
| Alias | Served by | Modality | Context | Price |
|---|---|---|---|---|
trellis-3dSelf-hosted, scale-to-zero — priced vs the cheapest hosted Trellis rate. Props, characters and game assets from a prompt or concept art. | Trellis 2 · in-cluster GPUs Fallbacks: PiAPI Trellis 2 → PiAPI Trellis → Meshy.ai | text/image-to-3D mesh (GLB/OBJ/FBX) | — | $0.04 /mesh |
hunyuan3d-proPremium tier: SOTA mesh quality with PBR textures, 8–20s generation. Hero characters and close-up props. | Hunyuan3D 2.5 Turbo (10B) · in-cluster GPUs Fallbacks: Trellis 2 self-hosted → PiAPI → Meshy.ai | image/text-to-3D with PBR textures | — | $0.6 /asset |
character-3dEnd-to-end: mesh, Mixamo-compatible auto-rig, animation presets (idle/walk/run/attack), turnaround sheets and engine-ready FBX. Animation clips included. | Mesh → UniRig auto-rig → HY-Motion/Cartwheel animation · pipeline Fallbacks: Meshy.ai rig+animate → sprite-sheet 2D mode | text/image-to-rigged-and-animated character (FBX + sprite sheets) | — | $0.8 /rigged character |
hy-motionDescribe a move ('sword slash with follow-through') and get a retargetable FBX animation for your rigged character. | HY-Motion-1.0 · in-cluster GPUs Fallbacks: Cartwheel text/video-to-motion → preset library | text-to-3D-animation (SMPL/SMPLH skeleton → FBX) | — | included with character-3d · à la carte on request |
HY-WorldPlay generates geometrically consistent, navigable environment video from a single concept image and a camera path — with one-click export to Unity as skyboxes, cubemaps and Cinemachine trajectories.
| Alias | Served by | Modality | Context | Price |
|---|---|---|---|---|
world-sceneGeometrically consistent environments from one concept image with WASD-style camera paths — establishing shots, flythroughs and transition B-roll. | HY-WorldPlay (HY-World 1.5, 8B) · in-cluster GPUs Fallbacks: WAN-based 5B backbone → clip tier | image/text + camera trajectory → navigable world video | ≤60s/scene | $0.1 /sec of video |
world-unity-exportDrop a generated environment straight into Unity: skybox sphere MP4, cubemap faces and a replayable camera path. | Unity World Exporter · pipeline Fallbacks: frame sequences + manifest for custom import | world video → Unity package (video skybox, 6-face cubemap, Cinemachine trajectory JSON) | — | $0.1 /sec of source scene |
Real-time conversational avatars (Whisper → LLM → TTS → live portrait), MuseTalk lip sync for talking-head video and dubbing, and interactive avatar bundles for mobile.
| Alias | Served by | Modality | Context | Price |
|---|---|---|---|---|
avatar-liveFull self-hosted loop: Whisper STT → LLM → Kokoro TTS → live portrait at 12.8ms/frame. NPCs, tutors and support agents that talk back. | FasterLivePortrait v2 · in-cluster GPUs Fallbacks: MuseTalk pre-render → Duix mobile bundle | real-time conversational avatar (audio-driven, 30+ FPS) | streaming | $0.78 /min streamed |
musetalkTalking-head videos and dubbing — premium hosted routes run 10x this. | MuseTalk 1.5 · in-cluster GPUs Fallbacks: hosted lip-sync (Sync Labs) also routed | audio-driven lip sync on existing video (multi-language) | — | $0.6 /min of video |
avatar-interactiveEducational tutors, presenters, storytellers and support avatars — playable in real time on mobile via the Duix SDK. | Interactive Avatar Pipeline (Duix SDK) · pipeline Fallbacks: pre-rendered talking-head video → offline mobile bundle | script + character → interactive avatar bundle (mobile real-time or pre-rendered) | — | $0.78 /min rendered |
Self-hosted ACE-Step for zero-marginal-cost background music, with Google Lyria 3 as the hosted route. Loopable BGM and per-scene scoring.
| Alias | Served by | Modality | Context | Price |
|---|---|---|---|---|
ace-stepSelf-hosted, open-weights — priced vs Lyria 3 Clip. Loopable game BGM and per-scene scoring. | ACE-Step 1.5 · in-cluster GPUs Fallbacks: Lyria 3 → hosted music tier | text-to-music (BGM, stems, vocals) | — | $0.08 /30s track |
lyria-3 | Google Lyria 3 Fallbacks: ACE-Step self-hosted | text-to-music | — | $0.08 / $0.16 /30s clip · /full song |
Speech-to-text and text-to-speech — same endpoint, same key.
| Alias | Served by | Modality | Context | Price |
|---|---|---|---|---|
whisper-1217x realtime — an hour of audio transcribed for $0.22. | Self-hosted Whisper Large V3 → Groq → OpenAI Fallbacks: in-cluster STT first (zero marginal cost), Groq Whisper V3 fallback | speech-to-text | — | $0.222 /hour audio |
gpt-4o-transcribe | OpenAI transcription Fallbacks: gpt-4o-mini-transcribe ($0.006/min) also routed | speech-to-text | — | $0.012 /min audio |
tts-1 | OpenAI TTS (+ gpt-4o-mini-tts) Fallbacks: direct route | text-to-speech | — | $30 /1M chars |
hexgrad/Kokoro-82M24x cheaper than tts-1 — used for our audiobooks. | Kokoro-82M TTS · DeepInfra Fallbacks: direct route | text-to-speech (82M, natural voices) | — | $1.24 /1M chars |
Qwen-family models served by vLLM on our own GPUs with scale-to-zero. Priced at 2x the cheapest comparable hosted rate — external fallbacks answer instantly while GPUs warm.
| Alias | Served by | Modality | Context | Price |
|---|---|---|---|---|
selfhosted-chatPrimary chat workhorse — $0.24/M input tokens. Also answers as qwen3-30b, qwen3-chat. | Qwen3-30B-A3B-AWQ · in-cluster vLLM Fallbacks: OpenAI bridge → Claude Haiku → Kimi K2 | text | 32K | $0.24 in · $1 out /1M tokens |
selfhosted-voiceLow-latency tier for chatboxes, games and voice. | Qwen3-30B-A3B · in-cluster vLLM (always-fast tier) Fallbacks: DeepInfra → OpenRouter → SiliconFlow → Haiku | text | 8K | $0.24 in · $1 out /1M tokens |
selfhosted-reasonerDeliberate chain-of-thought reasoning. | QwQ-32B · in-cluster vLLM Fallbacks: OpenAI pro bridge → Claude → Kimi K2 | text (reasoning) | 32K | $0.24 in · $1 out /1M tokens |
qwen3-omniMultimodal omni model, warm-on-demand. | Qwen3-Omni-30B · in-cluster vLLM Fallbacks: OpenAI pro bridge → Claude → Kimi K2 | text + audio + vision | 32K | $0.24 in · $1 out /1M tokens |
qwen3-coder | Qwen3-Coder · in-cluster vLLM Fallbacks: OpenAI bridge → Kimi K2 | text (code) | 32K | $0.24 in · $1 out /1M tokens |
Large mixture-of-experts models for frontier-class quality at open-model prices.
| Alias | Served by | Modality | Context | Price |
|---|---|---|---|---|
selfhosted-chat-proFrontier-class open model. | Qwen3.5-122B-A10B (122B MoE) Fallbacks: OpenAI pro bridge → Claude → Haiku → Kimi K2 | text (thinking + tools, 201 languages) | 262K | $0.58 in · $4.8 out /1M tokens |
minimax-coder-pro1M-token context. | MiniMax-M3 (428B MoE, 23B active) Fallbacks: OpenAI bridge → Kimi K2 | text (code + agents) | 1M | $0.6 in · $2.4 out /1M tokens |
Anthropic Claude with automatic prompt caching (cache reads billed at 0.1x the input rate).
| Alias | Served by | Modality | Context | Price |
|---|---|---|---|---|
claude-sonnetDaily-driver frontier model with prompt caching (cache reads at 0.1x). | Claude Sonnet 4.6 · AWS Bedrock Fallbacks: Opus → Haiku → Kimi K2 → OpenAI | text + vision + tools | 200K | $6 in · $30 out /1M tokens |
claude-opusTop-tier reasoning. | Claude Opus 4.6 · AWS Bedrock Fallbacks: Haiku → Kimi K2 → OpenAI | text + vision + tools | 200K | $10 in · $50 out /1M tokens |
claude-haikuFast frontier tier. | Claude Haiku 4.5 · AWS Bedrock Fallbacks: Kimi K2 → OpenAI | text + vision | 200K | $2 in · $10 out /1M tokens |
Independent providers used both directly and as cold-start bridges for the GPU tiers.
| Alias | Served by | Modality | Context | Price |
|---|---|---|---|---|
kimi-k2Native multimodal, thinking + non-thinking. | Kimi K2.6 · Moonshot AI Fallbacks: terminal tier (always answers) | text + image + video input | 256K | $1.9 in · $8 out /1M tokens |
deepseek/deepseek-chat1M context at commodity pricing. | DeepSeek V4-Flash Fallbacks: wildcard route — any deepseek/* model id | text (thinking optional) | 1M | $0.28 in · $0.56 out /1M tokens |
gemini/gemini-3-flash-preview | Gemini 3 Flash · Google AI Fallbacks: wildcard route — any gemini/* model id | text + image + video + audio | 1M | $1 in · $6 out /1M tokens |
openrouter/*Escape hatch to virtually every hosted open model. | 400+ models · OpenRouter Fallbacks: wildcard route — any openrouter/* model id | varies | varies | priced per routed model |
Direct OpenAI routes, including the guaranteed terminal fallback for every chain.
| Alias | Served by | Modality | Context | Price |
|---|---|---|---|---|
gpt-5.4gpt-4.1 family also routed. | OpenAI GPT-5.4 Fallbacks: direct route | text + vision + tools | 400K | $5 in · $30 out /1M tokens |
Vector embeddings for search and RAG — give your app a knowledge base on the same key.
| Alias | Served by | Modality | Context | Price |
|---|---|---|---|---|
text-embedding-3-small | OpenAI embeddings Fallbacks: direct route | embeddings (1536-dim) | 8K | $0.04 /1M tokens |
text-embedding-3-large | OpenAI embeddings Fallbacks: direct route | embeddings (3072-dim) | 8K | $0.26 /1M tokens |
Pay-as-you-go, billed per request. Prices include routing, automatic fallbacks, tracing and budget caps.
Models are stateless; products shouldn't be. Combine embeddings, cheap summarization and your own vector store on one key, and every kind of memory an LLM understands is available to your app or brand.
Remember each user's preferences, goals and history across sessions — so your app greets them where they left off instead of starting cold. Store facts as embeddings, recall them into the prompt at question time.
Summarize long conversations with the cheap self-hosted tier and carry the summary forward — infinite-feeling chat history without paying for infinite context.
Embed your docs, catalog or lore once ($0.04/M tokens) and every answer is grounded in your content — an in-app knowledge base your users can ask anything.
Locked brand kits, recurring characters and style systems — the memory layer ContentX uses so every asset looks like you, across formats and over time.
Playbooks and lessons learned from past runs feed back into the next one — the system gets better at your use case instead of repeating mistakes.
Route the same request to cheaper or stronger models per user tier, tone and language — 201 languages on the pro tier, one key for all of it.
The whole loop runs on rows from this catalog: to store, to summarize and answer, frontier models when it matters. Your vectors stay in your own store.
Same models, same pricing rule — pick the interface that fits your team.
One OpenAI-compatible key for every row on this page — drop it into your app, game or pipeline. Scoped keys, hard daily budget caps, full request tracing and automatic fallbacks included.
Get your API keyDon't want to write API calls? The ContentX studio drives these exact models end to end — films, commercials, music videos, documentaries, 3D characters and game trailers — produced, approved and auto-published on your brand.
What builders and studios ask us about the catalog.
Most of the catalog (Wan, FLUX, SkyReels, Trellis, Hunyuan3D, MuseTalk, ACE-Step, Whisper, the Qwen chat tier) runs on our own GPU fleet — the same fleet that produces ContentX's own commercial work — so you're paying for output, not for a vendor's margin stack. There are no seats, no minimums and no platform fee: a $25 budget really buys ~4 minutes of video or ~100M chat tokens. Frontier models are there when you need them, on the same key.
Two ways. Builders: grab an API key from the IntelliVerse AI Gateway — one OpenAI-compatible endpoint for every model on this page, with per-key budgets and full tracing. Teams that want finished content instead of API calls: book a call and the ContentX studio runs these same models for you, end to end.
Every alias has an automatic fallback chain (shown in each table). If a self-hosted GPU is cold or a provider has an outage, the request is re-routed to the next model in the chain in the same API call — no code changes, no retries on your side.
Yes. Trellis 2 and Hunyuan3D 2.5 Turbo turn prompts or concept art into GLB/OBJ/FBX meshes from $0.04; the character pipeline auto-rigs them (Mixamo-compatible) and adds animation presets for $0.80 per character. HY-WorldPlay generates navigable 3D environments from one image and exports straight to Unity as skyboxes, cubemaps and Cinemachine camera paths.
Yes — that's the memory layer. Use embeddings (from $0.04/M tokens) to store user preferences, session summaries and your own knowledge base, then recall them into any model's prompt at question time. Your app remembers who each user is, what they said last week and what your brand sounds like — the same architecture ContentX uses internally (brand profiles, playbooks and lessons) to stay consistent across thousands of production runs. Vectors live in your own store; nothing is used to train models.
Everything: smart routing, automatic fallback chains, full request tracing, per-key budget caps and spend alerts. The price in each row is the all-in price — pay-as-you-go, no seats, no minimums, no surprise line items on your invoice. Set a hard daily cap per key and a leaked key can never burn more than one day's budget.
No. On self-hosted routes your prompts are processed on our own AWS GPUs and never leave our infrastructure. Proxied requests go only to the provider you selected, under that provider's API terms.
Start with a free key, or let the studio produce for you — either way your first video, character or knowledge base costs less than lunch.