Text-only LLM serving is a solved problem, at least in the sense that the tooling, playbooks, and failure modes are now well-documented. You pick vLLM or SGLang, size your GPU nodes, tune your batch size, and wire it into your gateway. I have been doing this long enough, twenty years of building infrastructure across every cloud shape that has existed, to know when a problem has become routine.
Multimodal serving is not routine. Not yet. The teams I talk to in 2026 are deploying vision-language models, audio processing pipelines, and video understanding models into production and discovering that every assumption they built on top of pure-text inference is wrong in some way. GPU memory math that worked for their LLM does not work when you add a vision encoder. Batch sizing logic that kept latency predictable breaks when images have wildly different token counts. The failure cascade in a tightly coupled audio-to-text-to-LLM pipeline hits differently than anything they tested.
This article is the guide I wish existed eighteen months ago when I helped the first team in my orbit go multimodal in production. I am going to cover the architectural patterns that actually work, the GPU memory traps that will bite you, the serving stack choices as of late 2026, and the cost reality you need to model before you commit hardware.
What “Multimodal” Actually Means for Your Infrastructure
Marketing uses “multimodal” to mean anything that accepts more than one type of input. For infrastructure purposes, I split it into three categories that have meaningfully different serving characteristics:
Vision-language models (VLMs) accept images alongside text. The canonical architecture connects a vision encoder (typically a CLIP-style ViT) to a language model backbone. Models like LLaVA, Qwen3-VL, and InternVL sit in this category. In 2026, the frontier models (GPT-4o, Gemini Flash Vision, Claude 3.5+ Sonnet) have moved toward early fusion where modality inputs are blended in the token stream from the start rather than encoded separately. Open-weight models are following the same pattern.
Audio models split into two directions: speech-to-text (Whisper-style transcription and speaker diarization) and text-to-speech/voice generation. These are typically separate model weights that you compose into a pipeline rather than a unified architecture. The real-time voice AI stack I covered in real-time AI voice infrastructure treats these as pipeline stages connected over WebRTC. In batch or near-real-time contexts, you run them differently.
Video understanding and generation is the newest category to reach production maturity. Understanding models process video frames as sequences of image tokens. Generation models (text-to-video) are a different beast entirely, consuming enormous GPU memory and compute, mostly still on dedicated infrastructure rather than shared serving pools.
Most production teams start with VLMs for document understanding or visual search, then later need to add audio. Very few teams are running video generation in the same cluster as interactive serving; those workloads are still typically isolated.

The GPU Memory Math That Will Break Your Plans
Here is the thing that catches teams flat-footed: the VRAM footprint of a VLM is not simply the base model weight plus some overhead. Image resolution determines token count in ways that can inflate your memory usage by an order of magnitude depending on what your users send.
When a model like LLaVA or Qwen-VL encodes an image using a CLIP ViT encoder with a 336x336 patch resolution, a single standard image becomes 576 vision tokens. Bump the input resolution to 672x672 for higher-quality document understanding, and you are looking at up to 2,880 tokens from that single image alone. If a user sends a multi-page PDF with six high-resolution images, you have added potentially 17,000 tokens to what would otherwise be a short prompt. Vision tokens are not free; they consume KV cache just like text tokens.
The consequence: the KV cache you pre-allocated for your LLM serving under text-only assumptions gets evicted or you hit out-of-memory errors when vision input volume spikes. The vision encoder runs its forward pass and the resulting vision token tensors compete directly with the pre-allocated KV cache blocks. This is not a theoretical concern. It is the most common production incident I see in multimodal deployments.
Practical mitigations:
Cap resolution and image count at the gateway. A hard limit on input resolution (max 1024x1024 before resizing) and image count per request (max four images) prevents the tail of your distribution from blowing up your serving pool. This belongs in your AI gateway layer; see the AI gateway architecture guide for where to put these guardrails. Do this before your first production deployment, not after your first incident.
Size KV cache for multimodal workloads, not text workloads. When you provision your GPU nodes, run load tests with representative multimodal traffic, not text prompts. The memory pressure is categorically different.
Consider separate vision encoder instances. The disaggregated approach, running vision encoding as a separate service stage before the LLM, trades network latency for memory isolation. A Kimi-VL study published in mid-2026 that ran the vision encoder on separate hardware from the language model measured a substantial improvement in throughput and mean time-to-first-token under load compared to tightly coupled serving. The tradeoff is operational complexity: you now have two services to manage, scale, and debug. Whether that is worth it depends on your traffic volume and latency requirements.
Serving Engine Choices in Late 2026
The LLM inference engine landscape has consolidated around vLLM and SGLang for open-weight serving, and both have matured their multimodal support considerably through 2026.
vLLM added PagedAttention-based KV cache management for vision tokens in its multimodal support. The key operational difference from text-only vLLM: you need to tune --max-model-len carefully because that limit now needs to account for vision tokens in your longest expected requests, not just text. The vLLM team has published guidance on image-to-token counts for common encoder configurations; those numbers should inform your capacity planning.
SGLang handles multimodal through its RadixAttention mechanism and has strong support for the Qwen-VL family and LLaVA variants. The SGLang approach to prefix caching works for shared image tokens, meaning if many requests share the same system image (a common pattern in document processing where you always include a template image), the prefix cache hit rate can substantially reduce prefill cost.
BentoML takes a different approach: it gives you a composable serving framework where you define a pipeline of models and it handles routing, batching, and scaling for each stage independently. For teams building pipelines that chain vision encoding, LLM generation, and audio TTS, BentoML’s architecture fits naturally. The operational overhead is higher than a single vLLM instance, but the flexibility to scale each stage independently is valuable.
TensorRT-LLM from NVIDIA has multimodal support for the models in NVIDIA’s supported roster. If you are locked to NVIDIA hardware and prioritize raw throughput above flexibility, TensorRT-LLM’s compiler optimizations can extract meaningful performance gains. The model coverage is narrower than vLLM and the build workflow adds friction, but for high-volume serving of a specific model, the numbers are compelling.
For most teams, I recommend starting with vLLM for VLM serving. It has the widest model coverage, the largest production community, and the best documentation for the corner cases you will hit. Reach for BentoML when you are building a true multi-model pipeline, and consider TensorRT-LLM only when you have a specific high-throughput VLM that it supports and you can justify the build complexity.
Audio Pipeline Architecture
Audio models follow a different pattern than VLMs because they are almost always used as pipeline stages rather than endpoints in their own right. The typical voice AI pipeline looks like: raw audio input, transcription (Whisper or a fine-tuned variant), LLM processing, TTS output. Each stage is a separate model with separate scaling characteristics.
Transcription (STT): Whisper-large-v3 remains the workhorse for offline transcription as of 2026. For real-time transcription, smaller models (Whisper-base, Whisper-medium) or streaming-optimized alternatives trade accuracy for latency. Faster-Whisper, the CTranslate2-based reimplementation, significantly reduces memory footprint and increases throughput for batch transcription workloads compared to the original PyTorch implementation.
GPU vs. CPU for audio: STT models at the smaller end (base, small, medium) can run efficiently on CPU for moderate throughput. This matters for capacity planning: you may not need to put your transcription pipeline on the same GPU nodes as your LLM. Separating audio preprocessing onto CPU-optimized instances and reserving GPU capacity for LLM serving often yields better overall utilization and lower cost.
TTS models: Text-to-speech has seen significant model quality improvements in 2026, with several open-weight models approaching commercial quality. The infrastructure challenge is that TTS is inherently streaming (you want to start playing audio before the full response is generated), which pushes you toward a chunked generation pattern rather than waiting for a complete response. VRAM requirements for production TTS models vary widely; size your serving instances based on your specific model choice and concurrent request targets.
The fractional GPU techniques covered in the Kubernetes MIG/time-slicing article become relevant here: if your TTS model is small enough to share a GPU with other workloads using MIG or time-slicing, that can significantly improve utilization.

Video Understanding in Production
Video models present the most resource-intensive workload in the multimodal family. The core challenge: a video is a sequence of images, and encoding it means running your vision encoder for every frame (or every sampled frame, at some interval). A 30-second video at 1 frame per second generates 30 sets of vision tokens. At high resolution with dense sampling, you can easily generate more tokens from the video than your model’s context window supports.
Production video understanding systems therefore make architectural choices before the model ever runs:
Frame sampling strategy. You almost never encode every frame. Common approaches sample at fixed intervals (1 fps, 0.5 fps), use scene detection to sample at visual change points, or use a cheap lightweight model to select the most relevant frames before running the expensive VLM. The right strategy depends on your content type. For surveillance footage where most frames are identical, aggressive sampling is appropriate. For instructional video where actions happen quickly, denser sampling preserves more signal.
Asynchronous processing. Interactive video understanding is rare in production today; most use cases accept batch latency. Design your pipeline as an asynchronous job queue with separate workers for video ingestion, frame extraction, vision encoding, and LLM processing. This decouples the stages and lets you scale bottlenecks independently. The durable execution patterns covered elsewhere apply here: video processing pipelines are a natural fit for workflow engines that handle retries, state, and partial completion.
Cost management. Video understanding is expensive. The cost per video processed can vary by orders of magnitude based on length, resolution, and sampling rate. Before you build video understanding into a product feature, model the cost per operation and build per-request limits into your gateway. Tie this into your broader AI FinOps strategy from the start, not as an afterthought.
The Architecture Shift Nobody Briefed You On
Through 2023 and 2024, the dominant VLM architecture was a pretrained LLM as the backbone with vision as a bolt-on: a CLIP encoder, a projection layer, and then text tokens from the LLM. This is the LLaVA family of architectures.
The 2025-2026 generation of frontier models has moved to early fusion, where visual and text tokens are interleaved in the token stream from very early in the architecture, rather than encoded separately and merged. The open-weight models are following: Qwen3-VL, InternVL 3, and several others now use variants of early or mixed fusion.
Why does this matter for your infrastructure decisions? Serving early-fusion models with an inference engine that was designed for the adapter pattern can produce unexpected behavior, particularly around KV cache management. When you evaluate a new model version, specifically test that your serving stack handles the architecture correctly before migrating production traffic. Do not assume that a model upgrade is drop-in compatible with your current vLLM or SGLang configuration just because it is in the same model family.
This is also why I recommend keeping your model upgrade pipeline distinct from your serving infrastructure upgrades. Changing both simultaneously makes failures harder to diagnose.
Building the Gateway Layer for Multimodal Traffic
Your AI gateway needs multimodal-aware policies that you probably did not need for text-only traffic:
Content validation before serving. Image inputs should be validated for format, dimensions, and file size before the request reaches your serving pool. An oversized TIFF that bypasses validation will exhaust memory in ways that take down more than one request.
Resolution normalization. Standardize to a maximum resolution at the gateway so your model never sees inputs that exceed what it was benchmarked on. This is both a security control and a reliability control.
Modality routing. Not every request needs the vision encoder path. If your application mixes text-only and vision requests, route them differently. Text-only requests to a text-optimized serving pool, vision requests to a multimodal pool. This avoids loading the vision encoder on every request and lets you scale each pool independently.
Cost attribution by modality. Image tokens cost more to process than text tokens because of the encoding step. Your cost attribution model in your FinOps stack should distinguish between text token costs and vision token costs. This gives you the signal to understand whether vision use cases are carrying their weight economically.
Observability for Multimodal Pipelines
Multimodal pipelines need richer observability than text-only LLM serving. The metrics that matter most beyond the standard latency/throughput set:
Vision token count distribution. Track the p50, p95, and p99 of vision tokens per request. Spikes in this metric are the early warning signal for memory pressure events.
Stage latency breakdown. In a pipeline with vision encoding, LLM generation, and optional TTS, track latency for each stage separately. When overall request latency spikes, you need to know which stage is the bottleneck.
Encoder utilization vs. LLM utilization. If you run separate vision encoder and LLM serving, track GPU utilization for each independently. Imbalanced utilization tells you where to add capacity.
Input validation rejection rate. A rising rejection rate at the gateway (requests rejected due to resolution limits, image count limits, etc.) can signal that your users are hitting limits you set too aggressively, or that you are seeing input abuse.
The observability patterns from the LLM observability guide apply here, extended with these multimodal dimensions. If you are using OpenTelemetry, add custom attributes to your spans for vision token count and stage latency so you can correlate these with your overall request latency.

Capacity Planning for Multimodal Workloads
The standard GPU capacity planning approach for LLMs, starting with model weight memory and adding KV cache budget, needs adjustment for multimodal workloads.
For VLMs: start with the LLM backbone weight in your VRAM calculation, add the vision encoder weight (typically 300MB to 1.5GB for CLIP ViT variants), then size your KV cache budget not based on text sequence length alone but on text plus expected vision tokens. If your median request includes one 512x512 image (roughly 1,024 vision tokens), and your text context averages 500 tokens, your effective sequence length for KV cache sizing is around 1,500 tokens. Use your p95 or p99 input length, not the median, for safety margins.
For audio pipelines: STT and TTS GPU memory is typically dwarfed by LLM serving. The question is whether to co-locate them on the same GPU instances. For workloads with predictable audio volume, co-location works if you use fractional GPU techniques to partition memory. For bursty audio volume, separate CPU-based STT instances and GPU-based TTS (if you need the quality) give you independent scaling.
For video understanding: treat video processing as a separate capacity pool entirely. Video jobs are inherently bursty and latency-tolerant, which makes them a good fit for GPU spot instances with a queue. The spot instance strategies covered in this guide apply directly, with the caveat that you need checkpoint support in your video processing pipeline to handle interruptions gracefully.
The Model Churn Problem
One reality of multimodal AI in 2026 that I cannot overstate: the model landscape is changing faster than your infrastructure abstractions keep up. The model you deploy today may be superseded by something with a fundamentally different architecture in six months. LLaVA-style adapter models and early-fusion models have different serving requirements. Whisper-v3 and its alternatives have different quantization and batching tradeoffs.
Build your serving layer to be model-agnostic where possible. Use vLLM’s model-as-configuration pattern rather than baking model-specific logic into your serving code. Abstract your pipeline stages so that swapping the STT model does not require rebuilding the downstream LLM integration. Keep your model registry up to date and plan for at least one architecture migration per year in the multimodal space.
This is the lesson teams learned in the pure-text LLM world and it applies with even more force in multimodal: the infrastructure investment is in the serving layer, the gateway, the monitoring, and the cost controls, not in any specific model. The models are commoditizing faster than the operational tooling.
Recommendations: Where to Start
If you are building multimodal serving from scratch in late 2026, here is the sequence I would follow:
Start with a single VLM, vLLM-backed, with resolution and image count limits enforced at the gateway. Instrument vision token count from day one. Run load tests with representative visual inputs before going to production.
Add audio as a separate pipeline service, not as a co-located sidecar to your LLM serving. Keep the scaling concerns isolated.
Defer video understanding until you have text and vision operating reliably and you have clear demand signal. Video is expensive and operationally complex; the reward needs to justify the investment.
Build your RAG architecture to be multimodal-aware from the start if you know you will need it. Multimodal RAG, retrieving relevant images or documents based on image queries, requires your embedding model and vector store to support multimodal embeddings. Retrofitting this onto a text-only RAG pipeline later is painful.
Multimodal AI is not just a richer version of text serving. The memory characteristics, the pipeline architecture, the cost model, and the failure modes are all distinct. Treat it as a new infrastructure problem that borrows concepts from text serving but does not inherit its assumptions.
Get Cloud Architecture Insights
Practical deep dives on infrastructure, security, and scaling. No spam, no fluff.
By subscribing, you agree to receive emails. Unsubscribe anytime.
