Fast, cheap omni‑modality model serving for text, image, audio, video and more
vLLM‑Omni builds on the high‑performance vLLM engine to support omni‑modality models, enabling text, image, audio, video, and action inference in a single framework. It adds non‑autoregressive support for diffusion transformers and heterogeneous outputs, and offers request‑level batching, async output, and quantization. The library is aimed at researchers and developers who need fast, low‑cost inference for multimodal workloads, and it outperforms vanilla vLLM by adding multimodal capabilities without sacrificing speed.
View on GitHub →sprag-ai/vllm-omni