AWS Announces vLLM-Omni DLC for Image and Video Generation on SageMaker AI
AWS has released a new Deep Learning Container (DLC) called vLLM-Omni, which extends the capabilities of vLLM beyond text generation to include models that process or generate text, audio, images, and video through OpenAI-compatible APIs.
In a recent blog post, AWS announced the availability of the vLLM-Omni Deep Learning Container (DLC) for Amazon SageMaker AI. This new DLC extends the functionality of vLLM to include models that process or generate text, audio, images, and video through OpenAI-compatible APIs. The announcement includes a detailed technical guide on how to generate images and video using vLLM-Omni on SageMaker AI.
The blog post, authored by Yadan Wei, Dmitry Soldatkin, Mona Mona, and Daniel Wirjo, demonstrates the deployment of two endpoints from the AWS vLLM-Omni DLC: a real-time endpoint for FLUX.2-klein-4B image generation and an asynchronous endpoint for Wan2.1-VACE-1.3B video generation. The workflow involves sending a text prompt to generate a still image, then passing the image and a motion prompt to the video endpoint. The video output is retrieved from Amazon Simple Storage Service (Amazon S3).
The AWS vLLM-Omni DLC packages tracked vLLM-Omni releases and adds routing middleware for SageMaker AI. This release continues a series of posts about specialized AWS DLCs, with Part 1 covering real-time speech generation and Part 2 focusing on real-time and asynchronous inference for image and video generation.
The solution overview highlights the deployment of the same pinned AWS vLLM-Omni DLC image to two SageMaker AI endpoints, with separate endpoints for each model to optimize instance type and inference options. The post concludes with resources for further exploration, including the vLLM-Omni DLC documentation and SageMaker Asynchronous Inference Developer Guide.
Source: Amazon Web Services
