vllm-project/vllm-omni
View on GitHub[RFC]: Support Wan-AI/Wan2.2-S2V-14B in vllm-omni
Open
#2,612 opened on Apr 8, 2026
good first issuehelp wantednew model
Repository metrics
- Stars
- (4,990 stars)
- PR merge metrics
- (PR metrics pending)
Description
Motivation.
SOTA audio-driven model which enables lip-synced video generation based on input image and audio. This RFC is to propose Wan2.2-S2V model support in vllm-omni
Example model weight: Wan-AI/Wan2.2-S2V-14B
Input:
db4810e2-a379-41e9-815f-bbb1a9cafbae.mp3
output:
https://github.com/user-attachments/assets/25e108b8-93cf-4575-b6cd-ff4a00ddd845
Proposed Change.
Please provide the detailed design document of the RFC using the template.
Feedback Period.
No response
CC List.
@hsliuustc0106 @Gaohan123 @gcanlin
Any Other Things.
No response
Before submitting a new issue...
- Make sure you already searched for relevant issues, and asked the chatbot living at the bottom right corner of the documentation page, which can answer lots of frequently asked questions.