vllm-project/vllm-omni

[Good First Issue]: Add vendor-organized community recipes for “run model X on hardware Y for task Z” in vLLM-Omni

Open

#2,645 opened on Apr 9, 2026

View on GitHub
 (59 comments) (4 reactions) (9 assignees)Python (1,067 forks)github user discovery
good first issuehelp wantedhigh priority

Repository metrics

Stars
 (4,990 stars)
PR merge metrics
 (PR metrics pending)

Description

Community Recipes To-Do List

All supported hardware platforms (CUDA, ROCm, NPU, XPU) are welcome for every entry.

Legend: 🙋 help wanted | ⏳ PR raised | ✅ merged

Vendor Model Task Status Claimer PR Notes
Omni-Modality
Qwen Qwen3-Omni omni chat / serving 30B MoE (3B active)
Qwen Qwen2.5-Omni omni chat / speech @allgather #3151 7B / 3B
ByteDance BAGEL-7B-MoT image gen + understanding @GrayMiao123 MoT, ~42 GiB VRAM
ByteDance MammothModa2-Preview text-to-image 🙋 AR + DiT pipeline
inclusionAI Ming-flash-omni-2.0 omni-modal understanding @yuanheng-zhao #2890 MoE, 4-GPU TP
SNU AIDAS Dynin-Omni omni-modal (t2t/i2t/s2t/t2i/v2t/t2s) 🙋 3-stage token-based
SenseNova SenseNova-U1-8B-MoT omni-modal (T2I/I2I/I2T/T2T) @princepride #3319 MoE, ~36 GiB; think mode
Text-to-Image / Image Gen
Qwen Qwen-Image text-to-image @Semmer2 #2729 1024x1024, ~60 GiB
Qwen Qwen-Image-2512 text-to-image 🙋 @AbelSara updated variant
Qwen Qwen-Image-Edit image editing @yixiaoer text-guided editing
Qwen Qwen-Image-Layered layered image gen 🙋
Zhipu AI (GLM) GLM-Image image gen / editing @friden-zhang 2-stage AR+DiT
Tongyi Z-Image-Turbo text-to-image (distilled) @MrDongsls #3283 4-9 steps, ~25 GiB; 2x RTX 5880 48GB
Stepfun NextStep-1.1 text-to-image @GrayMiao123 dual-level CFG, 512x512
Meituan LongCat-Image text-to-image 🙋 @Mashirona 1024x1024
Meituan LongCat-Image-Edit image editing 🙋
OvisAI Ovis-Image text-to-image 🙋 1024x1024
OmniGen2 OmniGen2 text-to-image 🙋 @Jerry2423 @Yuyi-Ao @jackywangno007-cyber ~20 GiB
Stability AI Stable-Diffusion-3.5 text-to-image 🙋 @yangyonggit 1024x1024
Black Forest Labs FLUX.1-dev text-to-image 🙋 ~78 GiB
Black Forest Labs FLUX.1-schnell text-to-image (fast) 🙋 fast distilled variant
Black Forest Labs FLUX.2-dev text-to-image 🙋 needs CPU offload on 80GiB
Black Forest Labs FLUX.2-klein-4B text-to-image 🙋 distilled, 4B params
Black Forest Labs FLUX.2-klein-9B text-to-image 🙋 distilled, 9B params
Tencent HunyuanImage-3.0 image gen / understanding @Bounty-hunter #2495 MoE, dual task
Baidu ERNIE-Image text-to-image @RuixiangMa #2861 1024x1024, ~23 GiB
Text-to-Video / Image-to-Video
Tencent-Hunyuan HunyuanVideo-1.5-T2V text-to-video @allgather #3152 480p path on 1xA100; needs FP8 + VAE tiling for 720p
Tencent-Hunyuan HunyuanVideo-1.5-I2V image-to-video @xldeng-chn 1xA100 80GB
Wan-AI Wan2.2-T2V-A14B text-to-video @lengrongfu #3018 720x1280, 81 frames, ~60 GiB
Wan-AI Wan2.2-TI2V-5B text+image-to-video 🙋 unified, 5B params
Wan-AI Wan2.2-I2V-A14B image-to-video MoE
Wan-AI Wan2.1-VACE video creation (T2V/I2V/FLF2V) 🙋 1.3B / 14B variants
Wan-AI Wan2.2-S2V-14B speech-to-video (image+audio→video) @xuechendi #2751 720p, 81 frames; TP=2
Lightricks LTX-2-T2V text-to-video @fywc #3294 512x768, 121 frames
Lightricks LTX-2-I2V image-to-video @fywc #3294 same model as T2V
Lightricks LTX-2.3 text-to-video + audio @oglok #2893 22B, ~62 GiB peak; BWE 48kHz vocoder
Helios Helios-Base T2V / I2V / V2V @lengrongfu stage 1 only; 8xH800 or 8xH200
Helios Helios-Mid video denoising 🙋 @lengrongfu multi-stage pyramid; 8xH800 or 8xH200
Helios Helios-Distilled video (few-step) 🙋 @lengrongfu DMD distillation; 8xH800 or 8xH200
SII-GAIR MagiHuman video+audio with lip sync 🙋 TP=4 for 80GB, DiT MoE+T5
XuGuo699 DreamID-Omni I2V with identity preservation 🙋 @fywc ~72 GiB
Text-to-Speech / Audio
Qwen Qwen3-TTS TTS serving @chzhang2021 #3130 12Hz variants (CustomVoice, VoiceDesign, Base)
FunAudioLLM CosyVoice3 TTS @Moore-Z #3486 2-stage: talker + flow-matching, 0.5B
FishAudio Fish Speech S2 Pro TTS / voice cloning @menjiantong #3193 4B dual-AR, 44.1kHz
Mistral AI Voxtral TTS TTS with voice presets 🙋 4B, voice cloning
K2-FSA OmniVoice multilingual TTS (600+ lang) 🙋 @fray1024 zero-shot, Qwen3-0.6B backbone; 1xL20 48GB
OpenBMB VoxCPM2 TTS (diffusion AR) 🙋 @wjinxu 2B, 48kHz, 30+ lang
Xiaomi (MiMo) MiMo-Audio audio gen / TTS / ASR / dialogue 🙋 multi-task audio 7B
Stability AI Stable-Audio-Open audio gen (music, SFX) 🙋 @Ronnie-Rui diffusion-based
HKUST AudioX T2A / V2A / TV2A / T2M / V2M / TV2M @zhangj1an #2077 diffusion maf-mmdit; ~97 GiB VRAM
Tencent Covo-Audio-Chat audio chat (audio in, text+audio out) @Dnoob #2293 7B LLM + BigVGAN; ~20 GiB

How to contribute (step by step)

  1. Claim a row — comment on this issue with the row + your hardware (e.g. Qwen3-TTS, 1xL40S 48GB). Maintainer will mark the row ⏳ and assign you.
  2. Copy the templatecp recipes/_template.md recipes/<Vendor>/<Model>.md (template merged in #2646).
  3. Fill the required sections — model + task, tested HW, env (CUDA/ROCm/CANN versions), launch command, verification command + expected output, important flags, known limitations, links back to docs/ and examples/.
  4. Test end-to-end on the hardware you claimed — recipes that aren't personally validated should not be merged.
  5. Open PR titled [Recipe] <Vendor>/<Model> (or [Recipe] <Vendor>/<Model> (<HW>) if adding a HW section to an existing recipe), link this issue, and paste the launch command + verification output in the PR body. Include a test plan describing what you verified (e.g. specific prompts, expected outputs, latency/throughput numbers) and the test results showing the actual output or logs.

The initial recipes/ template has been merged #2646. We now welcome community contributions for practical "run model X on hardware Y for task Z" recipes in vllm-omni.

Please use the merged template as the starting point and organize recipes by model vendor, for example:

recipes/
  Qwen/
    Qwen3-Omni.md
    Qwen3-TTS.md
    Qwen2.5-Omni.md
  Tencent-Hunyuan/
    HunyuanVideo.md
  GLM/
    GLM-Image.md
  MiMo/
    MiMo-Audio.md
  Wan-AI/
    Wan2.2-T2V.md
  ...

Each recipe should include:

  • Model name and task
  • Tested hardware configuration
  • Required environment
  • Launch command
  • Verification command or expected output
  • Important flags or stage configs
  • Known limitations
  • Links back to canonical docs/ and runnable examples/

High Priority Starter Recipes

  • recipes/Qwen/Qwen3-TTS.md

    • Suggested scope: text-to-speech serving with Qwen3-TTS
    • Suggested hardware sections: CUDA GPU, NPU if available
    • Link to existing Qwen3-TTS docs/examples where possible
  • recipes/Qwen/Qwen2.5-Omni.md

    • Suggested scope: omni-modal chat or speech interaction
    • Suggested hardware sections: CUDA GPU, NPU if available
    • Link to existing Qwen2.5-Omni docs/examples where possible
  • recipes/Tencent-Hunyuan/HunyuanVideo.md

    • Suggested scope: text-to-video or image-to-video generation
    • Include tested VRAM requirements and performance notes where possible
  • recipes/GLM/GLM-Image.md

    • Suggested scope: image generation with GLM-Image
    • Include recommended generation settings and validation output
  • recipes/MiMo/MiMo-Audio.md

    • Suggested scope: audio generation or speech-related serving
    • Include model-specific setup and verification notes

Hardware Coverage We Want

  • NVIDIA CUDA recipes

    • Example targets: A100 80GB, H100, L40S, RTX 4090 where applicable
  • AMD ROCm recipes

    • Include ROCm version, GPU model, and any known caveats
  • Intel XPU recipes

    • Suggested devices from community feedback:
    • Intel Arc Pro B50, 16GB
    • Intel Arc Pro B60, 24GB
    • Intel Arc Pro B70, 32GB
  • Huawei NPU recipes

    • Include CANN/runtime versions and model-specific limitations

Good First Contributions

  • Add a recipe for a model you have successfully run locally
  • Add another hardware section to an existing recipe
  • Add verification output to an existing recipe
  • Link an existing recipe to the relevant docs/ and examples/
  • Improve clarity around memory usage, flags, or known limitations

Contribution Guidelines

When opening a recipe PR, please include:

  • The exact command you used
  • Hardware details
  • Software/runtime versions
  • Whether the recipe was personally tested
  • Test plan: what you verified (specific prompts, expected outputs, latency/throughput numbers)
  • Test results: actual output or logs from your run
  • Any limitations or assumptions
  • Links to relevant examples or docs

Recipes should not replace canonical documentation. They should act as practical, community-maintained runbooks that point users to the right docs and runnable examples.


Motivation.

We'd like to propose a new community-maintained recipes/ area in vllm-project/vllm-omni to answer a recurring user question:

How do I run model X on hardware Y for task Z?

Today, users often struggle to find the best path for a concrete deployment scenario in vLLM-Omni.

There are a few reasons:

  1. We currently do not have a dedicated place for operational runbooks that map a specific model, hardware target, and task to a known-good setup.
  2. The current examples/ layout mixes model-specific and task-specific directories, which can be confusing for discovery.
  3. We want to align the user experience with vllm-project/recipes, which already provides this style of practical guide for upstream vLLM.

Proposed Change.

Introduce a top-level recipes/ directory in vLLM-Omni for community-maintained runbooks.

To align with upstream vllm-project/recipes, recipes should be grouped by model vendor at the top level, with one Markdown file per model family by default.

Example direction:

recipes/
  Qwen/
    Qwen3-Omni.md
    Qwen3-TTS.md
    Qwen2.5-Omni.md
  Tencent-Hunyuan/
    HunyuanVideo.md
  GLM/
    GLM-Image.md
  MiMo/
    MiMo-Audio.md

Within each model doc, we should include multiple hardware-specific sections in the same Markdown file, following the structure used by the upstream DeepSeek guide: https://github.com/vllm-project/recipes/blob/main/DeepSeek/DeepSeek-V3.md

For example, a single recipe doc could contain sections such as:

  • 1x A100 80GB
  • 2x L40S
  • 4x H100
  • ROCm / NPU variants where applicable

Each section would describe:

  • supported task(s)
  • tested hardware
  • required environment
  • launch commands
  • important flags / stage configs
  • verification steps
  • known limitations

Design Principles

  • Align the top-level organization with vllm-project/recipes where practical, so discovery feels familiar across vLLM and vLLM-Omni.
  • Keep one Markdown file per model family by default.
  • Keep multiple hardware configurations inside the same recipe document unless the document becomes too large or hard to maintain.
  • Use recipes for practical "known-good setup" guidance, not as the canonical source of product documentation.

Relationship to examples/

The current examples/ folder is still valuable, but it serves a different purpose:

  • examples/: runnable code and scripts
  • docs/: canonical documentation
  • recipes/: practical "known-good setup" guides for concrete user scenarios

One benefit of adding recipes/ is that it may reduce pressure to make examples/ itself carry all discovery and onboarding needs.

This also gives us a cleaner answer to users who ask for a concrete deployment path without requiring them to infer it from a mix of task-oriented and model-oriented example folders.

Scope

This proposal is for community recipes only.

It is not intended to replace:

  • canonical product documentation under docs/
  • source-of-truth runnable examples under examples/

Instead, recipes would act as a user-oriented entry point that links back to those canonical materials.

Open Questions

  1. Should recipe ownership be fully community-maintained, or should each recipe have one or two named maintainers?
  2. Should recipes live only in the repo at first, or also be surfaced in the documentation site under a Community section?
  3. Should we add a lightweight recipe template from the beginning to keep structure consistent?
  4. Should any future cleanup of examples/ be handled in a separate RFC after recipes/ is established?

Initial Success Criteria

  • A new user can quickly find a concrete recipe for a target model, hardware, and task.
  • Recipes follow a consistent structure.
  • Recipes are organized by vendor in a way that feels familiar to users of upstream vllm-project/recipes.
  • Recipes link to examples/ and docs/ instead of duplicating canonical content.
  • The approach improves user experience without adding confusion about where canonical documentation lives.

Suggested First Recipes

  • recipes/Qwen/Qwen3-Omni.md
  • recipes/Qwen/Qwen3-TTS.md
  • recipes/Qwen/Qwen2.5-Omni.md
  • recipes/Tencent-Hunyuan/HunyuanVideo.md
  • recipes/GLM/GLM-Image.md
  • recipes/MiMo/MiMo-Audio.md

Feedback welcome on structure, ownership, and how tightly we should align with upstream vllm-project/recipes.

Feedback Period.

CC List.

vllm-omni maintainer team @ywang96 @Gaohan123 @ZJY0516 @princepride @lishunyang12 .....

Any Other Things.

Contributor guide