[Good First Issue]: Add vendor-organized community recipes for “run model X on hardware Y for task Z” in vLLM-Omni
#2,645 opened on Apr 9, 2026
Repository metrics
- Stars
- (4,990 stars)
- PR merge metrics
- (PR metrics pending)
Description
Community Recipes To-Do List
All supported hardware platforms (CUDA, ROCm, NPU, XPU) are welcome for every entry.
Legend: 🙋 help wanted | ⏳ PR raised | ✅ merged
| Vendor | Model | Task | Status | Claimer | PR | Notes |
|---|---|---|---|---|---|---|
| Omni-Modality | ||||||
| Qwen | Qwen3-Omni | omni chat / serving | ✅ | 30B MoE (3B active) | ||
| Qwen | Qwen2.5-Omni | omni chat / speech | ⏳ | @allgather | #3151 | 7B / 3B |
| ByteDance | BAGEL-7B-MoT | image gen + understanding | ⏳ | @GrayMiao123 | MoT, ~42 GiB VRAM | |
| ByteDance | MammothModa2-Preview | text-to-image | 🙋 | AR + DiT pipeline | ||
| inclusionAI | Ming-flash-omni-2.0 | omni-modal understanding | ✅ | @yuanheng-zhao | #2890 | MoE, 4-GPU TP |
| SNU AIDAS | Dynin-Omni | omni-modal (t2t/i2t/s2t/t2i/v2t/t2s) | 🙋 | 3-stage token-based | ||
| SenseNova | SenseNova-U1-8B-MoT | omni-modal (T2I/I2I/I2T/T2T) | ✅ | @princepride | #3319 | MoE, ~36 GiB; think mode |
| Text-to-Image / Image Gen | ||||||
| Qwen | Qwen-Image | text-to-image | ✅ | @Semmer2 | #2729 | 1024x1024, ~60 GiB |
| Qwen | Qwen-Image-2512 | text-to-image | 🙋 | @AbelSara | updated variant | |
| Qwen | Qwen-Image-Edit | image editing | ⏳ | @yixiaoer | text-guided editing | |
| Qwen | Qwen-Image-Layered | layered image gen | 🙋 | |||
| Zhipu AI (GLM) | GLM-Image | image gen / editing | ⏳ | @friden-zhang | 2-stage AR+DiT | |
| Tongyi | Z-Image-Turbo | text-to-image (distilled) | ⏳ | @MrDongsls | #3283 | 4-9 steps, ~25 GiB; 2x RTX 5880 48GB |
| Stepfun | NextStep-1.1 | text-to-image | ⏳ | @GrayMiao123 | dual-level CFG, 512x512 | |
| Meituan | LongCat-Image | text-to-image | 🙋 | @Mashirona | 1024x1024 | |
| Meituan | LongCat-Image-Edit | image editing | 🙋 | |||
| OvisAI | Ovis-Image | text-to-image | 🙋 | 1024x1024 | ||
| OmniGen2 | OmniGen2 | text-to-image | 🙋 | @Jerry2423 @Yuyi-Ao @jackywangno007-cyber | ~20 GiB | |
| Stability AI | Stable-Diffusion-3.5 | text-to-image | 🙋 | @yangyonggit | 1024x1024 | |
| Black Forest Labs | FLUX.1-dev | text-to-image | 🙋 | ~78 GiB | ||
| Black Forest Labs | FLUX.1-schnell | text-to-image (fast) | 🙋 | fast distilled variant | ||
| Black Forest Labs | FLUX.2-dev | text-to-image | 🙋 | needs CPU offload on 80GiB | ||
| Black Forest Labs | FLUX.2-klein-4B | text-to-image | 🙋 | distilled, 4B params | ||
| Black Forest Labs | FLUX.2-klein-9B | text-to-image | 🙋 | distilled, 9B params | ||
| Tencent | HunyuanImage-3.0 | image gen / understanding | ✅ | @Bounty-hunter | #2495 | MoE, dual task |
| Baidu | ERNIE-Image | text-to-image | ✅ | @RuixiangMa | #2861 | 1024x1024, ~23 GiB |
| Text-to-Video / Image-to-Video | ||||||
| Tencent-Hunyuan | HunyuanVideo-1.5-T2V | text-to-video | ⏳ | @allgather | #3152 | 480p path on 1xA100; needs FP8 + VAE tiling for 720p |
| Tencent-Hunyuan | HunyuanVideo-1.5-I2V | image-to-video | ⏳ | @xldeng-chn | 1xA100 80GB | |
| Wan-AI | Wan2.2-T2V-A14B | text-to-video | ⏳ | @lengrongfu | #3018 | 720x1280, 81 frames, ~60 GiB |
| Wan-AI | Wan2.2-TI2V-5B | text+image-to-video | 🙋 | unified, 5B params | ||
| Wan-AI | Wan2.2-I2V-A14B | image-to-video | ✅ | MoE | ||
| Wan-AI | Wan2.1-VACE | video creation (T2V/I2V/FLF2V) | 🙋 | 1.3B / 14B variants | ||
| Wan-AI | Wan2.2-S2V-14B | speech-to-video (image+audio→video) | ✅ | @xuechendi | #2751 | 720p, 81 frames; TP=2 |
| Lightricks | LTX-2-T2V | text-to-video | ✅ | @fywc | #3294 | 512x768, 121 frames |
| Lightricks | LTX-2-I2V | image-to-video | ✅ | @fywc | #3294 | same model as T2V |
| Lightricks | LTX-2.3 | text-to-video + audio | ✅ | @oglok | #2893 | 22B, ~62 GiB peak; BWE 48kHz vocoder |
| Helios | Helios-Base | T2V / I2V / V2V | ✅ | @lengrongfu | stage 1 only; 8xH800 or 8xH200 | |
| Helios | Helios-Mid | video denoising | 🙋 | @lengrongfu | multi-stage pyramid; 8xH800 or 8xH200 | |
| Helios | Helios-Distilled | video (few-step) | 🙋 | @lengrongfu | DMD distillation; 8xH800 or 8xH200 | |
| SII-GAIR | MagiHuman | video+audio with lip sync | 🙋 | TP=4 for 80GB, DiT MoE+T5 | ||
| XuGuo699 | DreamID-Omni | I2V with identity preservation | 🙋 | @fywc | ~72 GiB | |
| Text-to-Speech / Audio | ||||||
| Qwen | Qwen3-TTS | TTS serving | ✅ | @chzhang2021 | #3130 | 12Hz variants (CustomVoice, VoiceDesign, Base) |
| FunAudioLLM | CosyVoice3 | TTS | ⏳ | @Moore-Z | #3486 | 2-stage: talker + flow-matching, 0.5B |
| FishAudio | Fish Speech S2 Pro | TTS / voice cloning | ✅ | @menjiantong | #3193 | 4B dual-AR, 44.1kHz |
| Mistral AI | Voxtral TTS | TTS with voice presets | 🙋 | 4B, voice cloning | ||
| K2-FSA | OmniVoice | multilingual TTS (600+ lang) | 🙋 | @fray1024 | zero-shot, Qwen3-0.6B backbone; 1xL20 48GB | |
| OpenBMB | VoxCPM2 | TTS (diffusion AR) | 🙋 | @wjinxu | 2B, 48kHz, 30+ lang | |
| Xiaomi (MiMo) | MiMo-Audio | audio gen / TTS / ASR / dialogue | 🙋 | multi-task audio 7B | ||
| Stability AI | Stable-Audio-Open | audio gen (music, SFX) | 🙋 | @Ronnie-Rui | diffusion-based | |
| HKUST | AudioX | T2A / V2A / TV2A / T2M / V2M / TV2M | ✅ | @zhangj1an | #2077 | diffusion maf-mmdit; ~97 GiB VRAM |
| Tencent | Covo-Audio-Chat | audio chat (audio in, text+audio out) | ✅ | @Dnoob | #2293 | 7B LLM + BigVGAN; ~20 GiB |
How to contribute (step by step)
- Claim a row — comment on this issue with the row + your hardware (e.g.
Qwen3-TTS, 1xL40S 48GB). Maintainer will mark the row ⏳ and assign you. - Copy the template —
cp recipes/_template.md recipes/<Vendor>/<Model>.md(template merged in #2646). - Fill the required sections — model + task, tested HW, env (CUDA/ROCm/CANN versions), launch command, verification command + expected output, important flags, known limitations, links back to
docs/andexamples/. - Test end-to-end on the hardware you claimed — recipes that aren't personally validated should not be merged.
- Open PR titled
[Recipe] <Vendor>/<Model>(or[Recipe] <Vendor>/<Model> (<HW>)if adding a HW section to an existing recipe), link this issue, and paste the launch command + verification output in the PR body. Include a test plan describing what you verified (e.g. specific prompts, expected outputs, latency/throughput numbers) and the test results showing the actual output or logs.
The initial recipes/ template has been merged #2646. We now welcome community contributions for practical "run model X on hardware Y for task Z" recipes in vllm-omni.
Please use the merged template as the starting point and organize recipes by model vendor, for example:
recipes/
Qwen/
Qwen3-Omni.md
Qwen3-TTS.md
Qwen2.5-Omni.md
Tencent-Hunyuan/
HunyuanVideo.md
GLM/
GLM-Image.md
MiMo/
MiMo-Audio.md
Wan-AI/
Wan2.2-T2V.md
...
Each recipe should include:
- Model name and task
- Tested hardware configuration
- Required environment
- Launch command
- Verification command or expected output
- Important flags or stage configs
- Known limitations
- Links back to canonical
docs/and runnableexamples/
High Priority Starter Recipes
-
recipes/Qwen/Qwen3-TTS.md- Suggested scope: text-to-speech serving with Qwen3-TTS
- Suggested hardware sections: CUDA GPU, NPU if available
- Link to existing Qwen3-TTS docs/examples where possible
-
recipes/Qwen/Qwen2.5-Omni.md- Suggested scope: omni-modal chat or speech interaction
- Suggested hardware sections: CUDA GPU, NPU if available
- Link to existing Qwen2.5-Omni docs/examples where possible
-
recipes/Tencent-Hunyuan/HunyuanVideo.md- Suggested scope: text-to-video or image-to-video generation
- Include tested VRAM requirements and performance notes where possible
-
recipes/GLM/GLM-Image.md- Suggested scope: image generation with GLM-Image
- Include recommended generation settings and validation output
-
recipes/MiMo/MiMo-Audio.md- Suggested scope: audio generation or speech-related serving
- Include model-specific setup and verification notes
Hardware Coverage We Want
-
NVIDIA CUDA recipes
- Example targets: A100 80GB, H100, L40S, RTX 4090 where applicable
-
AMD ROCm recipes
- Include ROCm version, GPU model, and any known caveats
-
Intel XPU recipes
- Suggested devices from community feedback:
- Intel Arc Pro B50, 16GB
- Intel Arc Pro B60, 24GB
- Intel Arc Pro B70, 32GB
-
Huawei NPU recipes
- Include CANN/runtime versions and model-specific limitations
Good First Contributions
- Add a recipe for a model you have successfully run locally
- Add another hardware section to an existing recipe
- Add verification output to an existing recipe
- Link an existing recipe to the relevant
docs/andexamples/ - Improve clarity around memory usage, flags, or known limitations
Contribution Guidelines
When opening a recipe PR, please include:
- The exact command you used
- Hardware details
- Software/runtime versions
- Whether the recipe was personally tested
- Test plan: what you verified (specific prompts, expected outputs, latency/throughput numbers)
- Test results: actual output or logs from your run
- Any limitations or assumptions
- Links to relevant examples or docs
Recipes should not replace canonical documentation. They should act as practical, community-maintained runbooks that point users to the right docs and runnable examples.
Motivation.
We'd like to propose a new community-maintained recipes/ area in vllm-project/vllm-omni to answer a recurring user question:
How do I run model X on hardware Y for task Z?
Today, users often struggle to find the best path for a concrete deployment scenario in vLLM-Omni.
There are a few reasons:
- We currently do not have a dedicated place for operational runbooks that map a specific model, hardware target, and task to a known-good setup.
- The current
examples/layout mixes model-specific and task-specific directories, which can be confusing for discovery. - We want to align the user experience with
vllm-project/recipes, which already provides this style of practical guide for upstream vLLM.
Proposed Change.
Introduce a top-level recipes/ directory in vLLM-Omni for community-maintained runbooks.
To align with upstream vllm-project/recipes, recipes should be grouped by model vendor at the top level, with one Markdown file per model family by default.
Example direction:
recipes/
Qwen/
Qwen3-Omni.md
Qwen3-TTS.md
Qwen2.5-Omni.md
Tencent-Hunyuan/
HunyuanVideo.md
GLM/
GLM-Image.md
MiMo/
MiMo-Audio.md
Within each model doc, we should include multiple hardware-specific sections in the same Markdown file, following the structure used by the upstream DeepSeek guide: https://github.com/vllm-project/recipes/blob/main/DeepSeek/DeepSeek-V3.md
For example, a single recipe doc could contain sections such as:
- 1x A100 80GB
- 2x L40S
- 4x H100
- ROCm / NPU variants where applicable
Each section would describe:
- supported task(s)
- tested hardware
- required environment
- launch commands
- important flags / stage configs
- verification steps
- known limitations
Design Principles
- Align the top-level organization with
vllm-project/recipeswhere practical, so discovery feels familiar across vLLM and vLLM-Omni. - Keep one Markdown file per model family by default.
- Keep multiple hardware configurations inside the same recipe document unless the document becomes too large or hard to maintain.
- Use recipes for practical "known-good setup" guidance, not as the canonical source of product documentation.
Relationship to examples/
The current examples/ folder is still valuable, but it serves a different purpose:
examples/: runnable code and scriptsdocs/: canonical documentationrecipes/: practical "known-good setup" guides for concrete user scenarios
One benefit of adding recipes/ is that it may reduce pressure to make examples/ itself carry all discovery and onboarding needs.
This also gives us a cleaner answer to users who ask for a concrete deployment path without requiring them to infer it from a mix of task-oriented and model-oriented example folders.
Scope
This proposal is for community recipes only.
It is not intended to replace:
- canonical product documentation under
docs/ - source-of-truth runnable examples under
examples/
Instead, recipes would act as a user-oriented entry point that links back to those canonical materials.
Open Questions
- Should recipe ownership be fully community-maintained, or should each recipe have one or two named maintainers?
- Should recipes live only in the repo at first, or also be surfaced in the documentation site under a Community section?
- Should we add a lightweight recipe template from the beginning to keep structure consistent?
- Should any future cleanup of
examples/be handled in a separate RFC afterrecipes/is established?
Initial Success Criteria
- A new user can quickly find a concrete recipe for a target model, hardware, and task.
- Recipes follow a consistent structure.
- Recipes are organized by vendor in a way that feels familiar to users of upstream
vllm-project/recipes. - Recipes link to
examples/anddocs/instead of duplicating canonical content. - The approach improves user experience without adding confusion about where canonical documentation lives.
Suggested First Recipes
recipes/Qwen/Qwen3-Omni.mdrecipes/Qwen/Qwen3-TTS.mdrecipes/Qwen/Qwen2.5-Omni.mdrecipes/Tencent-Hunyuan/HunyuanVideo.mdrecipes/GLM/GLM-Image.mdrecipes/MiMo/MiMo-Audio.md
Feedback welcome on structure, ownership, and how tightly we should align with upstream vllm-project/recipes.
Feedback Period.
CC List.
vllm-omni maintainer team @ywang96 @Gaohan123 @ZJY0516 @princepride @lishunyang12 .....