#143 · Primary category: Inference & Local Deploy

vLLM-2080Ti-Definitive

The definitive vLLM runtime for dual RTX 2080 Ti 22GB + NVLink, delivering Qwen 27B local inference with maximum 100+ tok/s single-request decode with support of FP8 weight ( Join Discord :https://discord.gg/VFqVVySdMS )

Project last updated:08/27/26

GitHub Stars

741

Forks

110

Contributors

9

License

Apache-2.0

Why we included this project

This fork patches vLLM so a pair of RTX 2080 Ti cards with 22GB mods and NVLink can run Qwen3.x 27B models locally, reporting over 100 tokens per second on single-request decode. It's a hardware-focused build that preserves the patched source, launch profiles, and runtime notes, so you can reproduce the working stack instead of wrestling with vLLM's default support for older Turing silicon. The repo validates specific checkpoints and long-context routes, including 256K text and 136K text-plus-image, and includes MTP decoding and FP8 weight support. If you have aging Turing GPUs and want a serious local inference setup for one strong model on a modest budget, this is a practical option, though it's tuned for single-concurrency personal-agent workloads rather than multi-tenant serving.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category