#110 · Primary category: MLOps & Evaluation

InferenceX

ai amd benchmark cuda deepseek gb200 gb300 glm kimi llm mi355x minimax nvidia pytorch rocm sglang vllm

Open-source continuous inference benchmark research platform comparing leading LLMs across NVIDIA, AMD, and future TPU hardware with live dashboards.

Project last updated:08/29/26

GitHub Stars

1.6K

Forks

277

Contributors

93

License

Apache-2.0

Why we included this project

The project's core idea is to keep re-running a standard inference benchmark across the major open-source serving stacks, including SGLang, vLLM, and TensorRT-LLM, on current data-center GPUs like the GB200 and GB300 NVL72, B200, and AMD's MI355X. Because the runs are continuous, results capture how kernels and schedulers improve over time, and a single snapshot benchmark goes stale within weeks as upstream software ships updates. InferenceX provides an Apache-2.0 benchmark runner and a public dashboard where you can compare results across stacks and hardware variants without running the workload yourself. That gives procurement, capacity-planning, and framework-selection teams a continuously refreshed source of numbers instead of one-off vendor claims.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category