#110 · Primary category: MLOps & Evaluation
InferenceX
Open-source continuous inference benchmark research platform comparing leading LLMs across NVIDIA, AMD, and future TPU hardware with live dashboards.
Project last updated:08/29/26
GitHub Stars
1.6K
Forks
277
Contributors
93
License
Apache-2.0
Why we included this project
The project's core idea is to keep re-running a standard inference benchmark across the major open-source serving stacks, including SGLang, vLLM, and TensorRT-LLM, on current data-center GPUs like the GB200 and GB300 NVL72, B200, and AMD's MI355X. Because the runs are continuous, results capture how kernels and schedulers improve over time, and a single snapshot benchmark goes stale within weeks as upstream software ships updates. InferenceX provides an Apache-2.0 benchmark runner and a public dashboard where you can compare results across stacks and hardware variants without running the workload yourself. That gives procurement, capacity-planning, and framework-selection teams a continuously refreshed source of numbers instead of one-off vendor claims.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
unsloth
Local UI to run and train LLMs and diffusion models. Supports GGUF, MLX, Qwen3.8, Kimi K3, MiniMax-H3, Gemma 4, FLUX and more.
LlamaFactory
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
airflow
Apache Airflow - A platform to programmatically author, schedule, and monitor workflows
langfuse
🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integrates with OpenTelemetry, LangChain, OpenAI SDK, LiteLLM, and more. 🍊YC W23
netron
Visualizer for neural network, deep learning and machine learning models