#22 · Primary category: AI Cloud Platforms & PaaS

pai

ai artificial-intelligence chainer cloud cluster-management cluster-manager gpu gpu-cluster gpu-computing gpu-scheduler jupyter kubernetes machine-learning model-training on-premise pytorch resource-management scheduling tensorflow

Resource scheduling and cluster management for AI

Project last updated:08/15/26

GitHub Stars

2.7K

Forks

552

Contributors

101

License

MIT

Why we included this project

OpenPAI is a practical choice for teams that want to run distributed AI training on a shared pool of GPU servers without assembling a stack of separate tools. It bundles Kubernetes-based scheduling, GPU-aware job queuing, and a web portal for submitting and monitoring training runs, so a small platform team can stand up an on-premise cluster on its own. Training jobs can be written in PyTorch, TensorFlow, or Chainer, and a REST API covers automation. The honest caveat is that the project entered stable mode after v1.8.1 and the repository is now read-only, so it suits organizations that want a proven, frozen platform rather than active upstream development.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category