#7 · Primary category: AI Cloud Platforms & PaaS

skypilot

cloud-computing cloud-management cost-optimization deep-learning distributed-training gpu hyperparameter-tuning job-queue job-scheduler llm-serving llm-training machine-learning ml-infrastructure ml-platform mlops multicloud slurm spot-instances tpu

The AI Compute Platform for frontier teams. SkyPilot turns fragmented AI compute into one AI supercomputer, so frontier AI teams build custom intelligence faster.

Project last updated:08/29/26

GitHub Stars

10.5K

Forks

1.2K

Contributors

277

License

Apache-2.0

Why we included this project

SkyPilot is for teams whose AI workloads spread across more than one cloud, Kubernetes cluster, or Slurm fleet, and who are tired of stitching that infrastructure together by hand. You get one CLI for launching jobs and scaling GPU capacity, with a scheduler that handles placement, spot instance preemption, and prioritized job queues across providers. People use it for everything from short-lived dev boxes and sandboxes for LLM-generated code to large pre-training runs and online RL jobs spanning thousands of GPUs. It runs on top of infrastructure you already have, whether that is Kubernetes, Slurm, on-prem VMs, or public clouds like AWS, GCP, Azure, and CoreWeave, so you are not locked into one vendor. For small teams, the CLI keeps the learning curve low compared with hand-rolling YAML and cluster wiring.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category