#42 · Primary category: AI Cloud Platforms & PaaS

examples

ai gpu llms ml mlops serverless serverless-functions

Examples for Cerebrium Serverless GPUs

Project last updated:07/16/26

GitHub Stars

526

Forks

77

Contributors

9

License

Other

Why we included this project

Serverless GPU platforms are easy to demo but harder to vet, which is what makes Cerebrium's example collection useful. Each scenario ships as a self-contained folder with its own deploy instructions, and the range is genuinely broad: vLLM-based OpenAI-compatible endpoints, Whisper transcription, ComfyUI and SDXL image generation, LangChain question-answering, WebSocket streaming, and realtime voice agents. It doubles as a practical tour of what production GPU serving involves, including cold-start optimization, multi-GPU inference, faster model weight loading, and migrating existing workloads off COG. Since everything runs through Cerebrium's own CLI and platform, treat it as vendor-specific starter material rather than portable infrastructure. For a team comparing serverless GPU providers, though, working examples like these give a concrete sense of realistic deployment patterns before you commit to paying for compute.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category