#33 · Primary category: Inference & Local Deploy

serve

cloud-native cncf deep-learning docker fastapi framework generative-ai grpc jaeger kubernetes llmops machine-learning microservice mlops multimodal neural-search opentelemetry orchestration pipeline prometheus

☁️ Build multimodal AI applications with cloud-native stack

Project last updated:03/24/25

GitHub Stars

21.9K

Forks

2.2K

Contributors

185

License

Apache-2.0

Why we included this project

Jina lets you package inference logic into Executors and serve it over gRPC, HTTP, or WebSockets without building the surrounding infrastructure yourself. The framework handles scaling, streaming, and dynamic batching, and its native gRPC support plus Docker, Kubernetes, and Docker Compose integration means it fits naturally into a cloud-native microservice stack. That orchestration layer is where it saves real time: teams can move from a local prototype to a deployed, load-balanced endpoint, and multimodal workloads are a particular strength. LLM endpoints get token-by-token streaming, which keeps chat-style applications responsive.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category