#63 · Primary category: Foundation Models

PaddleMIX

aigc clip controlnet deepseek-vl dit eva-clip got-ocr20 image-to-text internvl2 llava minicpm-v multimodal ppdiffusers qwen2-vl sd-xl sora stable-diffusion stablevideodiffusion text-to-image text-to-video

Paddle Multimodal Integration and eXploration, supporting mainstream multi-modal tasks, including end-to-end large-scale multi-modal pretrain models and diffusion model toolbox. Equipped with high performance and flexibility.

Project last updated:03/06/26

GitHub Stars

724

Forks

223

Contributors

96

License

Apache-2.0

Why we included this project

Teams already working in the PaddlePaddle ecosystem get the shortest path to modern multimodal models without switching frameworks. The project bundles a wide range of vision-language checkpoints, including LLaVA, InternVL, Qwen2-VL, DeepSeek-VL, MiniCPM-V, and Janus, each with inference and training examples ready to use, so you can fine-tune or serve a specific model instead of building everything from scratch. The bundled diffusion toolbox covers text-to-image and text-to-video generation with Stable Diffusion and FLUX, which helps teams that need both understanding and generation in one place. It also includes specialized extras like a document-understanding model, a video-generation control model, and utilities for cleaning and tagging multimodal datasets. If you are evaluating multimodal capabilities and prefer Paddle over PyTorch, this is a practical place to start prototyping.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category