#63 · Primary category: Foundation Models
PaddleMIX
Paddle Multimodal Integration and eXploration, supporting mainstream multi-modal tasks, including end-to-end large-scale multi-modal pretrain models and diffusion model toolbox. Equipped with high performance and flexibility.
Project last updated:03/06/26
GitHub Stars
724
Forks
223
Contributors
96
License
Apache-2.0
Why we included this project
Teams already working in the PaddlePaddle ecosystem get the shortest path to modern multimodal models without switching frameworks. The project bundles a wide range of vision-language checkpoints, including LLaVA, InternVL, Qwen2-VL, DeepSeek-VL, MiniCPM-V, and Janus, each with inference and training examples ready to use, so you can fine-tune or serve a specific model instead of building everything from scratch. The bundled diffusion toolbox covers text-to-image and text-to-video generation with Stable Diffusion and FLUX, which helps teams that need both understanding and generation in one place. It also includes specialized extras like a document-understanding model, a video-generation control model, and utilities for cleaning and tagging multimodal datasets. If you are evaluating multimodal capabilities and prefer Paddle over PyTorch, this is a practical place to start prototyping.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
transformers
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
CLIP
CLIP (Contrastive Language-Image Pretraining), Predict the most relevant text snippet given an image
MiniCPM-V
A Pocket-Sized MLLM for Ultra-Efficient Image and Video Understanding on Your Phone
generative-models
Generative Models by Stability AI
unilm
Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalities