#70 · Primary category: Foundation Models
mPLUG-Owl
mPLUG-Owl: The Powerful Multi-modal Large Language Model Family
Project last updated:04/02/25
GitHub Stars
2.5K
Forks
189
Contributors
11
License
MIT
Why we included this project
mPLUG-Owl is a family of multimodal large language models that has grown through three releases: the original model, the CVPR 2024 highlighted mPLUG-Owl2, and mPLUG-Owl3, which focuses on understanding long image sequences. That range lets you match a version to your task, whether you need simple visual question answering or handling multiple images and video. The repository ships inference and training code plus pretrained weights on Hugging Face, and there is a Chinese-enhanced variant, mPLUG-Owl2.1, if localization matters. Because the vision module is separated from the language backbone, the architecture is a useful reference for researchers fine-tuning on their own data and a solid starting point for engineers building multimodal dialogue systems.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
transformers
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
CLIP
CLIP (Contrastive Language-Image Pretraining), Predict the most relevant text snippet given an image
MiniCPM-V
A Pocket-Sized MLLM for Ultra-Efficient Image and Video Understanding on Your Phone
generative-models
Generative Models by Stability AI
unilm
Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalities