#70 · Primary category: Foundation Models

mPLUG-Owl

alpaca chatbot chatgpt damo dialogue gpt gpt4 gpt4-api huggingface instruction-tuning large-language-models llama mplug mplug-owl multimodal pretraining pytorch transformer video visual-recognition

mPLUG-Owl: The Powerful Multi-modal Large Language Model Family

Project last updated:04/02/25

GitHub Stars

2.5K

Forks

189

Contributors

11

License

MIT

Why we included this project

mPLUG-Owl is a family of multimodal large language models that has grown through three releases: the original model, the CVPR 2024 highlighted mPLUG-Owl2, and mPLUG-Owl3, which focuses on understanding long image sequences. That range lets you match a version to your task, whether you need simple visual question answering or handling multiple images and video. The repository ships inference and training code plus pretrained weights on Hugging Face, and there is a Chinese-enhanced variant, mPLUG-Owl2.1, if localization matters. Because the vision module is separated from the language backbone, the architecture is a useful reference for researchers fine-tuning on their own data and a solid starting point for engineers building multimodal dialogue systems.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category