#67 · Primary category: Foundation Models
InternLM-XComposer
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions
Project last updated:05/26/25
GitHub Stars
2.9K
Forks
175
Contributors
17
License
Apache-2.0
Why we included this project
InternLM-XComposer is a practical base for teams that want a vision-language model without building one from scratch. The 7B checkpoint holds long 24K-to-96K interleaved image-text contexts, reads ultra-high-resolution images and fine-grained video, and carries on multi-turn, multi-image conversations; the repository ships the pretrained weights alongside finetuning code and evaluation scripts. Around the core model sit companions worth knowing about, including a multimodal reward model, an OmniLive system built for sustained streaming video and audio interaction, and datasets such as ShareGPT4V and MMDU. That combination makes it a sensible starting point to adapt for visual question answering, document and webpage understanding, or content generation. Treat it as an open foundation model with training recipes rather than a turnkey product.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
transformers
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
CLIP
CLIP (Contrastive Language-Image Pretraining), Predict the most relevant text snippet given an image
MiniCPM-V
A Pocket-Sized MLLM for Ultra-Efficient Image and Video Understanding on Your Phone
generative-models
Generative Models by Stability AI
unilm
Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalities