#47 · Primary category: Foundation Models
XPretrain
Multi-modality pre-training
Project last updated:08/28/26
GitHub Stars
510
Forks
36
Contributors
7
License
Other
Why we included this project
Researchers and ML engineers building joint vision-and-language models will find this a useful, structured collection of pre-training code from Microsoft Research's multimedia group. It bundles peer-reviewed implementations: HD-VILA and LF-VILA for high-resolution and long-form video-language understanding, CLIP-ViP for adapting image-language pretraining to video, and the HD-VILA-100M video-language dataset. Instead of a single deployable application, you get reproducible code and checkpoints to benchmark against published results, handy for video-text retrieval or multimodal pre-training baselines. The image-language entries Pixel-BERT, SOHO, and VisualParsing add coverage, so you can compare approaches across video and still-image settings. Note the Microsoft Research license restricts use to non-commercial purposes, so this.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
transformers
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
CLIP
CLIP (Contrastive Language-Image Pretraining), Predict the most relevant text snippet given an image
MiniCPM-V
A Pocket-Sized MLLM for Ultra-Efficient Image and Video Understanding on Your Phone
generative-models
Generative Models by Stability AI
unilm
Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalities