#47 · Primary category: Foundation Models

XPretrain

computer-vision multimedia multimodal-learning nlp pre-training

Multi-modality pre-training

Project last updated:08/28/26

GitHub Stars

510

Forks

36

Contributors

7

License

Other

Why we included this project

Researchers and ML engineers building joint vision-and-language models will find this a useful, structured collection of pre-training code from Microsoft Research's multimedia group. It bundles peer-reviewed implementations: HD-VILA and LF-VILA for high-resolution and long-form video-language understanding, CLIP-ViP for adapting image-language pretraining to video, and the HD-VILA-100M video-language dataset. Instead of a single deployable application, you get reproducible code and checkpoints to benchmark against published results, handy for video-text retrieval or multimodal pre-training baselines. The image-language entries Pixel-BERT, SOHO, and VisualParsing add coverage, so you can compare approaches across video and still-image settings. Note the Microsoft Research license restricts use to non-commercial purposes, so this.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category