#469 · Primary category: Education & Research

VLM_survey

clip computer-vision deep-learning knowledge-distillation multi-modal-model survey transfer-learning vision-language-model

Collection of AWESOME vision-language models for vision tasks

Project last updated:10/14/25

GitHub Stars

3.1K

Forks

234

Contributors

7

License

Other

Why we included this project

This repo is the companion page to a peer-reviewed TPAMI survey on vision-language models for vision tasks, so its paper tables follow the same taxonomy you would find in the published literature. It covers image classification, object detection, semantic segmentation, and related areas, with each entry linking straight to the paper and official project page. Maintainers actively merge pull requests, which keeps the listings current with recent conference papers. If you are scoping a research area, building a reading list, or picking a VLM baseline to try, browsing these curated tables beats searching from scratch.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category