#336 · Primary category: AI Tool Directories & Curated Lists

awesome-foundation-and-multimodal-models

blip clip computer-vision foundational-models grounding-dino image-captioning llava multimodal nlp open-vocabulary-detection open-vocabulary-segmentation segment-anything zero-shot-detection

👁️ + 💬 + 🎧 = 🤖 Curated list of top foundation and multimodal models! [Paper + Code + Examples + Tutorials]

Project last updated:02/29/24

GitHub Stars

637

Forks

46

Contributors

3

License

Other

Why we included this project

This index collects notable foundation and multimodal models, with each entry linking to the original paper, the code, and often a live demo or tutorial. For anyone working in vision-language AI, it saves the chore of hunting through arXiv and GitHub separately: models like YOLO-World and Depth Anything appear in one place with their modalities and the tasks they handle. Entries pair each model with runnable resources, so you can quickly judge whether an approach is worth trying for your own detection, segmentation, or captioning work. It is a reference and learning resource rather than software you deploy, which is what makes it useful for evaluation and research.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category