#161 · Primary category: AI Tool Directories & Curated Lists

awesome-vlm-architectures

ai-agents architecture-diagrams arxiv awesome awesome-list computer-vision deep-learning foundation-models large-multimodal-models mllm model-architecture multimodal multimodal-ai multimodal-llm natural-language-processing release-timeline research-papers transformers vision-language-models vlm

Curated visual catalog of 155+ vision-language model (VLM/MLLM) architectures: papers, diagrams, training recipes, datasets, and a release timeline for multimodal AI agents.

Project last updated:07/31/26

GitHub Stars

1.3K

Forks

55

Contributors

1

License

Other

Why we included this project

Researchers and engineers often struggle to see how the many vision-language model architectures relate to one another, and this repository addresses that by organizing 155+ VLM and multimodal LLM designs into a visual, citation-first index. Each entry links to the primary paper and explains the architecture and how it aligns vision and language. It also covers training stages and datasets, along with the design choices that set the model apart. The release timeline is handy for tracing how early systems like CLIP and Flamingo gave way to today's multimodal reasoning and agentic models. Since this is a curated index rather than runnable code, its value is as a grounded map of the literature: a good place to compare model families and find the right paper to cite before you start building.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category