#591 · Primary category: Education & Research
vlms-zero-to-hero
This series will take you on a journey from the fundamentals of NLP and Computer Vision to the cutting edge of Vision-Language Models.
Project last updated:01/23/25
GitHub Stars
1.2K
Forks
104
Contributors
1
License
Apache-2.0
Why we included this project
Most courses on vision-language models expect you to already know the field. This series instead builds up from the ground, pairing each landmark paper, from Word2Vec and Attention Is All You Need to BERT, CLIP, and LLaVA, with a PyTorch notebook you can open straight in Colab. Reading the original research and then following code that implements it makes the ideas concrete rather than abstract. The author's open-source computer vision background shows in the teaching, and because every concept ships as a runnable notebook, you can experiment instead of just reading. For developers and students who want genuine intuition about how image and text representations combine, this is a coherent path through the field's milestones, not a scattered collection of tutorials.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
prompts.chat
f.k.a. Awesome ChatGPT Prompts. Share, discover, and collect prompts from the community. Free and open source — self-host for your organization with complete privacy.
JavaGuide
Java Interview & Backend General Interview Guide, covering computer fundamentals, databases, distributed systems, high concurrency, system design, and AI application development.
system-prompts-and-models-of-ai-tools
A curated collection of system prompts, internal tools, and AI models from popular AI assistants and coding agents.
30-seconds-of-code
Coding articles to level up your development skills
generative-ai-for-beginners
21 Lessons, Get Started Building with Generative AI