#274 · Primary category: AI Tool Directories & Curated Lists

Transformer-in-Vision

computer-vision deep-learning multi-modal paper self-attention transformer vision-transformers visual-language

Recent Transformer-based CV and related works.

Project last updated:08/22/23

GitHub Stars

1.3K

Forks

140

Contributors

2

License

Other

Why we included this project

Transformer-in-Vision is a curated reading list for anyone trying to keep up with how transformer architectures spread from NLP into computer vision and beyond. The entries span CLIP and DALL·E 2, video generation, 3D synthesis, and large multimodal models, and each one points back to the original paper or official code instead of a rehosted summary. That design makes it handy for two different readers: newcomers can browse the list to get oriented, while researchers and engineers can quickly locate the canonical source behind a method they want to explore. It is more of a map of the field than a tutorial, so it works best for people who know the basics and want to decide which directions deserve a closer look.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category