#184 · Primary category: AI Chatbots
Multimodal-GPT
Multimodal-GPT
Project last updated:06/04/23
GitHub Stars
1.5K
Forks
128
Contributors
11
License
Apache-2.0
Why we included this project
Multimodal-GPT is an early, readable example of how a vision-language chatbot gets built. The OpenMMLab team took OpenFlamingo and trained it on visual instruction data they generated from existing VQA, captioning, reasoning, OCR, and dialogue datasets, then added language-only instruction training for the language component. That joint training is the interesting part: the README reports it noticeably improves how the model responds, so the repo works as a template if you want to build your own assistant on open weights. A technical report and training code accompany the project, which helps if you are new to adapting an OpenFlamingo-style model for dialogue. Note that development has been quiet since mid-2023, so treat it as a learning resource rather than a maintained dependency.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
open-webui
User-friendly AI Interface (Supports Ollama, OpenAI API, ...)
NextChat
✨ Light and Fast AI Assistant. Support: Web | iOS | MacOS | Android | Linux | Windows
gpt4free
The official gpt4free repository | various collection of powerful language models | opus 4.6 gpt 5.3 kimi 2.5 deepseek v3.2 gemini 3
system_prompts_leaks
Extracted system prompts from Anthropic - Claude Fable 5, Opus 5, Claude Design, Claude Code. OpenAI - ChatGPT GPT-5.6-Sol, Codex. Google - Gemini 3.5 Flash, 3.1 Pro, Antigravity. xAI - Grok, Cursor, Copilot, VS Code, Perplexity, and more. Updated regularly.
cherry-studio
AI productivity studio with smart chat, autonomous agents, and 300+ assistants. Unified access to frontier LLMs