#579 · Primary category: Education & Research
multimodal-agents-course
An MCP Multimodal AI Agent with eyes and ears!
Project last updated:01/05/26
GitHub Stars
575
Forks
148
Contributors
5
License
Apache-2.0
Why we included this project
This is a free, open-source course rather than a single tool, and it walks you through building a complete multimodal AI agent from scratch instead of just wiring an existing MCP server to a desktop app. You construct the whole stack yourself: a video processing pipeline with Pixeltable, an MCP server on FastMCP, a Groq-powered agent with its own MCP client, and Opik for observability and prompt versioning. The five modules pair written lessons with runnable code, each with a short video summary, so by the end you have a working Kubrick agent that understands images, video, audio, and text. It is written for ML, software, and data engineers who want to move past toy examples and see how production agentic systems are actually assembled, including the API layer and LLMOps practices. Because the course is free and the model providers offer free tiers, following along costs little, which makes it a practical path for teams weighing how to build similar systems.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
prompts.chat
f.k.a. Awesome ChatGPT Prompts. Share, discover, and collect prompts from the community. Free and open source — self-host for your organization with complete privacy.
JavaGuide
Java Interview & Backend General Interview Guide, covering computer fundamentals, databases, distributed systems, high concurrency, system design, and AI application development.
system-prompts-and-models-of-ai-tools
A curated collection of system prompts, internal tools, and AI models from popular AI assistants and coding agents.
30-seconds-of-code
Coding articles to level up your development skills
generative-ai-for-beginners
21 Lessons, Get Started Building with Generative AI