#579 · Primary category: Education & Research

multimodal-agents-course

agent embeddings groq mcp mcp-client mcp-server multimodal openai opik pixeltable

An MCP Multimodal AI Agent with eyes and ears!

Project last updated:01/05/26

GitHub Stars

575

Forks

148

Contributors

5

License

Apache-2.0

Why we included this project

This is a free, open-source course rather than a single tool, and it walks you through building a complete multimodal AI agent from scratch instead of just wiring an existing MCP server to a desktop app. You construct the whole stack yourself: a video processing pipeline with Pixeltable, an MCP server on FastMCP, a Groq-powered agent with its own MCP client, and Opik for observability and prompt versioning. The five modules pair written lessons with runnable code, each with a short video summary, so by the end you have a working Kubrick agent that understands images, video, audio, and text. It is written for ML, software, and data engineers who want to move past toy examples and see how production agentic systems are actually assembled, including the API layer and LLMOps practices. Because the course is free and the model providers offer free tiers, following along costs little, which makes it a practical path for teams weighing how to build similar systems.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category