#75 · Primary category: MLOps & Evaluation
mteb
MTEB: State-of-the-art evaluation of embeddings across languages and modalities
Project last updated:08/29/26
GitHub Stars
3.4K
Forks
680
Contributors
331
License
Apache-2.0
Why we included this project
Choosing an embedding model for retrieval or search usually comes down to trusting vendor claims, and MTEB gives you a way to check them. It collects hundreds of evaluation tasks across text, image, and audio in many languages behind a small API, so you can run the same model through retrieval, reranking, semantic-similarity, and clustering tests with a few lines of Python or a CLI command rather than building your own harness. A public leaderboard shows how recent releases compare, and you can add your own tasks or datasets when the standard suite does not cover your use case. Teams that ship or publish embeddings benefit the most from it, because the results are reproducible and comparable.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
unsloth
Local UI to run and train LLMs and diffusion models. Supports GGUF, MLX, Qwen3.8, Kimi K3, MiniMax-H3, Gemma 4, FLUX and more.
LlamaFactory
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
airflow
Apache Airflow - A platform to programmatically author, schedule, and monitor workflows
langfuse
🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integrates with OpenTelemetry, LangChain, OpenAI SDK, LiteLLM, and more. 🍊YC W23
netron
Visualizer for neural network, deep learning and machine learning models