#75 · Primary category: MLOps & Evaluation

mteb

benchmark bitext-mining clustering embeddings evaluation information-retrieval low-resource-nlp mteb multilingual-nlp multimodal neural-search reranking retrieval sbert semantic-search sentence-transformers sts text-classification text-embedding

MTEB: State-of-the-art evaluation of embeddings across languages and modalities

Project last updated:08/29/26

GitHub Stars

3.4K

Forks

680

Contributors

331

License

Apache-2.0

Why we included this project

Choosing an embedding model for retrieval or search usually comes down to trusting vendor claims, and MTEB gives you a way to check them. It collects hundreds of evaluation tasks across text, image, and audio in many languages behind a small API, so you can run the same model through retrieval, reranking, semantic-similarity, and clustering tests with a few lines of Python or a CLI command rather than building your own harness. A public leaderboard shows how recent releases compare, and you can add your own tasks or datasets when the standard suite does not cover your use case. Teams that ship or publish embeddings benefit the most from it, because the results are reproducible and comparable.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category