#3 · Primary category: Data Catalog & Metadata Management

CLUEDatasetSearch

chinese corpus datasets knowledge-graph machine-reading-comprehension machine-translation match ner nlp qa sentiment-analysis text-classification text-similarity text-summarization

Search all Chinese NLP datasets, with common English NLP datasets included

Project last updated:11/21/22

GitHub Stars

4.5K

Forks

626

Contributors

4

License

Other

Why we included this project

For anyone building or evaluating Chinese-language NLP systems, finding the right dataset is often the slowest part of the work. CLUEDatasetSearch collects hundreds of Chinese datasets plus a set of commonly used English ones, organized by task: named entity recognition, question answering, sentiment analysis, text matching, machine translation, and more. Each table row links to the original source and includes a short note on what the data contains and where it came from, so you can assess fit before downloading anything. There's also a web-based search tool for filtering interactively, and the project accepts community uploads to keep the catalog current. Treat it as a discovery resource rather than a piece of software to run, and it's a solid first stop for anyone hunting for Chinese NLP data.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category