#3 · Primary category: Data Catalog & Metadata Management
CLUEDatasetSearch
Search all Chinese NLP datasets, with common English NLP datasets included
Project last updated:11/21/22
GitHub Stars
4.5K
Forks
626
Contributors
4
License
Other
Why we included this project
For anyone building or evaluating Chinese-language NLP systems, finding the right dataset is often the slowest part of the work. CLUEDatasetSearch collects hundreds of Chinese datasets plus a set of commonly used English ones, organized by task: named entity recognition, question answering, sentiment analysis, text matching, machine translation, and more. Each table row links to the original source and includes a short note on what the data contains and where it came from, so you can assess fit before downloading anything. There's also a web-based search tool for filtering interactively, and the project accepts community uploads to keep the catalog current. Treat it as a discovery resource rather than a piece of software to run, and it's a solid first stop for anyone hunting for Chinese NLP data.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.