#125 · Primary category: NLP Tools & Text Processing
poetry
The most comprehensive corpus of modern Chinese poetry on Earth: 3k+ poets, 80K+ poems, 15M+ characters.
Project last updated:09/12/25
GitHub Stars
738
Forks
90
Contributors
4
License
MIT
Why we included this project
Chinese NLP work often stalls on the same problem: there is little clean, large-scale literary text to train on. This repository collects roughly 81,000 modern Chinese poems from about 3,500 poets into one structured dataset, with a documented format and scripts for contributing or adapting the data. That makes it a solid starting point for fine-tuning language models, building style-analysis tools, or studying contemporary verse computationally. Teams working on text generation, sentiment analysis, or literary research can skip the scraping and cleaning phase and get straight to modeling.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
flair
A very simple framework for state-of-the-art Natural Language Processing (NLP)
compromise
modest natural-language processing
tokenizers
💥 Fast State-of-the-Art Tokenizers optimized for Research and Production
CoreNLP
CoreNLP: A Java suite of core NLP tools for tokenization, sentence segmentation, NER, parsing, coreference, sentiment analysis, etc.
Chinese-Word-Vectors
100+ Chinese Word Vectors 上百种预训练中文词向量