NLP Tools & Text Processing
Libraries and utilities for core natural-language processing tasks such as language detection, text classification, tokenization, and text normalization.
124 projects
See methodology for ranking rules; order uses public GitHub metrics within this scenario.
| Rank | Project | Stars | Forks | Updated | License |
|---|---|---|---|---|---|
| 81 |
mt-dnn
Multi-Task Deep Neural Networks for Natural Language Understanding |
2.3K | 408 | 03/07/24 | MIT |
| 82 |
Rasa_NLU_Chi
Turn Chinese natural language into structured data 中文自然语言理解 |
1.5K | 416 | 07/30/24 | Apache-2.0 |
| 83 |
ecco
Explain, analyze, and visualize NLP language models. Ecco creates interactive visualizations directly in Jupyter notebooks explaining the behavior of Transformer-based language models (like GPT2, BERT, RoBERTA, T5, and T0). |
2.1K | 176 | 08/15/24 | BSD-3-Clause |
| 84 |
fast-bert
Super easy library for BERT based NLP models |
1.9K | 339 | 08/19/24 | Apache-2.0 |
| 85 |
nlpcda
一键中文数据增强包 ; NLP数据增强、bert数据增强、EDA:pip install nlpcda |
1.9K | 171 | 03/18/25 | Apache-2.0 |
| 86 |
jieba-php
"結巴"中文分詞:做最好的 PHP 中文分詞、中文斷詞組件。 / "Jieba" (Chinese for "to stutter") Chinese text segmentation: built to be the best PHP Chinese word segmentation module. |
1.4K | 257 | 12/16/25 | MIT |
| 87 |
textacy
NLP, before and after spaCy |
2.2K | 247 | 09/22/23 | Other |
| 88 |
Information-Extraction-Chinese
Chinese Named Entity Recognition with IDCNN/biLSTM+CRF, and Relation Extraction with biGRU+2ATT 中文实体识别与关系提取 |
2.3K | 798 | 02/01/24 | Other |
| 89 |
mq
A jq-like Markdown query language for command-line processing |
1.0K | 21 | 08/29/26 | MIT |
| 90 |
transformers_tasks
⭐️ NLP Algorithms with transformers lib. Supporting Text-Classification, Text-Generation, Information-Extraction, Text-Matching, RLHF, SFT etc. |
2.4K | 398 | 09/29/23 | Other |
| 91 |
extractous
Fast and efficient unstructured data extraction. Written in Rust with bindings for many languages. |
1.8K | 95 | 12/21/24 | Apache-2.0 |
| 92 |
lingua-rs
The most accurate natural language detection library for Rust, suitable for short text and mixed-language text |
1.1K | 63 | 03/26/26 | Apache-2.0 |
| 93 |
this-word-does-not-exist
This Word Does Not Exist |
1.0K | 85 | 06/17/26 | MIT |
| 94 |
whatlang-rs
Natural language detection library for Rust. Try demo online: https://whatlang.org/ |
1.1K | 119 | 12/24/25 | MIT |
| 95 |
clean-text
🧹 Python package for text cleaning |
1.0K | 83 | 05/15/26 | Other |
| 96 |
contextualized-topic-models
A python package to run contextualized topic modeling. CTMs combine contextualized embeddings (e.g., BERT) with topic models to get coherent topics. Published at EACL and ACL 2021 (Bianchi et al.). |
1.3K | 155 | 07/24/25 | MIT |
| 97 |
Keras-TextClassification
Keras-based toolkit for Chinese text classification, multi-label classification, and sentence similarity with diverse models (FastText, TextCNN, RCNN, BERT, etc.). |
1.8K | 397 | 06/17/24 | MIT |
| 98 |
CLUECorpus2020
Large-scale Pre-training Corpus for Chinese 100G 中文预训练语料 |
1.0K | 83 | 02/06/26 | MIT |
| 99 |
lingua-go
The most accurate natural language detection library for Go, suitable for short text and mixed-language text |
1.4K | 81 | 02/06/25 | Apache-2.0 |
| 100 |
DeepMoji
State-of-the-art deep learning model for analyzing sentiment, emotion, sarcasm etc. |
1.6K | 307 | 08/02/24 | MIT |