NLP Tools & Text Processing
Libraries and utilities for core natural-language processing tasks such as language detection, text classification, tokenization, and text normalization.
124 projects
See methodology for ranking rules; order uses public GitHub metrics within this scenario.
| Rank | Project | Stars | Forks | Updated | License |
|---|---|---|---|---|---|
| 1 |
flair
A very simple framework for state-of-the-art Natural Language Processing (NLP) |
14.4K | 2.1K | 10/27/25 | Other |
| 2 |
compromise
modest natural-language processing |
12.1K | 667 | 08/23/26 | MIT |
| 3 |
tokenizers
💥 Fast State-of-the-Art Tokenizers optimized for Research and Production |
11.0K | 1.2K | 08/27/26 | Apache-2.0 |
| 4 |
CoreNLP
CoreNLP: A Java suite of core NLP tools for tokenization, sentence segmentation, NER, parsing, coreference, sentiment analysis, etc. |
10.1K | 2.7K | 08/29/26 | GPL-3.0 |
| 5 |
Chinese-Word-Vectors
100+ Chinese Word Vectors 上百种预训练中文词向量 |
12.2K | 2.3K | 10/30/23 | Apache-2.0 |
| 6 |
TextBlob
Simple, Pythonic, text processing--Sentiment analysis, part-of-speech tagging, noun phrase extraction, translation, and more. |
9.5K | 1.2K | 08/25/26 | MIT |
| 7 |
minbpe
Minimal, clean code for the Byte Pair Encoding (BPE) algorithm commonly used in LLM tokenization. |
10.7K | 1.1K | 07/01/24 | MIT |
| 8 |
nlp_chinese_corpus
大规模中文自然语言处理语料 Large Scale Chinese Corpus for NLP |
9.9K | 1.6K | 02/06/26 | MIT |
| 9 |
pattern
Web mining module for Python, with tools for scraping, natural language processing, machine learning, network analysis and visualization. |
8.9K | 1.6K | 08/05/26 | BSD-3-Clause |
| 10 |
BERTopic
Leveraging BERT and c-TF-IDF to create easily interpretable topics. |
7.8K | 917 | 08/28/26 | MIT |
| 11 |
stanza
Stanford NLP Python library for tokenization, sentence segmentation, NER, and parsing of many human languages |
7.9K | 959 | 08/29/26 | Other |
| 12 |
trafilatura
Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML |
6.7K | 425 | 08/28/26 | Apache-2.0 |
| 13 |
text_classification
all kinds of text classification models and more with deep learning |
7.9K | 2.5K | 09/28/23 | MIT |
| 14 |
GPT2-Chinese
Chinese version of GPT2 training code, using BERT tokenizer. |
7.6K | 1.7K | 04/25/24 | MIT |
| 15 |
nlp.js
An NLP library for building bots, with entity extraction, sentiment analysis, automatic language identify, and so more |
6.6K | 631 | 01/09/25 | MIT |
| 16 |
Parsr
Transforms PDF, Documents and Images into Enriched Structured Data |
6.2K | 318 | 03/20/26 | Apache-2.0 |
| 17 |
sensitive-word
👮♂️The sensitive word tool for java. |
6.0K | 806 | 03/23/26 | Apache-2.0 |
| 18 |
openmed
Local-first healthcare AI: clinical NER & HIPAA PII de-identification that runs 100% on-device. 2,200+ medical models, 21 languages, Apple MLX + Python, no cloud, no patient data leaving your network. Apache-2.0 |
5.2K | 655 | 08/26/26 | Apache-2.0 |
| 19 |
grobid
A machine learning software for extracting information from scholarly documents |
5.1K | 568 | 08/29/26 | Apache-2.0 |
| 20 |
ansj_seg
ansj分词.ict的真正java实现.分词效果速度都超过开源版的ict. |
6.5K | 2.3K | 11/19/23 | Apache-2.0 |