NLP Tools & Text Processing

Libraries and utilities for core natural-language processing tasks such as language detection, text classification, tokenization, and text normalization.

124 projects

See methodology for ranking rules; order uses public GitHub metrics within this scenario.

41–60 of 124

Rank Project Stars Forks
41 prose

:book: A Golang library for text processing, including tokenization, part-of-speech tagging, and named-entity extraction.

3.1K 170
42 rust-bert

Rust native ready-to-use NLP pipelines and transformer-based models (BERT, DistilBERT, GPT2,...)

3.1K 249
43 setfit

Efficient few-shot learning with Sentence Transformers

2.8K 264
44 SimCSE

[EMNLP 2021] SimCSE: Simple Contrastive Learning of Sentence Embeddings https://arxiv.org/abs/2104.08821

3.7K 537
45 MITIE

MITIE: library and tools for information extraction

3.0K 532
46 gse

Go efficient multilingual NLP and text segmentation; support English, Chinese, Japanese and others.

2.8K 232
47 snips-nlu

Snips Python library to extract meaning from text

4.0K 505
48 Recognizers-Text

Cross-platform library for recognizing and resolving numbers, units, and date/time in multiple languages.

1.8K 434
49 BERT-BiLSTM-CRF-NER

Tensorflow solution of NER task Using BiLSTM-CRF model with Google BERT Fine-tuning And private Server services

4.9K 1.2K
50 scattertext

Beautiful visualizations of how language differs among document types.

2.3K 286
51 model2vec

Fast State-of-the-Art Static Embeddings

2.2K 124
52 pytextrank

Python implementation of TextRank algorithms ("textgraphs") for phrase extraction

2.2K 334
53 tika-python

Tika-Python is a Python binding to the Apache Tika™ REST services allowing Tika to be called natively in the Python community.

1.7K 249
54 opennlp

Apache OpenNLP

1.6K 505
55 docext

An on-premises, OCR-free unstructured data extraction, markdown conversion and benchmarking toolkit. (https://idp-leaderboard.org/)

2.1K 155
56 nltk_data

NLTK Data

1.8K 1.1K
57 scispacy

A full spaCy pipeline and models for scientific/biomedical documents.

2.0K 257
58 budoux

A compact, standalone machine learning tool for line breaking that supports multiple languages and HTML inputs.

1.8K 45
59 lingua-py

The most accurate natural language detection library for Python, suitable for short text and mixed-language text

1.8K 61
60 InvoiceNet

Deep neural network to extract intelligent information from invoice documents.

2.7K 411