#19 · Primary category: NLP Tools & Text Processing
grobid
A machine learning software for extracting information from scholarly documents
Project last updated:08/28/26
GitHub Stars
5.1K
Forks
568
Contributors
75
License
Apache-2.0
Why we included this project
GROBID turns messy scholarly PDFs into structured TEI/XML, using trained models that recognize headers, reference lists, citation contexts, full-text sections, figures, and tables instead of leaving you to scrape raw text. That structured output is what teams building academic search engines, citation databases, or literature knowledge graphs feed into their own systems. The project has been open source since 2011, is used in production by well-known research platforms, and ships with a web service API, Docker images, and batch processing, so it works for small integration projects as well as large-scale ingestion. It also includes its own training and benchmarking framework if you want to retune the models for your own document domain.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
flair
A very simple framework for state-of-the-art Natural Language Processing (NLP)
compromise
modest natural-language processing
tokenizers
💥 Fast State-of-the-Art Tokenizers optimized for Research and Production
CoreNLP
CoreNLP: A Java suite of core NLP tools for tokenization, sentence segmentation, NER, parsing, coreference, sentiment analysis, etc.
Chinese-Word-Vectors
100+ Chinese Word Vectors 上百种预训练中文词向量