#19 · Primary category: NLP Tools & Text Processing

grobid

bibliographical-references crf deep-learning fulltext hamburger-to-cow machine-learning metadata pdf rnn scientific-articles transformers

A machine learning software for extracting information from scholarly documents

Project last updated:08/28/26

GitHub Stars

5.1K

Forks

568

Contributors

75

License

Apache-2.0

Why we included this project

GROBID turns messy scholarly PDFs into structured TEI/XML, using trained models that recognize headers, reference lists, citation contexts, full-text sections, figures, and tables instead of leaving you to scrape raw text. That structured output is what teams building academic search engines, citation databases, or literature knowledge graphs feed into their own systems. The project has been open source since 2011, is used in production by well-known research platforms, and ships with a web service API, Docker images, and batch processing, so it works for small integration projects as well as large-scale ingestion. It also includes its own training and benchmarking framework if you want to retune the models for your own document domain.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category