#381 · Primary category: AI Tool Directories & Curated Lists

NLP_bahasa_resources

bahasa-indonesia corpus corpus-linguistics dataset indonesian indonesian-language library natural-language-processing nlp nlp-bahasa-resources packages sentiment-analysis sentiment-analysis-dataset

A Curated List of Dataset and Usable Library Resources for NLP in Bahasa Indonesia

Project last updated:02/17/23

GitHub Stars

574

Forks

144

Contributors

3

License

MIT

Why we included this project

Building an Indonesian-language NLP pipeline usually means digging through dozens of scattered GitHub repos and Hugging Face datasets. This index groups those resources by task, including named entity recognition, POS tagging, question answering, text summarization, and hate-speech detection, so you can find the corpus or library that fits the problem at hand. The dictionary section is a standout: sentiment lexicons, slang words, root words, stop words, and name-based gender and region lists that are hard to track down elsewhere. It also links to pre-trained models, usable libraries, spelling-correction tools, and Twitter scraping utilities, which makes it a practical starting point for research and production alike. Since it is a curated index rather than a single tool, use it as a map to compare candidate datasets and packages before you commit.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category