#104 · Primary category: NLP Tools & Text Processing
insuranceqa-corpus-zh
:helicopter: Insurance industry corpus, chatbot
Project last updated:05/26/25
GitHub Stars
1.1K
Forks
340
Contributors
1
License
Other
Why we included this project
This repository offers a Chinese translation of the InsuranceQA dataset, pairing questions asked by real users with answers written by insurance professionals, in both Chinese and English. That origin matters for training: the phrasing is genuinely how people ask, not synthetic examples. The corpus comes in two forms, raw translated question-answer text and cleaned, tokenized pairs with labels that are ready for supervised tasks like answer selection. A bundled Python package loads the train, validation, and test splits directly, which suits teams working on insurance chatbots, FAQ retrieval, or customer support NLP. The main hurdle is access: downloading the data requires a license certificate from the maintainer, so factor that step into your evaluation.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
flair
A very simple framework for state-of-the-art Natural Language Processing (NLP)
compromise
modest natural-language processing
tokenizers
💥 Fast State-of-the-Art Tokenizers optimized for Research and Production
CoreNLP
CoreNLP: A Java suite of core NLP tools for tokenization, sentence segmentation, NER, parsing, coreference, sentiment analysis, etc.
Chinese-Word-Vectors
100+ Chinese Word Vectors 上百种预训练中文词向量