#193 · Primary category: AI Chatbots
Dialog_Corpus
用于训练中英文对话系统的语料库 Datasets for Training Chatbot System
Project last updated:09/23/20
GitHub Stars
2.1K
Forks
489
Contributors
2
License
Other
Why we included this project
Training a chatbot in Chinese or English usually stalls on the same problem: there is no obvious place to find real dialogue data. This repository gathers the scattered public corpora that do exist into one index, linking each entry back to its original source. The notes are what make it useful, because quality varies a lot across the sets. Chinese movie dialogue is noisy and poorly paired, the ChatterBot Chinese data is small but clean, and the English NLP collections may need machine translation before they fit a Chinese bot. The list even flags one corpus, Microsoft Xiaoice, that is rumored to exist but was never released, saving you from chasing it. This is a map of what training data is out there, not a ready-to-use dataset.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
open-webui
User-friendly AI Interface (Supports Ollama, OpenAI API, ...)
NextChat
✨ Light and Fast AI Assistant. Support: Web | iOS | MacOS | Android | Linux | Windows
gpt4free
The official gpt4free repository | various collection of powerful language models | opus 4.6 gpt 5.3 kimi 2.5 deepseek v3.2 gemini 3
system_prompts_leaks
Extracted system prompts from Anthropic - Claude Fable 5, Opus 5, Claude Design, Claude Code. OpenAI - ChatGPT GPT-5.6-Sol, Codex. Google - Gemini 3.5 Flash, 3.1 Pro, Antigravity. xAI - Grok, Cursor, Copilot, VS Code, Perplexity, and more. Updated regularly.
cherry-studio
AI productivity studio with smart chat, autonomous agents, and 300+ assistants. Unified access to frontier LLMs