#34 · Primary category: AI Data Infrastructure & Storage
ai-data-extraction
extract all your personal data history from cursor, codex, claude-code, windsurf, and trae
Project last updated:08/19/26
GitHub Stars
1.3K
Forks
106
Contributors
3
License
Other
Why we included this project
Most people never look at what their coding assistant has saved about them, but that history is exactly the kind of data you need to build a fine-tuning corpus. This toolkit digs through the local storage of Cursor, Codex, Claude Code, Windsurf, Trae, Continue, Gemini CLI, and OpenCode and normalizes the chat messages, code context, diffs, tool calls, and timestamps into clean JSONL. The records keep the full conversation structure rather than just the text, which matters if you want to study how models behave on real developer workflows or assemble training data from your own machine. It runs on the Python standard library alone, so there is no dependency setup to fight through. If you want to recover or archive your own assistant history, this is a practical starting point.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
ClickHouse
ClickHouse® is a real-time analytics database management system
simdjson
Parsing gigabytes of JSON per second : used by Facebook/Meta Velox, the Node.js runtime, ClickHouse, WatermelonDB, Apache Doris, Milvus, StarRocks
gun
An open source cybersecurity protocol for syncing decentralized graph data.
emqx
The most scalable and reliable MQTT broker for AI, IoT, IIoT and connected vehicles
server
MariaDB server is a community developed fork of MySQL server. Started by core members of the original MySQL team, MariaDB actively works with outside developers to deliver the most featureful, stable, and sanely licensed open SQL server in the industry.