AI Data Infrastructure & Storage
Open-source storage engines, data fabrics, and infrastructure layers engineered for high-throughput AI pipelines, covering checkpointing, dataloading, and KV-cache systems.
42 projects
See methodology for ranking rules; order uses public GitHub metrics within this scenario.
| Rank | Project | Stars | Forks | Updated | License |
|---|---|---|---|---|---|
| 1 |
ClickHouse
ClickHouse® is a real-time analytics database management system |
49.5K | 8.9K | 08/29/26 | Apache-2.0 |
| 2 |
simdjson
Parsing gigabytes of JSON per second : used by Facebook/Meta Velox, the Node.js runtime, ClickHouse, WatermelonDB, Apache Doris, Milvus, StarRocks |
24.2K | 1.3K | 08/27/26 | Apache-2.0 |
| 3 |
gun
An open source cybersecurity protocol for syncing decentralized graph data. |
19.1K | 1.2K | 08/01/26 | Other |
| 4 |
emqx
The most scalable and reliable MQTT broker for AI, IoT, IIoT and connected vehicles |
16.7K | 2.5K | 08/29/26 | Other |
| 5 |
server
MariaDB server is a community developed fork of MySQL server. Started by core members of the original MySQL team, MariaDB actively works with outside developers to deliver the most featureful, stable, and sanely licensed open SQL server in the industry. |
8.2K | 2.1K | 08/29/26 | GPL-2.0 |
| 6 |
seatunnel
SeaTunnel is a multimodal, high-performance, distributed, massive data integration tool. |
9.6K | 2.4K | 08/29/26 | Apache-2.0 |
| 7 |
gel
Gel supercharges Postgres with a modern data model, graph queries, Auth & AI solutions, and much more. |
14.2K | 452 | 12/24/25 | Apache-2.0 |
| 8 |
databend
Data Agent Ready Warehouse : One for Analytics, Search, AI, Python Sandbox. — rebuilt from scratch. Unified architecture on your S3. |
9.4K | 895 | 08/29/26 | Other |
| 9 |
risingwave
Event streaming platform for agentic AI. Continuously ingest, transform, and serve event streams in real time, at scale. |
9.3K | 825 | 08/29/26 | Apache-2.0 |
| 10 |
crawlee-python
Python library for web scraping and browser automation to extract data for AI, LLMs, RAG, and GPTs. |
9.5K | 800 | 08/28/26 | Apache-2.0 |
| 11 |
3FS
A high-performance distributed file system designed to address the challenges of AI training and inference workloads. |
10.2K | 1.1K | 05/07/26 | MIT |
| 12 |
deeplake
Deeplake is AI Data Runtime for Agents. It provides serverless postgres with a multimodal datalake, enabling scalable retrieval and training. |
9.2K | 723 | 05/21/26 | Apache-2.0 |
| 13 |
lance
Open Lakehouse Format for Multimodal AI. Convert from Parquet in 2 lines of code for 100x faster random access, vector index, and data versioning. Compatible with Pandas, DuckDB, Polars, Pyarrow, and PyTorch with more integrations coming.. |
7.0K | 826 | 08/29/26 | Apache-2.0 |
| 14 |
unstract
LLM-Driven Extraction of Unstructured Data — Built for API Deployments & ETL Pipeline Workflows |
7.2K | 709 | 08/29/26 | AGPL-3.0 |
| 15 |
materialize
The live data layer for apps and AI agents. Create up-to-the-second views into your business, just using SQL |
6.4K | 511 | 08/29/26 | Other |
| 16 |
Daft
High-performance data engine for AI and multimodal workloads. Process images, audio, video, and structured data at any scale |
5.7K | 549 | 08/26/26 | Apache-2.0 |
| 17 |
datasets
TFDS is a collection of datasets ready to use with TensorFlow, Jax, ... |
4.6K | 1.6K | 08/21/26 | Apache-2.0 |
| 18 |
crate
CrateDB is a distributed and scalable SQL database for storing and analyzing massive amounts of data in near real-time, even with complex queries. It is PostgreSQL-compatible, and based on Lucene. |
4.4K | 612 | 08/28/26 | Apache-2.0 |
| 19 |
memgraph
High-performance open-source in-memory graph database for GraphRAG, AI memory, agentic AI, and real-time graph analytics. Cypher-compatible, built in C++. |
4.4K | 261 | 08/29/26 | Other |
| 20 |
vortex
An extensible, state-of-the-art framework for columnar compression, and the fastest FOSS columnar file format. Formerly at @spiraldb, now an Incubation Stage project at LFAI&Data, part of the Linux Foundation. |
3.2K | 214 | 08/29/26 | Apache-2.0 |