AI Data Infrastructure & Storage

Open-source storage engines, data fabrics, and infrastructure layers engineered for high-throughput AI pipelines, covering checkpointing, dataloading, and KV-cache systems.

42 projects

See methodology for ranking rules; order uses public GitHub metrics within this scenario.

1–20 of 42

Rank Project Stars Forks
1 ClickHouse

ClickHouse® is a real-time analytics database management system

49.5K 8.9K
2 simdjson

Parsing gigabytes of JSON per second : used by Facebook/Meta Velox, the Node.js runtime, ClickHouse, WatermelonDB, Apache Doris, Milvus, StarRocks

24.2K 1.3K
3 gun

An open source cybersecurity protocol for syncing decentralized graph data.

19.1K 1.2K
4 emqx

The most scalable and reliable MQTT broker for AI, IoT, IIoT and connected vehicles

16.7K 2.5K
5 server

MariaDB server is a community developed fork of MySQL server. Started by core members of the original MySQL team, MariaDB actively works with outside developers to deliver the most featureful, stable, and sanely licensed open SQL server in the industry.

8.2K 2.1K
6 seatunnel

SeaTunnel is a multimodal, high-performance, distributed, massive data integration tool.

9.6K 2.4K
7 gel

Gel supercharges Postgres with a modern data model, graph queries, Auth & AI solutions, and much more.

14.2K 452
8 databend

Data Agent Ready Warehouse : One for Analytics, Search, AI, Python Sandbox. — rebuilt from scratch. Unified architecture on your S3.

9.4K 895
9 risingwave

Event streaming platform for agentic AI. Continuously ingest, transform, and serve event streams in real time, at scale.

9.3K 825
10 crawlee-python

Python library for web scraping and browser automation to extract data for AI, LLMs, RAG, and GPTs.

9.5K 800
11 3FS

A high-performance distributed file system designed to address the challenges of AI training and inference workloads.

10.2K 1.1K
12 deeplake

Deeplake is AI Data Runtime for Agents. It provides serverless postgres with a multimodal datalake, enabling scalable retrieval and training.

9.2K 723
13 lance

Open Lakehouse Format for Multimodal AI. Convert from Parquet in 2 lines of code for 100x faster random access, vector index, and data versioning. Compatible with Pandas, DuckDB, Polars, Pyarrow, and PyTorch with more integrations coming..

7.0K 826
14 unstract

LLM-Driven Extraction of Unstructured Data — Built for API Deployments & ETL Pipeline Workflows

7.2K 709
15 materialize

The live data layer for apps and AI agents. Create up-to-the-second views into your business, just using SQL

6.4K 511
16 Daft

High-performance data engine for AI and multimodal workloads. Process images, audio, video, and structured data at any scale

5.7K 549
17 datasets

TFDS is a collection of datasets ready to use with TensorFlow, Jax, ...

4.6K 1.6K
18 crate

CrateDB is a distributed and scalable SQL database for storing and analyzing massive amounts of data in near real-time, even with complex queries. It is PostgreSQL-compatible, and based on Lucene.

4.4K 612
19 memgraph

High-performance open-source in-memory graph database for GraphRAG, AI memory, agentic AI, and real-time graph analytics. Cypher-compatible, built in C++.

4.4K 261
20 vortex

An extensible, state-of-the-art framework for columnar compression, and the fastest FOSS columnar file format. Formerly at @spiraldb, now an Incubation Stage project at LFAI&Data, part of the Linux Foundation.

3.2K 214
< Previous
/ 3
Next >