AI Data Infrastructure & Storage

Open-source storage engines, data fabrics, and infrastructure layers engineered for high-throughput AI pipelines, covering checkpointing, dataloading, and KV-cache systems.

42 projects

See methodology for ranking rules; order uses public GitHub metrics within this scenario.

41–42 of 42

Rank Project Stars Forks
41 petastorm

Petastorm library enables single machine or distributed training and evaluation of deep learning models from datasets in Apache Parquet format. It supports ML frameworks such as Tensorflow, Pytorch, and PySpark and can be used from pure Python code.

1.9K 286
42 ffcv

FFCV: Fast Forward Computer Vision (and other ML workloads!)

3.0K 180
< Previous
/ 3
Next >