#26 · Primary category: AI Data Infrastructure & Storage
alluxio
Alluxio, data orchestration for analytics and machine learning in the cloud
Project last updated:04/29/25
GitHub Stars
7.2K
Forks
2.9K
Contributors
1.4K
License
Apache-2.0
Why we included this project
Alluxio is a distributed caching layer that sits between your compute engines and durable storage. It keeps hot data in memory across a cluster, so jobs that rerun against the same large datasets read from local memory instead of hammering object storage or HDFS. Because it exposes one virtual namespace and a common interface over many backends, you can add or swap storage systems without touching application code. The project grew out of UC Berkeley's AMPLab (it was called Tachyon there) and is used widely with Presto, Spark, and Trino for structured analytics. One honest caveat: this open-source edition targets analytics workloads, while very large-scale AI training and inference acceleration lives in the commercial product.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
ClickHouse
ClickHouse® is a real-time analytics database management system
simdjson
Parsing gigabytes of JSON per second : used by Facebook/Meta Velox, the Node.js runtime, ClickHouse, WatermelonDB, Apache Doris, Milvus, StarRocks
gun
An open source cybersecurity protocol for syncing decentralized graph data.
emqx
The most scalable and reliable MQTT broker for AI, IoT, IIoT and connected vehicles
server
MariaDB server is a community developed fork of MySQL server. Started by core members of the original MySQL team, MariaDB actively works with outside developers to deliver the most featureful, stable, and sanely licensed open SQL server in the industry.