#12 · Primary category: AI Data Infrastructure & Storage
deeplake
Deeplake is AI Data Runtime for Agents. It provides serverless postgres with a multimodal datalake, enabling scalable retrieval and training.
Project last updated:05/21/26
GitHub Stars
9.2K
Forks
724
Contributors
141
License
Apache-2.0
Why we included this project
Deeplake puts multimodal data and the vectors built from it in one queryable store, so teams building LLM applications don't have to keep a file store, a vector database, and a training pipeline in sync by hand. Its storage format is tuned for deep-learning workloads, and large datasets can be streamed straight into training runs rather than copied to local disk first. Versioning and lineage track dataset changes the way code is tracked, which matters when a model's behavior depends on exactly which data it saw. The same dataset handles retrieval in LLM apps and model training, and LangChain and LlamaIndex integrations fit teams already on those stacks. If your data moves between search and training without constant reformatting, Deeplake is worth a close look.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
ClickHouse
ClickHouse® is a real-time analytics database management system
simdjson
Parsing gigabytes of JSON per second : used by Facebook/Meta Velox, the Node.js runtime, ClickHouse, WatermelonDB, Apache Doris, Milvus, StarRocks
gun
An open source cybersecurity protocol for syncing decentralized graph data.
emqx
The most scalable and reliable MQTT broker for AI, IoT, IIoT and connected vehicles
server
MariaDB server is a community developed fork of MySQL server. Started by core members of the original MySQL team, MariaDB actively works with outside developers to deliver the most featureful, stable, and sanely licensed open SQL server in the industry.