#12 · Primary category: AI Data Infrastructure & Storage

deeplake

agent agentic-rag ai clawbot computer-vision datalake deep-learning filesystem large-language-models llm memory mlops multimodal openclaw postgres pytorch rag skill vector-database

Deeplake is AI Data Runtime for Agents. It provides serverless postgres with a multimodal datalake, enabling scalable retrieval and training.

Project last updated:05/21/26

GitHub Stars

9.2K

Forks

724

Contributors

141

License

Apache-2.0

Why we included this project

Deeplake puts multimodal data and the vectors built from it in one queryable store, so teams building LLM applications don't have to keep a file store, a vector database, and a training pipeline in sync by hand. Its storage format is tuned for deep-learning workloads, and large datasets can be streamed straight into training runs rather than copied to local disk first. Versioning and lineage track dataset changes the way code is tracked, which matters when a model's behavior depends on exactly which data it saw. The same dataset handles retrieval in LLM apps and model training, and LangChain and LlamaIndex integrations fit teams already on those stacks. If your data moves between search and training without constant reformatting, Deeplake is worth a close look.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category