#6 · Primary category: AI Data Infrastructure & Storage
seatunnel
SeaTunnel is a multimodal, high-performance, distributed, massive data integration tool.
Project last updated:08/29/26
GitHub Stars
9.6K
Forks
2.4K
Contributors
523
License
Apache-2.0
Why we included this project
Teams moving serious volumes of data between systems will find SeaTunnel useful long after its Apache ETL roots stop being the point. It is a distributed integration layer that reads from databases, message queues, object stores, and files, applies light transforms in between, and lands the results in warehouses, lakehouses, search systems, and vector stores, all defined in one config file that runs unchanged on its own Zeta engine, Flink, or Spark. For AI teams, the interesting part is where it is headed: embedding transforms, vector-store sinks like Milvus and Qdrant, and work toward keeping vector indexes consistent as source documents change. Anyone building RAG or multimodal pipelines who does not want to hand-roll document-to-vector movement gets the same machinery for ordinary batch, streaming, and CDC jobs too. It is a data plumbing tool that now also feeds AI applications, not a model or agent framework.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
ClickHouse
ClickHouse® is a real-time analytics database management system
simdjson
Parsing gigabytes of JSON per second : used by Facebook/Meta Velox, the Node.js runtime, ClickHouse, WatermelonDB, Apache Doris, Milvus, StarRocks
gun
An open source cybersecurity protocol for syncing decentralized graph data.
emqx
The most scalable and reliable MQTT broker for AI, IoT, IIoT and connected vehicles
server
MariaDB server is a community developed fork of MySQL server. Started by core members of the original MySQL team, MariaDB actively works with outside developers to deliver the most featureful, stable, and sanely licensed open SQL server in the industry.