Show HN: Rsync.ai – self-hosted CDC and ETL/ELT, source-available (ELv2) Rsync.ai launched as a self-hosted, source-available data platform under the Elastic License 2.0, offering batch and change-data-capture pipelines, scheduled SQL models, and table-level lineage with 21 connectors. The tool lets users describe a pipeline in plain English, which an agent converts into a staged plan that pauses for approval before executing on Temporal, and it is unrelated to the rsync(1) file-synchronization utility. The Elastic License 2.0 permits free internal use and modification but prohibits reselling it as a hosted service. Self-hosted, source-available AI data platform for batch pipelines, CDC, scheduled models, and lineage. Describe a pipeline in plain English, approve the plan, and see exactly what ran, failed, or became stale. You describe the pipeline in plain English. It resolves the plan, then stops for your approval — batch, CDC, or changes-only — before a single row moves. rsync.ai moves data between databases, warehouses, object stores and APIs. You describe the job in a sentence; an agent turns it into an explicit, staged plan, pauses for you when something is ambiguous, and executes it on Temporal so a long sync survives restarts. Batch and change-data-capture are both first-class. Twenty-one connectors ship in the box. It is unrelated to rsync 1 https://rsync.samba.org/ , the file-synchronisation tool — this moves rows between systems, not files between hosts. It is source-available under the Elastic License 2.0 https://github.com/rsync-ai/rsync/blob/main/LICENSE : run it, modify it, and use it internally for free — you just cannot resell it as a hosted service. The full summary is below license . - Batch and CDC pipelines. Batch loads between the connectors below, plus Debezium-backed change data capture from PostgreSQL, MySQL, SQL Server, Oracle and MongoDB. A run pauses for your decision where the request is ambiguous, and each stage reports what it did. → PostgreSQL CDC https://github.com/rsync-ai/rsync/blob/main/docs/solutions/self-hosted-postgresql-cdc-pipeline.md · PostgreSQL to MySQL https://github.com/rsync-ai/rsync/blob/main/docs/solutions/postgresql-to-mysql-data-sync.md · Shopify to PostgreSQL https://github.com/rsync-ai/rsync/blob/main/docs/solutions/shopify-to-postgresql-data-pipeline.md - Scheduled, dependency-aware SQL models. Save a query as a model and rebuild it on a cron, an interval, or after the pipeline or model it reads from finishes; edits to scheduled SQL need an admin's approval, and a freshness deadline flags a table that stopped moving. → Scheduled SQL models https://github.com/rsync-ai/rsync/blob/main/docs/solutions/scheduled-sql-models-with-dependency-triggers.md - Data Explorer and lineage. Query what you connected in English or SQL, and see which pipelines write which tables and which models read them. Lineage is table-level, and the lineage view is recent — its page states how far it has been verified. → Data Explorer https://github.com/rsync-ai/rsync/blob/main/docs/explorer/README.md · Lineage and observability https://github.com/rsync-ai/rsync/blob/main/docs/solutions/data-lineage-and-pipeline-observability.md - Versioned MCP connectors. Each of the 21 connectors runs as its own versioned container, so you can upgrade or pin one without touching the rest. → Connector reference https://github.com/rsync-ai/rsync/blob/main/docs/connectors/reference.md More guides: all solutions https://github.com/rsync-ai/rsync/blob/main/docs/solutions/README.md . Ask in plain English, review the SQL it wrote, run it against a connected source. Here: how many swipes went left versus right. A finished run, stage by stage: what each one did, the rows it moved, and how stale the destination has become since. Nothing here was typed in by hand. Every pipeline in a workspace on one screen: batch or CDC, source to destination, and whether it is running right now. Lineage across pipelines and scheduled SQL models: which table a model writes, which model reads it, and which one runs after which. Managed ELT tools move data well but hand off at the warehouse door. Orchestrators and automation tools are general-purpose and leave the data semantics to you. rsync.ai aims at the middle: get the data moving and keep it modelled, on hardware you control. | Instead of | What it does well | What rsync.ai does differently | |---|---|---| | Fivetran | Managed and reliable, hundreds of connectors, someone else is on call | Runs on your infrastructure with your keys. A connector you need is a container you can write, not a support ticket. | | Airbyte | Large connector ecosystem, self-hostable, mature ELT | You describe the pipeline in a sentence and approve a plan instead of configuring each sync by hand, and batch and CDC are the same product rather than separate paths. | | dbt | The standard for SQL transformation, with deep testing and a large package ecosystem | Scheduled, dependency-aware SQL models are built in, so moving and modelling data is one tool instead of two. dbt's testing and packages are considerably deeper. | | Debezium on its own | Best-in-class change data capture | rsync.ai runs Debezium and adds the provisioning, sinks, retries and UI around it, so you are not assembling Kafka Connect by hand. | | Airflow / n8n | General orchestration and automation, enormously flexible | A pipeline is a first-class object with row counts, lineage and CDC built in, rather than something you assemble from operators or nodes. | Where it is honestly weaker. There is no managed option — every install is yours to run. The catalogue is 21 connectors, not hundreds. Data-quality assertions are not built yet. And the Kubernetes path is younger than the Docker one see Project status project-status . If you want someone else carrying the pager, use a managed tool. 1. Install with one command — Install install below. Docker is the only requirement. 2. Open http://localhost:3000 and click Start with sample data . The stack bundles a sample-data source and a throwaway demo-warehouse PostgreSQL, so this needs no credential of your own. 3. In /chat , ask for "sync customers and orders from sample data to the demo warehouse" , pick the tables, and confirm. That path is a batch pipeline. CDC, Shopify and your own databases need a source of your own — see the quickstart https://github.com/rsync-ai/rsync/blob/main/docs/getting-started/quickstart.md try-it-in-5-minutes-with-no-credentials and the self-hosting guide https://github.com/rsync-ai/rsync/blob/main/docs/deployment/self-hosting.md . php flowchart LR U "You, in plain English" -- FE "Frontend