Skip to main content
How it works

Three steps.
One sentence to a running pipeline.

No SQL. No YAML. No DAG editors. Describe what you want — the AI agent handles schema discovery, PII detection, and pipeline setup, then waits for your approval before anything runs.

Start free
rsync.ai table statistics for a streaming CDC pipeline: 4 tables, 500 changes captured and 500 applied
~30 sec
Describe it in plain EnglishPick tables, PII checked firstEvery change captured and appliedAsk in plain English, get the SQL

The three steps, expanded

Here's exactly what happens from the moment you type to the moment rows are flowing.

Step 01

Describe it in plain English

Type what you want — the same way you'd ask a teammate in Slack. The LLM agent parses your intent and extracts every parameter it needs:

  • Source: Which system to read from (Shopify, MySQL, SQL Server…)
  • Destination: Where to write (Postgres, BigQuery, S3…)
  • Cadence: Hourly, daily, real-time CDC, or on-demand
  • PII rules: "mask emails", "exclude phone numbers"
  • Table filters: "only the orders and products tables"

If the source isn't in the built-in catalog, the AI Tool Generator (on rsync.ai Cloud) builds a connector from an OpenAPI, GraphQL or docs URL in minutes — auth, schema discovery and pagination included. Review it before you run it.

rsync.ai agent

YOU

Sync my Shopify orders and products to Postgres hourly. Mask customer emails.

ASSISTANT

Understood: Shopify Admin (GraphQL) → PostgreSQL, hourly batch, masking customer.email. Create and run this pipeline?

Confirm Pipeline

Shopify (GraphQL)PostgreSQL
Start pipelineCancel
Step 02

Review the plan, then confirm

rsync.ai never creates or starts a pipeline until you confirm the plan. Along the way the agent checks the connection, reads the schema and scans for PII, and stops to ask whenever your request leaves something open:

1

Connection test

The agent checks that it can reach source and destination with the credentials you supplied. You are asked only if something is missing.

2

Schema discovery

Agent scans source tables, infers column types, and proposes the destination schema.

3

PII scan

Columns are scanned for likely emails, phones, IDs, and addresses. You choose: mask, hash (SHA-256 or HMAC), drop, or pass-through — per field.

4

Confirm the plan

You confirm the table list, cadence, and rules. Only then does the pipeline start.

Pipeline setup

Connectionpassed
Schema discoveredready
PII scan completereviewed
customer.emailmasked
customer.phonedropped
order_idpass-through
Final approvalawaiting you →
Step 03

Pipeline runs — you watch

Once you approve, rsync.ai kicks off a Temporal workflow. No servers to babysit, no cron jobs to monitor.

  • Live row counts

    Watch rows land in your destination in real time from the executions dashboard.

  • Checkpointed cursors

    Temporal persists state at every checkpoint — a crash mid-sync resumes from the last checkpoint instead of starting over.

  • Debezium CDC for databases

    Postgres, MySQL, SQL Server, and MongoDB changes propagate to your destination continuously via log-based CDC.

  • Run history and traces

    Every run keeps its history and OpenTelemetry traces, so you can see what happened in any past run.

Executionsrunning
rows synced48,293
last checkpointjust now
sourceShopify (GraphQL)
destinationPostgreSQL
cadencehourly batch
PII maskcustomer.email ✓
workflow idwf_shopify_pg_8af2
retries0
tracesOpenTelemetry →

Under the hood

Built on four components you can inspect: Debezium for change capture, Temporal for durable runs, Kafka for transport, OpenTelemetry for traces.

MCP-based connectors

Connectors implement the Model Context Protocol, so the connector that feeds a pipeline is also a tool an AI client can call. To let your own AI tools work in a workspace, use the Agent gateway.

Debezium CDC

Log-based change data capture for Postgres, MySQL, SQL Server, and MongoDB. Row-level inserts and updates are streamed to your destination in near real time.

Temporal workflows

Every pipeline is a durable Temporal workflow with checkpointed state, automatic retries, and a full event history.

OpenTelemetry

End-to-end traces for every sync run. Inspect span timings, row counts, and errors without reading logs.

Self-hosting

Your VPC. Your data.
One command.

rsync.ai runs as a managed Cloud you can start free, or self-hosted — run the entire stack inside your own infrastructure — credentials AES-256 encrypted at rest with a key you control, data never leaving your network.

  • Docker Compose on one VM (Kubernetes via Helm is in preview) — Temporal, Debezium, Kafka, API and UI
  • AES-256 encrypted credentials at rest
  • Ollama support for fully local LLM inference
  • Source-available under the rsync.ai Source-Available License
  • No per-row or per-MAR pricing. Self-hosted installs run with billing switched off.
Read the self-hosting guide

Install & run

$ curl -sSL https://raw.githubusercontent.com/rsync-ai/rsync.ai/main/install.sh | bash

# Kubernetes (Helm chart, in preview):

$ curl -sSL https://raw.githubusercontent.com/rsync-ai/rsync.ai/main/install-k8s.sh | bash

Services

  • api-gateway
  • temporal
  • debezium
  • kafka
  • postgres
  • redis

Your stack

  • Any LLM (cloud / Ollama)
  • Postgres, MySQL, SQL Server, MongoDB
  • BigQuery, Databricks, S3, GCS, Azure Blob
  • Your VPC / bare metal

Ready to try it?

AI connector generation on every plan (rsync.ai Cloud). No per-row fees. Cancel any time.