Skip to main content
Databricks · lakehouse destination

Load your databases into Databricks, live.

Describe the sync in plain English, approve it once, and rsync.ai streams every change from SQL Server, Postgres, MySQL, or MongoDB into Databricks — as Delta tables, registered in Unity Catalog.

Real-time CDCYou approve every syncNo per-row pricing
Delta tables with incremental MERGE You approve before anything moves
Sources

Bring every source into the lakehouse.

Point rsync.ai at your databases and apps. It builds and runs the pipeline into Databricks — you approve it before anything moves.

Databases

Log-based CDC keeps operational databases mirrored as Delta tables continuously, not in nightly batches. Deletes are applied as soft-deletes.

SQL ServerPostgreSQLMySQLMongoDB

SaaS & APIs

Pull business objects from the tools your team runs on, on a schedule you set and approve.

StripeGitHub

Files & storage

Ingest Parquet, JSON, or CSV from object storage into Delta tables.

Amazon S3GCSAzure Blob
What lands

Delta tables, ready for SQL and Spark.

Delta tables
Unity Catalog
Incremental MERGE
Soft deletes
Schema evolution
catalog.schema.table
TIMESTAMP / DATE
Documents → JSON
How it works

Four steps. Nothing runs until you approve.

01

Describe it

Say which source and tables to load into Databricks in plain English — no connector config.

02

Review the plan

rsync.ai discovers the schema, proposes the Unity Catalog path and Delta types, and scans columns for likely personal data before a row moves.

03

Approve

Nothing runs until you say yes. Later schema changes are applied and recorded; only a dropped table waits for you.

04

Stream

CDC keeps Delta tables up to date; incremental MERGE loads send only the rows that changed.

After it lands

Then work with the data without leaving rsync.ai

Synced tables become something you can query, schedule and share in the same workspace.

SQL Explorer

Run SQL against your connections, with personal columns masked in previews; writes are gated by role and every write is audited.

SQL Explorer docs→

Models

Turn a query into a table that rebuilds on a schedule or when its inputs refresh, with version history and data checks.

Models docs→

Charts and Dashboards

Draw a model or saved query as a chart, then put charts on one live dashboard.

Charts→ · Dashboards→

Workflows and Ask rsync.ai

Schedule reports and data alerts to Slack or email. Describe one in plain English and Ask rsync.ai drafts it.

Scheduled reports→ · Ask rsync.ai→

Compare

Why teams load Databricks this way.

What you care aboutNotebooks + jobsFivetranHand-built ETLrsync.ai
Set up without an engineerSpark requiredConnector formsNoPlain English
Real-time, log-based CDCBatchYesVariesYes
You approve before it runsn/aNoNoYes
Cost as volume growsCluster timePer MARYour timePer GB, not per row
Run in your own VPCYesNoVariesYes

Straight about status: the Databricks destination and the SQL Server, MongoDB, Postgres, and MySQL sources are live in production. The Stripe and GitHub sources are in preview. Self-hosting in your own VPC is available now — source-available under the rsync.ai Source-Available License — or run rsync.ai as a managed cloud.

FAQ

Loading Databricks — common questions

How fresh is the data in Databricks?
You choose the schedule, from a few times a day down to near real time. rsync.ai reads each source's change log (CDC) and applies an incremental MERGE into the Delta table, so it writes only the rows that changed — and a failed run resumes where it stopped.
Do tables register in Unity Catalog?
Yes. rsync.ai lands Delta tables at the catalog.schema.table path you approve and proposes the types in the review step. Nothing is created until you approve it, so there are no silent conversions.
Do I need to write Spark or configure a connector?
No. Describe the sync in plain English. rsync.ai discovers the source schema, maps it to Delta types, scans columns for likely personal data, and shows you the plan. Nothing runs until you approve it.
How are you priced?
On data volume (GB moved), not per row or Monthly Active Rows. Each plan includes a monthly GB allowance (10 GB on the free month), and pipelines pause at your cap instead of running up charges.
Who controls what moves?
You do. No pipeline is created or started until you confirm the plan, PII rules apply before data lands, and later schema changes are applied and recorded (only a dropped table waits for a person). Self-hosting entirely in your own VPC is available now for teams that need data to never leave their network.

Get your data into Databricks.

Connect a source free and land your first live Delta tables in Databricks today — registered in Unity Catalog.

Live on rsync.ai Cloud today · self-hosted (rsync.ai Source-Available License)