Skip to main content
Azure Blob Storage integration

Azure Blob Storage pipelines,
described in plain English.

Land PostgreSQL, MySQL, and SaaS data in Azure Blob as Parquet, JSON, or CSV — CDC streaming or scheduled snapshots — and move raw files byte-for-byte between Azure, S3, and GCS. Live on rsync.ai Cloud, no per-row fees.

TL;DR

rsync.ai writes to Azure Blob Storage two ways: structured output (Parquet from Postgres, MySQL, SQL Server or MongoDB CDC and snapshots, in Hive-style dt= date folders for Synapse and Fabric) and byte-identical blob passthrough (copy any object between Azure, S3, and GCS). Authenticate with an access key, SAS token, connection string, or service principal; PII rules apply before data lands.

  • Parquet, JSON, or CSV — Hive-style partitioning
  • CDC streaming or scheduled snapshots — or both
  • Access key, SAS token, or service-principal auth
  • Blob passthrough between Azure Blob, S3, and GCS

What the Azure Blob connector does

Structured exports for analytics, and raw blob passthrough for everything else.

Structured exports

Postgres, MySQL, SQL Server, and MongoDB tables to Parquet in dt= date folders, with a manifest per load.

Blob passthrough

Copy any object byte-for-byte between Azure, S3, and GCS — SHA-256 verified.

PII-safe

Mask or hash sensitive columns before a single byte lands in your container.

Parquet, JSON, or CSV outputHive-style dt= date foldersAccess key, SAS token, or service-principal authPII masking before the Azure writeResumable block-blob uploadsSHA-256 checksums on blob passthroughBlob passthrough: Azure ↔ S3 ↔ GCSNo per-row or per-MAR pricing

rsync.ai vs. Fivetran, Airbyte, custom scripts for Azure Blob

What you give up — and gain — choosing rsync.ai for pipelines into Azure Blob Storage.

Featurersync.aiyouFivetranAirbyteCustom scripts
Plain-English pipeline setup
CDC streaming to Azure Blob (Postgres, MySQL, SQL Server, and MongoDB)
Parquet output with a load manifest
Blob passthrough (Azure ↔ S3 ↔ GCS)
PII masking before write
No per-row / per-MAR pricing
Resumable snapshots (no restart on failure)

Azure Blob Storage pipelines — frequently asked

What can rsync.ai write to Azure Blob Storage?

Structured data and raw files. PostgreSQL, MySQL, SQL Server, and MongoDB tables (via CDC or snapshot) land as Parquet, other sources can write Parquet, JSON, or CSV, and each batch load gets a manifest — ready for Azure Synapse, Microsoft Fabric, or Databricks. Separately, blob passthrough copies any object byte-for-byte from S3 or GCS into Azure Blob without re-encoding.

How does rsync.ai authenticate to Azure Blob?

Use a storage-account access key, a scoped SAS token, a connection string, or a service principal (tenant ID, client ID, and secret). You point rsync.ai at the storage account and container and choose the path prefix.

Can Synapse and Microsoft Fabric read the files?

Yes. Database tables land as Parquet in Hive-style date folders (<table>/dt=2026-06-24/), a layout Synapse serverless SQL pools and Microsoft Fabric can query. Manifest and _SUCCESS markers sit in a separate _rsync/ folder, so they never mix with table data.

Is Azure Blob a source or a destination?

Both. Azure Blob is typically a destination for data-lake and archival workloads, but rsync.ai can also read objects from Azure Blob and move them to another store (S3, GCS, or another container) with byte-identical blob passthrough. Blob → relational database is intentionally rejected — a raw binary can't be written to a table row without parsing.

Do I have to deploy anything to use the Azure Blob connector?

Either. rsync.ai Cloud is live at app.rsync.ai — sign up free and build an Azure Blob pipeline in minutes, nothing to provision. Or run the whole stack inside your own VPC: self-hosting is available now, source-available under the rsync.ai Source-Available License.