Skip to main content
AWS S3 integration

AWS S3 data pipelines,
described in plain English.

Land PostgreSQL, MySQL, and SaaS data in S3 as Parquet, JSON, or CSV — CDC streaming or scheduled snapshots — and move raw files byte-for-byte between S3, GCS, and Azure. Live on rsync.ai Cloud, no per-row fees.

TL;DR

rsync.ai writes to AWS S3 two ways: structured output (Parquet from Postgres, MySQL, SQL Server or MongoDB CDC and snapshots, in Hive-style dt= date folders) and byte-identical blob passthrough (copy any object between S3, GCS, and Azure). PII rules apply before data lands. Works with S3-compatible stores like MinIO and R2.

  • Parquet, JSON, or CSV — Hive-style dt= date folders
  • CDC streaming or scheduled snapshots — or both
  • IAM role or access-key auth — MinIO & R2 supported
  • Blob passthrough between S3, GCS, and Azure Blob

What the S3 connector does

Structured exports for analytics, and raw blob passthrough for everything else.

Structured exports

Postgres, MySQL, SQL Server, and MongoDB tables to Parquet in dt= date folders, with a manifest per load.

Blob passthrough

Copy any object byte-for-byte between S3, GCS, and Azure — SHA-256 verified.

PII-safe

Mask or hash sensitive columns before a single byte lands in your bucket.

Parquet, JSON, or CSV output with gzip, Snappy, LZ4 or zstd compressionHive-style dt= date foldersIAM role or access-key authPII masking before the S3 writeResumable multi-part uploadsSHA-256 checksums on blob passthroughBlob passthrough: S3 ↔ GCS ↔ AzureNo per-row or per-MAR pricing

rsync.ai vs. Fivetran, Airbyte, custom scripts for S3

What you give up — and gain — choosing rsync.ai for pipelines into AWS S3.

Featurersync.aiyouFivetranAirbyteCustom scripts
Plain-English pipeline setup
CDC streaming to S3 (Postgres, MySQL, SQL Server, and MongoDB)
Parquet output with a load manifest
Blob passthrough (S3 ↔ GCS ↔ Azure)
PII masking before write
No per-row / per-MAR pricing
Resumable snapshots (no restart on failure)

AWS S3 pipelines — frequently asked

What can rsync.ai write to AWS S3?

Two things. First, relational and SaaS data as structured files — PostgreSQL, MySQL, SQL Server, and MongoDB tables (via CDC or snapshot) land as Parquet, other sources can write Parquet, JSON, or CSV, and each batch load gets a manifest. Second, raw files via blob passthrough — copy any object byte-for-byte from GCS or Azure Blob into S3 without re-encoding.

Is AWS S3 a source or a destination?

Both. S3 is most commonly a destination for data-lake ingestion and compliance archiving, but rsync.ai can also read objects from S3 and move them to another store (GCS, Azure Blob, or a different S3 bucket) using byte-identical blob passthrough. Blob → relational database is intentionally rejected — a raw binary can't be written to a table row without parsing.

How does S3 path partitioning work?

Database tables land in a Hive-style date layout: s3://bucket/<prefix>/<pipeline prefix>/<database>/[<schema>/]<table>/dt=YYYY-MM-DD/. Initial and batch loads are written as LOAD00000001.parquet files and CDC changes as timestamped Parquet files in the same folder, in files of up to about 32 MB or every 60 seconds, whichever comes first. The dt= folders work with Athena partition projection and AWS Glue crawlers, and manifest and _SUCCESS markers sit in a separate _rsync/ folder so they never mix with table data. You choose the pipeline prefix when you pick the tables and approve the layout before anything moves.

IAM roles or access keys — which should I use?

IAM roles are preferred. Give rsync.ai an IAM role ARN to assume (role_arn); it generates an external ID for the trust policy, so no long-lived keys are stored. Otherwise use access keys for a scoped IAM user with s3:PutObject, s3:GetObject, and s3:ListBucket on your specific bucket and prefix. An instance profile or task role applies only when you self-host rsync.ai on EC2, ECS, or EKS.

Does rsync.ai support MinIO and other S3-compatible stores?

Yes. Set a custom endpoint (endpoint_url) to point at an S3-compatible store such as MinIO or Cloudflare R2. The same output formats, folder layout, and PII rules apply.

Do I have to deploy anything to use the S3 connector?

Either. rsync.ai Cloud is live at app.rsync.ai — sign up free and build an S3 pipeline in minutes, nothing to provision. Or run the whole stack inside your own VPC: self-hosting is available now, source-available under the rsync.ai Source-Available License.