Skip to main content
Google Cloud Storage integration

Google Cloud Storage pipelines,
described in plain English.

Land PostgreSQL, MySQL, and SaaS data in GCS as Parquet, JSON, or CSV — CDC streaming or scheduled snapshots — and move raw files byte-for-byte between GCS, S3, and Azure. Live on rsync.ai Cloud, no per-row fees.

TL;DR

rsync.ai writes to Google Cloud Storage two ways: structured output (Parquet from Postgres, MySQL, SQL Server or MongoDB CDC and snapshots, in Hive-style dt= date folders for BigQuery external tables) and byte-identical blob passthrough (copy any object between GCS, S3, and Azure). Authenticate with a service-account key; PII rules apply before data lands.

  • Parquet, JSON, or CSV — Hive-style dt= folders for BigQuery
  • CDC streaming or scheduled snapshots — or both
  • Service-account key auth
  • Blob passthrough between GCS, S3, and Azure Blob

What the GCS connector does

Structured exports for analytics, and raw blob passthrough for everything else.

Structured exports

Postgres, MySQL, SQL Server, and MongoDB tables to Parquet in dt= date folders, with a manifest per load.

Blob passthrough

Copy any object byte-for-byte between GCS, S3, and Azure — SHA-256 verified.

PII-safe

Mask or hash sensitive columns before a single byte lands in your bucket.

Parquet, JSON, or CSV outputHive-style dt= date foldersService-account key authPII masking before the GCS writeResumable uploadsSHA-256 checksums on blob passthroughBlob passthrough: GCS ↔ S3 ↔ AzureNo per-row or per-MAR pricing

rsync.ai vs. Fivetran, Airbyte, custom scripts for GCS

What you give up — and gain — choosing rsync.ai for pipelines into Google Cloud Storage.

Featurersync.aiyouFivetranAirbyteCustom scripts
Plain-English pipeline setup
CDC streaming to GCS (Postgres, MySQL, SQL Server, and MongoDB)
Parquet output with a load manifest
Blob passthrough (GCS ↔ S3 ↔ Azure)
PII masking before write
No per-row / per-MAR pricing
Resumable snapshots (no restart on failure)

Google Cloud Storage pipelines — frequently asked

What can rsync.ai write to Google Cloud Storage?

Structured data and raw files. PostgreSQL, MySQL, SQL Server, and MongoDB tables (via CDC or snapshot) land as Parquet, other sources can write Parquet, JSON, or CSV, and each batch load gets a manifest — ready for BigQuery external tables, Dataproc, or DuckDB. Separately, blob passthrough copies any object byte-for-byte from S3 or Azure Blob into GCS without re-encoding.

How does rsync.ai authenticate to GCS?

Use a Google Cloud service-account key (JSON) with the Storage Object Admin role on your bucket. The bucket, path prefix, and object lifecycle are yours to configure.

Can I use GCS files as BigQuery external tables?

Yes. Database tables land as Parquet in Hive-style date folders (<table>/dt=2026-06-24/) that BigQuery external tables with hive partitioning read directly. Manifest and _SUCCESS markers sit in a separate _rsync/ folder, so an external table over <table>/* reads only Parquet.

Is GCS a source or a destination?

Both. GCS is typically a destination for data-lake and archival workloads, but rsync.ai can also read objects from GCS and move them to another store (S3, Azure Blob, or another GCS bucket) with byte-identical blob passthrough. Blob → relational database is intentionally rejected — a raw binary can't be written to a table row without parsing.

Do I have to deploy anything to use the GCS connector?

No. rsync.ai Cloud is live at app.rsync.ai — sign up free and build a GCS pipeline in minutes, nothing to provision. If you'd rather run the whole stack inside your own VPC, self-hosting (source-available under the rsync.ai Source-Available License) is available now.