Skip to main content

How a healthcare data vendor cut BigQuery ETL costs by 75% at petabyte scale

A healthcare data vendor was paying BigQuery over $4,000 a month to run petabyte-scale ETL queries over data in GCS. The same queries now run on ParaQuery for $1,000. Nothing moved, and twenty cycles later the pipeline is bigger than it has ever been.

Cost per monthly cycle
$1,000 $4,000 on BigQuery −75% · same queries, same output
Effective price
<$2/ TiB per TiB scanned, all-in BigQuery on-demand list: $6.25 / TiB
Largest single query
500TB 1.2 trillion rows, bytes scanned another: GROUP BY over 3T+ rows
Production cycles on schedule
20of 20 every month since January 2025 latest cycle was the largest yet

The customer

This customer builds monthly datasets for the healthcare and health-insurance markets through many large SQL queries over data in Google Cloud Storage (GCS). Their data is their core product.

These queries are sizable. The largest one processes 500 TB (bytes read) across 1.2 trillion rows, and another aggregates more than 3 trillion rows. On BigQuery, such queries didn't fit the team's Standard-edition slot allocation (capacity pricing), so the team was forced to split them, with certain queries being split into over 20 parts, maintained manually.

Then the credits ran out

The pipeline ran on BigQuery, and for as long as GCP credits were covering it, nobody had a reason to worry too much. Credits are a fantastic tool for growth. But when the credits ended, the real number arrived hard and fast: even after optimizing their BigQuery pipeline for cost, they faced over $4,000 per ETL run, with data volumes still climbing.

No one chooses an expensive warehouse. They grow into one, and then they have to maintain it.

What we changed

We swapped out the engine.

INPUT

GCS

UNCHANGED

ENGINE

BigQuery ParaQuery

SWAPPED

OUTPUT

GCS

UNCHANGED

Nothing about the customer's data or logic had to change. Their tables were already in GCS. Their pipeline was already Python-driven SQL queries. We simply ran a quick, automated rewrite from GoogleSQL to the equivalent, vendor-neutral Spark SQL, then checked them by hand. We pointed ParaQuery at the same buckets and ran the same queries on our GPU-accelerated clusters, writing results back to the same storage, scaling to almost a hundred GPUs in minutes.

No external table setup, no data migration, no ingestion, and no rewrite of business logic.

The customer's ETL ultimately feeds into downstream systems, including Snowflake. ParaQuery simply replaced BigQuery at the compute layer, maintaining compatibility with downstream consumers, so our customer could cut costs without redesigning the rest of their data stack.

It took less than a month to go from a committed evaluation to the first production cycle.

What they got

The result: 75% lower cost for the core ETL versus their previous BigQuery Standard Edition capacity pricing setup. The same monthly run fell from over $4,000 on BigQuery to $1,000 on ParaQuery. Our customer saved a minimum of $36,000 in 2025 and now has a lower-maintenance pipeline.

Cost per monthly cycle

BigQuery
$4,000
ParaQuery
$1,000

−75%

$3,000 kept per cycle · $36,000 across 2025

Each query could now be run without splitting into tens of fragments, simplifying logic and letting the whole pipeline finish faster. They didn't need over 10,000 BigQuery Enterprise slots at 1.5 times the price, nor did they need to worry about BigQuery's 6-hour query time limit.

ON BIGQUERY

One 500 TB query, split into 20+ pieces to fit slot allocation

ON PARAQUERY

1 query 500 TB · 1.2 trillion rows

Same SQL, run once. No splits to maintain

Twenty+ cycles later

ParaQuery has run this customer's monthly production cycle since January 2025, with data volume still growing across 2026. The most recent ETL was the largest yet, with a single query on over 500 TB across more than 1.2 trillion rows. It was delivered, simply, silently, and on schedule, like every cycle before it.

"We’ve been using ParaQuery to process petabytes of data each month for almost two years and have had a great experience. The service has been a reliable, more scalable and affordable alternative to BigQuery, and the ParaQuery team has always been responsive and helpful whenever we need anything."

CTO, healthcare data vendor (name withheld)


Workload details shared with customer permission. Names withheld at their request.

If your data warehouse bill needs some compaction, we'd like to see it. Get in touch for a savings estimate, a quick pilot, or any questions you may have. Bring a billing export to the call and we can show you what the same workload may cost on ParaQuery.