Prompt

Why are my BigQuery costs suddenly so high?

Technology / Data Infrastructure1 observationLast seen Oct 8, 2026

Latest observation

Oct 8, 2026OpenAI APIWeb search: off

BigQuery bills usually spike for a few common reasons. The fastest way to find the cause is to look for what changed recently in queries, storage, or data movement.

Most common causes

  1. More bytes scanned by queries

    • Someone ran a query without partition filters.
    • A dashboard started refreshing more often.
    • A query changed from scanning a few partitions to the whole table.
    • SELECT * on large tables.
    • Repeated ad hoc queries on large datasets.
  2. New or increased scheduled jobs

    • Scheduled queries, Dataform, dbt, or Airflow jobs may have multiplied.
    • A loop or retry issue can run the same job many times.
  3. Storage costs increased

    • Large table growth.
    • Increased time travel / fail-safe related storage.
    • Long-term storage not applying because data is being updated frequently.
  4. Streaming ingestion / CDC / materialized view refreshes

    • Streaming inserts or frequent updates can raise costs indirectly.
    • Materialized views can refresh more than expected.
  5. Cross-region or egress charges

    • Queries joining data across regions.
    • Exporting data out of BigQuery or to another cloud/service.
  6. BI tool behavior

    • Looker, Tableau, Power BI, etc. may run lots of queries behind the scenes.
    • Cache may be disabled or invalidated.
  7. Reservations / slots misconfiguration

    • On-demand query costs changed because reservations were removed or reduced.
    • Flat-rate slots might have been overprovisioned.

How to diagnose quickly

1) Check the Billing page

In Google Cloud Billing:

  • Go to Billing → Reports
  • Group by:
    • Service
    • SKU
    • Project
    • Region
  • Look for whether the spike is:
    • Query processing
    • Storage
    • Streaming
    • Data transfer

2) Inspect BigQuery job history

In BigQuery:

  • Go to Query history / Job history
  • Sort by:
    • Bytes processed
    • Start time
  • Find the largest jobs and see:
    • Who ran them
    • Whether they are repeated
    • Which tables were scanned

3) Review query patterns

Look for:

  • Missing partition filters
  • SELECT *
  • Unintended full table scans
  • Cross joins
  • Repeated dashboard queries

4) Check table growth and partitions

  • Which tables grew recently?
  • Are partitioned tables being queried without _PARTITIONTIME or partition column filters?
  • Are clustering keys being used effectively?

5) Check for retries or automation loops

  • ETL jobs repeatedly failing and retrying
  • Scheduled jobs accidentally duplicated
  • Orchestrators launching extra runs

Useful BigQuery metrics/queries

If you have access to audit/job logs, look at recent expensive queries and totals by user/project. Example approach:

  • Examine INFORMATION_SCHEMA.JOBS* views
  • Summarize total_bytes_processed, total_bytes_billed, and user_email
  • Identify the top offenders in the last 24h / 7d

Example:

SELECT
  user_email,
  COUNT(*) AS jobs,
  SUM(total_bytes_billed) / POW(1024, 4) AS billed_TiB
FROM `region-us`.INFORMATION_SCHEMA.JOBS_BY_PROJECT
WHERE creation_time >= TIMESTAMP_SUB(CURRENT_TIMESTAMP(), INTERVAL 7 DAY)
  AND job_type = 'QUERY'
GROUP BY user_email
ORDER BY billed_TiB DESC;

How to reduce costs

  • Partition tables and always filter on partition columns
  • Avoid SELECT *
  • Use clustering for common filters
  • Use materialized views or summary tables for frequent dashboard queries
  • Turn on query cost controls / maximum bytes billed
  • Cache results where possible
  • Review BI tool query frequency
  • Use reservations only if they match steady demand

If you want, I can help you pinpoint it

If you share any of these, I can help narrow it down:

  • the billing SKU breakdown
  • a few expensive queries
  • your top tables
  • whether you use Looker/Tableau/dbt/Airflow
  • whether the spike is in query, storage, or data transfer costs

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.