Prompt

Why are my incremental loads creating duplicates?

Technology / Data Infrastructure1 observationLast seen Oct 8, 2026

Latest observation

Oct 8, 2026OpenAI APIWeb search: off

Incremental loads usually create duplicates when the process isn’t truly idempotent or when the “new rows” filter doesn’t reliably exclude already-loaded records. Common causes:

1) No unique key / merge key

If you’re only appending rows, the same record can be loaded again if it appears in a later run.

Fix: Use a stable business key or surrogate key and do an upsert/merge instead of pure append.


2) Watermark based on the wrong column

If you use created_at or another non-update timestamp as the incremental cutoff, rows that are updated later may be re-picked up, or late-arriving rows may be missed and retried.

Fix: Use a reliable watermark such as:

  • updated_at for change capture
  • source CDC/LSN/offset if available
  • a combination of timestamp + tie-breaker key

3) Timestamp precision / boundary issues

If the last loaded timestamp is stored with lower precision than the source, rows at the boundary can be loaded again.

Example:

  • last run saved 2026-10-08 10:15:30
  • source has microseconds, and a record at 10:15:30.123
  • next run reprocesses it because of >= logic

Fix:

  • use > instead of >= where appropriate
  • store full precision
  • add a deterministic tie-breaker

4) Late-arriving data

Records may arrive after the watermark has advanced, so you re-run a window and pick up overlaps. If you’re appending, overlaps become duplicates.

Fix:

  • load with a lookback window and deduplicate in the target
  • or merge on key

5) Re-runs or retries without deduplication

If a job fails after partially loading and then retries, the same batch may be written again.

Fix:

  • make loads idempotent
  • stage then merge
  • track batch/job execution IDs

6) Source system changes existing rows

If the source updates a row but keeps the same key, append-only logic inserts another copy instead of replacing the old one.

Fix: MERGE on the primary/business key.


7) Joins or transformations multiplying rows

Duplicates may not come from the incremental filter itself, but from upstream logic:

  • many-to-many joins
  • exploding arrays
  • non-unique dimension joins

Fix: validate row counts at each step and ensure join keys are unique.


Best practice

For reliable incremental loads:

  1. Extract using a watermark or CDC offset
  2. Land into a staging table
  3. Deduplicate staging if needed
  4. MERGE into target on a unique key
  5. Store last successful checkpoint only after commit

If you want, I can help you debug your specific incremental load logic if you paste:

  • the incremental SQL/query
  • the watermark column
  • target table key
  • whether you append or merge

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.