Prompt
How do I build a time-series system that supports updates?
Latest observation
To build a time-series system that supports updates, you need to decide first what “updates” mean, because time-series data is usually append-heavy and immutable by default. Updates can mean:
- Late arrivals: new data for an old timestamp
- Corrections: replace a previously ingested value
- Deletions: remove bad points
- Recomputations: change derived aggregates after raw data changes
A good design depends on how often updates happen and how fresh your queries need to be.
1. Start with the core model
Store each point with at least:
series_idor tags/dimensionstimestampvalue- optional
version,ingested_at,source_id
Example schema:
series_id STRING
ts TIMESTAMP
value FLOAT
ingested_at TIMESTAMP
version INT
A key idea: treat timestamp as the event time, and ingested_at as the write time.
2. Choose an update strategy
Option A: Upsert by primary key
Use (series_id, ts) as the logical key and overwrite the row on update.
Best when:
- each timestamp has only one current value
- updates are common
- you don’t need history
Pros:
- simple
- fast reads of “latest” state
Cons:
- hard to audit changes
- expensive if storage engine rewrites many rows
Option B: Immutable append + versioning
Instead of overwriting, append a new record for the same (series_id, ts) with a higher version or later ingested_at.
Query the latest record for each key.
Best when:
- you need auditability
- updates happen but are not extremely frequent
- you want simpler ingestion
Pros:
- preserves history
- ingestion is append-only
Cons:
- reads are more complex
- compaction/dedup needed
Option C: Delta log + periodic compaction
Write every update to a log, then periodically merge into base storage.
Best when:
- high ingest rate
- frequent updates
- you want efficient writes and eventually consistent reads
Pros:
- scalable
- common in modern lakehouse/LSM designs
Cons:
- more complex
- query semantics may be “eventually updated” unless you read the delta too
3. Use a storage engine that fits updates
Good choices include:
- PostgreSQL + TimescaleDB for simpler systems
- ClickHouse if mostly append-heavy, with some update support
- Apache Druid / Pinot for analytical workloads with streaming ingest
- LSM-based stores like Cassandra, RocksDB-backed systems, or custom KV stores for high write throughput
- Lakehouse tables like Delta Lake / Apache Iceberg / Hudi for upserts and batch compaction
If updates are a first-class feature, avoid pure append-only file layouts unless you also build a compaction/update layer.
4. Design for deduplication
If multiple records can arrive for the same point, define a rule:
- latest
versionwins - latest
ingested_atwins - highest source priority wins
- compare
(source_timestamp, sequence_number)
Example conflict resolution:
same series_id + ts:
choose max(version)
if tie, choose latest ingested_at
This is critical for deterministic behavior.
5. Separate raw data from serving data
A robust system often has:
Raw/event store
- append-only
- keeps every received update
- used for replay and auditing
Serving store
- deduplicated, query-optimized
- contains the latest canonical value per point
This lets you:
- reprocess if business rules change
- support backfills and corrections
- keep query performance fast
6. Handle aggregates carefully
If you maintain rollups like:
- per minute average
- daily max
- moving window metrics
then updates to raw points can invalidate aggregates.
You need one of these approaches:
Recompute on change
Simple, but expensive for large ranges.
Maintain incremental aggregates with correction
Store enough metadata to reverse previous contributions.
Use lazy invalidation
Mark affected windows dirty and recompute asynchronously.
For example, if a point at 2026-10-08 10:03 changes, you may need to refresh:
- the 1m bucket for
10:03 - the 5m bucket containing it
- maybe the daily rollup
7. Partition for efficient writes and updates
Partition by time and possibly series hash:
dateorhour- hash of
series_id
This helps with:
- pruning old data
- parallel ingestion
- limiting update scope
But be careful: if updates frequently target old partitions, you’ll need a storage engine that can update old partitions efficiently or a compaction process.
8. Make ingestion idempotent
Updates often arrive more than once.
Use one of:
- natural key
(series_id, ts, version) - unique event ID
- sequence number per series/source
Then apply “at least once” ingestion safely.
Example:
- if
event_idalready exists, ignore it - if same
(series_id, ts)with larger version, replace it
9. Decide your consistency model
Common models:
Strong consistency
Readers always see the latest update.
Good for:
- operational monitoring
- control systems
Harder at scale.
Eventual consistency
Readers may temporarily see old values.
Good for:
- analytics
- high throughput systems
Usually easier and cheaper.
If updates matter to downstream alerts or decisions, you likely need stronger guarantees for the serving path.
10. A practical architecture
A common pattern:
-
Ingest API
- validates data
- assigns event IDs
- writes to a durable log
-
Raw append store
- immutable history
-
Update processor
- deduplicates by key
- resolves conflicts
- writes latest canonical state
-
Query store
- optimized for time-range queries
- supports latest-value and historical queries
-
Compaction/rebuild jobs
- refresh rollups
- clean stale versions
11. Example data flow
Suppose you get:
{"series_id":"temp.room1","ts":"2026-10-08T10:00:00Z","value":21.4,"version":1}
{"series_id":"temp.room1","ts":"2026-10-08T10:00:00Z","value":21.7,"version":2}
Your system should:
- store both in raw history
- mark version 2 as the current value
- update any affected aggregates if needed
12. Recommended implementation choices
If you want something simple
Use:
- PostgreSQL/TimescaleDB
- table with unique key
(series_id, ts) - upserts via
INSERT ... ON CONFLICT DO UPDATE
If you want very high scale
Use:
- append-only ingest log
- dedup/merge pipeline
- LSM or lakehouse backend
- background compaction
If you need audit + updates
Use:
- immutable raw store
- versioned records
- latest-materialized view
13. Common pitfalls
- Overwriting without history when you later need auditability
- Updating rollups without correcting them
- No deterministic conflict resolution
- Using timestamp alone as identity when duplicates exist
- Not handling late data
- Making queries read from raw append-only data directly without dedup
14. Rule of thumb
If updates are:
- rare: use append-only + periodic merge
- moderate: use versioned upserts
- frequent and latency-sensitive: use a storage engine designed for mutable time-series or a fast KV/LSM serving layer
If you want, I can also give you:
- a PostgreSQL/TimescaleDB schema + upsert examples,
- a Kafka + stream processing architecture, or
- a cloud-scale design for billions of points with updates.