Prompt
Can you help me choose a time-series database for millions of points per second and long retention?
Latest observation
Yes. For millions of points/sec + long retention, the best choice depends less on “time-series” as a label and more on your workload shape:
- write-heavy ingestion
- query style (dashboards vs ad hoc SQL vs alerts)
- retention horizon (months vs years)
- cardinality (number of unique series/tags)
- operational burden you can tolerate
Quick recommendation
If you want the shortest answer:
- Best open-source / self-hosted for very high ingest + long retention: ClickHouse
- Best managed / easiest at scale for metrics: BigQuery/Redshift/Snowflake are usually too slow/expensive for this use case, so not ideal
- Best purpose-built TSDBs: VictoriaMetrics or Cortex/Mimir for metrics, TimescaleDB if scale is moderate
- Best for extremely high ingest + analytics: ClickHouse often wins in practice
- Best for operational simplicity in a metrics stack: Prometheus + remote storage (Mimir/VictoriaMetrics)
Strong candidates by use case
1) ClickHouse
Best when: you need massive ingest, long retention, and fast analytical queries.
Pros
- Handles very high write rates
- Excellent compression
- Great for long retention at lower storage cost
- Powerful SQL analytics
- Good for mixing metrics with logs/events
Cons
- Not a “native metrics TSDB” in the Prometheus sense
- Requires good schema design and partitioning
- Real-time upserts/delete patterns are not its strength
Choose it if:
You want to store billions/trillions of time-stamped rows and run complex queries.
2) VictoriaMetrics
Best when: you want a TSDB specifically for metrics, with simple ops and high ingest.
Pros
- Very efficient on storage
- Handles high ingestion well
- Good long-term retention
- Compatible with Prometheus ecosystem
- Simpler than many distributed systems
Cons
- Primarily metrics-focused
- Less flexible than SQL-first systems
- Some advanced analytics use cases are limited
Choose it if:
Your data is mostly Prometheus-style metrics and you want strong performance with low operational overhead.
3) Grafana Mimir / Cortex
Best when: you need horizontally scalable Prometheus metrics storage.
Pros
- Designed for distributed metrics at scale
- Works well in Kubernetes/cloud-native setups
- Good for multi-tenant setups
- Integrates with Prometheus and Grafana
Cons
- Operationally complex
- More moving parts
- Usually more about metrics than arbitrary time-series analytics
Choose it if:
You’re already in the Prometheus ecosystem and need multi-tenant distributed scale.
4) TimescaleDB
Best when: you want SQL and Postgres compatibility, and scale is substantial but not extreme.
Pros
- Very familiar SQL/Postgres experience
- Great for hybrid relational + time-series data
- Good tooling and ecosystem
- Easier to adopt for teams already using Postgres
Cons
- At “millions of points/sec” you may hit limits unless carefully engineered and scaled
- Long retention at huge scale can get expensive
- Not usually the top pick for the highest ingest tiers
Choose it if:
You need SQL, joins, and relational integration more than raw TSDB throughput.
5) InfluxDB
Best when: you want a specialized TSDB with decent ergonomics.
Pros
- Purpose-built for time-series
- Easy to get started
- Good ecosystem
Cons
- At very large scale, cost/ops/performance tradeoffs can be less attractive
- Long retention and very high ingest can become challenging depending on edition and architecture
Choose it if:
Your scale is serious but not “absolute extreme,” and you value usability.
My practical ranking for your requirement
For millions of points/sec + long retention:
- ClickHouse — best overall if you can design schemas well
- VictoriaMetrics — best if data is metrics-centric
- Mimir/Cortex — best if you need distributed Prometheus compatibility
- TimescaleDB — best if SQL/Postgres integration matters
- InfluxDB — good, but less compelling at the very high end
What matters most in the decision
Choose ClickHouse if:
- You need long retention cheaply
- You need ad hoc analytics and aggregations
- Your data includes metrics + events + logs
- You can manage partitioning, sampling, and schema design
Choose VictoriaMetrics/Mimir if:
- The data is mostly metrics
- You already use Prometheus/Grafana
- You want efficient storage and strong ingest
- You want a more “TSDB-native” solution
Choose TimescaleDB if:
- You need SQL joins and relational modeling
- Your ingest is high but not the absolute upper limit
- Your team already knows Postgres well
A rough “don’t pick this unless…” guide
- Pick ClickHouse unless your data is strictly Prometheus-style metrics and you want a purpose-built TSDB.
- Pick VictoriaMetrics unless you need heavier analytics or non-metrics use cases.
- Pick TimescaleDB unless you need relational features and your ingest is manageable.
- Avoid using a generic OLTP DB like Postgres/MySQL alone for millions/sec + long retention.
Architecture tip for long retention
At this scale, many teams use:
- hot storage for recent data
- cold object storage for older data
- rollups/downsampling for older windows
This matters as much as the database choice.
If you want one direct answer
- For broadest recommendation: ClickHouse
- For metrics-only Prometheus-compatible storage: VictoriaMetrics
- For SQL + Postgres-style convenience: TimescaleDB
If you want, I can give you a decision matrix based on:
- metrics vs events/logs,
- self-hosted vs managed,
- retention length, and
- query patterns.