Prompt
How do I build a search result collection system that is stable?
Latest observation
To build a stable search result collection system, design for predictable retrieval, fault tolerance, idempotency, and observability. A stable system doesn’t just “work”; it keeps working under retries, partial failures, rate limits, and changing result sets.
1) Define “stable”
Decide what stability means for your use case:
- No missed results for a query/time window
- No duplicate results
- Consistent ordering across runs
- Graceful handling of failures
- Stable throughput under load
- Repeatable outputs for the same input
These goals affect your architecture.
2) Use an incremental, checkpointed collector
If you’re pulling search results from an API or scraping engine, do it in small, resumable steps:
- Store a checkpoint per query:
- last page fetched
- last result timestamp
- cursor/token
- last successful run time
- Resume from the checkpoint after crashes or restarts
- Use pagination/cursors instead of relying on page numbers when possible
This prevents redoing work and helps recover from interruptions.
3) Make writes idempotent
Search results often repeat across pages or runs. Your storage layer should handle duplicates safely.
- Create a unique key per result, such as:
- source + result_id
- source + canonical_url
- source + hash(title + url + published_at)
- Use upserts rather than blind inserts
- Keep an ingestion log so retries don’t create duplicates
This is one of the biggest stability improvements.
4) Normalize and deduplicate early
Search data is messy. Normalize before saving:
- Canonicalize URLs
- Normalize whitespace, casing, encodings
- Extract stable identifiers
- Deduplicate across:
- pages
- runs
- queries
- equivalent URLs
Consider a two-layer dedupe:
- Exact match on stable ID/URL
- Near-duplicate detection using title similarity or content hashing if needed
5) Handle rate limits and failures gracefully
Search systems frequently fail due to throttling, network issues, and upstream instability.
Use:
- Exponential backoff with jitter
- Retry limits
- Circuit breakers for repeated upstream failures
- Timeouts on every request
- Concurrency caps to avoid overload
Important: only retry safe/idempotent operations, or ensure retries won’t duplicate records.
6) Separate fetching, processing, and storage
A stable architecture usually has clear stages:
- Fetcher: retrieves search results
- Parser/Normalizer: transforms raw results
- Deduper: removes duplicates
- Storage writer: persists clean records
- Monitor: tracks errors and completeness
This separation makes failures easier to isolate and recover from.
A queue-based pipeline is often best:
- producer fetches pages
- workers process results
- writer stores them with idempotency
7) Store raw and processed data
Keep both:
- Raw response: original API payload or HTML snapshot
- Processed record: normalized fields used by your app
Why:
- Raw data helps debug parser changes
- Processed data is efficient for query and analytics
- You can reprocess later if your schema changes
8) Build for changing result sets
Search results can change between runs:
- new items appear
- items disappear
- ranking changes
- pagination shifts
To stabilize collection:
- Prefer timestamp-based harvesting where possible
- Use sliding windows with overlap
- Re-scan a small recent window to catch late-arriving items
- Accept that ranking is dynamic; don’t assume page 2 stays page 2
For example, collect results from the last 24 hours with a 1–2 hour overlap on each run.
9) Add completeness checks
A stable collector should know whether it likely got everything.
Examples:
- Expected count vs collected count
- Max page reached
- Cursor exhaustion
- Hashing raw pages to detect changes
- Alert if result count drops unexpectedly
If a query usually returns 1,000 results and suddenly returns 12, you want to know.
10) Use observability from day one
Track metrics such as:
- requests per minute
- success/failure rate
- retry count
- latency
- duplicate rate
- results collected per query
- checkpoint lag
- time since last successful run
Log:
- query parameters
- request IDs
- cursor/page number
- response size
- error type
Add alerts for:
- sustained failures
- empty-result anomalies
- backlog growth
- storage write errors
11) Make the system restart-safe
Assume any component can crash.
Design rules:
- Every step can be retried
- Checkpoints are persisted before/after safe boundaries
- Workers are stateless where possible
- Unfinished jobs can be reclaimed
- Partial runs don’t corrupt final output
A common pattern is:
- fetch page
- write raw payload
- parse and upsert results
- persist checkpoint only after successful write
12) Version your schemas and parsers
Search result formats change.
- Version your raw payload schema
- Version your parsed record schema
- Keep parser versions tied to source versions
- Use migrations for downstream consumers
This prevents silent breakage when sources change fields or HTML structure.
13) Use backpressure and queueing
If input rate exceeds processing rate, stability drops.
Add:
- bounded queues
- worker autoscaling
- backpressure to slow fetchers
- dead-letter queues for bad payloads
This avoids runaway memory use and cascading failures.
14) Test for failure modes
Don’t only test the happy path.
Test:
- timeouts
- partial pages
- duplicate pages
- malformed responses
- delayed results
- upstream rate limiting
- restart mid-run
- storage failures
Synthetic tests and replaying real payloads are especially useful.
15) A practical reference design
A stable search collection system often looks like this:
- Scheduler decides which queries to run
- Fetcher workers retrieve pages/cursors with retries
- Raw store saves response snapshots
- Normalizer extracts canonical result objects
- Deduper/upserter writes to DB using unique keys
- Checkpoint store records progress
- Monitoring alerts on anomalies
16) Minimal stability checklist
If you want the shortest actionable list:
- Use cursors/checkpoints
- Make all writes idempotent
- Normalize and dedupe
- Retry with exponential backoff + jitter
- Set timeouts everywhere
- Separate raw and processed data
- Monitor result counts and failures
- Design for restarts and partial progress
If you want, I can also give you:
- a reference architecture diagram,
- a database schema for search result storage, or
- a sample implementation in Python/Node.js.