Prompt
I'm building an internal research tool that turns web pages into datasets, what stack should I use?
Latest observation
Oct 4, 2026
OpenAI APIWeb search: off
A good stack for a “web pages → structured dataset” internal research tool usually has 5 parts:
- Crawler / fetcher
- Page rendering
- Extraction / normalization
- Storage
- Workflow / review UI
Recommended stack
1) Backend API
- Python + FastAPI
- Why: strong ecosystem for scraping, NLP, data processing, and easy API development.
2) Crawling and fetching
- Scrapy for large-scale crawling
- httpx / aiohttp for custom fetching jobs
- Add Playwright for JavaScript-heavy sites and login flows
A common pattern:
- Use plain HTTP first
- Fall back to Playwright only when needed
- Cache responses to avoid refetching
3) Extraction
- BeautifulSoup / lxml for HTML parsing
- trafilatura or readability-lxml for article-like pages
- Custom rules + heuristics for structured fields
- Optional: LLM-based extraction for messy pages
If your pages vary a lot, use:
- Pydantic models to enforce schema
- An extraction pipeline that can combine:
- CSS/XPath selectors
- regex
- DOM heuristics
- LLM fallback
4) Storage
Use separate stores for different needs:
- PostgreSQL: metadata, job state, extracted records
- S3-compatible object storage: raw HTML, screenshots, PDFs
- Redis: queue, caching, short-lived state
If you need search:
- OpenSearch / Elasticsearch for full-text search over pages and records
If you need analytics:
- DuckDB for local/interactive analysis
- or BigQuery/Snowflake if datasets get large
5) Task orchestration
- Celery + Redis/RabbitMQ for background jobs
- Or RQ if you want simpler setup
- If workflows become complex: Temporal or Prefect
6) Review / labeling UI
- React + Next.js frontend
- Include:
- dataset preview
- extraction rule editor
- diff between raw page and extracted data
- manual correction interface
- audit trail
7) Authentication / access control
- If internal: Auth0 / Okta / Microsoft Entra ID
- Or integrate with your company SSO
- Add role-based access control for:
- who can crawl
- who can edit schemas
- who can export data
A practical “default” stack
If you want the safest starting point:
- Frontend: Next.js
- API: FastAPI
- Crawler: Scrapy + Playwright
- Parsing: lxml + BeautifulSoup + trafilatura
- Jobs: Celery + Redis
- DB: PostgreSQL
- Raw storage: S3
- Search: OpenSearch
- Deploy: Docker + Kubernetes or ECS
- Observability: Prometheus + Grafana + Sentry
If you want a simpler MVP
Start with:
- Python
- FastAPI
- Playwright
- PostgreSQL
- S3
- Redis
- Next.js
This is enough for:
- fetching pages
- extracting into tables
- storing raw snapshots
- reviewing and exporting datasets
Key design advice
Make the pipeline modular
You’ll want clear stages:
- URL ingestion
- Fetch/render
- Parse/extract
- Validate
- Store
- Review/export
Keep raw and derived data separate
Always store:
- raw HTML
- rendered DOM snapshot if relevant
- extracted structured output
- extraction version / schema version
This helps with reprocessing later.
Version everything
Track:
- extraction rules
- schema
- site-specific configs
- crawler version
Web pages change constantly, so versioning matters a lot.
Expect failure modes
Plan for:
- rate limiting
- CAPTCHAs
- login/session expiration
- layout changes
- duplicate pages
- encoding issues
When to use an LLM
Use an LLM for:
- messy or semi-structured pages
- mapping content to a target schema
- classifying page types
Don’t rely on it as the only extractor for everything. Best practice is:
- deterministic extraction first
- LLM only as fallback or assistive layer
- validate outputs with schemas
My recommendation by team size
Solo / small team
- FastAPI
- Playwright
- PostgreSQL
- Redis
- Next.js
- S3
Medium team / production
- Add Scrapy
- Add Celery/Temporal
- Add OpenSearch
- Add strong observability and schema versioning
Large scale / enterprise
- Temporal
- Kubernetes
- OpenSearch
- Data warehouse
- Strong RBAC / audit logs / secrets management
If you want, I can also give you:
- a concrete architecture diagram,
- a minimal MVP stack, or
- a production-grade stack with exact services and deployment choices.