Prompt
How do I collect job listings from multiple websites?
Latest observation
There are a few common ways to collect job listings from multiple websites, depending on your scale, budget, and how often you need updates.
1. Use official APIs when available
Many job boards and company sites expose APIs or feeds.
Examples:
- LinkedIn and Indeed have limited or restricted access
- Greenhouse, Lever, Workday, Ashby, and SmartRecruiters often provide structured job data or endpoints
- Some sites offer RSS feeds or XML job feeds
Pros:
- Reliable
- Structured data
- Less likely to break
Cons:
- Not available everywhere
- May require authentication or approvals
2. Scrape the websites
If no API exists, you can build a scraper that extracts job data from pages.
Typical stack:
- Python
requestsorhttpxfor fetching pagesBeautifulSouporlxmlfor parsing HTMLPlaywrightorSeleniumfor JavaScript-heavy sitespandasor a database for storing results
Data fields to collect:
- Job title
- Company
- Location
- Salary
- Job URL
- Posting date
- Description
- Employment type
- Source website
Pros:
- Flexible
- Works for many sites
Cons:
- Breaks when site layout changes
- Needs careful rate limiting
- Must respect robots.txt and site terms
3. Use aggregators or data providers
There are services that already collect job postings and provide search APIs or exports.
Examples:
- Adzuna
- SerpApi
- RapidAPI job search endpoints
- Data vendors in recruiting space
Pros:
- Fastest to implement
- Less maintenance
Cons:
- Cost
- Less control
- Coverage may vary
4. Use structured data embedded in pages
Many job pages include schema.org JobPosting JSON-LD.
You can parse:
<script type="application/ld+json">
{
"@context": "https://schema.org",
"@type": "JobPosting",
"title": "Software Engineer",
"hiringOrganization": {"name": "Acme"},
"jobLocation": {"address": {"addressLocality": "New York"}}
}
</script>
This is often the cleanest scraping approach because the data is already structured.
5. Build a multi-source pipeline
For multiple sites, a practical setup is:
-
Discovery layer
Keep a list of target sites and job pages. -
Collection layer
- API clients for sites with APIs
- Scrapers for HTML pages
- Browser automation for dynamic sites
-
Normalization layer
Convert all sources into one schema:titlecompanylocationdescriptionposted_atsourcesource_url
-
Deduplication layer
Compare by:- job title
- company
- location
- source URL
- text similarity
-
Storage layer
Use:- PostgreSQL for structured storage
- Elasticsearch/OpenSearch for search
- CSV/Parquet for exports
-
Scheduler Run daily/hourly with:
- cron
- Airflow
- Prefect
- Celery/queues for scale
6. Respect legal and technical limits
Before collecting:
- Check robots.txt
- Read site terms of service
- Use rate limiting
- Identify your scraper with a user-agent if appropriate
- Avoid aggressive crawling
- Prefer official feeds/APIs
Example approach in Python
If you want a simple start:
- Use
requeststo fetch pages - Parse HTML with
BeautifulSoup - Look for JSON-LD job metadata first
- Fall back to CSS selectors
- Save results to a database or CSV
Recommended strategy
If your goal is practical and maintainable:
- Start with APIs and structured feeds
- Use JSON-LD extraction wherever possible
- Use scraping only for missing sources
- Normalize everything into one schema
- Deduplicate and schedule updates
If you want, I can also help you with:
- a Python scraper template for multiple job sites,
- a database schema for storing listings, or
- a site-by-site extraction strategy.