\n```\n\nThis is often the cleanest scraping approach because the data is already structured.\n\n---\n\n## 5. Build a multi-source pipeline\nFor multiple sites, a practical setup is:\n\n1. **Discovery layer** \n Keep a list of target sites and job pages.\n\n2. **Collection layer** \n - API clients for sites with APIs \n - Scrapers for HTML pages \n - Browser automation for dynamic sites\n\n3. **Normalization layer** \n Convert all sources into one schema:\n - `title`\n - `company`\n - `location`\n - `description`\n - `posted_at`\n - `source`\n - `source_url`\n\n4. **Deduplication layer** \n Compare by:\n - job title\n - company\n - location\n - source URL\n - text similarity\n\n5. **Storage layer** \n Use:\n - PostgreSQL for structured storage\n - Elasticsearch/OpenSearch for search\n - CSV/Parquet for exports\n\n6. **Scheduler**\n Run daily/hourly with:\n - cron\n - Airflow\n - Prefect\n - Celery/queues for scale\n\n---\n\n## 6. Respect legal and technical limits\nBefore collecting:\n- Check robots.txt\n- Read site terms of service\n- Use rate limiting\n- Identify your scraper with a user-agent if appropriate\n- Avoid aggressive crawling\n- Prefer official feeds/APIs\n\n---\n\n## Example approach in Python\nIf you want a simple start:\n\n- Use `requests` to fetch pages\n- Parse HTML with `BeautifulSoup`\n- Look for JSON-LD job metadata first\n- Fall back to CSS selectors\n- Save results to a database or CSV\n\n---\n\n## Recommended strategy\nIf your goal is practical and maintainable:\n\n- **Start with APIs and structured feeds**\n- **Use JSON-LD extraction wherever possible**\n- **Use scraping only for missing sources**\n- **Normalize everything into one schema**\n- **Deduplicate and schedule updates**\n\nIf you want, I can also help you with:\n1. a **Python scraper template** for multiple job sites, \n2. a **database schema** for storing listings, or \n3. a **site-by-site extraction strategy**."}}}

Prompt

How do I collect job listings from multiple websites?

Technology / API Platforms1 observationLast seen Oct 4, 2026

Latest observation

Oct 4, 2026OpenAI APIWeb search: off

There are a few common ways to collect job listings from multiple websites, depending on your scale, budget, and how often you need updates.

1. Use official APIs when available

Many job boards and company sites expose APIs or feeds.

Examples:

  • LinkedIn and Indeed have limited or restricted access
  • Greenhouse, Lever, Workday, Ashby, and SmartRecruiters often provide structured job data or endpoints
  • Some sites offer RSS feeds or XML job feeds

Pros:

  • Reliable
  • Structured data
  • Less likely to break

Cons:

  • Not available everywhere
  • May require authentication or approvals

2. Scrape the websites

If no API exists, you can build a scraper that extracts job data from pages.

Typical stack:

  • Python
  • requests or httpx for fetching pages
  • BeautifulSoup or lxml for parsing HTML
  • Playwright or Selenium for JavaScript-heavy sites
  • pandas or a database for storing results

Data fields to collect:

  • Job title
  • Company
  • Location
  • Salary
  • Job URL
  • Posting date
  • Description
  • Employment type
  • Source website

Pros:

  • Flexible
  • Works for many sites

Cons:

  • Breaks when site layout changes
  • Needs careful rate limiting
  • Must respect robots.txt and site terms

3. Use aggregators or data providers

There are services that already collect job postings and provide search APIs or exports.

Examples:

  • Adzuna
  • SerpApi
  • RapidAPI job search endpoints
  • Data vendors in recruiting space

Pros:

  • Fastest to implement
  • Less maintenance

Cons:

  • Cost
  • Less control
  • Coverage may vary

4. Use structured data embedded in pages

Many job pages include schema.org JobPosting JSON-LD.

You can parse:

<script type="application/ld+json">
{
  "@context": "https://schema.org",
  "@type": "JobPosting",
  "title": "Software Engineer",
  "hiringOrganization": {"name": "Acme"},
  "jobLocation": {"address": {"addressLocality": "New York"}}
}
</script>

This is often the cleanest scraping approach because the data is already structured.


5. Build a multi-source pipeline

For multiple sites, a practical setup is:

  1. Discovery layer
    Keep a list of target sites and job pages.

  2. Collection layer

    • API clients for sites with APIs
    • Scrapers for HTML pages
    • Browser automation for dynamic sites
  3. Normalization layer
    Convert all sources into one schema:

    • title
    • company
    • location
    • description
    • posted_at
    • source
    • source_url
  4. Deduplication layer
    Compare by:

    • job title
    • company
    • location
    • source URL
    • text similarity
  5. Storage layer
    Use:

    • PostgreSQL for structured storage
    • Elasticsearch/OpenSearch for search
    • CSV/Parquet for exports
  6. Scheduler Run daily/hourly with:

    • cron
    • Airflow
    • Prefect
    • Celery/queues for scale

6. Respect legal and technical limits

Before collecting:

  • Check robots.txt
  • Read site terms of service
  • Use rate limiting
  • Identify your scraper with a user-agent if appropriate
  • Avoid aggressive crawling
  • Prefer official feeds/APIs

Example approach in Python

If you want a simple start:

  • Use requests to fetch pages
  • Parse HTML with BeautifulSoup
  • Look for JSON-LD job metadata first
  • Fall back to CSS selectors
  • Save results to a database or CSV

Recommended strategy

If your goal is practical and maintainable:

  • Start with APIs and structured feeds
  • Use JSON-LD extraction wherever possible
  • Use scraping only for missing sources
  • Normalize everything into one schema
  • Deduplicate and schedule updates

If you want, I can also help you with:

  1. a Python scraper template for multiple job sites,
  2. a database schema for storing listings, or
  3. a site-by-site extraction strategy.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.