Prompt

What's the best way to scrape blog articles?

Technology / Data Infrastructure1 observationLast seen Jul 15, 2026

Latest observation

Jul 15, 2026OpenAI APIWeb search: off

The best way is usually not to “scrape” in the broad sense, but to collect article content in a way that’s reliable, polite, and allowed by the site.

Best approach

  1. Check for an RSS/Atom feed

    • Many blogs expose full or partial feeds.
    • This is the cleanest and most stable way to get new posts.
  2. Check for an official API

    • Some platforms (WordPress, Ghost, Medium-like platforms, custom CMSs) provide APIs.
    • APIs are far more reliable than scraping HTML.
  3. If you must scrape HTML, use the site’s structure carefully

    • Fetch the article pages directly.
    • Parse the main content with an HTML parser like:
      • Python: BeautifulSoup, lxml, readability-lxml, trafilatura
      • JavaScript: cheerio, playwright if JS-rendered
    • Avoid brittle selectors that depend on random CSS classes.

Good scraping practices

  • Respect robots.txt and terms of service
  • Throttle requests so you don’t hammer the site
  • Set a clear User-Agent
  • Handle pagination and canonical URLs
  • Deduplicate articles
  • Store metadata like title, author, publish date, tags, URL
  • Expect changes and build error handling/retries

Choosing the tool

  • Static HTML blog: requests + BeautifulSoup or trafilatura
  • JavaScript-heavy site: Playwright or Selenium
  • Large-scale crawling: Scrapy is a strong choice
  • Just extracting article text: trafilatura is often excellent

A practical recommendation

If your goal is to gather blog posts for research, monitoring, or a content archive:

  • Use RSS/API first
  • Fall back to HTML extraction
  • Use Playwright only when necessary

If you want, I can show you:

  • a Python example
  • a Node.js example
  • or a full scraper design for a blog site.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.