Prompt

How to crawl a site list automatically in Python

Technology / Cloud Infrastructure1 observationLast seen Oct 4, 2026

Latest observation

Oct 4, 2026OpenAI APIWeb search: off

If by “crawl a site list” you mean visit a list of URLs automatically and collect data from each page, you can do that in Python with requests + BeautifulSoup for simple HTML pages, or Scrapy for larger crawls.

1) Simple crawler for a list of URLs

Install

pip install requests beautifulsoup4

Example

import requests
from bs4 import BeautifulSoup

urls = [
    "https://example.com",
    "https://example.org",
]

for url in urls:
    try:
        response = requests.get(url, timeout=10, headers={
            "User-Agent": "Mozilla/5.0"
        })
        response.raise_for_status()

        soup = BeautifulSoup(response.text, "html.parser")
        title = soup.title.text.strip() if soup.title else "No title"

        print(f"{url} -> {title}")

    except requests.RequestException as e:
        print(f"Failed to fetch {url}: {e}")

2) Crawl from a file of URLs

If your site list is in urls.txt:

https://example.com
https://example.org

Use:

import requests
from bs4 import BeautifulSoup

with open("urls.txt", "r", encoding="utf-8") as f:
    urls = [line.strip() for line in f if line.strip()]

for url in urls:
    try:
        r = requests.get(url, timeout=10, headers={"User-Agent": "Mozilla/5.0"})
        r.raise_for_status()
        soup = BeautifulSoup(r.text, "html.parser")
        print(url, soup.title.text.strip() if soup.title else "No title")
    except Exception as e:
        print(url, "error:", e)

3) If you want to crawl links discovered on each page

This is closer to a real crawler: start with seed URLs, extract links, then continue.

import requests
from bs4 import BeautifulSoup
from urllib.parse import urljoin, urlparse

visited = set()
to_visit = ["https://example.com"]

def same_domain(base, target):
    return urlparse(base).netloc == urlparse(target).netloc

while to_visit:
    url = to_visit.pop(0)
    if url in visited:
        continue

    visited.add(url)

    try:
        r = requests.get(url, timeout=10, headers={"User-Agent": "Mozilla/5.0"})
        r.raise_for_status()
        soup = BeautifulSoup(r.text, "html.parser")

        print("Crawled:", url)

        for a in soup.find_all("a", href=True):
            link = urljoin(url, a["href"])
            if link not in visited and same_domain(url, link):
                to_visit.append(link)

    except Exception as e:
        print("Error:", url, e)

4) For larger crawling jobs, use Scrapy

Scrapy is better if you need:

  • many URLs
  • speed
  • retry handling
  • rate limiting
  • structured output

Install:

pip install scrapy

Create a project:

scrapy startproject mycrawler

Then define a spider that starts from a list of URLs.


Important tips

  • Respect robots.txt
  • Add delays between requests
  • Use a proper User-Agent
  • Avoid overloading the site
  • Check the site’s terms of service

Example delay:

import time
time.sleep(1)

If you want, I can give you:

  1. a crawler that reads URLs from a CSV/Excel file,
  2. a crawler that follows links automatically, or
  3. a Scrapy version.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as Circuit. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.