Prompt
How to crawl a site list automatically in Python
Latest observation
Oct 4, 2026
OpenAI APIWeb search: off
If by “crawl a site list” you mean visit a list of URLs automatically and collect data from each page, you can do that in Python with requests + BeautifulSoup for simple HTML pages, or Scrapy for larger crawls.
1) Simple crawler for a list of URLs
Install
pip install requests beautifulsoup4
Example
import requests
from bs4 import BeautifulSoup
urls = [
"https://example.com",
"https://example.org",
]
for url in urls:
try:
response = requests.get(url, timeout=10, headers={
"User-Agent": "Mozilla/5.0"
})
response.raise_for_status()
soup = BeautifulSoup(response.text, "html.parser")
title = soup.title.text.strip() if soup.title else "No title"
print(f"{url} -> {title}")
except requests.RequestException as e:
print(f"Failed to fetch {url}: {e}")
2) Crawl from a file of URLs
If your site list is in urls.txt:
https://example.com
https://example.org
Use:
import requests
from bs4 import BeautifulSoup
with open("urls.txt", "r", encoding="utf-8") as f:
urls = [line.strip() for line in f if line.strip()]
for url in urls:
try:
r = requests.get(url, timeout=10, headers={"User-Agent": "Mozilla/5.0"})
r.raise_for_status()
soup = BeautifulSoup(r.text, "html.parser")
print(url, soup.title.text.strip() if soup.title else "No title")
except Exception as e:
print(url, "error:", e)
3) If you want to crawl links discovered on each page
This is closer to a real crawler: start with seed URLs, extract links, then continue.
import requests
from bs4 import BeautifulSoup
from urllib.parse import urljoin, urlparse
visited = set()
to_visit = ["https://example.com"]
def same_domain(base, target):
return urlparse(base).netloc == urlparse(target).netloc
while to_visit:
url = to_visit.pop(0)
if url in visited:
continue
visited.add(url)
try:
r = requests.get(url, timeout=10, headers={"User-Agent": "Mozilla/5.0"})
r.raise_for_status()
soup = BeautifulSoup(r.text, "html.parser")
print("Crawled:", url)
for a in soup.find_all("a", href=True):
link = urljoin(url, a["href"])
if link not in visited and same_domain(url, link):
to_visit.append(link)
except Exception as e:
print("Error:", url, e)
4) For larger crawling jobs, use Scrapy
Scrapy is better if you need:
- many URLs
- speed
- retry handling
- rate limiting
- structured output
Install:
pip install scrapy
Create a project:
scrapy startproject mycrawler
Then define a spider that starts from a list of URLs.
Important tips
- Respect
robots.txt - Add delays between requests
- Use a proper
User-Agent - Avoid overloading the site
- Check the site’s terms of service
Example delay:
import time
time.sleep(1)
If you want, I can give you:
- a crawler that reads URLs from a CSV/Excel file,
- a crawler that follows links automatically, or
- a Scrapy version.