Prompt
How do I extract listings from multiple pages with pagination and save them as JSON?
Latest observation
Oct 4, 2026
OpenAI APIWeb search: off
To extract listings from multiple pages with pagination and save them as JSON, the usual workflow is:
- Request the first page
- Parse the listings on that page
- Find the link/parameter for the next page
- Repeat until there are no more pages
- Save all collected listings to a JSON file
Here’s a simple Python example using requests and BeautifulSoup:
import requests
from bs4 import BeautifulSoup
import json
base_url = "https://example.com/listings"
all_listings = []
page = 1
while True:
url = f"{base_url}?page={page}"
response = requests.get(url)
response.raise_for_status()
soup = BeautifulSoup(response.text, "html.parser")
# Adjust these selectors to match the site structure
listing_cards = soup.select(".listing-card")
if not listing_cards:
break
for card in listing_cards:
title = card.select_one(".title").get_text(strip=True) if card.select_one(".title") else None
price = card.select_one(".price").get_text(strip=True) if card.select_one(".price") else None
link_tag = card.select_one("a")
link = link_tag["href"] if link_tag and link_tag.has_attr("href") else None
all_listings.append({
"title": title,
"price": price,
"link": link
})
# Stop if there is no next page
next_button = soup.select_one("a.next")
if not next_button:
break
page += 1
# Save to JSON
with open("listings.json", "w", encoding="utf-8") as f:
json.dump(all_listings, f, ensure_ascii=False, indent=2)
print(f"Saved {len(all_listings)} listings to listings.json")
If the pagination uses a “Next” button instead of ?page=
You can follow the href from the next button directly:
import requests
from bs4 import BeautifulSoup
import json
from urllib.parse import urljoin
start_url = "https://example.com/listings"
all_listings = []
url = start_url
while url:
response = requests.get(url)
response.raise_for_status()
soup = BeautifulSoup(response.text, "html.parser")
for card in soup.select(".listing-card"):
all_listings.append({
"title": card.select_one(".title").get_text(strip=True) if card.select_one(".title") else None,
"price": card.select_one(".price").get_text(strip=True) if card.select_one(".price") else None,
"link": urljoin(url, card.select_one("a")["href"]) if card.select_one("a") else None
})
next_link = soup.select_one("a.next")
url = urljoin(url, next_link["href"]) if next_link and next_link.has_attr("href") else None
with open("listings.json", "w", encoding="utf-8") as f:
json.dump(all_listings, f, ensure_ascii=False, indent=2)
Tips
- Use browser dev tools to inspect the HTML and find the correct selectors
- Add
time.sleep()between requests to avoid overloading the site - Check the site’s robots.txt and terms of service before scraping
- If the content is loaded by JavaScript, you may need Selenium or Playwright instead of
requests
If you want, I can also give you:
- a version for requests + pandas
- a Selenium/Playwright version
- or help you adapt this to a specific website’s HTML