Prompt
How do I extract tables from web pages and export them as JSON?
Latest observation
Oct 4, 2026
OpenAI APIWeb search: off
You can do this in a few different ways depending on the site and how many tables you need.
Option 1: Python with pandas.read_html() for simple HTML tables
If the table is in the page’s HTML, this is the easiest approach.
Install
pip install pandas lxml
Example
import pandas as pd
url = "https://example.com/page-with-table"
tables = pd.read_html(url)
print(f"Found {len(tables)} tables")
# Export the first table to JSON
tables[0].to_json("table.json", orient="records", indent=2)
Common JSON formats
orient="records"→ list of row objectsorient="split"→ rows/columns structureorient="index"→ keyed by row index
Example:
tables[0].to_json("table.json", orient="records", indent=2)
Produces:
[
{"Name":"Alice","Age":30},
{"Name":"Bob","Age":25}
]
Option 2: requests + BeautifulSoup + custom parsing
Use this if pandas.read_html() fails or you need more control.
Install
pip install requests beautifulsoup4
Example
import requests
from bs4 import BeautifulSoup
import json
url = "https://example.com/page-with-table"
html = requests.get(url).text
soup = BeautifulSoup(html, "html.parser")
table = soup.find("table")
headers = [th.get_text(strip=True) for th in table.find_all("th")]
rows = []
for tr in table.find_all("tr"):
cells = [td.get_text(strip=True) for td in tr.find_all("td")]
if cells:
rows.append(dict(zip(headers, cells)))
with open("table.json", "w", encoding="utf-8") as f:
json.dump(rows, f, indent=2, ensure_ascii=False)
Option 3: Selenium or Playwright for JavaScript-rendered tables
If the table is loaded dynamically by JavaScript, use a browser automation tool.
Example with Selenium
pip install selenium pandas
Then load the page, wait for the table to render, grab page_source, and use pandas.read_html() or parse it with BeautifulSoup.
from selenium import webdriver
import pandas as pd
driver = webdriver.Chrome()
driver.get("https://example.com/dynamic-table")
html = driver.page_source
tables = pd.read_html(html)
tables[0].to_json("table.json", orient="records", indent=2)
driver.quit()
Option 4: Scrapy for large-scale crawling
If you need to extract tables from many pages, Scrapy is a good framework.
Important notes
- Check the site’s robots.txt and terms of service.
- Add delays to avoid overwhelming servers.
- Some sites block scraping; use headers or a browser-based approach if needed.
- Tables may have:
- merged cells (
rowspan,colspan) - nested tables
- hidden rows/columns
- pagination
- merged cells (
If you want a one-liner
For many pages with standard HTML tables:
import pandas as pd
pd.read_html("https://example.com")[0].to_json("table.json", orient="records", indent=2)
If you want, I can also give you:
- a script that extracts all tables from a page, or
- a version that handles JavaScript-rendered tables.