Prompt
How do I extract phone numbers and emails from business directories?
Latest observation
Oct 4, 2026
OpenAI APIWeb search: off
To extract phone numbers and emails from business directories, the safest approach is to use the site’s allowed access methods first and avoid scraping anything prohibited by the site’s terms or robots rules.
Good options
-
Check for an official API
- Many directories offer search APIs or partner feeds.
- This is usually the most reliable and compliant method.
-
Export / download features
- Some directories let you export leads or business listings as CSV/XLSX.
-
Web scraping, if permitted
- If the directory allows automated access, you can parse the listing pages and collect:
- phone numbers
- emails
- website URLs
- Be careful with rate limits and legal restrictions.
- If the directory allows automated access, you can parse the listing pages and collect:
Typical extraction approach
1) Fetch the page
Use a request library to get the HTML.
2) Parse the HTML
Look for:
- visible text like
Phone: (123) 456-7890 mailto:links for emails- structured data like JSON-LD
- schema.org markup
3) Extract with patterns
Use regex for:
- emails
- phone numbers
4) Clean and normalize
- Remove duplicates
- Normalize phone numbers to E.164 if possible
- Validate emails
Example in Python
import re
import requests
from bs4 import BeautifulSoup
url = "https://example.com/business-directory-page"
html = requests.get(url, timeout=15).text
soup = BeautifulSoup(html, "html.parser")
# Emails from mailto links and page text
emails = set()
for a in soup.find_all("a", href=True):
href = a["href"]
if href.startswith("mailto:"):
emails.add(href.replace("mailto:", "").split("?")[0])
text = soup.get_text(" ", strip=True)
email_regex = r'[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}'
emails.update(re.findall(email_regex, text))
# Phone numbers (basic pattern; adjust for your region)
phone_regex = r'(\+?\d[\d\-\s\(\)]{7,}\d)'
phones = set(re.findall(phone_regex, text))
print("Emails:", emails)
print("Phones:", phones)
Important notes
- Respect robots.txt and terms of service.
- Avoid bypassing protections like CAPTCHAs, logins, or obfuscation.
- Some directories hide contact details specifically to prevent automated extraction.
- If you need data at scale, consider licensed datasets or lead providers.
Better alternatives
If your goal is sales or outreach, you can use:
- official business listings APIs
- CRM enrichment tools
- licensed business data providers
- contact-finding services with compliance features
If you want, I can show you:
- a Python scraper template for a directory page
- how to extract from JSON-LD
- how to normalize phone numbers and emails safely
Brands