Prompt

How do I extract phone numbers and emails from business directories?

Technology / API Platforms1 observationLast seen Oct 4, 2026

Latest observation

Oct 4, 2026OpenAI APIWeb search: off

To extract phone numbers and emails from business directories, the safest approach is to use the site’s allowed access methods first and avoid scraping anything prohibited by the site’s terms or robots rules.

Good options

  1. Check for an official API

    • Many directories offer search APIs or partner feeds.
    • This is usually the most reliable and compliant method.
  2. Export / download features

    • Some directories let you export leads or business listings as CSV/XLSX.
  3. Web scraping, if permitted

    • If the directory allows automated access, you can parse the listing pages and collect:
      • phone numbers
      • emails
      • website URLs
    • Be careful with rate limits and legal restrictions.

Typical extraction approach

1) Fetch the page

Use a request library to get the HTML.

2) Parse the HTML

Look for:

  • visible text like Phone: (123) 456-7890
  • mailto: links for emails
  • structured data like JSON-LD
  • schema.org markup

3) Extract with patterns

Use regex for:

  • emails
  • phone numbers

4) Clean and normalize

  • Remove duplicates
  • Normalize phone numbers to E.164 if possible
  • Validate emails

Example in Python

import re
import requests
from bs4 import BeautifulSoup

url = "https://example.com/business-directory-page"
html = requests.get(url, timeout=15).text

soup = BeautifulSoup(html, "html.parser")

# Emails from mailto links and page text
emails = set()

for a in soup.find_all("a", href=True):
    href = a["href"]
    if href.startswith("mailto:"):
        emails.add(href.replace("mailto:", "").split("?")[0])

text = soup.get_text(" ", strip=True)
email_regex = r'[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}'
emails.update(re.findall(email_regex, text))

# Phone numbers (basic pattern; adjust for your region)
phone_regex = r'(\+?\d[\d\-\s\(\)]{7,}\d)'
phones = set(re.findall(phone_regex, text))

print("Emails:", emails)
print("Phones:", phones)

Important notes

  • Respect robots.txt and terms of service.
  • Avoid bypassing protections like CAPTCHAs, logins, or obfuscation.
  • Some directories hide contact details specifically to prevent automated extraction.
  • If you need data at scale, consider licensed datasets or lead providers.

Better alternatives

If your goal is sales or outreach, you can use:

  • official business listings APIs
  • CRM enrichment tools
  • licensed business data providers
  • contact-finding services with compliance features

If you want, I can show you:

  • a Python scraper template for a directory page
  • how to extract from JSON-LD
  • how to normalize phone numbers and emails safely

Brands

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.