Prompt
How do I scrape a website that requires login credentials?
Latest observation
Oct 7, 2026
OpenAI APIWeb search: off
If you have permission to access the site, the usual approach is:
-
Check for an official API first
Many sites provide APIs that are easier and more reliable than scraping. -
Automate the login flow
- Use a browser automation tool like Playwright or Selenium if the site uses a form-based login, CSRF tokens, or JavaScript-heavy auth.
- Or use
requests.Session()if the login is a simple POST form and the site doesn’t rely heavily on client-side behavior.
-
Preserve session cookies
- After logging in, keep the session cookies so subsequent requests stay authenticated.
- In
requests, this is whatSession()does. - In browser automation, the browser context stores them automatically.
-
Handle anti-bot protections
- Some sites use CAPTCHA, MFA, or bot detection.
- Don’t try to bypass these protections. If they appear, use the site’s intended access methods or get authorized programmatic access.
-
Respect robots.txt, terms of service, and rate limits
- Even if login is required, the site may still restrict automated access.
Simple example with requests (for a basic form login)
import requests
from bs4 import BeautifulSoup
login_url = "https://example.com/login"
protected_url = "https://example.com/account"
session = requests.Session()
# 1) Load login page to get cookies / CSRF token if needed
resp = session.get(login_url)
soup = BeautifulSoup(resp.text, "html.parser")
csrf = soup.find("input", {"name": "csrf_token"})
csrf_token = csrf["value"] if csrf else None
# 2) Submit credentials
payload = {
"username": "your_username",
"password": "your_password",
}
if csrf_token:
payload["csrf_token"] = csrf_token
login_resp = session.post(login_url, data=payload)
# 3) Access protected page
page = session.get(protected_url)
print(page.status_code)
print(page.text[:500])
If the site is JS-heavy, use Playwright
from playwright.sync_api import sync_playwright
with sync_playwright() as p:
browser = p.chromium.launch(headless=True)
page = browser.new_page()
page.goto("https://example.com/login")
page.fill('input[name="username"]', "your_username")
page.fill('input[name="password"]', "your_password")
page.click('button[type="submit"]')
page.wait_for_load_state("networkidle")
page.goto("https://example.com/account")
print(page.content())
browser.close()
Good practice
- Store credentials securely in environment variables or a secret manager.
- Add delays and backoff to avoid hammering the site.
- Log only what you need; avoid saving sensitive data.
If you tell me the site type (basic form login, OAuth, SSO, 2FA, etc.) and your preferred language, I can show a more specific example.