Prompt
How do I troubleshoot 502 errors on a hosted web app?
Latest observation
A 502 Bad Gateway usually means a proxy, load balancer, or gateway couldn’t get a valid response from the upstream app server. For a hosted web app, the fastest way to troubleshoot is to work from the edge inward.
1) Confirm where the 502 is coming from
Check whether the error is generated by:
- CDN / WAF
- Load balancer / reverse proxy
- App platform router
- Your app server
Look at response headers if possible:
servervia- platform-specific headers like
x-served-by,x-request-id,x-cache, etc.
This helps identify the layer failing.
2) Check whether the app is actually healthy
Verify:
- App process is running
- Instances/containers are up
- Health checks are passing
- No recent deploy failed or rolled back partially
On hosted platforms, a 502 often happens when the platform routes traffic to:
- a crashed container
- an app that isn’t listening on the expected port
- a process that started but isn’t ready yet
3) Inspect logs around the time of the error
Look at:
- application logs
- web server logs
- platform/router logs
- container logs
- load balancer logs
Search for:
- crashes
- startup failures
- timeout errors
- connection refused
- upstream reset
- out-of-memory kills
- stack traces
A 502 is often paired with an upstream failure message like:
connect() failedupstream timed outbad gatewayECONNRESET503 from upstreamprocess exited with code 1
4) Check recent deployments and config changes
Common causes after a deployment:
- Wrong port exposed
- App binding to
localhostinstead of0.0.0.0 - Missing environment variables
- Bad secrets or credentials
- Dependency install failure
- Build artifacts not created
- Migration failure causing startup crash
If the issue started right after a deploy, roll back and compare.
5) Validate timeouts and request size limits
A 502 can be caused by an upstream timing out or closing the connection:
- slow database queries
- long-running requests
- cold starts
- too-small proxy timeout
- large payloads or headers
Check and compare:
- proxy timeout
- app server timeout
- DB timeout
- load balancer idle timeout
6) Check capacity and resource pressure
If the app is overloaded, the gateway may fail to get a response:
- CPU pegged
- memory exhausted
- container OOM kills
- too few app instances
- connection pool exhaustion
- thread pool exhaustion
Look at:
- CPU, memory, latency, error rate
- autoscaling events
- saturation of DB connections and queues
7) Test the app directly, bypassing the gateway if possible
From inside the same network or via a direct internal endpoint:
- curl the app service directly
- hit the health endpoint
- compare direct response vs public endpoint
If direct access works but public access fails, the problem is likely:
- proxy config
- load balancer config
- TLS termination
- routing rules
- CDN/WAF
8) Check DNS, TLS, and networking
Sometimes 502s are caused by connectivity issues between layers:
- DNS points to wrong backend
- expired/invalid TLS cert
- firewall/security group blocks backend
- broken service discovery
- upstream IP changed
Validate:
- DNS resolution
- certificate validity
- backend port is open
- security rules allow traffic from proxy to app
9) Review health check configuration
A bad health check can make healthy instances look unhealthy and cause 502s. Check:
- path is correct
- expected status code matches
- timeout is reasonable
- startup grace period is long enough
- check endpoint doesn’t depend on downstream systems unnecessarily
Best practice: health endpoint should be fast and not require database access unless you truly want readiness to fail when DB is down.
10) Reproduce and isolate
Try to reproduce with:
- the same URL
- same method
- same headers/body
- same user/account
- same region/zone if applicable
See whether the problem is:
- all traffic
- one route
- one instance
- one region
- one client type
Common root causes
- App crashed on startup
- Wrong port or host binding
- Proxy timeout too short
- Upstream overloaded
- OOM kill / memory leak
- Bad deploy or bad config
- Broken health checks
- Network/security group issue
- TLS mismatch between proxy and app
- Dependency or DB outage causing app to hang
Quick triage checklist
- Check platform status page
- Check logs
- Confirm app is running and healthy
- Roll back latest deploy if needed
- Verify port, host binding, and env vars
- Check timeouts and resource usage
- Test direct-to-app connectivity
- Inspect health checks and network rules
If you want, tell me what platform you’re using, like AWS ALB, Nginx, Cloudflare, Heroku, Render, Vercel, Kubernetes, etc., and I can give you a more specific troubleshooting flow.