Prompt
How do I run server-side A/B tests?
Latest observation
Server-side A/B testing means the experiment logic runs on your backend, not in the browser. That gives you more control, better performance, and cleaner measurement for things like pricing, recommendations, APIs, and onboarding flows.
Basic flow
-
Pick a unit to assign
- Usually a user ID, account ID, or anonymous visitor ID.
- Keep it stable so the same person always gets the same variant.
-
Choose variants
- Example:
controlvstreatment - Define the traffic split, e.g. 50/50 or 90/10.
- Example:
-
Assign on the server
- When a request comes in, your backend decides which variant the user gets.
- Persist the assignment in a database, cache, cookie, or feature flag system.
-
Serve different logic
- Return different content, pricing, ranking, API responses, or workflows based on the assigned variant.
-
Track exposures and conversions
- Log when a user was exposed to a variant.
- Log downstream outcomes, such as clicks, purchases, retention, or errors.
-
Analyze results
- Compare conversion rates, revenue, retention, etc. between variants.
- Use statistical testing or a Bayesian approach.
A simple server-side example
Pseudocode
def get_variant(user_id, experiment_name):
# deterministic assignment
bucket = hash(user_id + experiment_name) % 100
if bucket < 50:
return "control"
else:
return "treatment"
def handle_request(user_id):
variant = get_variant(user_id, "new_checkout")
log_exposure(user_id, "new_checkout", variant)
if variant == "control":
return render_checkout(old_flow=True)
else:
return render_checkout(old_flow=False)
This ensures the same user gets the same variant every time.
Common implementation patterns
1. Deterministic hashing
Best for:
- Stable assignment
- No need to store every assignment
- Large-scale systems
How it works:
- Hash
user_id + experiment_name - Map the hash to a bucket from 0–99 or 0–9999
- Use bucket ranges to assign variants
2. Centralized feature flag / experimentation service
Best for:
- Multiple experiments
- Rollouts and kill switches
- Real-time changes without redeploying
Examples:
- LaunchDarkly
- Optimizely
- Statsig
- Split.io
- GrowthBook
3. Server-stored assignment
Best for:
- Anonymous users
- Auditable assignments
- Complex eligibility rules
Store the variant in:
- Redis
- Database
- Session store
- Signed cookie
Important design rules
Keep assignment sticky
A user should not bounce between variants, or your results will be noisy.
Randomize only eligible users
Filter out ineligible traffic first:
- Employees
- bots
- unsupported regions
- users without required permissions
Log exposure, not just assignment
A user can be assigned but never actually see the treatment. For clean analysis, track actual exposure.
Avoid interfering experiments
If you run multiple tests at once, make sure they don’t overlap in ways that bias results. Use:
- namespaces
- mutually exclusive buckets
- layered assignment
Watch for caching issues
If your pages or API responses are cached, ensure the variant is part of the cache key or you may serve the wrong version.
Respect privacy and compliance
Don’t use sensitive attributes for assignment unless necessary and allowed.
What to measure
Typical metrics:
- Conversion rate
- Revenue per user
- Retention
- Click-through rate
- Latency
- Error rate
- Bounce rate
Also watch guardrails:
- Page load time
- Server cost
- Support tickets
- Crash rate
How to analyze
For a simple conversion metric:
- Compute conversion rate for each variant
- Compare lift:
(treatment - control) / control - Use confidence intervals or Bayesian credible intervals
- Check sample size and power before judging results
Recommended stack
A typical setup:
- Experiment assignment: backend middleware or feature flag service
- Exposure logging: event pipeline like Kafka, Segment, Snowplow, or your own logs
- Analysis: SQL, dbt, Python/R, or BI tools
- Decisioning: feature flags or config service for rollout control
Minimal production checklist
- Stable user identifier
- Deterministic or persisted assignment
- Exposure logging
- Metric logging
- Guardrail metrics
- Ability to stop or roll back the test
- Cache-safe response handling
- Analysis plan before launch
If you want, I can also give you:
- a Node.js / Express example,
- a Python / Django or FastAPI example,
- or a database schema + SQL setup for analyzing server-side experiments.