Prompt

How do I run server-side A/B tests?

Technology / AB Testing & Experimentation1 observationLast seen Oct 7, 2026

Latest observation

Oct 7, 2026OpenAI APIWeb search: off

Server-side A/B testing means the experiment logic runs on your backend, not in the browser. That gives you more control, better performance, and cleaner measurement for things like pricing, recommendations, APIs, and onboarding flows.

Basic flow

  1. Pick a unit to assign

    • Usually a user ID, account ID, or anonymous visitor ID.
    • Keep it stable so the same person always gets the same variant.
  2. Choose variants

    • Example: control vs treatment
    • Define the traffic split, e.g. 50/50 or 90/10.
  3. Assign on the server

    • When a request comes in, your backend decides which variant the user gets.
    • Persist the assignment in a database, cache, cookie, or feature flag system.
  4. Serve different logic

    • Return different content, pricing, ranking, API responses, or workflows based on the assigned variant.
  5. Track exposures and conversions

    • Log when a user was exposed to a variant.
    • Log downstream outcomes, such as clicks, purchases, retention, or errors.
  6. Analyze results

    • Compare conversion rates, revenue, retention, etc. between variants.
    • Use statistical testing or a Bayesian approach.

A simple server-side example

Pseudocode

def get_variant(user_id, experiment_name):
    # deterministic assignment
    bucket = hash(user_id + experiment_name) % 100
    if bucket < 50:
        return "control"
    else:
        return "treatment"

def handle_request(user_id):
    variant = get_variant(user_id, "new_checkout")
    
    log_exposure(user_id, "new_checkout", variant)
    
    if variant == "control":
        return render_checkout(old_flow=True)
    else:
        return render_checkout(old_flow=False)

This ensures the same user gets the same variant every time.


Common implementation patterns

1. Deterministic hashing

Best for:

  • Stable assignment
  • No need to store every assignment
  • Large-scale systems

How it works:

  • Hash user_id + experiment_name
  • Map the hash to a bucket from 0–99 or 0–9999
  • Use bucket ranges to assign variants

2. Centralized feature flag / experimentation service

Best for:

  • Multiple experiments
  • Rollouts and kill switches
  • Real-time changes without redeploying

Examples:

  • LaunchDarkly
  • Optimizely
  • Statsig
  • Split.io
  • GrowthBook

3. Server-stored assignment

Best for:

  • Anonymous users
  • Auditable assignments
  • Complex eligibility rules

Store the variant in:

  • Redis
  • Database
  • Session store
  • Signed cookie

Important design rules

Keep assignment sticky

A user should not bounce between variants, or your results will be noisy.

Randomize only eligible users

Filter out ineligible traffic first:

  • Employees
  • bots
  • unsupported regions
  • users without required permissions

Log exposure, not just assignment

A user can be assigned but never actually see the treatment. For clean analysis, track actual exposure.

Avoid interfering experiments

If you run multiple tests at once, make sure they don’t overlap in ways that bias results. Use:

  • namespaces
  • mutually exclusive buckets
  • layered assignment

Watch for caching issues

If your pages or API responses are cached, ensure the variant is part of the cache key or you may serve the wrong version.

Respect privacy and compliance

Don’t use sensitive attributes for assignment unless necessary and allowed.


What to measure

Typical metrics:

  • Conversion rate
  • Revenue per user
  • Retention
  • Click-through rate
  • Latency
  • Error rate
  • Bounce rate

Also watch guardrails:

  • Page load time
  • Server cost
  • Support tickets
  • Crash rate

How to analyze

For a simple conversion metric:

  • Compute conversion rate for each variant
  • Compare lift: (treatment - control) / control
  • Use confidence intervals or Bayesian credible intervals
  • Check sample size and power before judging results

Recommended stack

A typical setup:

  • Experiment assignment: backend middleware or feature flag service
  • Exposure logging: event pipeline like Kafka, Segment, Snowplow, or your own logs
  • Analysis: SQL, dbt, Python/R, or BI tools
  • Decisioning: feature flags or config service for rollout control

Minimal production checklist

  • Stable user identifier
  • Deterministic or persisted assignment
  • Exposure logging
  • Metric logging
  • Guardrail metrics
  • Ability to stop or roll back the test
  • Cache-safe response handling
  • Analysis plan before launch

If you want, I can also give you:

  1. a Node.js / Express example,
  2. a Python / Django or FastAPI example,
  3. or a database schema + SQL setup for analyzing server-side experiments.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.