01

Why before and after can mislead

A test run this month and another next month may differ because of inventory, creative, tags, devices, connection quality, promotions, or site releases. The optimization is not the only changing factor.

What changes between a before test and an after test, other than the change you made
VariableWhy it moves on its ownHow a controlled test handles it
Traffic mixA paid campaign starting or ending changes device and network mixBoth groups draw from the same traffic at the same time
InventoryThe pages themselves change as vehicles arrive and sellBoth groups see the same inventory
Third-party tagsA vendor ships an update without telling the dealershipBoth groups carry the same tag stack
Time of day and day of weekNetwork conditions and device mix vary by hourBoth groups are sampled across the same period
02

Keep a comparable control

A split test retains the original experience for part of eligible traffic while the optimized experience serves another part. Assignment, exclusions, bot handling, cache behavior, and page eligibility should be defined before reading results.

A controlled dealership speed study over time. The original and optimized experiences run concurrently on split traffic, across the same days and the same conditions, instead of one after the other.
Concurrent, not sequential. Running the two versions at the same time is what removes traffic mix, inventory and vendor updates as explanations for the difference.
03

Report the test honestly

  • Website and included page templates
  • Start and end dates
  • Eligible and excluded traffic
  • Sample sizes and device mix
  • Metric definitions and aggregation
  • Known concurrent changes and limitations
04

Use lab tests for diagnosis

Laboratory traces remain valuable for explaining what changed. Keep them labeled as lab evidence; do not blend them with field or controlled-traffic observations into one number.

05

How do I run a defensible speed test?

  1. 01

    Declare the plan first

    Primary metric, sample target, exclusion rules and analysis method, written down before any result is visible.

  2. 02

    Validate with an A/A test

    Run the same experience against itself. If it shows a difference, your instrument is broken and no later result can be trusted.

  3. 03

    Assign traffic stably

    A returning visitor must stay in the same group. Visitors bouncing between experiences contaminate both.

  4. 04

    Run to the target

    Not until the number looks good. Early stopping on a favorable result is the most common way a false positive gets published.

  5. 05

    Report what happened

    Including unfavorable and inconclusive outcomes, with the confidence interval and the limitations.

06

Why is before-and-after so unreliable at a dealership?

Because almost everything that determines dealership website performance changes month to month for reasons unrelated to any speed project. Inventory turns over, and with it the media weight of the average vehicle page. Manufacturer incentives start and stop, changing which pages get traffic. Campaigns launch, changing the device and connection mix of the visitors arriving. The platform ships releases. An agency adds a tag.

A before-and-after comparison attributes every one of those movements to whatever the dealership changed in between, which is why the measurement methodology requires a concurrent control instead. Sometimes that flatters the vendor and sometimes it buries a genuine improvement, and there is no way to tell which from the comparison itself. That is not a subtle statistical objection; the month-to-month variance on a script-heavy dealership site is frequently larger than the effect being measured.

A concurrent control removes the problem by construction. Both groups experience the same inventory, the same incentives, the same campaigns and the same releases, because they are experiencing them at the same time.

07

What has to be decided before the test starts?

Fixing those in advance is what separates a test from a search for a favorable number. Stopping when the figure looks good is the most common way a false positive gets published, and it is invisible in the write-up unless the stopping rule was declared first. The full methodology sets out how each of these is handled here.

  • Which traffic is eligible, and what excludes a session: decided without knowing the result
  • How assignment works and how long it stays stable for a returning visitor
  • Which page templates are in scope, and reported separately instead of pooled
  • The primary metric, the guardrail metrics, and what result would stop the test early for safety
  • The minimum sample and the minimum duration, which for field data is at least 28 complete days
  • How fallbacks are analyzed: an intention-to-treat view keeps them in their assigned group

Sources and further reading

External sources support the general technical guidance on this page. They do not represent a DealerSpeed Engine performance result.