This guide is part of the dealer website speed resource library.
Why before and after can mislead
A test run this month and another next month may differ because of inventory, creative, tags, devices, connection quality, promotions, or site releases. The optimization is not the only changing factor.
| Variable | Why it moves on its own | How a controlled test handles it |
|---|---|---|
| Traffic mix | A paid campaign starting or ending changes device and network mix | Both groups draw from the same traffic at the same time |
| Inventory | The pages themselves change as vehicles arrive and sell | Both groups see the same inventory |
| Third-party tags | A vendor ships an update without telling the dealership | Both groups carry the same tag stack |
| Time of day and day of week | Network conditions and device mix vary by hour | Both groups are sampled across the same period |
Keep a comparable control
A split test retains the original experience for part of eligible traffic while the optimized experience serves another part. Assignment, exclusions, bot handling, cache behavior, and page eligibility should be defined before reading results.
Report the test honestly
- Website and included page templates
- Start and end dates
- Eligible and excluded traffic
- Sample sizes and device mix
- Metric definitions and aggregation
- Known concurrent changes and limitations
Use lab tests for diagnosis
Laboratory traces remain valuable for explaining what changed. Keep them labeled as lab evidence; do not blend them with field or controlled-traffic observations into one number.
How do I run a defensible speed test?
- 01
Declare the plan first
Primary metric, sample target, exclusion rules and analysis method, written down before any result is visible.
- 02
Validate with an A/A test
Run the same experience against itself. If it shows a difference, your instrument is broken and no later result can be trusted.
- 03
Assign traffic stably
A returning visitor must stay in the same group. Visitors bouncing between experiences contaminate both.
- 04
Run to the target
Not until the number looks good. Early stopping on a favorable result is the most common way a false positive gets published.
- 05
Report what happened
Including unfavorable and inconclusive outcomes, with the confidence interval and the limitations.
Why is before-and-after so unreliable at a dealership?
Because almost everything that determines dealership website performance changes month to month for reasons unrelated to any speed project. Inventory turns over, and with it the media weight of the average vehicle page. Manufacturer incentives start and stop, changing which pages get traffic. Campaigns launch, changing the device and connection mix of the visitors arriving. The platform ships releases. An agency adds a tag.
A before-and-after comparison attributes every one of those movements to whatever the dealership changed in between, which is why the measurement methodology requires a concurrent control instead. Sometimes that flatters the vendor and sometimes it buries a genuine improvement, and there is no way to tell which from the comparison itself. That is not a subtle statistical objection; the month-to-month variance on a script-heavy dealership site is frequently larger than the effect being measured.
A concurrent control removes the problem by construction. Both groups experience the same inventory, the same incentives, the same campaigns and the same releases, because they are experiencing them at the same time.
What has to be decided before the test starts?
Fixing those in advance is what separates a test from a search for a favorable number. Stopping when the figure looks good is the most common way a false positive gets published, and it is invisible in the write-up unless the stopping rule was declared first. The full methodology sets out how each of these is handled here.
- Which traffic is eligible, and what excludes a session: decided without knowing the result
- How assignment works and how long it stays stable for a returning visitor
- Which page templates are in scope, and reported separately instead of pooled
- The primary metric, the guardrail metrics, and what result would stop the test early for safety
- The minimum sample and the minimum duration, which for field data is at least 28 complete days
- How fallbacks are analyzed: an intention-to-treat view keeps them in their assigned group
Sources and further reading
External sources support the general technical guidance on this page. They do not represent a DealerSpeed Engine performance result.