01

Predeclare the study before reading the outcome

The measurement plan is written down and locked, a predeclared method, before anyone looks at how either side did: who is eligible, which page templates are tested, how traffic is split, what the main and supporting metrics are, the safety guardrails, what gets excluded, the minimum number of visits needed, how big a change has to be before it counts, how the numbers get analyzed, and when the test is allowed to stop.

The default page matrix includes representative homepage, search-results, vehicle-detail and conversion experiences. Mobile and desktop are segmented; device, network, warm or cold visit state and material site changes are recorded instead of blended without explanation.

02

Establish baseline and validate assignment

Before anything is changed, the measurement is run against itself: two groups, both receiving the original site. If it reports a difference where none exists, the measurement is wrong and nothing proceeds until that is resolved. That check has to pass first.

What is established before the optimized experience is enabled
StepMinimumExtends toWhat it is for
Field baseline14 complete days28 days when traffic or variability requires itContext for what the site was doing before anything changed
Historical field coverageLabeled as URL-level, origin-level or unavailableNo extension: it is context either wayChrome UX Report data is context, not the result
Laboratory runs9 valid runs per URL and deviceAcross 3 days, under documented conditionsA repeatable diagnosis rather than one run
A/A validation3 complete days7 complete daysRunning the measurement against itself: two groups, both the original site
Four stages in order: a field baseline of at least 14 complete days extending to 28; an A/A validation of 3 to 7 complete days in which both groups receive the original site; an A/B comparison of at least 28 complete days that does not stop early; and publication gates requiring technical, data, safety, independent-review and written dealer approval. The four stages are drawn at the same width and no duration is encoded in the drawing.
The order a study runs in, and the minimum time fixed for each stage before it starts.
03

Compare against untouched traffic

Traffic is divided between the original site and the changed one. A given shopper stays on the same side of that line for the whole test, so nobody is moved mid-visit in a way that would flatter the result.

A visit stays counted under whichever experience it was assigned to, even when the change didn't end up applying for that visit: what the field calls an intention-to-treat analysis. A version counting only the visits where the change actually took effect may be shown separately, clearly labeled, but it never replaces the primary result.

04

Use one fixed main metric, and keep evidence types separate

If Interaction to Next Paint data is missing for a visit, it is left out of that metric rather than counted as zero. Results are broken out by page type, so a page that happens to get more traffic can't quietly outweigh the page mix that was agreed on beforehand.

  • Primary field metric: Largest Contentful Paint on mobile, at the 75th percentile, on a fixed mix of page types set before the test runs and never changed afterward
  • Supporting field metrics: Interaction to Next Paint and First Contentful Paint; Cumulative Layout Shift is a safety guardrail and Time to First Byte is context only
  • Operational guardrails: JavaScript errors, failed functions, fallback triggers and the protected-flow checks run on the dealership's own site
  • Laboratory evidence: repeated test runs and their medians, used to diagnose a cause, never relabeled as what real visitors experienced
05

Fix how the results get analyzed, and when the test can stop

The default analysis is a bootstrap: it resamples the visit data itself, grouped by visitor and broken out by page type, 10,000 resamples in total, or an equivalent documented method, to work out how different the two groups really are and how confident that difference is. How big a change has to be before it counts as meaningful is fixed before anyone looks at the result.

After the A/A check passes, the comparison runs for at least 28 complete days and until the sample size set in advance is reached. It does not stop early just because a favorable number shows up. The maximum planned window is 56 complete days; if there still isn't enough data by then to be confident either way, the result is reported as inconclusive rather than as a finding.

06

Keep what gets excluded fair to both sides, and documented

Bots, internal test traffic, invalid telemetry and conditions ruled out in advance may be excluded, but only by rules applied without knowing which group came out ahead. A poor day, a slower device, paid traffic or a session that didn't go well is never removed just because it weakens a result.

Every change gets logged. Site releases, campaign changes, inventory shifts, vendor incidents, outages and measurement changes are recorded in the change log and evaluated against the validity rules.

07

Classify the result without marketing inflation

Every finding remains scoped to the named site, dates, page set, device segment, sample, metric, confidence interval and product configuration. Groups reporting across several rooftops should read how to report website speed across a dealer group before aggregating anything. A performance result does not establish search-ranking, lead, conversion or revenue impact.

  • Verified improvement
  • Detectable but not materially large
  • No material difference
  • Inconclusive
  • Regression
  • Invalid study
08

Require publication and privacy gates

The company behind this standard publishes it in full so a dealership, or anyone checking this claim, can read exactly how a result would be measured before trusting one. A result is published only after three things happen. The technical, data-quality and safety checks pass. The analysis and claim ledger get independent review. And the dealership approves the exact public wording in writing. There are currently no approved public performance results or case studies.

Measurement payloads exclude form or chat content, names, email addresses, phone numbers and customer records. The production data flow, retention and subprocessors are set out in the agreement covering a specific implementation.

Sources and further reading

External sources support the general technical guidance on this page. They do not represent a DealerSpeed Engine performance result.