Most CRO case studies are written on successes only; an uplift gets its place in the title, gets a green arrow in the chart, and the story ends here.
This story did not end in a success for us; we did a pricing display experiment for a menswear retail client (Oliver Brown), did two runs of three-week test each time, chased a very promising early signal in between, but got a zero result anyway. We are still going to share this story, though, as the real story behind all those "we got a 34% uplift" posts is a much better insight into CRO than just a success story.
It's a true story of one of our CRO retainer clients, the experiment hypothesis, deceptive early results, decisions that were taken based on them, and the conclusion our client got despite of zero changes implemented.
The Hypothesis
On Oliver Brown's collections pages, pricing was displayed in a light grey, non-serif font easy to overlook, and doing little to signal value or urgency. The hypothesis was simple: make the price visually stronger and see if it improves visibility, add-to-basket behaviour, and conversion.
Control: standard price styling, light grey, non-serif, low contrast.
Variant: enhanced price styling, darker, bolder, serif typeface.
This was about as low-risk a test as CRO gets. It needed no development work it could be built and run entirely inside the visual testing tool. Time estimate on the ticket: one hour to set up, one hour to monitor and report. The kind of test you'd expect to call quickly, one way or the other.
It took three weeks to actually get an answer.
What the First Run Told Us
The test launched across all devices on 28 May 2025, planned to run two weeks.
At the mid-test checkpoint 7 days in, 11,000+ sessions the topline read was flat. No statistically significant improvement across the full traffic pool.
But splitting the data by device told a different story in the desktop-only segment (3,633+ sessions):
Mobile sessions made up the majority of the traffic and were dragging the blended average down, potentially masking a real desktop-specific effect underneath.
The Decision Point
This is the moment that actually matters in the story, more than any single number.
We had two easy, wrong options available: ship the change based on the promising desktop cut and call it a win, or write the whole test off as a failure because the blended number was flat. Both would have been defensible on a surface read. Neither would have been honest.
Instead, we stopped the original test and restarted it isolated to desktop only specifically to find out whether the early desktop signal was a real effect or just noise from a small, early sample. The reasoning we gave the client at the time: mobile hadn't shown any positive movement, so continuing to blend it into the topline number would only keep hiding whatever was actually happening on desktop, in either direction.
Giving It a Proper Second Run
The desktop-only test restarted on 9 June, built as a clean, isolated read this time.
Eight days in (9–17 June, 3,936 sessions), the data was genuinely mixed worth showing, because this is closer to what real CRO data usually looks like before there's enough of it to trust:
- Product detail views: +0.25% uplift in the variant
- Average order value: +£49.01 uplift in the variant
- Conversion rate / add-to-basket rate: no uplift observed at that stage
Our recommendation at that point was to keep running it. There wasn't yet enough data to call it either way, and the early AOV and engagement signals were worth waiting on.
We let it run the full 14 days: 9–24 June, 7,000+ desktop sessions.
The final result: no significant improvement in transactions. The test was stopped and reverted to the original design.
What We Actually Learned
The end-of-test findings didn't just say "no effect" they came with a reason, which is what makes this worth more than a null result.
The read was that desktop shoppers on this site preferred the original, brand-consistent collections layout over the stronger price emphasis. Making pricing more prominent didn't just fail to help conversion it appeared to introduce friction, working against the calmer, premium browsing experience this audience expected from the brand.
That distinction a test that fails for a reason versus a test that just fails is the difference between wasted testing cycles and a useful one. This particular test gave the client:
- Confirmation, not assumption, that their current design instinct restrained, premium, uncluttered matches what their desktop audience actually responds to.
- A concrete redirect for future optimisation: focus on product storytelling and brand messaging on collections pages, not pricing prominence.
- Two process lessons for how we'd handle the next ambiguous read: extend test duration to 28+ days to reduce volatility from weekly shopping cycles, and pair quantitative test data with qualitative research (session recordings, heatmaps) to understand why a metric isn't moving, not just whether it is.
The Agency Perspective
A test that ends in "no change" isn't the same as a test that was wasted. It's only wasted if you weren't disciplined enough to run it properly before deciding what it meant.
We could have shipped the desktop pricing change off the first mid-test report the numbers looked good enough to justify it on paper. We didn't, because a positive read on 3,600 sessions across a few days isn't the same confidence level as a positive read on 7,000 sessions across two full weeks. Only the second one was reliable enough to act on, and it said something different.
The uncomfortable version of this story: we spent three weeks testing a pricing change that never shipped. The useful version: the client now knows, with real evidence rather than a guess, that their brand positioning is working as intended and we know exactly what to test next instead of pricing.
A Note on Scope
Everything above describes one test, on one client's collections pages, at one point in time. It isn't a universal claim about pricing display, serif fonts, or desktop shopper psychology it's what happened on this specific site, with this specific audience, in May and June 2025. That's intentional. The generic version of this article would be "how to run an A/B test properly." The actual value is in showing the mess: the segment cut that looked promising, the decision to isolate it rather than ship it, the eight-day report that still wasn't conclusive enough to call, and the discipline it took to wait for day fourteen anyway.
Want a CRO Program That Reports the Losses Too?
If your current testing partner only ever shows you wins, you're not getting the full picture and you're probably not learning as much as you think you are from your test program. We run CRO retainers that report honestly on what didn't move the needle, because those results usually tell you more about your customers than the ones that did. If you want a second opinion on your current testing roadmap, or you're not sure why your last few tests came back inconclusive, get in touch and we'll walk through what we'd test next.
Conclusion
The most valuable outcome of this test wasn't a conversion lift it was clarity. Oliver Brown now has evidence, not a guess, that their existing collections design is doing its job, and a clear next direction for where to focus optimisation effort instead. That's the actual point of running tests properly: not every test needs to win to be worth running, but every test needs to be run long enough, and honestly enough, to be worth trusting.

.png)

.png)
