Skip to content

Insights · Growth · Oct 5, 2025 · 7 min read

CRO beyond button colours: how to run experiments that matter

Button-colour tests are not CRO. A practical guide to research-led conversion optimisation: session replays and interviews, hypothesis discipline, test math for smaller sites, and the form and trust fixes that matter.

Attribution · this quarter

GA4 · server-side
Direct
Paid
Social
Organic
Email

312

Booked calls

3.1x

ROAS

−31%

CPL

Conversion rate optimisation is a research discipline, not a colour palette. The programmes that actually move revenue start with evidence — session replays, customer interviews and analytics that explain why visitors hesitate — and only then test the changes that evidence supports. Everything else, including the famous button-colour experiment, is decoration.

Key takeaways

  • Serious CRO begins with research: watch real sessions, read support tickets and interview customers before touching the page.
  • A hypothesis worth testing names the observed problem, the proposed change, the metric it should move and the reasoning that connects them.
  • Test math is unforgiving on smaller sites: modest traffic can only detect large effects, so test bold changes or do not test at all.
  • Obvious friction — broken fields, surprise costs, missing trust signals — should simply be fixed, not queued behind an experiment.
  • Forms and trust cues decide more conversions than headlines do, and they deserve first-class attention.

Why do most CRO programmes stall at surface tweaks?

The typical pattern looks like this: a team reads a case study about a colour change or a headline swap, copies the tactic, runs it for two weeks, sees nothing conclusive and quietly loses faith in testing. The problem is not the tool or the traffic. It is that the test was never connected to a diagnosed problem on that specific site, for that specific audience.

Copied tactics fail because context does not copy. A change that worked for a marketplace with enormous traffic tells you almost nothing about a regional B2B site that gets a handful of enquiries a month. Their visitors arrive with different questions, different objections and different levels of trust. Until you know what your visitors are actually struggling with, every test is a guess wearing a lab coat.

The fix is to invert the order of work. Research first, hypothesis second, test (or fix) third. That sequence sounds obvious written down, yet most optimisation effort still starts at step three.

What does real conversion research look like?

Good research triangulates at least three sources, because each one lies in its own way.

  • Session replays and heatmaps show behaviour: where people scroll, hesitate, rage-click and abandon. Watch replays of sessions that reached a key page and left without converting. Patterns emerge quickly — a filter nobody finds, a field that eats input on mobile, a price revealed too late.
  • Analytics and funnel data show where the losses concentrate. A funnel report cannot tell you why step two bleeds visitors, but it tells you where to point the replays and which pages deserve interviews. Clean measurement is a precondition here; if your events are unreliable, fix analytics and CRO foundations before trusting any number.
  • Customer interviews and support tickets supply the language and the objections. A handful of honest conversations with recent buyers — and, harder but more valuable, with people who almost bought — will surface doubts no heatmap can show: unclear delivery terms, missing reassurance about cancellation, uncertainty about whether the product fits their situation.

When all three sources point at the same page or the same doubt, you have found something worth acting on. When they disagree, you have found something worth investigating further. Either outcome beats guessing.

How do you write a hypothesis worth testing?

A disciplined hypothesis has four parts: the evidence, the change, the expected effect and the metric. A useful template: because we observed X in research, we believe changing Y will improve Z for this audience, and we will judge it by this metric over this period.

The evidence clause is the one teams skip, and it is the one that matters most. It forces every proposed test to trace back to something a real visitor did or said. It also makes losing tests useful: when a change fails, you learn that the diagnosis or the remedy was wrong, and you can revise one without discarding the other.

A test without a hypothesis is not an experiment. It is a coin flip with a dashboard.

Prioritise hypotheses by three qualitative questions: how strong is the evidence, how large could the effect plausibly be, and how expensive is the change to build? A weakly evidenced, cheap, potentially large change may still be worth trying; a weakly evidenced, expensive, marginal one never is.

What does test math mean for a smaller site?

Here is the reality most CRO content avoids. The sample size an A/B test needs depends on your baseline conversion rate and the smallest effect you want to detect, and the relationship is punishing: roughly speaking, halving the minimum detectable effect quadruples the sample you need. Sites with modest traffic and modest conversion volume can therefore only detect large effects in any reasonable time.

Run the numbers through any standard sample-size calculator before you build anything. If the answer says a test needs to run for most of a year to detect the lift you hope for, that is not a failure — it is a decision. It tells you that for that page, controlled testing is the wrong instrument, and evidence-led fixing is the right one.

Two related traps deserve a mention. Peeking — checking results daily and stopping the moment a variant pulls ahead — inflates false positives badly, because early leads are usually noise. And running many small tests at once on overlapping traffic muddies attribution. For smaller sites the honest playbook is: fewer tests, bolder changes, longer patience, and a written stopping rule agreed before launch.

When should you skip the test and just fix it?

Not every improvement needs a control group. Testing exists to resolve genuine uncertainty about visitor behaviour; where there is no genuine uncertainty, testing is theatre that delays a fix. The table below is the decision guide we use in practice.

SituationEvidence you typically haveRight moveWhy
Broken or confusing element (field errors, dead links, layout faults on mobile)Replays, error logs, support complaintsFix immediatelyThere is no plausible world where the broken version wins
Clear friction against well-established practice (surprise costs late in checkout, forced account creation)Funnel drop-off plus interview objectionsFix, then monitor before/afterUncertainty is low; a test mostly delays the benefit
Contested change where reasonable people disagree (new value proposition, restructured pricing page)Mixed research signalsTest, if traffic allowsGenuine uncertainty is exactly what experiments are for
High-stakes change on a revenue-critical, high-traffic pageStrong hypothesis, meaningful volumeTest with a pre-agreed stopping ruleThe downside of shipping a loser untested is real
Minor copy polish on a low-traffic pageEditorial judgementJust ship itThe test could never reach a conclusion anyway

Monitoring before-and-after is weaker evidence than a controlled test — seasonality and campaigns muddy it — but for unambiguous friction it is evidence enough, and it costs nothing in delay.

Why do forms and trust decide the sale?

Most conversion journeys end at a form, and forms are where good intentions go to die. The recurring offenders we see in audits: too many fields asked too early, labels that vanish on focus, error messages that scold without explaining, phone-number fields that reject valid formats, and mobile keyboards that do not match the input type. Each one is invisible in a design review and obvious in a session replay.

Trust is the quieter half of the same problem. Visitors near the point of commitment look for reasons to believe: a real business address, plain-language privacy and refund terms, recognisable payment reassurance, photographs of actual people and work rather than stock imagery, and a promise about response time that the business visibly keeps. None of this is glamorous, and much of it is build work rather than marketing work — which is why form and checkout fixes often route through web development rather than another round of copy edits.

Accessibility belongs here too. Forms that meet WCAG 2.2 AA — proper labels, visible focus, errors announced to assistive technology — are easier for everyone, and the fixes overlap almost entirely with conversion fixes.

How OlDevs approaches CRO, and how to start

OlDevs has built and optimised conversion journeys since 2014, from our studio in Vancouver, for clients across Canada and beyond. Analytics and CRO is one of the disciplines inside our performance marketing practice, and because we are one accountable team of developers and marketers, the same engagement can diagnose the friction, rebuild the form and measure the result — no hand-offs between agencies. You see a working demo every week, we follow the research-first sequence described above, and you own all code, designs, accounts and IP outright.

If your site gets traffic that does not turn into enquiries or orders, we would rather examine your evidence than sell you a test. Request a quote through our contact page and tell us what your funnel looks like today; we reply to every enquiry within one business day.

FAQ

Questions on this topic.

For most sites the honest answer is: as long as a sample-size calculation says, decided before launch. Set a stopping rule in advance and resist checking daily, because early leads are usually noise. If the calculation says the test would take many months, fix the page based on research evidence instead of testing.

Yes, but differently. Modest traffic can only detect large effects, so small sites should run fewer, bolder experiments and rely more on research: session replays, customer interviews and before-and-after monitoring. Clear friction such as broken fields or surprise costs should simply be fixed without a test.

Start with your forms and trust signals. Watch session replays of visitors who reached the form and left, then remove unnecessary fields, repair confusing error messages, and add plain-language reassurance about delivery, refunds and privacy. These fixes usually beat another round of headline or colour changes.

Still have a question? Ask us when you request a quote

Let’s connect

Want this applied to your business?

Tell us what you’re building. We’ll reply within one business day with next steps and a tailored quote.

We’ll only use your details to prepare your quote. No lists, no spam.

Call us Request a quote