L6 · Task completion and revenue

Conversion Rate Optimization: Tests That Start From the Query

Conversion rate optimization raises the share of visitors who finish the task their query expressed, using hypotheses from intent and tests sized before they launch.

CONVERSION PATHARRIVECOMPLETE

Key takeaways

  • Conversion rate optimization starts from the query, because the query states the task a page has to complete.
  • Each page role has one macro conversion and a few micro conversions, reported by query cluster rather than site-wide.
  • A conversion hypothesis names the cluster, the task, the failure point, the change, and the metric expected to move.
  • Sample size is fixed before launch from the baseline rate, the minimum detectable effect, significance, and power, and tests run in full weeks.
  • A test result is read once at the planned sample, against the primary metric, and checked against guardrails such as lead quality.

Conversion rate optimization is the practice of raising the share of visitors who complete a page’s task, starting from the query that brought them. This article defines what conversion rate optimization is, then covers measuring task completion with macro and micro conversions per query cluster, hypotheses from query intent, test design and sample size, and reading a test result. It then shows conversion rate optimization on holisticradar.com, where every conversion path is measured as a task. It closes with conversion rate optimization in CRO and analytics services, followed by CRO tools, low-traffic sites, and attribution.

What conversion rate optimization is

Conversion rate optimization is the practice of raising the share of visitors who complete a page’s task, where the conversion rate equals completed tasks divided by visitors for a defined segment and period.

A conversion rate is only meaningful when both halves of the fraction are defined. The numerator is a specific action, such as a submitted audit request or a purchase. The denominator is a specific population, such as organic sessions that landed on service pages in one month. A site-wide conversion rate mixes visitors who came to learn with visitors who came to buy, so it moves when the traffic mix changes even if no page got better.

Conversion rate optimization starts from the query because the query states the task. A visitor who searched “pricing” arrived to learn a price; a visitor who searched a definition arrived to understand a subject. A page converts when it completes the task the query expressed, and fails when it asks for an action the visitor did not come to take. Testing button colors on a page that answers the wrong question optimizes the least important variable.

In the Six-Layer Pyramid, conversion rate optimization works at L6, task completion and revenue. A weak layer caps every layer above it, so a page with weak page semantics at L4, or a topical map at L3 that sends the wrong queries to it, cannot be tested into converting. The test finds the limit; the fix may sit lower in the pyramid.

Measuring task completion

Task completion is measured with one macro conversion per page role and a small set of micro conversions that show progress toward it, all reported by query cluster.

A macro conversion is the action that proves the page finished its job. A micro conversion is a smaller action that shows the visitor is moving toward it: reaching the pricing section, starting a form, or opening a case study. Macro conversions measure outcome; micro conversions locate where progress stops.

Page roleMacro conversionMicro conversions
Outer articleBridge click into the core pageScroll past the summary, next-read click
Service pageAudit, call, or demo request submittedForm start, proof section viewed, pricing viewed
Case studyClick to the service the case provesResults table viewed
Pricing pageCheckout started or proposal requestedPlan comparison opened
Request formForm submitted and confirmedEach field completed, errors shown

In Google Analytics, the actions a business treats as completions are marked as key events, the name Google adopted in March 2024 for what were previously called conversions in Analytics. A standard property can mark up to 30 events as key events, and an Analytics 360 property up to 50. Key events can also be imported into Google Ads as conversions for bidding, which is why the same definition should be used in both places.

Events per query cluster require one join. Google Analytics does not report the organic search query behind each session, while Search Console reports queries by landing page. Tagging every page with its query cluster, as a content group or custom dimension, lets the conversion rate be reported per cluster: the share of sessions on a cluster’s pages that completed the cluster’s task. That report shows which part of the topical map produces qualified demand and which part produces traffic without action.

Hypotheses from query intent

A conversion hypothesis from query intent names the query cluster, the task its queries express, the point where the page fails that task, the change, and the metric expected to move.

  1. Cluster: the group of queries landing on the page, taken from Search Console.
  2. Task: what those queries ask the page to do, such as compare two options or state a price.
  3. Failure point: where qualified visitors stop, found in scroll depth, form analytics, or session review.
  4. Change: one change that addresses the failure point.
  5. Metric: the macro conversion expected to move, and the minimum effect worth detecting.

A written example: “Visitors from the comparison cluster land on the service page, scroll to the feature list, and leave before the form. The page answers what the service is but never compares it with the alternative they searched. Adding a comparison table above the feature list will raise audit requests from that cluster.” Every part can be checked, and a failed test still teaches something specific about the cluster.

Intent class predicts the likely failure. Definitional queries landing on a service page usually signal a reader not yet ready for the primary action, so the fix is often routing rather than persuasion. Comparative queries fail when the page never names the alternative. Transactional queries fail when the price, the next step, or the time to delivery is missing near the action. A hypothesis that ignores the intent of the arriving traffic tends to test changes the visitors were never going to notice.

Test design and sample size

An A/B test needs a sample size fixed before launch from three inputs: the baseline conversion rate, the minimum detectable effect, and the significance and power levels.

The minimum detectable effect is the smallest difference the test is designed to find. Significance is the accepted chance of declaring a difference when none exists, conventionally 5 percent. Power is the chance of detecting a real difference of the planned size, conventionally 80 percent. Evan Miller’s guidance on how not to run an A/B test gives a rule of thumb for those conventional levels: the sample per variant is about 16 times the variance divided by the square of the minimum detectable effect, where the variance of a conversion rate p is p times (1 minus p).

InputWhat it setsExample value
Baseline conversion rateThe variance of the metric3%
Minimum detectable effectThe smallest change worth finding0.6 percentage points (a 20% relative lift)
Significance levelAccepted false positive rate5%
PowerChance of detecting a real effect of that size80%
Sample per variantWhen the test is readAbout 12,900 visitors

The example works out as 16 times 0.0291, divided by 0.000036, which is about 12,900 visitors per variant and about 25,900 for a two-variant test. Halving the minimum detectable effect quadruples the sample, which is why small sites cannot detect small changes.

Test duration is set in full weeks. Conversion rates differ by day of week, so a test that runs ten days over-represents some days in both variants and reflects that week’s mix rather than a normal one. A test runs at least one full week and continues in whole weeks until the planned sample is reached in every variant.

The sample size is fixed in advance because looking at results repeatedly and stopping at the first significant reading invalidates the significance level. Miller shows that in one such scenario, a nominal 5 percent significance level becomes an actual false positive rate of 26.1 percent. Teams that need to look early use a sequential or Bayesian design built for interim checks, rather than checking a fixed-sample test.

Reading a test result

A test result is read once, at the planned sample size, against the primary metric named in the hypothesis, and then checked against guardrail metrics and segments.

  1. Check the traffic split. If a planned 50/50 split arrived as noticeably uneven, assignment or tracking is broken, and the result is not read until that is fixed.
  2. Read the primary metric. Record the difference, its confidence interval, and whether it clears the significance level set before launch.
  3. Check guardrails. A rise in form submissions with a fall in qualified leads, or in revenue per visitor, is not a win.
  4. Check segments without hunting. Confirm the effect holds on the device types and clusters that carry most traffic; do not search dozens of segments for one that wins.
  5. Apply the result where it generalizes. A winning change to one template is rolled out to every page in that role and recorded in the brief.

A flat result is a finding, not a failure. It means any effect is smaller than the minimum detectable effect, so the change is not worth its complexity at that traffic level, or the failure point was elsewhere. A loss is equally useful: it shows the visitors of that cluster needed the element the variant removed.

Results also decay. An early lift can come from novelty, as returning visitors react to anything new. Comparing the first and second halves of the test shows whether the effect held or faded.

Conversion rate optimization on holisticradar.com

Holistic Radar measures every conversion path on holisticradar.com as a task with a defined start and a defined completion: the Radar Scan request, the client questionnaire, and the contact enquiry.

Conversion pathTask startTask completionQualified outcome
Radar Scan requestRequest form startedRequest submittedScan delivered within five business days and walked through
Client questionnaireQuestionnaire openedQuestionnaire submittedEnough site and goal detail for a strategist to start work
Contact enquiryContact page reachedEnquiry submittedA conversation with a business the team can serve

Each path is tracked once, at template level, so every page that offers the Radar Scan reports the same event, and the completion rate can be compared across page roles and query clusters. The outer articles follow the same rule: an article completes its task when the reader takes the single bridge link into the core page, and that click is the article’s macro conversion.

Defining the task first changes what gets tested. A request form that collects more submissions but fewer complete site details has not improved, because the qualified outcome, a scan a strategist can actually write, has fallen. The questionnaire is judged by whether it gives the strategist what is needed, not by how fast it is finished.

Conversion rate optimization in CRO and analytics services

Running one well-designed test is a technique; running a program where every result changes a template, a brief, or the topical map is a measurement system. The system needs a definition of task completion per page role, events implemented once across the site, and reporting that ties each completion back to the query cluster that produced the visit.

Holistic Radar’s CRO and analytics services build that system: a measurement architecture, a funnel and event map, landing page diagnostics, an experiment backlog ranked by the demand each test affects, and query-to-revenue reporting. Results are reported monthly by query cluster, with the Holistic Authority Score recalculated quarterly beside revenue.

→ CRO and analytics services: task completion and revenue measured by the query cluster that produced them

Conversion rate optimization tools, low-traffic sites, and attribution

Which tools does conversion rate optimization need?

Conversion rate optimization needs an analytics platform with event tracking, a tag manager, Search Console for query data, a testing tool, and session recording or form analytics for diagnosis. The tools matter less than the definitions behind them: an event that means different things on different templates makes every tool report noise.

How does conversion rate optimization work on a low-traffic site?

A low-traffic site tests larger changes, measures micro conversions with higher base rates, and leans on qualitative research. Because sample size grows with the square of the effect, a site with a few thousand monthly visitors can detect a change in page structure but not a change in a button label. Before-and-after comparisons are a fallback, read with care because seasonality and traffic mix change at the same time.

Which attribution model should conversion rate optimization use?

Conversion rate optimization uses the attribution model the business already reports revenue with, applied consistently, and judges tests on the page where the task completed. Google Analytics retired its first-click, linear, time-decay, and position-based models in 2023, leaving data-driven and last-click models. Attribution decides which channel gets credit; it does not change whether a page completed its task.

→ Conversion Copywriting: Frameworks That Answer the Buyer’s Next Question: how the copy that tests change is written in the first place

→ The Holistic Authority Score: the coverage measure reported beside conversion and revenue

→ Email Lifecycle Stages: From First Signup to Repeat Revenue: how completed tasks become repeat revenue after the first conversion

Questions readers ask

Frequently asked questions

What is a good conversion rate?

A good conversion rate is one that improves for a defined page role and query cluster over time, because site-wide averages mix visitors with different tasks and cannot be compared across businesses.

What is the difference between a macro and a micro conversion?

A macro conversion is the action that proves a page finished its task, such as a submitted request, while a micro conversion is a smaller step toward it, such as starting the form.

How long should an A/B test run?

An A/B test runs in whole weeks, at least one, until every variant reaches the sample size calculated before launch.

Can an A/B test be stopped early when it looks significant?

A fixed-sample A/B test should not be stopped at the first significant reading, because repeated checking inflates the false positive rate far above the stated level.

What are key events in Google Analytics?

Key events are the events a business marks as important completions in Google Analytics, the name used since March 2024 for what were previously called conversions.

Sourcing

Sources

Measurement definitions are drawn from Google Analytics Help documentation, test statistics from Evan Miller's published guidance, and the task model from Holistic Radar engagement practice.

  1. Google Analytics Help, About key events
  2. Evan Miller, How Not To Run an A/B Test (2010)

Author

Founder · Semantic SEO & Conversion Architecture

Bilal Sameer founded Holistic Radar and leads its topical map strategy: central entity, source context, and the query network behind every map. His work sits inside the Holistic Radar team, alongside the engineers, reviewers, and designers who deliver each engagement.

View the full profile

Get my free Radar ScanFree Radar Scan