ASO course · lesson 37 of 60

ASO A/B testing: metrics, guardrails and decision rules

Define what the experiment is meant to improve before looking at results.

By Gabriel Machuret · Editorial methodology

Before you start: A creative hypothesis and your metric dictionary.

Module 7: Experiment design and interpretation

Plan about 50 minutes for reading, practice and review

You will produce: A metric and decision-rule appendix to the experiment brief.

Understand the decision

Match the metric to the decision

Choose a primary outcome that the experiment can measure for the eligible population. Record the precise source definition and denominator. A metric that is easy to observe may still be the wrong decision metric: install-button clicks and completed acquisitions are not interchangeable. Keep the terminology used by the actual report and state what behavior it captures.

Use guardrails without inventing linkage

A guardrail protects against a harmful tradeoff, such as attracting people who cannot use the product. Activation, retained users or support issues may be useful when they can be observed appropriately. If variant-level downstream measurement is unavailable, say so. Aggregate product evidence can provide context but should not be presented as a precise treatment-level causal estimate.

Write the rule before the outcome

A decision rule explains the evidence needed to adopt, reject or investigate a treatment, including practical value and uncertainty. Do not simply say “whichever is higher.” Include operational stop conditions and the assumptions that would invalidate interpretation. Locking these choices before launch reduces the temptation to find a winning metric after the primary result disappoints.

Apply the method in practice

Choose the decision before the dashboard

Write what action the team will take if the treatment provides useful evidence. Then select the primary measure that best represents that decision within the platform's available reporting. A screenshot test may use a native acquisition outcome, while activation remains a downstream guardrail if it can be measured appropriately. Avoid selecting several coequal primary metrics and later promoting whichever looks favorable. Record numerator, denominator, eligible population and time window. A metric name without these details leaves room for incompatible interpretations.

Distinguish practical value from statistical evidence

A tiny improvement can be estimated precisely and still be commercially unimportant. A large observed improvement can be too uncertain to rely on. Before launch, discuss what magnitude would matter given traffic, production effort and maintenance. Use scenarios with explicit assumptions to understand scale, not as a promise. Keep the platform's statistical method separate from the business decision. Native confidence or probability displays should be read according to their documentation rather than translated into unsupported certainty language.

Define guardrails and exceptions

Name outcomes that could make an apparent acquisition gain undesirable, such as a verified increase in misleading expectations or a deterioration in product value. Some guardrails require longer observation or cannot be linked directly to experiment assignment; state that limitation. Decide who can pause a test for a factual error or technical problem. Operational safety checks differ from stopping because a result looks temporarily favorable. Preserve the agreed interpretation plan and log any justified deviation.

Platform references for this work: Apple — Product page optimization results · Google Play — Run store listing experiments

Your step-by-step procedure

  1. Name the primary metric, exact definition and eligible population.
  2. Select useful guardrails and document measurement limits.
  3. Define the minimum business-relevant effect as a planning assumption.
  4. Write adopt, reject, inconclusive and operational-stop decisions before launch.
1. Name the primary metric, exact definition and eligible population. 2. Select useful guardrails and document measurement limits. 3. Define the minimum business-relevant effect as a planning assumption. 4. Write adopt, reject, inconclusive and operational-stop decisions before launch.
Graphic 1. The procedure. Select the graphic to view or save the full-size version.

Worked example and interpretation

Trail Notes and numerical research scenarios are fictional teaching examples. Platform limits, where shown, come from the linked official references.

Trail Notes chooses the native experiment’s supported acquisition-related outcome and records its current definition. The team would prefer more activated journal users, but cannot link that event reliably to each variant. They therefore keep activation as contextual evidence, disclose the gap and avoid a promise that the experiment proves subscriber value. A malformed screenshot remains an immediate operational stop regardless of early numbers.

Measure type: Primary; Purpose: Supports the main decision; Example question: Did eligible acquisition improve?. Measure type: Guardrail; Purpose: Detects an important tradeoff; Example question: Did suitable users still reach value?. Measure type: Diagnostic; Purpose: Explains possible mechanisms; Example question: Did interpretation differ by audience?
Graphic 2. The worked example. Select the graphic to view or save the full-size version.
Example details
Measure typePurposeExample question
PrimarySupports the main decisionDid eligible acquisition improve?
GuardrailDetects an important tradeoffDid suitable users still reach value?
DiagnosticExplains possible mechanismsDid interpretation differ by audience?

The three types should not be substituted after results arrive. A diagnostic metric can generate a new hypothesis without proving the primary outcome improved. A guardrail with incomplete attribution can still inform caution, but its limitation belongs in the readout. The brief should say how much evidence the team requires for each decision and what remains unresolved if a measure is unavailable.

Build a handover someone can use

Decision: State the adoption or learning decision the experiment is intended to support. Primary definition: Specify exact metric, denominator, population and reporting source. Practical threshold: Explain what improvement would matter under explicit business assumptions. Guardrails: Record tradeoffs, measurement limitations and any operational stop conditions. Interpretation plan: Assign the reviewer and actions for supported improvement, decline or uncertainty.
Graphic 3. The handover. Select the graphic to view or save the full-size version.

Decision

State the adoption or learning decision the experiment is intended to support.

Primary definition

Specify exact metric, denominator, population and reporting source.

Practical threshold

Explain what improvement would matter under explicit business assumptions.

Guardrails

Record tradeoffs, measurement limitations and any operational stop conditions.

Interpretation plan

Assign the reviewer and actions for supported improvement, decline or uncertainty.

Your ASO assignment

Use your own app and evidence, or work through the teaching case. Keep your observations separate from assumptions and explain the reasoning behind your decisions.

  1. Complete the metric and decision-rule fields in the planner.
  2. Explain why one tempting secondary metric is not primary.
  3. Write a case where the result should be considered invalid or incomplete.

Scenario challenge

A treatment increases install clicks but the team cannot observe completed acquisitions by assignment. Write a conclusion that respects the measured outcome and identify the additional evidence needed before making a stronger claim.

Assess your work

Common mistakes to catch

Knowledge check

The primary metric is flat but a secondary metric rises. Can you declare the test won?

Read the answer and reasoning

Report the primary result honestly and treat the secondary movement according to the prespecified plan. It may generate a follow-up hypothesis, but should not silently replace the original decision rule. Consider uncertainty and multiple comparisons before making a broader claim.

Sources and further reading

Platform references reviewed 29 September 2026. These sources support platform capabilities and constraints; the teaching frameworks, assignments and illustrative cases are original course material.

Your ASO lesson notes

Use this space for your assignment, evidence and handover. Select Save to keep a draft in this browser, or download a copy to take with you.

Notes stay on this device and are not submitted to ASO Agency.

Complete your assignment with the ASO experiment planner.

Loading your progress…

Progress is saved only in this browser. It does not sync across devices. Mark a lesson complete after doing its exercise.