ASO course · lesson 37 of 60
ASO A/B testing: metrics, guardrails and decision rules
Define what the experiment is meant to improve before looking at results.
By Gabriel Machuret · Editorial methodology
Before you start: A creative hypothesis and your metric dictionary.
Module 7: Experiment design and interpretation
Plan about 50 minutes for reading, practice and review
You will produce: A metric and decision-rule appendix to the experiment brief.
Understand the decision
Match the metric to the decision
Choose a primary outcome that the experiment can measure for the eligible population. Record the precise source definition and denominator. A metric that is easy to observe may still be the wrong decision metric: install-button clicks and completed acquisitions are not interchangeable. Keep the terminology used by the actual report and state what behavior it captures.
Use guardrails without inventing linkage
A guardrail protects against a harmful tradeoff, such as attracting people who cannot use the product. Activation, retained users or support issues may be useful when they can be observed appropriately. If variant-level downstream measurement is unavailable, say so. Aggregate product evidence can provide context but should not be presented as a precise treatment-level causal estimate.
Write the rule before the outcome
A decision rule explains the evidence needed to adopt, reject or investigate a treatment, including practical value and uncertainty. Do not simply say “whichever is higher.” Include operational stop conditions and the assumptions that would invalidate interpretation. Locking these choices before launch reduces the temptation to find a winning metric after the primary result disappoints.
Apply the method in practice
Choose the decision before the dashboard
Write what action the team will take if the treatment provides useful evidence. Then select the primary measure that best represents that decision within the platform's available reporting. A screenshot test may use a native acquisition outcome, while activation remains a downstream guardrail if it can be measured appropriately. Avoid selecting several coequal primary metrics and later promoting whichever looks favorable. Record numerator, denominator, eligible population and time window. A metric name without these details leaves room for incompatible interpretations.
Distinguish practical value from statistical evidence
A tiny improvement can be estimated precisely and still be commercially unimportant. A large observed improvement can be too uncertain to rely on. Before launch, discuss what magnitude would matter given traffic, production effort and maintenance. Use scenarios with explicit assumptions to understand scale, not as a promise. Keep the platform's statistical method separate from the business decision. Native confidence or probability displays should be read according to their documentation rather than translated into unsupported certainty language.
Define guardrails and exceptions
Name outcomes that could make an apparent acquisition gain undesirable, such as a verified increase in misleading expectations or a deterioration in product value. Some guardrails require longer observation or cannot be linked directly to experiment assignment; state that limitation. Decide who can pause a test for a factual error or technical problem. Operational safety checks differ from stopping because a result looks temporarily favorable. Preserve the agreed interpretation plan and log any justified deviation.
Platform references for this work: Apple — Product page optimization results · Google Play — Run store listing experiments
Your step-by-step procedure
- Name the primary metric, exact definition and eligible population.
- Select useful guardrails and document measurement limits.
- Define the minimum business-relevant effect as a planning assumption.
- Write adopt, reject, inconclusive and operational-stop decisions before launch.
Worked example and interpretation
Trail Notes and numerical research scenarios are fictional teaching examples. Platform limits, where shown, come from the linked official references.
Trail Notes chooses the native experiment’s supported acquisition-related outcome and records its current definition. The team would prefer more activated journal users, but cannot link that event reliably to each variant. They therefore keep activation as contextual evidence, disclose the gap and avoid a promise that the experiment proves subscriber value. A malformed screenshot remains an immediate operational stop regardless of early numbers.
| Measure type | Purpose | Example question |
|---|---|---|
| Primary | Supports the main decision | Did eligible acquisition improve? |
| Guardrail | Detects an important tradeoff | Did suitable users still reach value? |
| Diagnostic | Explains possible mechanisms | Did interpretation differ by audience? |
The three types should not be substituted after results arrive. A diagnostic metric can generate a new hypothesis without proving the primary outcome improved. A guardrail with incomplete attribution can still inform caution, but its limitation belongs in the readout. The brief should say how much evidence the team requires for each decision and what remains unresolved if a measure is unavailable.
Build a handover someone can use
Decision
State the adoption or learning decision the experiment is intended to support.
Primary definition
Specify exact metric, denominator, population and reporting source.
Practical threshold
Explain what improvement would matter under explicit business assumptions.
Guardrails
Record tradeoffs, measurement limitations and any operational stop conditions.
Interpretation plan
Assign the reviewer and actions for supported improvement, decline or uncertainty.
Your ASO assignment
Use your own app and evidence, or work through the teaching case. Keep your observations separate from assumptions and explain the reasoning behind your decisions.
- Complete the metric and decision-rule fields in the planner.
- Explain why one tempting secondary metric is not primary.
- Write a case where the result should be considered invalid or incomplete.
Scenario challenge
A treatment increases install clicks but the team cannot observe completed acquisitions by assignment. Write a conclusion that respects the measured outcome and identify the additional evidence needed before making a stronger claim.
Assess your work
- The primary outcome is defined before results are observed.
- Guardrail claims match available measurement.
- Higher numerical performance alone is not the entire rule.
Common mistakes to catch
- Switching metrics until one appears positive.
- Claiming variant-level retention from an unlinked aggregate report.
Knowledge check
The primary metric is flat but a secondary metric rises. Can you declare the test won?
Read the answer and reasoning
Report the primary result honestly and treat the secondary movement according to the prespecified plan. It may generate a follow-up hypothesis, but should not silently replace the original decision rule. Consider uncertainty and multiple comparisons before making a broader claim.
Sources and further reading
Platform references reviewed 29 September 2026. These sources support platform capabilities and constraints; the teaching frameworks, assignments and illustrative cases are original course material.
- Apple — Product page optimization results — Native result interpretation and Bayesian analysis.
- Google Play — Run store listing experiments — Experiment configuration and native interpretation guidance.
Your ASO lesson notes
Use this space for your assignment, evidence and handover. Select Save to keep a draft in this browser, or download a copy to take with you.
Notes stay on this device and are not submitted to ASO Agency.
Complete your assignment with the ASO experiment planner.
Loading your progress…
Progress is saved only in this browser. It does not sync across devices. Mark a lesson complete after doing its exercise.