# Lesson 37: ASO A/B testing: metrics, guardrails and decision rules

Define what the experiment is meant to improve before looking at results.

Prerequisite: A creative hypothesis and your metric dictionary.

Teaching examples are fictional unless explicitly identified as platform definitions.

## Foundations

### Match the metric to the decision

Choose a primary outcome that the experiment can measure for the eligible population. Record the precise source definition and denominator. A metric that is easy to observe may still be the wrong decision metric: install-button clicks and completed acquisitions are not interchangeable. Keep the terminology used by the actual report and state what behavior it captures.

### Use guardrails without inventing linkage

A guardrail protects against a harmful tradeoff, such as attracting people who cannot use the product. Activation, retained users or support issues may be useful when they can be observed appropriately. If variant-level downstream measurement is unavailable, say so. Aggregate product evidence can provide context but should not be presented as a precise treatment-level causal estimate.

### Write the rule before the outcome

A decision rule explains the evidence needed to adopt, reject or investigate a treatment, including practical value and uncertainty. Do not simply say “whichever is higher.” Include operational stop conditions and the assumptions that would invalidate interpretation. Locking these choices before launch reduces the temptation to find a winning metric after the primary result disappoints.

## Apply the method

### Choose the decision before the dashboard

Write what action the team will take if the treatment provides useful evidence. Then select the primary measure that best represents that decision within the platform's available reporting. A screenshot test may use a native acquisition outcome, while activation remains a downstream guardrail if it can be measured appropriately. Avoid selecting several coequal primary metrics and later promoting whichever looks favorable. Record numerator, denominator, eligible population and time window. A metric name without these details leaves room for incompatible interpretations.

### Distinguish practical value from statistical evidence

A tiny improvement can be estimated precisely and still be commercially unimportant. A large observed improvement can be too uncertain to rely on. Before launch, discuss what magnitude would matter given traffic, production effort and maintenance. Use scenarios with explicit assumptions to understand scale, not as a promise. Keep the platform's statistical method separate from the business decision. Native confidence or probability displays should be read according to their documentation rather than translated into unsupported certainty language.

### Define guardrails and exceptions

Name outcomes that could make an apparent acquisition gain undesirable, such as a verified increase in misleading expectations or a deterioration in product value. Some guardrails require longer observation or cannot be linked directly to experiment assignment; state that limitation. Decide who can pause a test for a factual error or technical problem. Operational safety checks differ from stopping because a result looks temporarily favorable. Preserve the agreed interpretation plan and log any justified deviation.

## Procedure

1. Name the primary metric, exact definition and eligible population.
2. Select useful guardrails and document measurement limits.
3. Define the minimum business-relevant effect as a planning assumption.
4. Write adopt, reject, inconclusive and operational-stop decisions before launch.

## Worked example

Trail Notes chooses the native experiment’s supported acquisition-related outcome and records its current definition. The team would prefer more activated journal users, but cannot link that event reliably to each variant. They therefore keep activation as contextual evidence, disclose the gap and avoid a promise that the experiment proves subscriber value. A malformed screenshot remains an immediate operational stop regardless of early numbers.

| Measure type | Purpose | Example question |
| --- | --- | --- |
| Primary | Supports the main decision | Did eligible acquisition improve? |
| Guardrail | Detects an important tradeoff | Did suitable users still reach value? |
| Diagnostic | Explains possible mechanisms | Did interpretation differ by audience? |

The three types should not be substituted after results arrive. A diagnostic metric can generate a new hypothesis without proving the primary outcome improved. A guardrail with incomplete attribution can still inform caution, but its limitation belongs in the readout. The brief should say how much evidence the team requires for each decision and what remains unresolved if a measure is unavailable.

## Your ASO assignment

1. Complete the metric and decision-rule fields in the planner.
2. Explain why one tempting secondary metric is not primary.
3. Write a case where the result should be considered invalid or incomplete.

A treatment increases install clicks but the team cannot observe completed acquisitions by assignment. Write a conclusion that respects the measured outcome and identify the additional evidence needed before making a stronger claim.

## Handover

### Decision

State the adoption or learning decision the experiment is intended to support.

### Primary definition

Specify exact metric, denominator, population and reporting source.

### Practical threshold

Explain what improvement would matter under explicit business assumptions.

### Guardrails

Record tradeoffs, measurement limitations and any operational stop conditions.

### Interpretation plan

Assign the reviewer and actions for supported improvement, decline or uncertainty.

## Check your work

- The primary outcome is defined before results are observed.
- Guardrail claims match available measurement.
- Higher numerical performance alone is not the entire rule.

## Common mistakes

- Switching metrics until one appears positive.
- Claiming variant-level retention from an unlinked aggregate report.

## Knowledge check

The primary metric is flat but a secondary metric rises. Can you declare the test won?

Report the primary result honestly and treat the secondary movement according to the prespecified plan. It may generate a follow-up hypothesis, but should not silently replace the original decision rule. Consider uncertainty and multiple comparisons before making a broader claim.

## Sources

- [Apple — Product page optimization results](https://developer.apple.com/help/app-store-connect-analytics/acquisition/product-page-optimization) — Native result interpretation and Bayesian analysis.
- [Google Play — Run store listing experiments](https://support.google.com/googleplay/android-developer/answer/12053285?hl=en) — Experiment configuration and native interpretation guidance.

Platform references reviewed 29 September 2026.
www.asoagency.com
