# Lesson 40: ASO A/B test results: uplift and statistical uncertainty

Distinguish a numerical difference from reliable and commercially useful evidence.

Prerequisite: Basic percentage arithmetic and a defined experiment outcome.

Teaching examples are fictional unless explicitly identified as platform definitions.

## Foundations

### Separate effect size from certainty

An observed rate difference is an estimate. Percentage-point change is the direct subtraction of rates; relative lift divides that change by the baseline. Neither calculation describes uncertainty. Small samples can produce large apparent differences by chance. A useful result report includes the observed effect, the uncertainty information supplied by the method and the business importance of plausible effects.

### Understand an illustrative interval

For two independent fixed samples with binary outcomes and adequate counts, a simple normal-approximation interval uses the difference plus or minus 1.96 times its estimated standard error. This is a teaching illustration, not a replacement for native Bayesian or sequential store methods. Assumptions, allocation and stopping behavior matter; the same numbers do not justify every statistical procedure.

### Interpret the decision range

If the uncertainty range includes both a meaningful loss and a useful gain, the current evidence may not support a confident adoption claim. Conversely, a tiny reliable effect may not justify expensive maintenance. Define practical importance before results arrive. Do not call “not decisive” the same as “no effect,” and do not read a confidence interval as a guarantee about every future audience.

## Apply the method

### Calculate absolute and relative changes separately

In a teaching example, 200 conversions from 1,000 observations is 20%, while 220 from another 1,000 is 22%. The absolute rate change is two percentage points. The relative lift is 2 ÷ 20 = 10%. The treatment produced 20 more conversions in these equally sized samples. All three statements are correct but answer different questions. Report original counts alongside rates so readers can see the sample size and avoid confusing percentage points with relative percentages.

### Understand what an uncertainty interval adds

For a simplified fixed-sample comparison of independent proportions, an approximate standard error for the difference is the square root of p1(1−p1)/n1 + p2(1−p2)/n2. Here it is about 1.82 percentage points. A normal-approximation 95% interval around the two-point difference is roughly −1.57 to +5.57 percentage points. The observed lift is compatible with both a decline and an improvement under this model. This calculation has assumptions; it does not reproduce Apple's Bayesian PPO analysis or every sequential experiment system.

### Connect the result to a business decision carefully

An interval spanning zero does not prove the treatments are identical. It means this analysis does not establish a clear direction at the stated level. Consider practical value, experiment validity and the cost of further learning. If the result is uncertain, choose a documented next action rather than rounding the point estimate into certainty. Do not repeatedly inspect a fixed-sample test and stop whenever a preferred threshold appears without an appropriate sequential design. Use the method that matches how the experiment was run.

## Procedure

1. Calculate baseline and treatment rates from explicit counts.
2. Report both percentage-point change and relative lift.
3. Read uncertainty using the actual experiment method and its assumptions.
4. Compare plausible outcomes with the prespecified business decision threshold.

## Worked example

In an illustrative fixed-sample example, control has 200 successes from 1,000 observations and treatment has 220 from 1,000. Rates are 20% and 22%: +2 percentage points, or +10% relative. The simple independent-proportion standard error is approximately 1.82 percentage points, giving an illustrative 95% interval of about −1.57 to +5.57 points. That range includes a loss. These arithmetic results must not be substituted for a native store’s different statistical analysis.

| Measure | Control | Treatment |
| --- | --- | --- |
| Conversions | 200 | 220 |
| Observations | 1,000 | 1,000 |
| Rate | 20% | 22% |
| Observed difference | 2 percentage points | 10% relative lift |

Using the stated approximation, the interval for the difference is about −1.57 to +5.57 percentage points. A report that says only “10% improvement” hides the uncertainty and the small absolute difference. A better readout includes counts, effect estimate, method, interval and decision. It also states whether assignment was valid and whether the calculation matches the platform's actual analysis. Arithmetic alone cannot repair a biased comparison.

## Your ASO assignment

1. Reproduce the example’s rate, point change and relative lift.
2. Explain what the illustrative range allows and does not establish.
3. Write a decision for an effect that is reliable but too small to justify its cost.

Explain why “the test proves a 10% lift” is too strong for the example. Then explain why “there is definitely no difference” is also unsupported. Use the interval and method assumptions in your answer.

## Handover

### Counts and effect

Show sample sizes, outcomes, rates, percentage-point change and relative lift.

### Method

Name the statistical approach and its assumptions, including whether it matches native reporting.

### Uncertainty

Preserve the interval or native probability measure without translating it into a false guarantee.

### Practical meaning

Explain the size of the effect in business terms under explicit assumptions.

### Decision

Choose rollout, rejection or further learning with the limitations visible.

## Check your work

- Percentage points and relative lift are correctly distinguished.
- The statistical method is matched to its assumptions.
- Inconclusive evidence is not mislabeled as proof of no effect.

## Common mistakes

- Declaring a winner because one rate is numerically larger.
- Repeatedly applying a fixed-sample threshold while stopping whenever it looks favorable.

## Knowledge check

Does +10% relative lift automatically justify rollout?

No. Assess the uncertainty, design validity, practical value, maintenance cost and guardrails. The example’s point estimate is positive, but its illustrative interval includes a negative effect. A native platform may use a different method, which must be interpreted on its own terms.

## Sources

- [NIST — Difference of proportions confidence interval](https://www.itl.nist.gov/div898/software/dataplot/refman1/auxillar/diffprop.htm) — Independent-proportion interval methods; distinct from native store experimentation.
- [Apple — Product page optimization results](https://developer.apple.com/help/app-store-connect-analytics/acquisition/product-page-optimization) — Native result interpretation and Bayesian analysis.

Platform references reviewed 29 September 2026.
www.asoagency.com
