Google Play Store Listing Experiments: A/B testing
On Google Play, you don't have to guess whether a new first screenshot converts better than the old one — you can test it on real Play Store traffic. Store Listing Experiments are built into Play Console under Test and release → Store listing experiments (Google has shuffled the menu location more than once; older guides call it "Grow → Store listing experiments"). They're free, they run on your live listing, and they handle the traffic split and the statistics for you.
If you ship to both stores, this is the counterpart to Apple's setup. We covered the iOS side in A/B testing App Store screenshots with Product Page Experiments — same idea, different mechanics. This post is the Play-side complement.
How Store Listing Experiments work
You create an experiment, choose which asset to test (app icon, screenshots, feature graphic, preview video, short description, or long description), and upload up to 3 variants alongside your current live listing — which acts as the control. Google then serves the variants to a slice of users who land on your store page and measures which one drives more installs.
Two metrics matter in the results:
- Acquisition (store listing conversion). The share of visitors who install after seeing a given variant. This is the headline number for screenshot tests.
- Retained first-time installers (1-day retention). Whether the people a variant attracted actually stuck around a day later. A variant that pulls more installs but worse retention may have over-promised.
Google computes a confidence interval per variant and, when a variant clears your chosen confidence threshold, declares it the winner. You can opt into an email alert so you don't have to babysit the dashboard. Applying a winner is one click — but you usually still want to push the winning asset through your normal release/review flow so the rest of your listing stays consistent.
Default graphics vs localized experiments
This is the part people trip over, and it changes how many tests you can run in parallel.
- Default graphics experiment. Tests the assets on your default store listing — the listing shown to anyone who isn't served a localized version. Users who are served localized assets are excluded from this experiment's audience. You can run only one default graphics experiment at a time.
- Localized store listing experiment. Tests assets (and/or text) for a specific language. You can run up to 5 localized experiments simultaneously — for example, one for German, one for Japanese, one for Brazilian Portuguese, and so on, all at once.
Practically: if your install base is concentrated in a handful of markets, localized experiments let you test those markets in parallel instead of queuing behind a single global test. And don't assume a language equals a country — selecting "English (United States)" doesn't restrict the audience to the US; it targets everyone served that localized listing, wherever they are.
Setting up a screenshot A/B test
- Open Store listing experiments and create one. Choose the default listing or a specific localized listing, and give it a name you'll still understand in three weeks ("Hero screenshot — benefit-led caption v2").
- Pick the asset. Select screenshots. Test one element per experiment — don't change the icon and the screenshots in the same test, or you won't know which moved the needle.
- Upload up to 3 variants. Your live listing is the control. Each variant is a full screenshot set, at the correct Play dimensions — see the Google Play screenshot size guide for the exact specs.
- Set the audience, confidence, and MDE. Allocate traffic across variants (commonly even — 50/50 for one variant vs control, or roughly a third each for three). Pick a confidence level (90%, 95%, 98%, or 99%) and a minimum detectable effect — the smallest improvement worth detecting, configurable roughly 0.5%–6%. Play Console shows the completion conditions so you can see what you're committing to before you start.
- Start it and leave it alone. Resist the urge to peek and stop early the moment a variant looks ahead. Let it run to the conditions you set.
Sample size, duration, and reaching significance
The honest answer to "how long should I run it?" is: until it reaches the significance conditions you set — not a fixed number of days. But there are real-world floors and ceilings.
- Run at least 7 days. Install behavior swings between weekdays and weekends. Anything shorter than a full week bakes day-of-week bias into your result. Two weeks (14 days) is a common, safer default, and low-traffic apps often need 28 days.
- Enough installs per variant to matter. Significance is a function of your install volume, the number of variants, your confidence level, and your MDE. As a rough working target, aim for on the order of 1,000+ installs per variant before you trust the call — more if you set a high confidence level or a small MDE.
- Tighter settings cost traffic. A 99% confidence level or a 0.5% MDE needs far more installers than 90% / 3%. If your app is low-volume, a demanding configuration may never reach significance — loosen the MDE or accept 90%.
A test that ends "inconclusive" is a real outcome, not a failure. It usually means the variants were too similar to separate at your traffic level, or you didn't run long enough. Both are fixable — make the variants more distinct, or give it more time.
The variants limit and what it implies
Three variants against the control is the hard cap per experiment. That's a feature, not a constraint to fight: every extra variant splits your traffic thinner and pushes significance further out. With three challengers you're already dividing your store traffic four ways (control + 3). For most apps, testing one bold, clearly-different challenger against the control reaches a conclusion faster than four mushy variations ever will.
Use the slots for genuinely different hypotheses — a different lead benefit, a different visual style, portrait vs a wide multi-panel layout — not for shades of the same idea.
Pitfalls to avoid
- Changing more than one variable. New screenshots and a new short description in one test = an uninterpretable result. Isolate the variable.
- Stopping early on a "winner". Early leads regress. Calling it on day 3 because a variant is up 8% is how you ship a worse listing with confidence.
- Running under 7 days. You'll catch one part of the weekly cycle and miss the rest.
- Over-tight stats on a low-traffic app. Demanding 99% / 0.5% MDE with modest installs guarantees an inconclusive, expensive-in-time test.
- Ignoring 1-day retention. A screenshot that over-promises can lift installs and tank retention. Watch both numbers.
- Forgetting the default/localized split. A default graphics experiment doesn't touch users on localized listings — if your audience is mostly localized, test there.
- Only ever testing the last screenshot. The first 2–3 frames do most of the convincing on the Play Store. Test those first. (More on how Play differs from Apple here: Play Store vs App Store screenshot differences.)
Where Mokbi fits
Mokbi doesn't run the A/B test — Play Console does that, and it's free. What slows most teams down is producing the variants in the first place: a credible challenger means a full screenshot set, at the right Play dimensions, with the caption and layout actually changed. In Mokbi you design one set in the browser, then duplicate it and swap the lead caption, the layout, or the background to make variant B — and one-click translate the captions if you're running localized experiments across markets. Design is free; a subscription (Solo €29.99/mo or Studio €49.99/mo) unlocks the final, store-ready sets to upload. It builds the contenders fast; Play Console decides which one wins.