What an A/B Test Can Prove
A/B testing proximity notifications differs from standard digital A/B testing in one fundamental respect: the environment is not controlled. On a website, every visitor who lands on a page sees one variant or the other under broadly similar conditions. In a physical space, the same zone might feel different at 9am on a Tuesday compared with 2pm on a Saturday, and the Bluetooth signal reaching a phone in a pocket behaves differently from one held in a hand.

The core idea remains the same. You serve two or more variants of a notification to comparable groups of visitors and compare outcomes. Where proximity testing diverges is in how you define those groups, how you account for physical variables, and how you interpret results when signal behaviour and footfall patterns introduce noise you cannot eliminate.
What you are actually testing
In a proximity context, the variables you can test fall into a few practical categories:
- Trigger logic. The same zone fires a notification at one metre in variant A and at three metres in variant B. This tests whether distance-to-trigger affects engagement, but it also changes who receives the notification, because a wider radius catches more passers-by.
- Timing and frequency. Variant A allows one notification per zone per visit; variant B permits a second notification if the visitor re-enters after ten minutes. This interacts directly with consent preferences and notification fatigue.
- Message framing. Different copy, different calls to action, different personalisation levels. This is the closest analogue to web A/B testing, but the physical context — what the visitor can see, hear and do at that moment — shapes whether the message makes sense.
- Zone boundaries. Two beacons with slightly different placement or transmit power, creating overlapping but distinct trigger areas. This tests whether a small physical shift changes which visitors are reached.
The critical point is that changing the trigger or the zone is not a clean test of the message. It changes the audience. If variant B fires at a wider radius, it reaches people who are walking past rather than stopping, which will almost always depress engagement rates. That is not a message failure; it is a targeting difference. Keeping trigger conditions identical while varying only the message content is the only way to isolate the copy effect.
Why sample sizes are harder to achieve
A retail store might see several hundred visitors per day who have Bluetooth enabled and have opted in to notifications. Of those, only a fraction will enter the specific zone during the test period, and only a fraction of those will interact with the notification. Unlike a high-traffic website where a test can reach statistical significance in hours, a physical pilot might need days or weeks to gather enough data — and by then, footfall patterns, staffing or stock levels may have shifted.
This does not make A/B testing pointless. It means you need to set realistic expectations, run tests for longer, and accept that some findings will be directional rather than conclusive.
Designing a Clean Proximity Experiment
Retail: entrance versus aisle triggers
A common retail test compares a notification fired at the store entrance with one fired deeper inside, near a specific category. The practical challenge is that entrance notifications reach everyone who opts in, while aisle notifications reach only those who walk that route. If you are testing message variants, hold the trigger point constant. If you are testing trigger points, accept that you are comparing two different audiences and frame the conclusion accordingly.
Seasonal stock, promotional displays and even queue lengths near the tested zone will affect results. Document what is physically present in the zone during the test window and note any changes.
Museums: exhibit content variants
Museums often want to test whether a short factual prompt or a question-driven prompt leads to more audio-guide starts. The physical variable here is dwell time: visitors who stop and read a physical label may be more receptive to a notification than those walking through. If you cannot control for dwell time, at least record it where your platform allows, so you can see whether engagement correlates with how long someone has been in the zone before the notification arrives.
Events: session-change notifications
At a conference, you might test two versions of a notification that fires when attendees leave a keynote and enter a lobby zone. Variant A lists the next three sessions; variant B highlights a single session with a brief speaker quote. The complication is that event schedules create synchronised footfall surges — hundreds of phones enter the zone within minutes of each other, which can strain beacon detection and create clustering effects in your data. Staggered session times across multiple rooms make cleaner tests than a single mass exodus.
Structuring the test
For any of these scenarios, a workable structure looks like this:
- Define one variable. Change only the element you want to learn about. If you change the message and the trigger distance simultaneously, you cannot attribute the result to either.
- Hold the environment steady. Run the test during comparable periods — same days of the week, similar times of day, no known events or stock changes that would skew footfall.
- Randomise by device, not by person. Most proximity platforms assign variants at the device level when the phone first enters a zone. Ensure your platform does not reassign variants on subsequent visits during the test, or you will contaminate the groups.
- Set a minimum run time. Decide your required sample size before starting and do not stop early because the numbers look favourable. Early stopping inflates false-positive rates.
- Record physical context. Note anything in or near the zone that changes during the test: a moved display case, a temporary sign, a closed entrance.
Interpreting Results Without False Certainty
Testing during atypical periods
Running an A/B test during a launch event, a school holiday week or a period of building works will produce data that does not generalise to normal operations. If you must test during an unusual period, label the results as specific to that context and retest during a standard week before making permanent changes.
Ignoring device and OS differences
Android and iOS handle Bluetooth scanning, background permissions and notification presentation differently. If your visitor split is 60:40 between operating systems and one variant happens to receive a disproportionate number of iOS devices, the OS difference can masquerade as a message difference. Check the device and OS distribution across your variants before drawing conclusions. Some platforms allow you to segment results by OS; if yours does not, the finding is less reliable.
Counting deliveries instead of interactions
A notification that is "delivered" to the operating system is not the same as one that is seen by the user. Lock-screen behaviour, notification grouping and silent delivery settings all affect visibility. Interaction — a tap, an expand, a swipe — is a more meaningful metric than delivery, though the numbers will be smaller. Be clear about which metric you are testing and do not present delivery rates as engagement.
Overlapping zones from neighbouring beacons
If a beacon in zone A has its transmit power set high enough that phones in zone B also detect it, visitors may receive notifications intended for a different area. This corrupts both the variant assignment and the relevance of the message. Before starting any A/B test, verify zone boundaries with a site survey and check that beacons are not bleeding into adjacent test areas.
Consent and frequency caps applying unevenly
If variant A is more aggressive — for example, allowing a second notification on re-entry — some users who previously opted in may revoke permissions during the test. This shrinks your test pool and introduces a self-selection bias: the remaining recipients are those who tolerate more notifications, which is not representative. Monitor opt-out rates by variant and stop the test if one variant causes a measurable increase in permission revocation.
Key checks before going live
- Confirm that variant assignment is random and sticky for the test duration.
- Verify that trigger conditions — distance, dwell time, frequency cap — are identical across variants unless the trigger itself is the variable under test.
- Check that zone boundaries do not overlap with other active zones or test areas.
- Ensure your analytics can segment results by variant, OS, time of day and day of week.
- Document the physical state of the zone at the start of the test.
- Agree a minimum run time and sample size before examining results.
- Have a rollback plan: if one variant causes noticeable visitor confusion or staff complaints, you need to be able to revert to a single notification immediately.
A/B testing in physical spaces will never offer the clean separation of a controlled web experiment. The value lies not in achieving statistical perfection but in learning whether a change in trigger logic, timing or message framing produces a directionally useful shift in behaviour — and understanding the physical and technical reasons why it might not.




