What must be ready before the work starts
Evaluating a proximity technology pilot means comparing what actually happened against what you set out to test. That sounds straightforward, but in practice pilots generate noisy data, and the temptation to read success into ambiguous results is strong—particularly when budget, timelines and internal expectations are pushing towards a rollout decision.

The first discipline is to return to the pilot objectives written before installation. If the objective was to test whether beacons could reliably trigger a notification at a specific entrance zone, the evaluation centres on detection consistency and trigger accuracy, not on how many people redeemed an offer. If the objective was to measure dwell-time changes near a museum exhibit, the evaluation needs a baseline and a statistically meaningful sample, not a handful of observations from a quiet Tuesday morning.
Pilot evaluation sits between two neighbouring activities. Monitoring, covered separately, is the live process of checking that hardware is broadcasting and the system is functioning during the pilot period. Procurement, also covered separately, is what follows if the evaluation produces a positive scale decision. Evaluation itself is the structured judgement: given what the pilot was designed to test, what did it actually prove, what remains uncertain, and what conditions would need to change for a full deployment to be justified?
A useful framing is to separate technical performance from behavioural and operational observations. Technical performance covers whether the signals behaved as expected in the physical environment. Behavioural observations cover how visitors or customers interacted with the system. Operational observations cover what your team had to do to keep the pilot running. All three matter, but conflating them leads to poor decisions.
Run the work in verifiable stages
Retail entrance-zone pilots
In a retail pilot, a common test is whether beacons can detect devices at a defined entrance zone and trigger a welcome notification with sufficient reliability. Evaluation here starts with the detection log: what proportion of known test devices were detected when they crossed the threshold, and how consistent was the RSSI reading at that point?
Next, examine the notification delivery chain. Detection is not the same as delivery. If the mobile app was backgrounded, did the operating system suppress the notification? What proportion of detections resulted in a visible alert on the test devices? This distinction matters because it is a platform constraint, not a beacon placement problem, and it will not be solved by adding more hardware.
Behavioural data in retail pilots is often thin because sample sizes are small and the pilot period is short. If the objective included measuring conversion from notification to purchase, check whether the tracking mechanism actually connected the notification event to the transaction. In many pilots this link is assumed rather than verified, and the evaluation should flag that gap honestly rather than presenting it as a result.
Museum exhibit-trigger pilots
Museum pilots typically test whether a visitor standing near a specific exhibit receives the correct content. The evaluation should examine trigger accuracy: did visitors at Exhibit A receive content for Exhibit A, or did signals from adjacent beacons cause misfires?
Look at the RSSI logs for the exhibit zone and compare them against the trigger threshold configured in the system. If the threshold was set at an assumed value rather than a calibrated one, the pilot results will tell you whether that assumption held. Document the variance: if readings at the same physical spot swung widely over the pilot period, the environment has an interference problem that calibration alone may not fix.
Also assess content-load timing. If the exhibit content took several seconds to appear after the trigger, note whether that delay affected visitor behaviour—did people walk away before the content loaded? This is an operational observation with direct implications for the content strategy if the system scales.
Event temporary-infrastructure pilots
Events present a particular evaluation challenge because conditions change rapidly. A beacon system tested during setup may behave differently when the venue is full of people, because bodies attenuate Bluetooth signals. If the pilot only ran during low-occupancy periods, the evaluation should explicitly state that peak-condition performance was not tested.
Check the physical condition of hardware after the event. Temporary installations are handled differently from fixed ones, and if beacons were knocked, repositioned or had their adhesive fail, that is a legitimate finding about installation method, not a reason to dismiss the technology.
Operational burden
Regardless of the use case, record what your team actually did during the pilot. How many site visits were needed to check or reposition beacons? Did any units need battery replacement during what was supposed to be a short test? Was the asset register accurate enough to locate specific units when a problem was flagged? These observations are easy to overlook when the focus is on detection rates, but they determine whether a scaled deployment is operationally viable.
Re-test, maintain and improve
Treating pilot reach as production reach
A pilot that detected 200 unique devices in a week does not imply that a full deployment will reach 200 devices per week in perpetuity. Pilot populations are often unrepresentative: they may include staff devices, repeat visitors who are curious about the new system, or people who opted in specifically because they were asked. When evaluating reach, note the opt-in rate among approached visitors and the proportion of total footfall that represents. If only a small fraction opted in, the pilot tested the technology, not the audience.
Ignoring consent-rate effects
In UK deployments, consent is a legal requirement under UK GDPR and the Privacy and Electronic Communications Regulations. A pilot that achieved a high opt-in rate because visitors were personally approached by a researcher will not replicate that rate when consent is sought through an app onboarding flow. Evaluate the consent mechanism separately from the technology performance, and be explicit about whether the pilot consent process was realistic for production.
Confusing correlation with causation
If dwell time increased near an exhibit during the pilot week, that does not mean the beacon system caused the increase. The exhibit may have been featured in external marketing, the weather may have driven more visitors indoors, or the pilot week may have coincided with a school holiday. Without a control zone or a comparable prior period, the evaluation should describe the observation without attributing causation.
Overlooking environmental changes mid-pilot
Were any physical changes made to the space during the pilot? New signage, moved fixtures, closed corridors, or even stacked stock can alter signal propagation. If the environment changed and the detection logs shifted at the same time, the evaluation needs to note that the results reflect two variables, not one.
Key checks before recommending a scale decision
- Objective alignment: Does the evidence directly address what the pilot was designed to test, or does it answer a different question?
- Sample adequacy: Is the volume of detections and interactions sufficient to draw a conclusion, or is the data too sparse to generalise from?
- Environmental representativeness: Did the pilot run under conditions that match normal operations, including peak occupancy and typical physical layout?
- Consent realism: Was the opt-in process comparable to what a production system would use, or was it artificially assisted?
- Operational feasibility: Can your team sustain the maintenance burden observed during the pilot at a larger scale, or would additional resources or processes be needed?
- Known unknowns: Has the evaluation clearly listed what the pilot did not test, so that the scale decision accounts for remaining risks?
If the evaluation produces a clear positive signal across these checks, the natural next step is to carry the findings into a structured procurement process. If the signal is mixed, the honest outcome may be a refined pilot with adjusted placement, calibration or consent mechanics rather than an immediate rollout or cancellation.


