The brief, at an XR agency in Paris where I led UX/UI design on this: capture the full interior of a user's mouth using only a smartphone camera, while they look into a bathroom mirror, with no dentist, no technician, and no training session.
The problem: no mental model for what "correct" looks like
The scanning method required filming specific angles and positions in a precise sequence, each held steady for several seconds at the right distance. Early tests made the core issue obvious: first-time users had no reference for what a correct scan angle even looked like from inside their own mouth. They moved too fast, missed entire zones, or held the phone at the wrong angle entirely. On-screen instructions got ignored the moment filming started. Voice prompts competed with running water or simply went unheard. Guidance had to work while users were focused entirely on their own reflection.
The solution: spatial markers plus haptic timing
I combined visual cues visible in the mirror with timed vibration feedback. The front camera captured the user's face and mirror reflection simultaneously, with semi-transparent geometric overlays that only aligned correctly once the phone was at the right distance and angle. When alignment locked, the marker changed color and the phone vibrated in a distinct pattern, a short pulse for hold steady, a double pulse for move to the next position, a long vibration for slow down.
![]()
The full scan broke into eight positions across upper arch, lower arch, and key bite angles, each with its own spatial target that only appeared once the previous position held for the required duration. I avoided anatomical diagrams or dental terminology entirely; people don't think in those terms while brushing their teeth, so the markers used simple shapes matching the general curvature of the arch, and the mirror made the task intuitive: make the shape fit what you see.
Why haptics carried the timing
People rush through unfamiliar tasks by default. Vibration acted as a metronome felt rather than watched, so users never had to look away from the mirror to know if they were moving too fast. A slow pulsing rhythm encouraged steady movement between positions; a sustained vibration told them to freeze. Testing showed users improving sharply after the first two positions, the combination of visual lock and haptic confirmation building confidence fast enough that precision was visibly better by position four.
![]()
Getting the marker visibility threshold right took several passes. Early versions showed faint outlines continuously, which created visual noise; the final version displayed only the active target for the current position, fading in as the user approached the correct zone, giving direction without clutter. Haptic intensity also needed calibration across devices, since a clear signal on a flagship phone read as weak on an older model, so intensity adapted to device capability while the meaning of each pattern stayed consistent.
The mirror itself introduced a left-right inversion problem. I solved it with symmetrical marker designs that worked regardless of the flip, communicating direction through shape rather than arrows that could be misread in reverse. Progress showed as a ring filling around the active marker, no numbers or percentages, just a marker completing and the next target appearing, which kept cognitive load low throughout.
What this requires beyond dental scanning
When users perform precise physical movements without training, pairing visual spatial feedback with haptic confirmation works because vision keeps them oriented in space while touch handles timing without pulling focus. Instruction should shrink to the absolute minimum: if an interface needs paragraphs of reading or a tutorial before use, most users are already lost. And any front-camera mirror experience needs its visual language built around the left-right inversion and split attention specifically, not adapted from a standard camera UI.
The scan quality in a diagnostic tool depends entirely on user compliance, not just the accuracy of the underlying algorithm. A precise model paired with a frustrating capture flow still produces poor data. This project held up because it respected exactly how someone behaves in front of a bathroom mirror, holding a phone at arm's length, doing something slightly awkward, and met them exactly there instead of asking them to adapt first.