Navigation aids for blind users
Blind navigation aids turn a camera or sensor feed into spoken obstacle warnings, where a missed obstacle is a physical safety failure, not an inconvenience.
- 11 min read
- 3 reading levels
- Published
Read these first
On this page 6
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
A navigation aid for blind users turns a camera or sensor feed into spoken warnings about what is nearby.
Think about a friend walking beside you in a crowded market, quietly saying "chair to your left" or "step down ahead." You never have to look down to know what is coming.
A navigation aid built for blind and low-vision users tries to give that same kind of steady, useful warning. It is generated from a phone camera or a dedicated sensor, instead of a friend's eyes.
Why it exists
A white cane and a guide dog remain the most trusted tools for independent travel for many blind people, and both have real limits. A cane detects what is directly in its sweep path, close to the ground. It cannot see a hanging obstacle at head height, or something several steps ahead.
Camera- and sensor-based navigation aids aim to extend that awareness further out and to head height. They are a complement to a cane or guide dog, not a replacement for either.
Building this well means solving a problem most computer vision systems never face this directly. The cost of missing something is not a lower accuracy score — it is a person walking into it.
How it works
Camera or sensor sees the space ahead
|
v
Estimate distance to nearby objects, across the whole field of view
|
v
Look for the CLOSEST point in each direction -- left, centre, right --
not the average, since a thin pole among open floor must not be averaged away
|
v
If something is close: "caution, obstacle ahead" or "stop, obstacle very close"Where you have already seen it
Microsoft's Seeing AI and several dedicated wearable devices offer real-time obstacle and scene description features aimed at exactly this use case. Some smart canes now include built-in sensors that vibrate or speak warnings about obstacles above cane height, working alongside the cane's own physical detection.
An honest warning
In this domain, a missed obstacle is not a statistic — it is someone walking into something they trusted the system to warn them about. A navigation aid that occasionally over-warns about something harmless is annoying. A navigation aid that misses something genuinely dangerous has failed at the one job that matters most.
Every serious system in this space is built, and should be evaluated, with that asymmetry front and centre. Even the best current systems are explicitly designed as a supplement to a cane or guide dog. None are a stand-alone replacement for one.
Remember this
- Navigation aids extend a blind traveller's awareness beyond what a cane or guide dog alone can detect, especially at head height and further ahead.
- Looking at the closest point in a region, not the average, is what catches a thin, dangerous obstacle among open floor.
- A missed obstacle is a real safety failure, not a lower accuracy number. This domain should always be designed and judged with that cost in mind.
What to learn next
- Reading the world aloud — a related real-time, camera-guided assistive technology, focused on text instead of obstacles.
- Monocular depth estimation — the technique that turns an ordinary camera image into distance estimates.
- Safety in physical AI systems — the broader engineering discipline of building AI systems where being wrong has physical consequences.
Developer — Code and libraries.
This example works from a synthetic depth map rather than a real camera feed, so the core decision — which region is dangerous, and why averaging is the wrong tool for that decision — stays fully visible and runs in an instant.
Setup
pip install numpyMinimal runnable code
depth_map simulates one estimated distance, in metres, per pixel of a camera frame — mostly open floor, with a thin pole close by on the right side. Real systems get this from a depth model (see monocular depth estimation) or a dedicated depth sensor.
import numpy as np
rng = np.random.default_rng(0)
# A depth map: one estimated distance (in metres) per pixel of a camera
# frame. Real systems get this from a monocular depth model (see
# monocular-depth-estimation) or a dedicated depth sensor. Here it is built
# by hand: mostly open floor (4-6m), with a pole close on the right side.
height, width = 60, 90
depth_map = rng.uniform(4.0, 6.0, size=(height, width))
depth_map[:, 65:80] = rng.uniform(0.8, 1.2, size=(height, 15)) # the pole
def zone_alert(depth_map):
third = depth_map.shape[1] // 3
zones = {
"left": depth_map[:, :third],
"centre": depth_map[:, third:2 * third],
"right": depth_map[:, 2 * third:],
}
alerts = []
for name, region in zones.items():
# the CLOSEST point in a zone is what matters, not the average --
# averaging would hide a thin pole among a lot of open floor
closest = region.min()
if closest < 1.5:
alerts.append((name, "STOP / obstacle very close", closest))
elif closest < 3.0:
alerts.append((name, "caution, obstacle ahead", closest))
return alerts
alerts = zone_alert(depth_map)
if not alerts:
print("path is clear")
for zone, message, distance in alerts:
print(f"{zone:8s} zone: {message} (closest point {distance:.1f} m)")right zone: STOP / obstacle very close (closest point 0.8 m)
Walkthrough
depth_map represents a frame that is almost entirely open floor, 4 to 6 metres away in every direction, except for a narrow band on the right side where a pole sits under 1.2 metres away — a small fraction of the total pixels in the frame.
zone_alert deliberately uses region.min(), the closest point in each third of the frame, rather than region.mean(), the average. This choice is the entire lesson: if the code averaged instead, the pole's close readings would be diluted by the much larger area of open floor around it, and the average distance for the right zone would land comfortably in "safe" territory — the pole would vanish from the warning entirely, even though it is the one thing in the frame someone could actually walk into.
Only the right zone triggers an alert here, correctly, because that is the only zone actually containing something close.
Common mistakes
Averaging depth across a region instead of taking the minimum. This is the single most important lesson in this example. A thin, dangerous obstacle is exactly the kind of thing an average silently erases.
Setting alert thresholds without walking-speed context. 1.5 metres might be plenty of warning while standing still, and barely enough time to react while walking at a normal pace. Real systems tune thresholds against typical walking speed and reaction time, not an arbitrary round number.
Dividing the frame into too few or too badly shaped zones. Three simple left/centre/right columns, as used here, are a reasonable starting point, but miss anything meaningfully identified only by height, such as a hanging obstacle overhead versus a curb underfoot — real systems typically use a finer grid.
Treating a single frame's reading as final. A single frame's depth estimate can be noisy or briefly wrong, from a person walking past or a reflective surface confusing the camera. Real systems smooth alerts across several consecutive frames, the same stability idea used in reading the world aloud, before committing to a spoken warning.
Try it yourself
Change depth_map[:, 65:80] to place the pole in the centre zone instead of the right, and rerun. Then add a second, closer obstacle in the left zone in the same way, and check that zone_alert correctly reports both zones at once, each with its own distance.
What to learn next
- Monocular depth estimation — the real technique that produces a depth map like the one used here.
- Object detection — identifying what an obstacle actually is, not only how far away it sits.
- Safety in physical AI systems — the general discipline of designing AI systems around the cost of being wrong.
Researcher — Mathematics and papers.
The formal setting
Let D(x, y) be an estimated depth value at pixel (x, y) in a camera frame, produced by a monocular depth model, stereo pair, or dedicated depth sensor. Partition the frame into regions R_1, ..., R_k (zones, in the developer example). A navigation aid computes an alert signal per region:
alert(R_i) = f( min_{(x,y) ∈ R_i} D(x, y) )f— a thresholding function mapping the closest detected distance in a region to a discrete alert level- The
minoperator, rather than mean or median, is the statistically correct choice for a worst-case, safety-relevant signal: the quantity that matters is the closest approach distance, not the central tendency of the region's depth distribution
This is a specific instance of a broader principle in safety-critical perception systems: aggregation statistics should match the cost structure of the decision, and mean-based aggregation is appropriate for estimation tasks, not for detecting the presence of a rare, dangerous minority of pixels within a region — a point closely related to the class-imbalance and cost-asymmetry themes explored in disability and dataset bias and evaluating assistive AI.
Monocular depth reliability
Monocular depth estimation (see monocular depth estimation) infers distance from a single 2D image without direct geometric triangulation, and is consequently less reliable than stereo or active depth sensing, particularly for reflective, transparent, or texture-poor surfaces — glass doors and still water are documented failure cases across the monocular depth literature (Godard et al., 2019, among others). Deployed navigation aids commonly fuse multiple sensing modalities — camera-based depth, ultrasonic or infrared time-of-flight sensors, and inertial data — precisely because no single monocular signal is reliable enough alone for a safety-relevant application, an instance of general sensor fusion practice (see sensor fusion and the Kalman filter).
Temporal smoothing and false-alarm cost
Given a raw per-frame alert stream, real systems apply temporal smoothing to suppress single-frame false positives while bounding added latency, structurally identical to the stability-filtering problem formalised in reading the world aloud:
alert_t (smoothed) = alert if alert(R_i) holds for ≥ k of the last n framesThe choice of k, n, and the underlying distance thresholds is fundamentally a cost-weighted decision, not a pure accuracy-maximisation one — see choosing a threshold from costs for the general treatment, and evaluating assistive AI for why raw accuracy is the wrong metric to optimise here.
Complexity and cost
For a depth map of h × w pixels divided into k zones:
| Component | Typical cost |
|---|---|
| Monocular depth estimation | O(hw), one forward pass through a depth model, per frame |
| Zone-wise minimum computation | O(hw), a single pass over all pixels |
Temporal smoothing over n frames | O(n · k), negligible relative to depth estimation |
Depth estimation dominates the compute budget, which is why on-device navigation aids typically run a smaller, distilled depth model at reduced resolution, trading some accuracy for the low, consistent latency a real-time safety application requires.
Papers
- Godard, C. et al. (2019). Digging Into Self-Supervised Monocular Depth Estimation. ICCV. A widely used monocular depth architecture and a direct discussion of its known failure modes on reflective and texture-poor surfaces.
- Ahmetovic, D., Gleason, C., Ruan, C., Kitani, K., Takagi, H. and Asakawa, C. (2016). NavCog: A Navigational Cognitive Assistant for the Blind. MobileHCI. A representative real-time, sensor-based navigation assistant for blind travellers.
- Kalman, R. (1960). A New Approach to Linear Filtering and Prediction Problems. Journal of Basic Engineering 82(1). The foundational filtering technique behind combining multiple noisy sensors, covered fully in sensor fusion and the Kalman filter.
Current state
Multi-sensor fusion, rather than any single camera-based signal alone, is the accepted approach for safety-relevant navigation assistance, mirroring standard practice in robotics and autonomous vehicle perception. Independent evaluation with real blind and low-vision users under real walking conditions — not benchmark accuracy alone — remains the field's accepted standard for validating a system before any claim of practical reliability, and the assistive technology research community is explicit that current systems remain a supplement to established mobility tools such as the white cane and guide dog, not a replacement for either.
What to learn next
- Monocular depth estimation — the core perception technique this lesson's depth map represents.
- Sensor fusion and the Kalman filter — combining multiple imperfect sensors into one more reliable estimate.
- Evaluating assistive AI with real users — why this domain's evaluation standard differs fundamentally from ordinary accuracy metrics.