AI for Accessibility

Navigation aids for blind users

Blind navigation aids turn a camera or sensor feed into spoken obstacle warnings, where a missed obstacle is a physical safety failure, not an inconvenience.

On this page 6
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. An honest warning
  5. Remember this
  6. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

A navigation aid for blind users turns a camera or sensor feed into spoken warnings about what is nearby.

Think about a friend walking beside you in a crowded market, quietly saying "chair to your left" or "step down ahead." You never have to look down to know what is coming.

A navigation aid built for blind and low-vision users tries to give that same kind of steady, useful warning. It is generated from a phone camera or a dedicated sensor, instead of a friend's eyes.

Why it exists

A white cane and a guide dog remain the most trusted tools for independent travel for many blind people, and both have real limits. A cane detects what is directly in its sweep path, close to the ground. It cannot see a hanging obstacle at head height, or something several steps ahead.

Camera- and sensor-based navigation aids aim to extend that awareness further out and to head height. They are a complement to a cane or guide dog, not a replacement for either.

Building this well means solving a problem most computer vision systems never face this directly. The cost of missing something is not a lower accuracy score — it is a person walking into it.

How it works

Camera or sensor sees the space ahead
        |
        v
Estimate distance to nearby objects, across the whole field of view
        |
        v
Look for the CLOSEST point in each direction -- left, centre, right --
not the average, since a thin pole among open floor must not be averaged away
        |
        v
If something is close: "caution, obstacle ahead"  or  "stop, obstacle very close"

Where you have already seen it

Microsoft's Seeing AI and several dedicated wearable devices offer real-time obstacle and scene description features aimed at exactly this use case. Some smart canes now include built-in sensors that vibrate or speak warnings about obstacles above cane height, working alongside the cane's own physical detection.

An honest warning

In this domain, a missed obstacle is not a statistic — it is someone walking into something they trusted the system to warn them about. A navigation aid that occasionally over-warns about something harmless is annoying. A navigation aid that misses something genuinely dangerous has failed at the one job that matters most.

Every serious system in this space is built, and should be evaluated, with that asymmetry front and centre. Even the best current systems are explicitly designed as a supplement to a cane or guide dog. None are a stand-alone replacement for one.

Remember this

  • Navigation aids extend a blind traveller's awareness beyond what a cane or guide dog alone can detect, especially at head height and further ahead.
  • Looking at the closest point in a region, not the average, is what catches a thin, dangerous obstacle among open floor.
  • A missed obstacle is a real safety failure, not a lower accuracy number. This domain should always be designed and judged with that cost in mind.

What to learn next

Developer — Code and libraries.

This example works from a synthetic depth map rather than a real camera feed, so the core decision — which region is dangerous, and why averaging is the wrong tool for that decision — stays fully visible and runs in an instant.

Setup

bash
pip install numpy

Minimal runnable code

depth_map simulates one estimated distance, in metres, per pixel of a camera frame — mostly open floor, with a thin pole close by on the right side. Real systems get this from a depth model (see monocular depth estimation) or a dedicated depth sensor.

navaid.py
import numpy as np

rng = np.random.default_rng(0)

# A depth map: one estimated distance (in metres) per pixel of a camera
# frame. Real systems get this from a monocular depth model (see
# monocular-depth-estimation) or a dedicated depth sensor. Here it is built
# by hand: mostly open floor (4-6m), with a pole close on the right side.
height, width = 60, 90
depth_map = rng.uniform(4.0, 6.0, size=(height, width))
depth_map[:, 65:80] = rng.uniform(0.8, 1.2, size=(height, 15))   # the pole

def zone_alert(depth_map):
    third = depth_map.shape[1] // 3
    zones = {
        "left":   depth_map[:, :third],
        "centre": depth_map[:, third:2 * third],
        "right":  depth_map[:, 2 * third:],
    }
    alerts = []
    for name, region in zones.items():
        # the CLOSEST point in a zone is what matters, not the average --
        # averaging would hide a thin pole among a lot of open floor
        closest = region.min()
        if closest < 1.5:
            alerts.append((name, "STOP / obstacle very close", closest))
        elif closest < 3.0:
            alerts.append((name, "caution, obstacle ahead", closest))
    return alerts

alerts = zone_alert(depth_map)
if not alerts:
    print("path is clear")
for zone, message, distance in alerts:
    print(f"{zone:8s} zone: {message}  (closest point {distance:.1f} m)")
Output
right    zone: STOP / obstacle very close  (closest point 0.8 m)

Walkthrough

depth_map represents a frame that is almost entirely open floor, 4 to 6 metres away in every direction, except for a narrow band on the right side where a pole sits under 1.2 metres away — a small fraction of the total pixels in the frame.

zone_alert deliberately uses region.min(), the closest point in each third of the frame, rather than region.mean(), the average. This choice is the entire lesson: if the code averaged instead, the pole's close readings would be diluted by the much larger area of open floor around it, and the average distance for the right zone would land comfortably in "safe" territory — the pole would vanish from the warning entirely, even though it is the one thing in the frame someone could actually walk into.

Only the right zone triggers an alert here, correctly, because that is the only zone actually containing something close.

Common mistakes

Averaging depth across a region instead of taking the minimum. This is the single most important lesson in this example. A thin, dangerous obstacle is exactly the kind of thing an average silently erases.

Setting alert thresholds without walking-speed context. 1.5 metres might be plenty of warning while standing still, and barely enough time to react while walking at a normal pace. Real systems tune thresholds against typical walking speed and reaction time, not an arbitrary round number.

Dividing the frame into too few or too badly shaped zones. Three simple left/centre/right columns, as used here, are a reasonable starting point, but miss anything meaningfully identified only by height, such as a hanging obstacle overhead versus a curb underfoot — real systems typically use a finer grid.

Treating a single frame's reading as final. A single frame's depth estimate can be noisy or briefly wrong, from a person walking past or a reflective surface confusing the camera. Real systems smooth alerts across several consecutive frames, the same stability idea used in reading the world aloud, before committing to a spoken warning.

Try it yourself

Change depth_map[:, 65:80] to place the pole in the centre zone instead of the right, and rerun. Then add a second, closer obstacle in the left zone in the same way, and check that zone_alert correctly reports both zones at once, each with its own distance.

What to learn next

Researcher — Mathematics and papers.

The formal setting

Let D(x, y) be an estimated depth value at pixel (x, y) in a camera frame, produced by a monocular depth model, stereo pair, or dedicated depth sensor. Partition the frame into regions R_1, ..., R_k (zones, in the developer example). A navigation aid computes an alert signal per region:

alert(R_i) = f( min_{(x,y) ∈ R_i} D(x, y) )
  • f — a thresholding function mapping the closest detected distance in a region to a discrete alert level
  • The min operator, rather than mean or median, is the statistically correct choice for a worst-case, safety-relevant signal: the quantity that matters is the closest approach distance, not the central tendency of the region's depth distribution

This is a specific instance of a broader principle in safety-critical perception systems: aggregation statistics should match the cost structure of the decision, and mean-based aggregation is appropriate for estimation tasks, not for detecting the presence of a rare, dangerous minority of pixels within a region — a point closely related to the class-imbalance and cost-asymmetry themes explored in disability and dataset bias and evaluating assistive AI.

Monocular depth reliability

Monocular depth estimation (see monocular depth estimation) infers distance from a single 2D image without direct geometric triangulation, and is consequently less reliable than stereo or active depth sensing, particularly for reflective, transparent, or texture-poor surfaces — glass doors and still water are documented failure cases across the monocular depth literature (Godard et al., 2019, among others). Deployed navigation aids commonly fuse multiple sensing modalities — camera-based depth, ultrasonic or infrared time-of-flight sensors, and inertial data — precisely because no single monocular signal is reliable enough alone for a safety-relevant application, an instance of general sensor fusion practice (see sensor fusion and the Kalman filter).

Temporal smoothing and false-alarm cost

Given a raw per-frame alert stream, real systems apply temporal smoothing to suppress single-frame false positives while bounding added latency, structurally identical to the stability-filtering problem formalised in reading the world aloud:

alert_t (smoothed) = alert  if alert(R_i) holds for ≥ k of the last n frames

The choice of k, n, and the underlying distance thresholds is fundamentally a cost-weighted decision, not a pure accuracy-maximisation one — see choosing a threshold from costs for the general treatment, and evaluating assistive AI for why raw accuracy is the wrong metric to optimise here.

Complexity and cost

For a depth map of h × w pixels divided into k zones:

ComponentTypical cost
Monocular depth estimationO(hw), one forward pass through a depth model, per frame
Zone-wise minimum computationO(hw), a single pass over all pixels
Temporal smoothing over n framesO(n · k), negligible relative to depth estimation

Depth estimation dominates the compute budget, which is why on-device navigation aids typically run a smaller, distilled depth model at reduced resolution, trading some accuracy for the low, consistent latency a real-time safety application requires.

Papers

  • Godard, C. et al. (2019). Digging Into Self-Supervised Monocular Depth Estimation. ICCV. A widely used monocular depth architecture and a direct discussion of its known failure modes on reflective and texture-poor surfaces.
  • Ahmetovic, D., Gleason, C., Ruan, C., Kitani, K., Takagi, H. and Asakawa, C. (2016). NavCog: A Navigational Cognitive Assistant for the Blind. MobileHCI. A representative real-time, sensor-based navigation assistant for blind travellers.
  • Kalman, R. (1960). A New Approach to Linear Filtering and Prediction Problems. Journal of Basic Engineering 82(1). The foundational filtering technique behind combining multiple noisy sensors, covered fully in sensor fusion and the Kalman filter.

Current state

Multi-sensor fusion, rather than any single camera-based signal alone, is the accepted approach for safety-relevant navigation assistance, mirroring standard practice in robotics and autonomous vehicle perception. Independent evaluation with real blind and low-vision users under real walking conditions — not benchmark accuracy alone — remains the field's accepted standard for validating a system before any claim of practical reliability, and the assistive technology research community is explicit that current systems remain a supplement to established mobility tools such as the white cane and guide dog, not a replacement for either.

What to learn next