Retail, Demand and Supply Chain

Messy addresses and geocoding

Two customers can describe the same building in completely different words, and a delivery system has to recognise they mean the same place before it can find it on a map.

Read these first

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Two people can describe the same building in completely different words — and a delivery system has to know they mean the same place.

Think about giving directions to your own home to a friend who has never visited. You probably do not read out a formal postal address. You say something like "the yellow building near the tea stall, second lane past the temple." Someone else describing the exact same building might say "opposite the temple, ask for the Sharma house."

Both descriptions point to the same front door. Neither looks anything like the other in writing.

Why it exists

Geocoding means converting a written address into a precise location — usually latitude and longitude — that a map or a delivery app can actually use. It sounds like it should be a lookup. In practice, it is closer to detective work.

Formal postal addressing is inconsistent across India, and even where it exists, people rarely type it precisely. The same building shows up in customer records as "MG Road," "M.G Rd," and "Mahatma Gandhi Road." Landmarks substitute for street names constantly, and the same landmark gets spelled several different ways by several different customers.

Before a system can even geocode an address, it often has to solve a separate problem first. It has to recognise that two very differently written addresses refer to the same real place. That is address matching, and it usually has to happen before geocoding can work reliably.

How it works

Customer typed at checkout:   "12 mg rd near city central mall blr 560001"

Address on file (canonical):  "12 MG Road, near City Central Mall,
                                Bengaluru 560001"

These do not match as exact text.
Both plainly describe the same place to a human reader.

The fix compares addresses by their meaningful parts — the words, the landmark, the pin code — rather than demanding an exact character-by-character match. Two addresses that share most of their important words, even in a different order or spelling, are treated as very likely to be the same place.

Where you have already seen it

  • A food delivery app recognising your saved "Home" address even when you type it slightly differently each time you order.
  • A courier company merging duplicate customer records that were typed in by different call-centre agents over the years, each spelling things their own way.
  • Google Maps understanding "near Cauvery Emporium" as a real, findable location, using nearby landmarks it has already geocoded.

Remember this

  • Two addresses can be worded completely differently and still describe the same place — exact text matching misses this constantly.
  • Comparing addresses by their shared words and landmarks, not their exact spelling, is what makes matching reliable.
  • This matching step usually has to happen before a location can be reliably placed on a map, not after.

What to learn next

Developer — Code and libraries.

Setup

bash
# standard library only
python3 --version

Minimal runnable code

Three canonical addresses on file, and three messy versions a customer might actually type. We compare exact string matching against a simple token-based fuzzy match.

address_matching.py
import re

canonical_addresses = {
    "addr_101": "12 MG Road, near City Central Mall, Bengaluru 560001",
    "addr_102": "45 Brigade Road, opposite St Marks Church, Bengaluru 560025",
    "addr_103": "7 Residency Road, next to Cauvery Emporium, Bengaluru 560025",
}

# What customers actually typed at checkout -- typos, abbreviations, reordering.
raw_inputs = [
    "12 mg rd near city central mall blr 560001",
    "flat near st. marks church, brigade rd, banglore",
    "7,Residency Rd,nxt 2 cauvery emporium",
]


def normalize(text):
    text = text.lower()
    text = re.sub(r"[^a-z0-9\s]", " ", text)
    return set(text.split())


ABBREVIATIONS = {"rd": "road", "blr": "bengaluru", "banglore": "bengaluru", "nxt": "next", "2": "to"}


def normalize_with_abbrev(text):
    words = normalize(text)
    return {ABBREVIATIONS.get(w, w) for w in words}


def jaccard(a, b):
    return len(a & b) / len(a | b) if (a | b) else 0.0


canonical_tokens = {aid: normalize_with_abbrev(addr) for aid, addr in canonical_addresses.items()}

for raw in raw_inputs:
    raw_exact_hit = raw in canonical_addresses.values()
    raw_tokens = normalize_with_abbrev(raw)
    scores = {aid: jaccard(raw_tokens, toks) for aid, toks in canonical_tokens.items()}
    best_id = max(scores, key=scores.get)
    print(f"raw input: {raw!r}")
    print(f"  exact string match found: {raw_exact_hit}")
    print(f"  best fuzzy match: {best_id} (score {scores[best_id]:.2f}) -> {canonical_addresses[best_id]}")
    print()
Output
raw input: '12 mg rd near city central mall blr 560001'
  exact string match found: False
  best fuzzy match: addr_101 (score 1.00) -> 12 MG Road, near City Central Mall, Bengaluru 560001

raw input: 'flat near st. marks church, brigade rd, banglore'
  exact string match found: False
  best fuzzy match: addr_102 (score 0.55) -> 45 Brigade Road, opposite St Marks Church, Bengaluru 560025

raw input: '7,Residency Rd,nxt 2 cauvery emporium'
  exact string match found: False
  best fuzzy match: addr_103 (score 0.78) -> 7 Residency Road, next to Cauvery Emporium, Bengaluru 560025

What actually happened

Every single raw input fails an exact match — raw_exact_hit is False all three times, despite each one plainly describing a real address on file.

normalize strips punctuation and lowercases everything, turning an address into an unordered bag of words — a set, where word order and exact formatting no longer matter. ABBREVIATIONS expands a handful of common shortenings before comparing, so "rd" and "road" are treated as the same word rather than as two different ones.

Jaccard similarity scores two sets by (shared words) / (all words either one uses) — a simple, effective way to measure how much two short pieces of text overlap. The first input scores a perfect 1.00 once abbreviations are expanded. The second scores lower, 0.55, because it genuinely shares fewer words with the canonical address ("flat" is not in the original, "banglore" is a real misspelling) — but 0.55 is still comfortably the highest score among the three candidates, so the match is still made correctly.

Common mistakes

Matching on the pin code alone. A pin code narrows things down to a neighbourhood, sometimes a few thousand households — nowhere near specific enough to identify one address.

Building an abbreviation dictionary once and never updating it. New shorthand appears constantly ("opp" for "opposite," "nr" for "near"). A fuzzy matcher is only as good as the normalisation feeding it.

Trusting a fuzzy match blindly at a low similarity score. A low best-score match, even if it is the highest among candidates, might genuinely be a new address that is not in the system at all. Set a minimum score threshold, and route anything below it to a human for review rather than guessing.

Ignoring the landmark entirely and only comparing street-level fields. In much of India, the landmark ("near City Central Mall") often carries more identifying weight than the street name, especially once abbreviations muddy the street name itself.

Try it yourself

Add a fourth raw input describing a place that is genuinely not in canonical_addresses, like "14 Church Street, near Amoeba Records". Re-run, and check the best score returned — a properly low score is the system correctly telling you "I don't actually know this address," rather than confidently guessing wrong.

What to learn next

Researcher — Mathematics and papers.

Address matching as entity resolution

Matching differently written addresses to the same real-world place is a specific case of entity resolution (also called record linkage or deduplication): deciding whether two records, with no shared unique identifier, refer to the same underlying entity. Fellegi and Sunter (1969), A Theory for Record Linkage, JASA, provide the classical probabilistic framework: for each candidate pair, compute agreement across several fields, and classify the pair as a match, non-match, or "possible match requiring review" based on likelihood ratios estimated from labelled examples.

Beyond Jaccard: edit distance and phonetic matching

Token-set overlap, used in the developer example, is fast and effective for word-level variation (abbreviation, reordering, dropped words) but blind to character-level typos within a single token ("Bengaluru" vs "Bengalore"). Production address-matching systems typically combine several signals:

  • Levenshtein (edit) distance — the minimum number of single-character insertions, deletions or substitutions to turn one string into another — catches typos within a token that set-based methods miss entirely.
  • Soundex or Metaphone — phonetic encodings that group words by how they sound rather than how they are spelled, useful for transliterated names with inconsistent spelling, common across Indian languages romanised inconsistently.
  • TF-IDF-weighted token similarity — rather than treating every word equally, as plain Jaccard does, downweight very common words ("road," "near," "bengaluru") and upweight rare, more identifying ones (a specific landmark or building name), since a shared rare word is much stronger evidence of a true match than a shared common one.

A learned classifier — typically gradient-boosted trees — combining several such similarity scores as features, trained on labelled matching and non-matching pairs, generally outperforms any single hand-tuned similarity metric in production.

Geocoding as a separate, downstream problem

Once an address is matched or parsed into structured components (building, street, landmark, area, pin code), geocoding proper — resolving it to coordinates — is typically solved by a combination of a structured address index (when postal addressing is reliable) and landmark-anchored search (when it is not): find the landmark's known coordinates, then apply a learned or rule-based offset for "near," "opposite," or "behind." This landmark-anchored approach is disproportionately important in markets with informal addressing systems, and is an active area of applied work at Indian mapping and logistics companies specifically because of how often it is the only reliable resolution strategy.

Evaluation and the cost of a wrong match

A false match — confidently resolving to the wrong building — is usually far more costly in a delivery context than a false non-match, since a package can be sent to a genuinely wrong address, unlike a non-match, which only escalates to manual review. This asymmetry argues for a deliberately conservative similarity threshold in production, evaluated with precision weighted more heavily than recall — directly analogous to the cost-asymmetric thinking in Turning a forecast into an order quantity, applied to a classification decision instead of an inventory one.

Key references

  • Fellegi, I. & Sunter, A. (1969). A Theory for Record Linkage. Journal of the American Statistical Association 64(328).
  • Christen, P. (2012). Data Matching: Concepts and Techniques for Record Linkage, Entity Resolution, and Duplicate Detection. Springer — the standard applied reference across the whole field.
  • Levenshtein, V. (1966). Binary Codes Capable of Correcting Deletions, Insertions, and Reversals. Soviet Physics Doklady 10(8).

What to learn next

What to learn next

These follow on from what you just read.

  • Manufacturing and Predictive Maintenance

    Industrial sensor and historian data

    A factory's sensors report on different schedules and occasionally freeze or drift, so the raw data has to be aligned and checked before any model can trust it.

  • Manufacturing and Predictive Maintenance

    Predictive maintenance

    Predictive maintenance uses sensor data to warn that a machine is about to fail, instead of fixing it on a fixed calendar or waiting for it to break.

  • Manufacturing and Predictive Maintenance

    Remaining useful life

    Remaining useful life predicts how many days a machine has left, not only whether it will fail soon, and the honest version of that number gets more confident as the machine gets closer to failing.