Lists and dictionaries
A list holds many values in order, found by position. A dictionary holds many values by label, found by name. Almost every dataset you will ever load is built from these two.
- 14 min read
- 3 reading levels
- Published
Read these first
On this page 7
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
A list holds many values in a fixed order. A dictionary holds many values under labels.
Picture the seats in a bus. Seat one, seat two, seat three, all the way to the back. To find someone you count along the row. If a new passenger boards, they take the next empty seat at the end. That is a list — order matters, and every value has a position number.
Now picture the contacts on your phone. You never think "open contact number forty-seven". You think "Amma" and search for that. That is a dictionary — each value is filed under a label, and the order you saved them in does not matter.
One lesson ago you had a name pointed at a single value. Real work needs a name pointed at hundreds of them.
Why you should care
Data does not arrive one value at a time. It arrives as a class list, a sales sheet, a folder of photos, a month of readings.
Every dataset you meet in AI is a pile of rows, and every row is a set of labelled fields. In Python that is a list of dictionaries. Learn these two containers and most real data stops looking strange.
Why they exist
You could give every value its own name. student1, student2, student3. People genuinely do this for a week, then stop.
It breaks the moment you have a hundred students, or a number of students you do not know in advance. You cannot write a hundred names, and you cannot write a name for a student who has not enrolled yet.
A list solves that. One name holds all of them, however many there turn out to be.
A dictionary solves a second problem. Position is a poor way to find things when the thing has a name. Remembering that field three is the price is how mistakes happen. Asking for "price" cannot go wrong the same way.
How they work
a list a dictionary
------ ------------
position 0 → "idli" "idli" → 40
position 1 → "vada" "vada" → 30
position 2 → "filter coffee" "filter coffee" → 25
found by counting found by label
order is kept order does not matter to youCounting starts at zero, not one. The first seat on the bus is seat zero. This feels wrong for about a week and then feels normal forever.
The two nest inside each other, and that is where they become powerful:
orders → [ row 0, row 1, row 2 ]
|
v
{ "name" → "Asha", "item" → "idli", "qty" → 2 }A list of rows, where each row is a set of labelled fields. That shape is the same shape as a spreadsheet, a database table, and every dataset you will train a model on.
A real example you have seen
Open any food delivery app and look at your cart. The cart is a list — items in the order you added them, and a new item goes at the bottom.
Tap one item. Name, price, quantity, customisation, restaurant. That is a dictionary. Nobody would want the price to be "the third thing".
When that cart travels to the company's servers, it travels as exactly this shape. Lists of dictionaries is not a teaching example. It is the format the internet actually uses.
An honest word
Choosing between them confuses people, so here is a rule that holds up.
If you would naturally say "the first one" or "the next one", you want a list. If you would naturally say "the one called X", you want a dictionary.
When a list gets long, searching it means checking every entry from the start. A dictionary finds its answer without searching. On a hundred items nobody notices. On a million it is the difference between instant and unusable.
Remember this
- A list keeps values in order and finds them by position, counting from zero.
- A dictionary keeps values under labels and finds them by name, without searching.
- Real data is almost always a list of dictionaries — rows, each with named fields.
What to learn next
- Loops and conditions — doing something to every item.
- Functions — naming a piece of work so you can reuse it.
- Pandas — a list of dictionaries, upgraded into a real table.
Developer — Code and libraries.
Setup
Core Python. Nothing to install, nothing to import.
Square brackets [ ] make a list. Curly brackets { } with label: value pairs make a dictionary. Those two symbols are most of the syntax.
The two containers side by side
# A list: ordered, positions count, duplicates allowed.
items = ["idli", "vada", "filter coffee"]
# A dictionary: each value filed under a label, and labels must be unique.
prices = {"idli": 40, "vada": 30, "filter coffee": 25}
print("first item :", items[0])
print("last item :", items[-1])
print("how many :", len(items))
print("a slice :", items[0:2])
items.append("pongal") # add to the end of the list
prices["pongal"] = 55 # add a new label to the dictionary
print("after adding:", items)
print("price of vada:", prices["vada"])
print("is dosa sold?:", "dosa" in prices)
print("price of dosa:", prices.get("dosa", "not sold"))
del items[1] # remove by position
print("after removing:", items)
print("all labels :", list(prices.keys()))
print("all values :", list(prices.values()))first item : idli last item : filter coffee how many : 3 a slice : ['idli', 'vada'] after adding: ['idli', 'vada', 'filter coffee', 'pongal'] price of vada: 30 is dosa sold?: False price of dosa: not sold after removing: ['idli', 'filter coffee', 'pongal'] all labels : ['idli', 'vada', 'filter coffee', 'pongal'] all values : [40, 30, 25, 55]
Line-by-line walkthrough
items[0] is the first item, because positions start at zero. items[-1] counts backwards from the end, so it is the last item. Negative indexing saves you from writing items[len(items) - 1] everywhere.
items[0:2] takes positions zero and one. The number after the colon is where to stop, and it is not included. Every off-by-one argument in Python traces back to this rule, and the rule is deliberate: items[0:2] and items[2:] fit together with no overlap and no gap.
.append(x) adds one item to the end and returns nothing. Writing items = items.append("pongal") sets items to None and destroys your list. This is a real and frequent accident.
prices["pongal"] = 55 creates the label if it is absent and overwrites it if it is present. There is no separate "add" and "update".
"dosa" in prices tests the labels, not the values. For a list, in tests the values.
.get(label, fallback) returns the fallback instead of crashing when the label is missing. Use it whenever the data might be incomplete, which is always.
.keys() and .values() hand back live views rather than lists, which is why list(...) is wrapped around them for printing. Since Python 3.7, a dictionary keeps the order you inserted things in, and that order is part of the language, not luck.
Rows of real data
# Real data almost always arrives in this exact shape.
orders = [
{"name": "Asha", "item": "idli", "qty": 2, "price": 40},
{"name": "Ravi", "item": "vada", "qty": 1, "price": 30},
{"name": "Meera", "item": "pongal", "qty": 3, "price": 55},
]
first = orders[0]
print("one whole row:", first)
print("one field :", first["name"], "wants", first["qty"], first["item"])
print("row 3, price :", orders[2]["price"])
first["qty"] = 4 # dictionaries can be edited in place
print("after editing:", orders[0])
print("rows in total:", len(orders))one whole row: {'name': 'Asha', 'item': 'idli', 'qty': 2, 'price': 40}
one field : Asha wants 2 idli
row 3, price : 55
after editing: {'name': 'Asha', 'item': 'idli', 'qty': 4, 'price': 40}
rows in total: 3Look at the last two prints carefully. first and orders[0] are two names for one dictionary. Editing through one name is visible through the other. Nothing was copied when first was created.
To do something with every row rather than one row, you need a loop. That is the next lesson, and it is where these containers start earning their keep.
Common mistakes
1. Asking for a position that is not there.
items = ["idli", "vada", "filter coffee"]
print(items[3])Traceback (most recent call last):
File "cart.py", line 2, in <module>
print(items[3])
IndexError: list index out of rangeThree items occupy positions 0, 1 and 2. The last valid position is always len(items) - 1.
2. Asking for a label that is not there.
prices = {"idli": 40, "vada": 30}
print(prices["dosa"])Traceback (most recent call last):
File "cart.py", line 2, in <module>
print(prices["dosa"])
KeyError: 'dosa'Fix: prices.get("dosa", 0) when a missing label is expected, or test with if "dosa" in prices when it is not.
3. Thinking a second name means a second list.
original = ["idli", "vada"]
backup = original # not a copy - a second name for the same list
backup.append("dosa")
print("original:", original)
print("backup :", backup)
real_copy = original.copy() # this one is independent
real_copy.append("pongal")
print("original after real copy:", original)original: ['idli', 'vada', 'dosa'] backup : ['idli', 'vada', 'dosa'] original after real copy: ['idli', 'vada', 'dosa']
Your "backup" was never a backup. Assignment binds a second name to the same object, exactly as described in variables and data types. Use .copy() for a flat list, and copy.deepcopy() when the list contains dictionaries you also intend to edit.
4. Keeping parallel lists instead of rows.
names = ["Asha", "Ravi", "Meera"]
qty = [2, 1, 3]
prices = [40, 30, 55]Three lists that must stay in the same order forever. Sort one and the data is silently corrupted — no error, no warning, wrong answers. Delete from one and zip will quietly truncate the rest, as python basics shows.
Fix: one list of dictionaries. The fields travel together and cannot drift apart.
Try it yourself
Start from the orders list above. Print the item ordered by the person in the last row, using negative indexing rather than orders[2].
Then add a fourth order for a person called Iqbal. Then use .get("discount", 0) on any row and confirm it returns zero rather than raising KeyError.
Finally, make copy_of_orders = orders.copy(), edit copy_of_orders[0]["qty"], and print orders[0]. The result will surprise you. .copy() on a list copies the list, not the dictionaries inside it. Work out why that is the sane default.
What to learn next
- Loops and conditions — visiting every item once.
- Functions — packaging work you repeat.
- Pandas — the same rows, with column names and real speed.
Researcher — Mathematics and papers.
What a list actually is
A CPython list is a PyListObject: a length, an allocated capacity, and a pointer to a heap array of PyObject*. It is a dynamic array of references, never of values. Two consequences follow.
Elements are boxed, so a list of a million floats holds a million pointers plus a million separate 24-byte float objects, scattered across the heap. Iterating it defeats the cache in a way a contiguous buffer does not.
Growth is amortised O(1) because list_resize over-allocates. The pattern in Objects/listobject.c is roughly new_allocated = size + (size >> 3) + 6, a growth factor near 1.125 rather than the 2.0 used by many other languages. Small factor, frequent but cheap reallocations, low memory slack.
| Operation | Complexity | Note |
|---|---|---|
lst[i] | O(1) | pointer arithmetic |
lst.append(x) | O(1) amortised | over-allocated backing array |
lst.insert(0, x), lst.pop(0) | O(n) | shifts every later pointer |
x in lst | O(n) | linear scan calling __eq__ |
lst.sort() | O(n log n) | Timsort, stable, adaptive |
lst[a:b] | O(b − a) | copies the pointers, not the objects |
Timsort (Peters, 2002) is worth knowing about: it detects existing runs and merges them, so it is O(n) on already-sorted or reverse-sorted input, and it is stable, which is what makes multi-key sorting work by sorting on the least significant key first.
What a dictionary actually is
A dict is an open-addressed hash table in the compact layout introduced by Raymond Hettinger for Python 3.6. It has two parts: a small indices array of integers, and a dense entries array of (hash, key, value) triples appended in insertion order.
Three properties fall out of that design:
- Insertion order is preserved, because the entries array is append-only. This became a language guarantee in 3.7, rather than remaining a CPython implementation detail.
- Memory drops by roughly 20 to 25 percent versus the old sparse table, because the sparse part now stores narrow indices — one byte each while the table is under 128 slots — rather than 24-byte triples.
- Iteration is a linear walk over a dense array, so it is cache-friendly.
Collisions are handled by open addressing with a perturbed probe sequence, not by chaining:
perturb = hash
i = hash & mask
loop:
i = (5 * i + 1 + perturb) & mask
perturb >>= 5Mixing the high bits of the hash back in via perturb is what stops clustered hashes from degenerating. The table is resized when it reaches two thirds full.
Lookup, insert and delete are O(1) average and O(n) worst case. The worst case is reachable by an adversary who controls the keys, which is why string hashing is randomised per process under PEP 456 (SipHash-1-3). A practical consequence: hash("abc") differs between runs unless PYTHONHASHSEED is fixed, so never persist a hash value or rely on set iteration order for reproducible output.
The hashability contract
A key must be hashable: it needs __hash__ and __eq__, and objects that compare equal must hash equal. Mutable built-ins deliberately set __hash__ = None.
from collections import Counter, defaultdict, deque
d = {"b": 2, "a": 1, "c": 3}
print("insertion order kept:", list(d))
print("unhashable key:", end=" ")
try:
{[1, 2]: "x"}
except TypeError as e:
print(type(e).__name__, "-", e)
dq = deque([1, 2, 3])
dq.appendleft(0) # O(1) at both ends; list.insert(0, x) is O(n)
print("deque:", list(dq))
print("counter:", Counter("mississippi").most_common(3))
groups = defaultdict(list)
for word in ["idli", "dosa", "vada", "dhokla"]:
groups[word[0]].append(word)
print("grouped:", dict(groups))insertion order kept: ['b', 'a', 'c']
unhashable key: TypeError - unhashable type: 'list'
deque: [0, 1, 2, 3]
counter: [('i', 4), ('s', 4), ('p', 2)]
grouped: {'i': ['idli'], 'd': ['dosa', 'dhokla'], 'v': ['vada']}defaultdict and Counter remove most hand-written grouping code, and deque is the correct structure for a queue or a sliding window — a replay buffer built on list.pop(0) is quadratic and this is a genuinely common bug in reinforcement-learning code.
Note that a dataclass with eq=True, frozen=True gets a generated __hash__, which makes it usable as a dictionary key. A plain class uses identity for both hash and equality, so two structurally identical instances are different keys.
Where these structures stop being appropriate
The accidental quadratic is the dominant failure. A loop over n items that tests x in some_list of length n is O(n²). Converting the inner container to a set makes it O(n). This appears constantly in vocabulary filtering, deduplication and stop-word removal, and it is invisible on a test file of 200 rows.
The deeper limit is representational. Numerical work wants one dtype, contiguous memory and no per-element Python object. That is precisely what an ndarray provides and what a list of dictionaries cannot. The standard pipeline is therefore: parse JSON into lists of dicts, validate, then convert once into a DataFrame or an array and never touch per-element Python again.
For fixed-schema rows, dataclasses with slots=True or NamedTuple beat a dict on both memory and attribute-access cost, because they drop the per-instance dictionary. For sparse feature maps in classical ML, scipy.sparse beats both.
References
- Hettinger, R. "Modern Python Dictionaries: A Confluence of a Dozen Great Ideas", PyCon US 2017.
- Peters, T.
Objects/listsort.txtin the CPython source — the Timsort design note. - PEP 456 — Secure and Interchangeable Hash Algorithm, Heimes, 2013.
- PEP 468 — Preserving the Order of **kwargs in a Function, Snow, 2014.
- PEP 372 — Adding an Ordered Dictionary to collections, Ronacher and Hettinger, 2008.
- Knuth, D. The Art of Computer Programming, Vol. 3, section 6.4 — open addressing and probe sequences.
- CPython source,
Objects/listobject.candObjects/dictobject.c.
What to learn next
- Loops and conditions — the iterator protocol these containers implement.
- NumPy — contiguous buffers and why boxing costs so much.
- Pandas — columnar storage over those buffers.