Chapter 17

Nested data — a list of dictionaries

The shape real data arrives in: dictionaries inside lists, lists inside dictionaries. Reaching in, grouping, and why copy() is not enough here.

35 minPython 3.12
  1. 1Encounter
  2. 2Understand
  3. 3Worked
  4. 4Predict
  5. 5Apply
  6. 6Stretch

The problem we are solving

The last three chapters taught four containers — lists, tuples, dictionaries, sets. Every example used one of them, one level deep.

Real data does not arrive that way. A list of orders looks like this:

text
customer   item     quantity   price
rafi       pen      3          15.0
ahmed      bag      1          850.0

There are two levels here. On the outside a collection — the number of orders is not known, so a list. Inside, each order is a record whose every part has a name — so a dictionary.

What comes back from an API, what you get from reading a CSV, what a database returns: almost always this shape, a list of dictionaries. Today's chapter introduces no new container. It is about putting the ones you have inside each other, and recognising the mistakes that follow.

By the end of this chapter you can

  • Reach into a list of dictionaries and walk over it
  • Read a deep KeyError or IndexError and tell which level went wrong
  • Group and summarise a list of dictionaries
  • Handle incomplete records with .get()'s default
  • Explain why copy() is not enough for nested data

Prerequisites: Sets — a collection without duplicates.


Recognising the shape

python
orders = [
    {"customer": "rafi", "item": "pen", "quantity": 3, "price": 15.0},
    {"customer": "ahmed", "item": "bag", "quantity": 1, "price": 850.0},
]

print(len(orders))
print(orders[0])
print(orders[0]["customer"])
print(orders[1]["price"])
text
2
{'customer': 'rafi', 'item': 'pen', 'quantity': 3, 'price': 15.0}
rafi
850.0

The line orders[0]["customer"] is the core skill of this chapter, and it is read left to right, one step at a time:

  • orders — a list
  • orders[0] — its first item, which is a dictionary
  • orders[0]["customer"] — that dictionary's customer value

When in doubt, stop and print the middle step. Writing print(orders[0]) to see that it is a dictionary is one habit that removes half the trouble with nested data.

Walking it

python
orders = [
    {"customer": "rafi", "item": "pen", "quantity": 3, "price": 15.0},
    {"customer": "ahmed", "item": "bag", "quantity": 1, "price": 850.0},
]

for order in orders:
    total = order["quantity"] * order["price"]
    print(f"{order['customer']:<8} {order['item']:<6} {total:>8.2f}")
text
rafi     pen       45.00
ahmed    bag      850.00

for order in orders gives a dictionary each time round, and inside it you reach things by name. There is no index arithmetic anywhere.

Quotes inside an f-string. f"{order['customer']}" uses single quotes inside because the f-string itself is written with double quotes. Use the same kind inside and out and Python cannot tell where the text ends. Keep one kind inside and the other outside.

Incomplete records

In real data, not every record has every key:

python
orders = [{"customer": "rafi"}]
print(orders[0]["discount"])
text
KeyError: 'discount'

The KeyError says the key is missing, but not which record is missing it. When that happens in one order out of twelve thousand it is painful to find — which is why walking with enumerate keeps the number to hand.

When a key being absent is normal, .get():

python
orders = [{"customer": "rafi"}]
print(orders[0].get("discount", 0))
text
0

Chapter fifteen's rule applies here too: square brackets when it should be there, .get() when its absence is normal. The temptation is stronger with nested data, because .get() can silence every error — and then incomplete data quietly becomes zero and walks into your arithmetic.

The other shape — lists inside a dictionary

python
by_customer = {
    "rafi": ["pen", "ink"],
    "ahmed": ["bag"],
}

by_customer["rafi"].append("bottle")
by_customer["dia"] = ["eraser"]

print(by_customer)

for customer, items in by_customer.items():
    print(f"{customer}: {len(items)} item(s) - {', '.join(items)}")
text
{'rafi': ['pen', 'ink', 'bottle'], 'ahmed': ['bag'], 'dia': ['eraser']}
rafi: 3 item(s) - pen, ink, bottle
ahmed: 1 item(s) - bag
dia: 1 item(s) - eraser

The line by_customer["rafi"].append("bottle") is two steps in one: by_customer["rafi"] pulls out a list, and then append adds to that list.

This shape answers a problem from chapter fifteen, where a key could only exist once and so one person could not have several records. Make the value a list and they can.

Grouping — the most useful pattern

Turning a list of dictionaries into "who bought what":

python
orders = [
    {"customer": "rafi", "item": "pen"},
    {"customer": "ahmed", "item": "bag"},
    {"customer": "rafi", "item": "ink"},
]

grouped = {}
for order in orders:
    customer = order["customer"]
    if customer not in grouped:
        grouped[customer] = []
    grouped[customer].append(order["item"])

print(grouped)
text
{'rafi': ['pen', 'ink'], 'ahmed': ['bag']}

The three middle lines are the whole technique: if the key is missing, put an empty list there first, then append without worrying. Without that, the very first order raises a KeyError.

The same thing fits on one line:

python
orders = [
    {"customer": "rafi", "item": "pen"},
    {"customer": "rafi", "item": "ink"},
]

grouped = {}
for order in orders:
    grouped.setdefault(order["customer"], []).append(order["item"])

print(grouped)
text
{'rafi': ['pen', 'ink']}

setdefault(key, []) says "give me this key's value, and if it has none put this empty list there and then give me that" — and whatever comes back is appended to. The difference from .get() matters: .get() puts nothing into the dictionary, setdefault does.

The three-line form reads more easily, the one-line form writes faster. You will meet both in real code.

Deeper still

python
shop = {
    "name": "Corner Store",
    "staff": [
        {"name": "rafi", "shifts": ["mon", "tue"]},
        {"name": "dia", "shifts": ["wed"]},
    ],
}

print(shop["staff"][0]["shifts"][1])
print(len(shop["staff"]))

for person in shop["staff"]:
    print(f"{person['name']}: {len(person['shifts'])} shift(s)")
text
tue
2
rafi: 2 shift(s)
dia: 1 shift(s)

shop["staff"][0]["shifts"][1] is four steps: dictionary → list → dictionary → list. Read left to right it is not complicated, only long.

Past three levels, stop and think. Breaking the chain and giving the middle step a name makes the code readable:

python
first_person = shop["staff"][0]
print(first_person["shifts"][1])

copy() is not enough for nested data

Chapter twelve showed that b = a leaves two names on one list, and that copy() is the fix. With nested data, copy() protects only one level:

python
import copy

template = {"name": "", "items": []}

a = template.copy()
b = copy.deepcopy(template)

a["items"].append("pen")
b["items"].append("bag")

print(template)
print(a)
print(b)
text
{'name': '', 'items': ['pen']}
{'name': '', 'items': ['pen']}
{'name': '', 'items': ['bag']}

a was made with copy() — a new outer dictionary, but the inner list is the same list. So a["items"].append("pen") changed the original template as well.

b was made with deepcopy(), which copies everything inside too. So changes to b spread nowhere.

copy is a module, so it needs import copy before use — modules in full arrive in chapter twenty-two.

When do you need deepcopy? When you are taking a copy of a nested structure in order to change it, and want the original left alone. Be aware that deepcopy is slow, and most of the time it is not needed — most code builds new data rather than modifying existing data.

A complete example

orders.py:

python
# The shape data actually arrives in: a list of records
orders = [
    {"customer": "rafi", "item": "pen", "quantity": 3, "price": 15.0},
    {"customer": "ahmed", "item": "bag", "quantity": 1, "price": 850.0},
    {"customer": "rafi", "item": "ink", "quantity": 2, "price": 120.0},
    {"customer": "dia", "item": "pen", "quantity": 5, "price": 15.0},
]

print("Customer  Item    Qty      Total")
grand_total = 0.0
for order in orders:
    total = order["quantity"] * order["price"]
    grand_total = grand_total + total
    print(f"{order['customer']:<9} {order['item']:<6} {order['quantity']:>3} {total:>10.2f}")

print(f"{'':<20}{grand_total:>10.2f}")
print()

# Group the spend by customer: the key may not exist yet, so .get supplies a start
spend = {}
for order in orders:
    customer = order["customer"]
    spend[customer] = spend.get(customer, 0.0) + order["quantity"] * order["price"]

# Sort by amount by putting the amount first in a tuple
pairs = []
for customer, amount in spend.items():
    pairs.append((amount, customer))
pairs.sort(reverse=True)

print("Spend per customer, highest first:")
for amount, customer in pairs:
    print(f"  {customer:<8} {amount:>8.2f}")

distinct = set()
for order in orders:
    distinct.add(order["item"])

print()
print("Distinct items:", sorted(distinct))
print("Best customer :", pairs[0][1])
text
Customer  Item    Qty      Total
rafi      pen      3      45.00
ahmed     bag      1     850.00
rafi      ink      2     240.00
dia       pen      5      75.00
                       1210.00

Spend per customer, highest first:
  ahmed      850.00
  rafi       285.00
  dia         75.00

Distinct items: ['bag', 'ink', 'pen']
Best customer : ahmed

This small program uses every container from the last four chapters, each exactly where it belongs.

A list for the orders, because the number is unknown and the order matters.

Dictionaries for each order, because every part needs a name; and for spend, because it accumulates by name.

Tuples for each pair in pairs. And there is a trick here: the amount is put first in the tuple, because Python sorts tuples by their first value. So a plain .sort(reverse=True) orders by money and no key= is needed.

A set for distinct, because the question was "how many different", and pen appears twice.

And spend.get(customer, 0.0) is chapter fifteen's counting pattern in another guise — adding an amount instead of adding one.


When it breaks

KeyError: 'discount' — but in which record? Walk with for i, order in enumerate(orders): so i is known when it fails. Or put print(order) inside the loop; the last one printed is the guilty one.

TypeError: string indices must be integers Almost always this means you thought you had a dictionary and you have a string. If orders is a list of names rather than a list of dictionaries, order["customer"] gives exactly this. Check with print(type(order)).

TypeError: list indices must be integers or slices, not str The opposite mistake — you have a list and are asking by name. Usually a [0] is missing: orders[0]["customer"], not orders["customer"].

AttributeError: 'list' object has no attribute 'get' The same family, different clothes. The thing you called .get() on is a list, not a dictionary.

I changed a copy and the original changed too copy() is shallow — inner lists and dictionaries stay shared. import copy and use copy.deepcopy(...).

A KeyError on the very first pass while grouping grouped[key].append(...) ran before the key existed. Create it with if key not in grouped: and an empty list, or use setdefault.