Nested data — a list of dictionaries
The shape real data arrives in: dictionaries inside lists, lists inside dictionaries. Reaching in, grouping, and why copy() is not enough here.
- 1Encounter
- 2Understand
- 3Worked
- 4Predict
- 5Apply
- 6Stretch
The problem we are solving
The last three chapters taught four containers — lists, tuples, dictionaries, sets. Every example used one of them, one level deep.
Real data does not arrive that way. A list of orders looks like this:
customer item quantity price
rafi pen 3 15.0
ahmed bag 1 850.0There are two levels here. On the outside a collection — the number of orders is not known, so a list. Inside, each order is a record whose every part has a name — so a dictionary.
What comes back from an API, what you get from reading a CSV, what a database returns: almost always this shape, a list of dictionaries. Today's chapter introduces no new container. It is about putting the ones you have inside each other, and recognising the mistakes that follow.
By the end of this chapter you can
- Reach into a list of dictionaries and walk over it
- Read a deep
KeyErrororIndexErrorand tell which level went wrong - Group and summarise a list of dictionaries
- Handle incomplete records with
.get()'s default - Explain why
copy()is not enough for nested data
Prerequisites: Sets — a collection without duplicates.
Recognising the shape
orders = [
{"customer": "rafi", "item": "pen", "quantity": 3, "price": 15.0},
{"customer": "ahmed", "item": "bag", "quantity": 1, "price": 850.0},
]
print(len(orders))
print(orders[0])
print(orders[0]["customer"])
print(orders[1]["price"])2
{'customer': 'rafi', 'item': 'pen', 'quantity': 3, 'price': 15.0}
rafi
850.0The line orders[0]["customer"] is the core skill of this chapter, and it is read left to right, one step at a time:
orders— a listorders[0]— its first item, which is a dictionaryorders[0]["customer"]— that dictionary'scustomervalue
When in doubt, stop and print the middle step. Writing print(orders[0]) to see that it is a dictionary is one habit that removes half the trouble with nested data.
Walking it
orders = [
{"customer": "rafi", "item": "pen", "quantity": 3, "price": 15.0},
{"customer": "ahmed", "item": "bag", "quantity": 1, "price": 850.0},
]
for order in orders:
total = order["quantity"] * order["price"]
print(f"{order['customer']:<8} {order['item']:<6} {total:>8.2f}")rafi pen 45.00
ahmed bag 850.00for order in orders gives a dictionary each time round, and inside it you reach things by name. There is no index arithmetic anywhere.
Quotes inside an f-string. f"{order['customer']}" uses single quotes inside because the f-string itself is written with double quotes. Use the same kind inside and out and Python cannot tell where the text ends. Keep one kind inside and the other outside.Incomplete records
In real data, not every record has every key:
orders = [{"customer": "rafi"}]
print(orders[0]["discount"])KeyError: 'discount'The KeyError says the key is missing, but not which record is missing it. When that happens in one order out of twelve thousand it is painful to find — which is why walking with enumerate keeps the number to hand.
When a key being absent is normal, .get():
orders = [{"customer": "rafi"}]
print(orders[0].get("discount", 0))0Chapter fifteen's rule applies here too: square brackets when it should be there, .get() when its absence is normal. The temptation is stronger with nested data, because .get() can silence every error — and then incomplete data quietly becomes zero and walks into your arithmetic.
The other shape — lists inside a dictionary
by_customer = {
"rafi": ["pen", "ink"],
"ahmed": ["bag"],
}
by_customer["rafi"].append("bottle")
by_customer["dia"] = ["eraser"]
print(by_customer)
for customer, items in by_customer.items():
print(f"{customer}: {len(items)} item(s) - {', '.join(items)}"){'rafi': ['pen', 'ink', 'bottle'], 'ahmed': ['bag'], 'dia': ['eraser']}
rafi: 3 item(s) - pen, ink, bottle
ahmed: 1 item(s) - bag
dia: 1 item(s) - eraserThe line by_customer["rafi"].append("bottle") is two steps in one: by_customer["rafi"] pulls out a list, and then append adds to that list.
This shape answers a problem from chapter fifteen, where a key could only exist once and so one person could not have several records. Make the value a list and they can.
Grouping — the most useful pattern
Turning a list of dictionaries into "who bought what":
orders = [
{"customer": "rafi", "item": "pen"},
{"customer": "ahmed", "item": "bag"},
{"customer": "rafi", "item": "ink"},
]
grouped = {}
for order in orders:
customer = order["customer"]
if customer not in grouped:
grouped[customer] = []
grouped[customer].append(order["item"])
print(grouped){'rafi': ['pen', 'ink'], 'ahmed': ['bag']}The three middle lines are the whole technique: if the key is missing, put an empty list there first, then append without worrying. Without that, the very first order raises a KeyError.
The same thing fits on one line:
orders = [
{"customer": "rafi", "item": "pen"},
{"customer": "rafi", "item": "ink"},
]
grouped = {}
for order in orders:
grouped.setdefault(order["customer"], []).append(order["item"])
print(grouped){'rafi': ['pen', 'ink']}setdefault(key, []) says "give me this key's value, and if it has none put this empty list there and then give me that" — and whatever comes back is appended to. The difference from .get() matters: .get() puts nothing into the dictionary, setdefault does.
The three-line form reads more easily, the one-line form writes faster. You will meet both in real code.
Deeper still
shop = {
"name": "Corner Store",
"staff": [
{"name": "rafi", "shifts": ["mon", "tue"]},
{"name": "dia", "shifts": ["wed"]},
],
}
print(shop["staff"][0]["shifts"][1])
print(len(shop["staff"]))
for person in shop["staff"]:
print(f"{person['name']}: {len(person['shifts'])} shift(s)")tue
2
rafi: 2 shift(s)
dia: 1 shift(s)shop["staff"][0]["shifts"][1] is four steps: dictionary → list → dictionary → list. Read left to right it is not complicated, only long.
Past three levels, stop and think. Breaking the chain and giving the middle step a name makes the code readable:
first_person = shop["staff"][0]
print(first_person["shifts"][1])copy() is not enough for nested data
Chapter twelve showed that b = a leaves two names on one list, and that copy() is the fix. With nested data, copy() protects only one level:
import copy
template = {"name": "", "items": []}
a = template.copy()
b = copy.deepcopy(template)
a["items"].append("pen")
b["items"].append("bag")
print(template)
print(a)
print(b){'name': '', 'items': ['pen']}
{'name': '', 'items': ['pen']}
{'name': '', 'items': ['bag']}a was made with copy() — a new outer dictionary, but the inner list is the same list. So a["items"].append("pen") changed the original template as well.
b was made with deepcopy(), which copies everything inside too. So changes to b spread nowhere.
copy is a module, so it needs import copy before use — modules in full arrive in chapter twenty-two.
When do you needdeepcopy? When you are taking a copy of a nested structure in order to change it, and want the original left alone. Be aware thatdeepcopyis slow, and most of the time it is not needed — most code builds new data rather than modifying existing data.
A complete example
orders.py:
# The shape data actually arrives in: a list of records
orders = [
{"customer": "rafi", "item": "pen", "quantity": 3, "price": 15.0},
{"customer": "ahmed", "item": "bag", "quantity": 1, "price": 850.0},
{"customer": "rafi", "item": "ink", "quantity": 2, "price": 120.0},
{"customer": "dia", "item": "pen", "quantity": 5, "price": 15.0},
]
print("Customer Item Qty Total")
grand_total = 0.0
for order in orders:
total = order["quantity"] * order["price"]
grand_total = grand_total + total
print(f"{order['customer']:<9} {order['item']:<6} {order['quantity']:>3} {total:>10.2f}")
print(f"{'':<20}{grand_total:>10.2f}")
print()
# Group the spend by customer: the key may not exist yet, so .get supplies a start
spend = {}
for order in orders:
customer = order["customer"]
spend[customer] = spend.get(customer, 0.0) + order["quantity"] * order["price"]
# Sort by amount by putting the amount first in a tuple
pairs = []
for customer, amount in spend.items():
pairs.append((amount, customer))
pairs.sort(reverse=True)
print("Spend per customer, highest first:")
for amount, customer in pairs:
print(f" {customer:<8} {amount:>8.2f}")
distinct = set()
for order in orders:
distinct.add(order["item"])
print()
print("Distinct items:", sorted(distinct))
print("Best customer :", pairs[0][1])Customer Item Qty Total
rafi pen 3 45.00
ahmed bag 1 850.00
rafi ink 2 240.00
dia pen 5 75.00
1210.00
Spend per customer, highest first:
ahmed 850.00
rafi 285.00
dia 75.00
Distinct items: ['bag', 'ink', 'pen']
Best customer : ahmedThis small program uses every container from the last four chapters, each exactly where it belongs.
A list for the orders, because the number is unknown and the order matters.
Dictionaries for each order, because every part needs a name; and for spend, because it accumulates by name.
Tuples for each pair in pairs. And there is a trick here: the amount is put first in the tuple, because Python sorts tuples by their first value. So a plain .sort(reverse=True) orders by money and no key= is needed.
A set for distinct, because the question was "how many different", and pen appears twice.
And spend.get(customer, 0.0) is chapter fifteen's counting pattern in another guise — adding an amount instead of adding one.
When it breaks
KeyError: 'discount' — but in which record? Walk with for i, order in enumerate(orders): so i is known when it fails. Or put print(order) inside the loop; the last one printed is the guilty one.
TypeError: string indices must be integers Almost always this means you thought you had a dictionary and you have a string. If orders is a list of names rather than a list of dictionaries, order["customer"] gives exactly this. Check with print(type(order)).
TypeError: list indices must be integers or slices, not str The opposite mistake — you have a list and are asking by name. Usually a [0] is missing: orders[0]["customer"], not orders["customer"].
AttributeError: 'list' object has no attribute 'get' The same family, different clothes. The thing you called .get() on is a list, not a dictionary.
I changed a copy and the original changed too copy() is shallow — inner lists and dictionaries stay shared. import copy and use copy.deepcopy(...).
A KeyError on the very first pass while grouping grouped[key].append(...) ran before the key existed. Create it with if key not in grouped: and an empty list, or use setdefault.
Step 4 of 6 — Predict
Check your understanding
Two levels deep. What is printed?
orders = [
{"customer": "rafi", "item": "pen"},
{"customer": "ahmed", "item": "bag"},
]
print(orders[1]["item"])- Abag
- Bpen
- Cahmed
- DA `KeyError`
An attempt to group. What happens on the very first pass?
orders = [
{"customer": "rafi", "item": "pen"},
]
grouped = {}
for order in orders:
grouped[order["customer"]].append(order["item"])
print(grouped)- A`KeyError: 'rafi'` — the key does not exist yet, so there is nothing to append to
- BIt works and prints `{'rafi': ['pen']}`
- CAn `AttributeError` — a dictionary has no `append`
- DIt prints `{}`
A copy() was taken and two changes made to it. What happened to the original?
import copy
template = {"name": "", "items": []}
a = template.copy()
a["name"] = "rafi"
a["items"].append("pen")
print(template)- A{'name': '', 'items': ['pen']}
- B{'name': '', 'items': []}
- C{'name': 'rafi', 'items': ['pen']}
- D{'name': 'rafi', 'items': []}
Answering needs an account
Sign in to check your answers
The questions are above, and working them out in your head is the part that matters. Sign in to see the answers, the explanations and the three-level hints.
Your turn
Write a file called library.py holding a list of at least six books, each one a dictionary with a title, an author, a year and a number of copies.
Then:
- Print every book in aligned columns
- Work out the total copies and the average year
- Group by author into
{"author": ["title", "title"]}and print how many books each has - Count the distinct authors with a set
- Print the title of the book with the most copies
Then break it the way real data does:
- Delete the
"year"key from one book entirely and run the program. Which line stopped it, and did the message tell you which book was at fault? - Now use
.get("year", 0)on that line and run it again. What is the average year now, and is it true?
That last question is the real lesson. .get() stopped the error, but treating a missing year as zero and averaging it is a wrong answer arrived at silently. Removing an error is not the same as solving a problem.
Step 6 of 6
Stretch — the chapter quiz
Ten questions from easy to hard. The last ones are difficult on purpose.
Sign in to take the quiz