Comprehensions — a loop on one line
List, dictionary and set comprehensions, the difference between a filtering if and a choosing if/else, and when a comprehension is the wrong choice.
- 1Encounter
- 2Understand
- 3Worked
- 4Predict
- 5Apply
- 6Stretch
The problem we are solving
One shape of code has kept reappearing over the last six chapters:
marks = [72, 45, 90, 33]
doubled = []
for mark in marks:
doubled.append(mark * 2)
print(doubled)
doubled = [mark * 2 for mark in marks]
print(doubled)[144, 90, 180, 66]
[144, 90, 180, 66]The four lines above and the one line below do exactly the same thing.
Three of those four lines are not the work, they are the bookkeeping — make an empty list, walk over the source, append each time. The only real decision is mark * 2. Everything else is the same every time.
A comprehension removes that repetition. It is not a new ability — it is what you could already write, in fewer words. Which is also where its limit comes from, and that is the subject of this chapter's last section: shorten what becomes clearer for being shortened, not everything.
By the end of this chapter you can
- Write and read a list comprehension
- Filter with
if, and choose withif/else - Write dictionary and set comprehensions
- Say when a comprehension is the wrong choice
- Take an unreadable comprehension apart back into a loop
Prerequisites: Nested data — a list of dictionaries.
The shape
[ what I want for each thing in where it comes from ]Read aloud it is an ordinary sentence: "`mark 2 for each mark in marks`"*.
One way to hold it: read the for part first, then the beginning. Knowing where things come from makes what is being built easy to see.
Filtering
An if on the end keeps only the things that match:
marks = [72, 45, 90, 33]
passed = [mark for mark in marks if mark >= 40]
print(passed)[72, 45, 90]Here the first part is just mark — no calculation, taken as it is. Filtering and transforming can be done together:
names = [" rafi ", "AHMED", " dia"]
clean = [name.strip().lower() for name in names if name.strip()]
print(clean)['rafi', 'ahmed', 'dia']The condition is name.strip() with no comparison at all. Recall chapter eight's truthiness rule: empty text is false. So somebody who typed only spaces is dropped.
if at the end versus if at the start
Two different things, and the position is the difference:
marks = [72, 33, 90]
labels = ["pass" if m >= 40 else "fail" for m in marks]
print(labels)['pass', 'fail', 'pass']An if at the end filters — some things never reach the result, so the result can be shorter.
An if ... else at the start chooses — every thing produces something, so the length is unchanged. It is the one-line condition from chapter fifteen, inside a comprehension.
Does the length change? That question is the easiest way to tell the two apart.
Dictionaries and sets too
The same shape, different brackets:
marks = {"rafi": 72, "ahmed": 45, "bilal": 90}
passed = {name: mark for name, mark in marks.items() if mark >= 50}
print(passed)
doubled = {name: mark * 2 for name, mark in marks.items()}
print(doubled){'rafi': 72, 'bilal': 90}
{'rafi': 144, 'ahmed': 90, 'bilal': 180}Curly brackets and a colon make a dictionary comprehension. marks.items() hands back a tuple each time round and name, mark unpacks it, exactly as in a loop.
The same curly brackets without a colon make a set:
orders = [
{"item": "pen"},
{"item": "bag"},
{"item": "pen"},
]
items = {order["item"] for order in orders}
print(sorted(items))['bag', 'pen']Chapter sixteen's "how many different" question in one line. pen arrives twice and appears once.
Numbers into text — which you have already seen
This line appeared in chapter thirteen with a promise that the explanation was coming:
marks = [72, 45, 90]
print(", ".join(str(m) for m in marks))
print([str(m) for m in marks])72, 45, 90
['72', '45', '90']Now it can be read in full. str(m) for m in marks is a comprehension without the square brackets — that form is called a generator expression, and rather than building the whole list at once it hands join one item at a time. The brackets can be dropped when it is a function's only argument, and on large data it saves memory.
When not to write one
Being able to write a comprehension does not mean you should. Look at this:
rows = [["1", "2"], ["3", "x"]]
numbers = [int(c) for row in rows for c in row if c.isdigit()]
print(numbers)[1, 2, 3]The line is correct. But it has two fors and an if, and you have to remember the order: the fors read outermost to innermost, the way nested loops would be written. Six months later that line is work to read.
As a loop:
rows = [["1", "2"], ["3", "x"]]
numbers = []
for row in rows:
for c in row:
if c.isdigit():
numbers.append(int(c))
print(numbers)[1, 2, 3]More lines, less thinking. The goal of code is not to be short but to be clear — and most of the time shortening makes it clearer, but not always.
As a working rule: one for and at most one if — a comprehension. More than that — a loop. And never put a print or an append inside a comprehension; it is a thing for building a value, not for doing work.
A complete example
views.py:
# One list of records in, three different views out
orders = [
{"customer": "rafi", "item": "pen", "quantity": 3, "price": 15.0},
{"customer": "ahmed", "item": "bag", "quantity": 1, "price": 850.0},
{"customer": "rafi", "item": "ink", "quantity": 2, "price": 120.0},
{"customer": "dia", "item": "pen", "quantity": 5, "price": 15.0},
]
totals = [order["quantity"] * order["price"] for order in orders]
print("Totals :", totals)
big = [order["item"] for order in orders if order["quantity"] * order["price"] > 100]
print("Over 100 :", big)
items = {order["item"] for order in orders}
print("Distinct :", sorted(items))
by_item = {order["item"]: order["price"] for order in orders}
print("Price list:", by_item)
print("Grand :", sum(totals))Totals : [45.0, 850.0, 240.0, 75.0]
Over 100 : ['bag', 'ink']
Distinct : ['bag', 'ink', 'pen']
Price list: {'pen': 15.0, 'bag': 850.0, 'ink': 120.0}
Grand : 1210.0The same data as chapter seventeen, with most of what was a loop there now one line.
Three things worth noticing.
totals and sum(totals) are separate lines. The total could have been sum(order["quantity"] * order["price"] for order in orders), but then the individual line values would not exist. An intermediate name is often worth more than a one-line trick.
The same calculation appears twice inside big — once in the condition, and it would be needed again in the result. Here the result only takes the name, so nothing is lost; but when you need both, that is a good reason to leave the comprehension and write a loop, where the value can be computed once and given a name.
by_item lost a record. Four orders, three pairs in the result — pen appears twice, and by chapter fifteen's rule the later one quietly replaced the earlier. Here both pens have the same price so no harm was done; with different prices it would be silently wrong. A comprehension does not create that problem, but being short gives it less chance to be noticed.
When it breaks
SyntaxError from putting if and else at the end [m for m in marks if m >= 40 else 0] is not valid. The filtering if goes at the end and has no else; the choosing if ... else goes at the front: [m if m >= 40 else 0 for m in marks].
The result is full of None Something inside the comprehension returns nothing — such as [items.append(x) for x in other]. append returns None, so the result is a list of None. A comprehension builds values; to do work, write an ordinary loop.
NameError — the comprehension's variable is gone outside it After [m * 2 for m in marks] you cannot use m. In chapter eleven a for loop's variable survived the loop; a comprehension's does not — it lives only inside.
A dictionary comprehension produced fewer pairs than expected The same key was produced more than once and the later one replaced the earlier. Compare with len().
I cannot read the line That is reason enough. Break it into a loop — the code will not run slower, and the next person to read it, probably you, will be grateful.
Step 4 of 6 — Predict
Check your understanding
There is an if at the end. What is printed?
marks = [72, 45, 90, 33]
passed = [mark for mark in marks if mark >= 40]
print(passed)- A[72, 45, 90]
- B[72, 45, 90, 33]
- C[True, True, True, False]
- D[33]
This time the if is at the front, with an else. What is printed?
marks = [72, 33, 90]
labels = ["pass" if m >= 40 else "fail" for m in marks]
print(labels)- A['pass', 'fail', 'pass']
- B['pass', 'pass']
- C[72, 90]
- D['pass', 'fail']
append is called inside a comprehension. What do the two lines print?
source = ["pen", "bag"]
target = []
result = [target.append(item) for item in source]
print(target)
print(result)- A['pen', 'bag'] [None, None]
- B['pen', 'bag'] ['pen', 'bag']
- C[] ['pen', 'bag']
- DA `TypeError`
Answering needs an account
Sign in to check your answers
The questions are above, and working them out in your head is the part that matters. Sign in to see the answers, the explanations and the three-level hints.
Your turn
Write a file called clean.py with a list of messy names — some with extra spaces, some capitalised, and at least one that is entirely blank.
Using comprehensions, build:
- A list of clean names — trimmed, lowercased, with the blanks dropped
- A list of the length of each name
- A list of the names longer than five characters
- A dictionary of
{name: length} - A set of the first letters
Then two experiments:
- Write number 1 again as an ordinary
forloop, and check the two results are equal with== - Write a comprehension with two
fors and anif, then take it apart into a loop. Which is easier to read?
That last question has no "right" answer, and that is the point. Learning to ask it is the real lesson of this chapter.
Step 6 of 6
Stretch — the chapter quiz
Ten questions from easy to hard. The last ones are difficult on purpose.
Sign in to take the quiz