Reading and writing files
Reading and writing with `with open`, what separates r/w/a/x, the newline at the end of every line, why `encoding` is always written, and pathlib for the small jobs.
- 1Encounter
- 2Understand
- 3Worked
- 4Predict
- 5Apply
- 6Stretch
The problem we are solving
Everything we have built across twenty-three chapters disappeared the moment the program ended.
basket = []
basket.append("pen")
print(basket)['pen']Run it again and the basket is empty again. Lists, dictionaries, functions — all of it lives in memory, and memory lasts exactly as long as the program.
The last chapter was about sharing code between programs. This one is about data — how one run leaves its work for the next run, or for somebody else.
with open("notes.txt", "w", encoding="utf-8") as fh:
fh.write("first line\n")
fh.write("second line\n")
with open("notes.txt", "r", encoding="utf-8") as fh:
print(fh.read(), end="")first line
second lineThe program has ended, and notes.txt is still there.
By the end of this chapter you can
- Read and write files with
with open(...), and say whywith - Say what separates
"r","w","a"and"x"— especially what"w"destroys - Read a file one line at a time, and handle the
\nat the end of each - Say why
encoding="utf-8"is always written - Use
pathlib.Pathfor the small jobs - Read a
FileNotFoundErrorand aUnicodeDecodeError
Prerequisites: Modules and imports.
Why with
Opening a file borrows something from the operating system, and a loan has to be repaid. Without closing, what you wrote may never reach the disk, and enough unclosed files will stop the program.
A with block makes that repayment itself — the file is closed as the block ends, even if something inside it raised.
with open("notes.txt", "r", encoding="utf-8") as fh:
text = fh.read()
print(fh.closed)
print(fh.read())True
ValueError: I/O operation on closed file.The name fh still exists after the block — scope works exactly as before — but the file is shut. In this course open() is always written with with.
Reading
The whole contents at once come from .read():
with open("notes.txt", "r", encoding="utf-8") as fh:
text = fh.read()
print(repr(text))
print(len(text))'first line\nsecond line\n'
23repr is used deliberately, because a plain print hides the line breaks. A file is one long piece of text, with its line endings sitting inside it as \n — remembering that is half of this chapter.
One line at a time
.read() is a poor choice for a large file, because it lifts the whole thing into memory. Looping over the file itself makes Python read a line at a time:
with open("orders.txt", "r", encoding="utf-8") as fh:
for line in fh:
print(repr(line))'pen,15.0,3\n'
'bag,850.0,1\n'
'ink,120.0,2\n'Every line ends with \n. This is the most common confusion of all — int(line) does not work, line == "pen" does not match, and the reason is invisible.
So the first move is almost always .strip():
with open("orders.txt", "r", encoding="utf-8") as fh:
for line in fh:
name, price, qty = line.strip().split(",")
print(f"{name:<5} {float(price) * int(qty):>8.2f}")pen 45.00
bag 850.00
ink 240.00.strip() then .split(",") then float() and int() — that sequence will come back again and again when reading files. Everything that comes out of a file is text, never numbers; chapter seven's conversion is needed here too.
When all the lines are wanted at once, there are two ways:
with open("orders.txt", "r", encoding="utf-8") as fh:
lines = fh.read().splitlines()
print(lines)['pen,15.0,3', 'bag,850.0,1', 'ink,120.0,2'].splitlines() drops the line endings. There is also .readlines(), but it keeps the \n — which is where most people trip.
Writing, and one dangerous letter
with open("notes.txt", "w", encoding="utf-8") as fh:
fh.write("new\n")The file previously contained old content. Now:
new"w" empties the file — at the moment of opening, before anything is written. There is no warning and no way back.
To keep what is there and add to the end, "a":
with open("notes.txt", "a", encoding="utf-8") as fh:
fh.write("new\n")old content
newAnd to refuse to write at all if the file already exists, "x":
FileExistsError: [Errno 17] File exists: 'notes.txt'The four modes in one line: "r" reads, "w" erases and writes, "a" adds to the end, "x" only creates.
Pause before typing"w". People lose real work to this one letter. If the file you are opening was not made by your program, think before writing"w".
The \n is yours to supply
print breaks the line for you. write does not.
rows = ["pen", "bag"]
with open("out.txt", "w", encoding="utf-8") as fh:
for row in rows:
fh.write(row)'penbag'So the cleanest way to write a list of lines is:
rows = ["pen", "bag"]
with open("out.txt", "w", encoding="utf-8") as fh:
fh.write("\n".join(rows) + "\n")'pen\nbag\n'"\n".join(rows) puts breaks between the lines, and the trailing + "\n" ends the file on a newline — which is the convention for text files, and what many tools expect.
encoding="utf-8" — always
Without encoding, Python uses the operating system's default, which differs from machine to machine. The result is a program that works where you are and breaks where somebody else is.
Reading with the wrong one gives:
UnicodeDecodeError: 'utf-8' codec can't decode byte 0xff in position 3: invalid start byteSo write encoding="utf-8" every time you write open() — for reading and for writing alike. The only exception is binary mode ("rb", "wb"), where the question does not arise because there is no encoding.
pathlib — a short path for small jobs
For a small file the whole with block can feel like a lot.
from pathlib import Path
path = Path("orders.txt")
print(path.exists())
print(len(path.read_text(encoding="utf-8").splitlines()))
out = Path("report.txt")
out.write_text("done\n", encoding="utf-8")
print(out.read_text(encoding="utf-8"), end="")True
3
doneread_text and write_text open and close the file internally, so no with is needed. Path does more besides — .exists(), .name, .suffix, .parent — and folder paths are joined with /, as in Path("data") / "orders.txt".
Which when? Path when the whole file is wanted at once, with open(...) when reading line by line. write_text has one more advantage: it replaces the file by default, so the mistake of typing "w" is at least never accidental.
A complete example
orders.txt:
pen,15.0,3
bag,850.0,1
ink,120.0,2main.py:
"""Read an order file, total it, and write a report beside it."""
from pathlib import Path
TAX_RATE = 0.15
def read_order(path):
"""Each line is name,price,quantity. Blank lines are skipped."""
lines = []
for raw in path.read_text(encoding="utf-8").splitlines():
if not raw.strip():
continue
name, price, quantity = raw.split(",")
lines.append((name, float(price), int(quantity)))
return lines
def report(lines):
rows = []
total = 0.0
for name, price, quantity in lines:
amount = round(price * quantity * (1 + TAX_RATE), 2)
total += amount
rows.append(f"{name:<6} {quantity:>3} {amount:>9.2f}")
rows.append("-" * 20)
rows.append(f"{'total':<6} {'':>3} {total:>9.2f}")
return rows
def main():
order = read_order(Path("orders.txt"))
rows = report(order)
Path("report.txt").write_text("\n".join(rows) + "\n", encoding="utf-8")
print(Path("report.txt").read_text(encoding="utf-8"), end="")
print()
print("lines read:", len(order))
if __name__ == "__main__":
main()pen 3 51.75
bag 1 977.50
ink 2 276.00
--------------------
total 1305.25
lines read: 3Four things worth looking at.
Three functions with three separate jobs. read_order reads from disk, report only calculates, and main writes. report is given no file at all — it is given a list and returns a list. So testing it needs no file, and in chapter twenty-eight we will be grateful for exactly that.
The blank line is handled in advance. A stray newline at the end of a file is entirely normal, and without handling it raw.split(",") would not get three things and would raise a ValueError. if not raw.strip(): continue settles it in one line.
float(price) and int(quantity) are not optional, because everything out of a file arrives as text. Forget them and price * quantity would quietly repeat the text instead — the worst kind of bug, the kind that does not break.
The report goes to a separate file. orders.txt is never opened with "w" — the input file stays intact. That is a habit worth forming: do not write to what you are reading.
When it breaks
FileNotFoundError: [Errno 2] No such file or directory: 'orders.txt' The file is not there, or you are running the program from a different folder. The path is worked out from where you ran it, not from the script's folder. Print Path("orders.txt").resolve() to see where it is looking.
UnicodeDecodeError: 'utf-8' codec can't decode byte ... The file was written in another encoding, or it is not a text file at all. Fix the encoding, or open it with "rb" if it is binary.
ValueError: I/O operation on closed file. The file was used outside its with block. Read what you need inside the block.
ValueError: not enough values to unpack Some line does not have the commas you expected — often the empty line at the end. Skip blank lines, and when in doubt print(repr(raw)).
I wrote the file but it is empty It is being read before the with block has finished. Writes reach the disk when the file is closed.
All my lines ran together into one write does not break lines. The "\n" is yours to supply.
Everything in my file is gone It was opened with "w". Use "r" to read and "a" to add to the end.
Step 4 of 6 — Predict
Check your understanding
repr is used, so nothing is hidden. What is printed?
# orders.txt
# pen,15.0,3
# bag,850.0,1
# ink,120.0,2
with open("orders.txt", "r", encoding="utf-8") as fh:
for line in fh:
print(repr(line))- A'pen,15.0,3\n' 'bag,850.0,1\n' 'ink,120.0,2\n'
- B'pen,15.0,3' 'bag,850.0,1' 'ink,120.0,2'
- C'pen,15.0,3\nbag,850.0,1\nink,120.0,2\n'
- D['pen,15.0,3', 'bag,850.0,1', 'ink,120.0,2']
The file is opened with "w" but nothing is written. What is printed?
# notes.txt already contains: important work
with open("notes.txt", "w", encoding="utf-8") as fh:
print("opened")
with open("notes.txt", "r", encoding="utf-8") as fh:
print(repr(fh.read()))- Aopened ''
- Bopened 'important work\n'
- Copened 'opened\n'
- DA `FileExistsError`
Two things are written separately. What is in the file?
rows = ["pen", "bag"]
with open("out.txt", "w", encoding="utf-8") as fh:
for row in rows:
fh.write(row)
with open("out.txt", "r", encoding="utf-8") as fh:
print(repr(fh.read()))- A'penbag'
- B'pen\nbag\n'
- C'pen bag'
- D'pen\nbag'
Answering needs an account
Sign in to check your answers
The questions are above, and working them out in your head is the part that matters. Sign in to see the answers, the explanations and the three-level hints.
Your turn
Make a file called people.txt with a name and a number on each line, separated by a comma — at least four lines, and deliberately leave one blank line in the middle.
Then write summary.py with three functions: one reading the file into a list of pairs, one turning that list into lines of text (who has the largest number, and the total), and main writing it to summary.txt. Do not forget if __name__ == "__main__":.
Then five experiments:
- Remove the
.strip(). Which error appears, and why? - Delete a number from one line so the comma is there but the number is not. Read the message.
- Change
"w"to"a"inmain, then run the program three times. What happened tosummary.txt? - Try opening
people.txtwith"w"... do not. Make a copy and do it to that instead, then look at what is left. - Print
Path("people.txt").resolve(), then run the program from a different folder. Does the path change?
That last one explains a mistake almost everybody makes once: the program worked, and then it was run from somewhere else and raised FileNotFoundError.
Step 6 of 6
Stretch — the chapter quiz
Ten questions from easy to hard. The last ones are difficult on purpose.
Sign in to take the quiz