Python for AI: A Crash Course

Nov 16 2024 · Python 3.12, JupyterLab 4.2.4

Lesson 04: Working with Local Data (File Operations & Data Handling)

File Handling Demo

Episode complete

Play next episode

Next
Transcript

Demo

In this demo, you’ll write code to read and write text files. Or, more accurately, you’ll write code that writes a text file and then write code that reads that file.

Writing To a Text File

Start by writing a text file to your computer’s filesystem. That way, you’ll have a text file you can read for the file-reading part of this exercise.

Open the working-wth-files-starter.ipynb notebook. You’ll see that its first code cell defines a variable named programming_languages. Run the programming_languages code cell. The variable contains an array of dictionaries. Each dictionary describes a programming language. You’ll use programming_languages as the basis for the text files you’ll create in this section and for the CSV and JSON files you’ll create later.

Writing To a Text File the Old Way

The old way of writing text files in Python—which happens to be the current way of writing text files in most other programming languages—is to open a file, write to that file, and then close it. Do that now using the open() function and file.close() method.

Scroll to the Working with Text Files section of the notebook and enter the following into a new code cell:

file = open("programming-languages.txt", "w")
file.write("Programming Languages")
file.write("=====================")
file.close()

Run the cell, open programming-languages.txt in your favorite text editor or JupyterLab, and confirm that it contains just one line:

Programming Languages=====================

Unlike the print() function, which automatically adds a newline character to the end of everything it outputs, you need to add the newline to the end of lines that you write to a file with file.write().

Also, note that the file programming-languages.txt didn’t exist until now. That’s what writing to a non-existent file in "w" mode does — it creates the file before writing to it.

Update the code in the cell so that each of the strings in the calls to file.write() ends with a newline (\n) character. It’ll end up looking like this:

file = open("programming-languages.txt", "w")
file.write("Programming Languages\n")
file.write("=====================\n")
file.close()

Run the cell, and then open programming-languages.txt. Its contents will now look like this:

Programming Languages
=====================

Note that when you ran the code with the corrections, it overwrote the original contents of programming-languages.txt, replacing them with the new content, where each line ends with a newline character. That’s what writing to an existing file in "w" mode does — overwrites the previous content.

Writing To a Text File the Preferred Way

In the preferred way, you perform file operations inside a with block, perform setup operations at the start of the block and clean up operations once the block has completed executing. In the case of files, with automatically closes the file at the end of the block.

The block approach is less error-prone because you no longer have to consider closing the file. It has the added benefit of making you think a little more about the file operations you’re performing because they all have to happen inside the with block.

Try it out by running the following in a new code cell:

with open("programming-languages.txt", "w") as file:
  file.write("Programming Languages\n")
  file.write("=====================\n")

  for language in programming_languages:
    line = (
      f"{language["name"]} was created by {language["creator"]} " +
      f"and first appeared in {language["year_appeared"]}.\n"
    )
    file.write(line)

Open programming-languages.txt. It should contain the following:

Programming Languages
=====================
Python was created by Guido van Rossum and first appeared in 1991.
Miranda was created by David Turner and first appeared in 1985.
Ruby was created by Yukihiro Matsumoto and first appeared in 1995.
Rebol was created by Carl Sassenrath and first appeared in 1997.
Swift was created by Chris Lattner and first appeared in 2014.
ActionScript was created by Gary Grossman and first appeared in 1998.
Kotlin was created by JetBrains and first appeared in 2011.
CoffeeScript was created by Jeremy Ashkenas and first appeared in 2009.

Appending To a Text File

Now, add to programming-languages.txt without overwriting any existing content. You can do this by writing to it in "a" (append) mode.

Enter the following into a code cell and run it:

old_school_languages = [
  "ALGOL\n"
  "BASIC\n",
  "COBOL\n",
  "FORTRAN\n",
  "Lisp\n"
]

with open("programming-languages.txt", "a") as file:
  file.write("Let's not forget the old guard:\n")
  file.writelines(old_school_languages)

The code above uses the writelines() method, which takes a list of strings and writes them to the file in the order in which they appear in the list.

If you open programming-languages.txt now, you’ll see that it has these contents:

Programming Languages
=====================
Python was created by Guido van Rossum and first appeared in 1991.
Miranda was created by David Turner and first appeared in 1985.
Ruby was created by Yukihiro Matsumoto and first appeared in 1995.
Rebol was created by Carl Sassenrath and first appeared in 1997.
Swift was created by Chris Lattner and first appeared in 2014.
ActionScript was created by Gary Grossman and first appeared in 1998.
Kotlin was created by JetBrains and first appeared in 2011.
CoffeeScript was created by Jeremy Ashkenas and first appeared in 2009.
Let's not forget the old guard:
ALGOL
BASIC
COBOL
FORTRAN
Lisp

Note that each string in old_school_languages, the variable containing the list passed to writelines(), ends with a newline character. Without that newline character, the old-school languages would been written to the file this way:

ALGOLBASICCOBOLFORTRANLisp

Reading From a Text File

You started this exercise by writing a text file so you’d have one to read. Not, it’s time to read it!

Enter the following into a code cell and run it:

with open("programming-languages.txt", "r") as file:
  data = file.read()
print(data)

You’ll see the contents of programming-languages.txt:

Programming Languages
=====================
Python was created by Guido van Rossum and first appeared in 1991.
Miranda was created by David Turner and first appeared in 1985.
Ruby was created by Yukihiro Matsumoto and first appeared in 1995.
Rebol was created by Carl Sassenrath and first appeared in 1997.
Swift was created by Chris Lattner and first appeared in 2014.
ActionScript was created by Gary Grossman and first appeared in 1998.
Kotlin was created by JetBrains and first appeared in 2011.
CoffeeScript was created by Jeremy Ashkenas and first appeared in 2009.
Let's not forget the old guard:
ALGOL
BASIC
COBOL
FORTRAN
Lisp

The code above reads the entire file into memory, putting its contents into a single string and assigns it to the variable data.

If you’d rather work with the contents of a file as a list of lines instead of as a single string, you can use the readlines() method. Try it out by running this in a new code cell:

with open("programming-languages.txt", "r") as file:
  data = file.readlines()
print(data)

This time, data contains a list of strings, each representing one line from the file, complete with the newline character at the end.

Note: When reading lines from a file, a “line” is considered to be a string that ends with a newline character. The newline character is part of the line.

In both examples above, the file’s entire contents are read into memory. This approach works when the file is small enough to fit into memory, but that won’t always be the case. AI thrives on large datasets, and the general rule seems to be “the larger the dataset, the better.”

Fortunately, there’s an alternate approach: you can run the file one line at a time using a for loop and the read() method. Run the following in a new code cell:

with open("programming-languages.txt", "r") as file:
  for line in file:
    print(line)

This works because the file object file does more than give you access to the file; it also acts as an iterator, allowing you to retrieve the contents of the file one line at a time. This approach requires considerably less memory than reading the entire fill simultaneously.

Note that the output of the code above is double-spaced:

Programming Languages

=====================

Python was created by Guido van Rossum and first appeared in 1991.

Miranda was created by David Turner and first appeared in 1985.

Ruby was created by Yukihiro Matsumoto and first appeared in 1995.

Rebol was created by Carl Sassenrath and first appeared in 1997.

Swift was created by Chris Lattner and first appeared in 2014.

ActionScript was created by Gary Grossman and first appeared in 1998.

Kotlin was created by JetBrains and first appeared in 2011.

CoffeeScript was created by Jeremy Ashkenas and first appeared in 2009.

Let's not forget the old guard:

ALGOL

BASIC

COBOL

FORTRAN

Lisp

That’s because there are two newline characters at the end of each line:

  1. There’s the newline at the end of each line in the file.
  2. There’s also the newline that the print() function automatically adds to the end of what it prints.

The simplest way to convert the code’s output to single-spaced is by removing the newline character at the end of each line using the string class’ rstrip() method, which removes any whitespace from the right side of the string. Try running the following in a new code cell to see the result:

with open("programming-languages.txt", "r") as file:
  for line in file:
    print(line.rstrip())

Handling Exceptions

For the sake of simplicity, the previous examples in this exercise have ignored exception handling. It’s generally a good idea to incorporate it into file-handling code.

Deliberately cause a “file not found” exception by trying to read a non-existent file. Enter the following into a new code cell and run it:

try:
  with open("does-not-exist.txt", "r") as file:
    data = file.read()
except FileNotFoundError as e:
  print(f"File not found! Details:\n{e}")
except OSError as e:
  print(f"I/O error (probably)! Details:\n{e}")
except Exception as e:
  print("An unexpected error occurred! Call the developer.")
  print(f"Details:\n{e}")

You’ll see this error message:

File not found! Details:
[Errno 2] No such file or directory: 'does-not-exist.txt'

Other types of exceptions are harder to recreate, but you can simulate them with the raise expression, which is like Ruby’s raise, or as it’s called in C#, Java, JavScript, Kotlin, and Swift, throw.

Here’s some code that will successfully display the contents of programming-languages.txt two-thirds of the time but end with an I/O error one-third of the time. Try it out by entering it into a new code cell and running it a few times:

import random

try:
  with open("programming-languages.txt", "r") as file:
    if random.randint(1, 3) == 1:
      raise OSError("Disk error")
    print(file.read())
except FileNotFoundError as e:
  print(f"File not found! Details:\n{e}")
except OSError as e:
  print(f"I/O error (probably)! Details:\n{e}")
except Exception as e:
  print("An unexpected error occurred! Call the developer.")
  print(f"Details:\n{e}")
else:
  print("Congratulations! No errors!")
finally:
  print("All done.")
See forum comments
Cinema mode Download course materials from Github
Previous: File Handling Next: CVS Files