All in One View

Content from Command Line


Last updated on 2026-10-07 | Edit this page

Estimated time: 25 minutes

Overview

Questions

  • How can I make my results easier to reproduce?

Objectives

  • Introduce Zipf’s Law project.
  • Introduce git tags and branches
  • Illustrate pain and fragility of building with editor and sequences of command line commands.

Let’s imagine that you are writing code for a research project to test Zipf’s Law.

Callout

Zipf’s Law

The most frequently-occurring word occurs approximately twice as often as the second most frequent word. This is [Zipf’s Law][zipfs-law].

We have setup a git repository with a sequence of branches to step through the tutorial today. You should start by cloning the repository

BASH

$ git clone repository work_dir

OUTPUT

Cloning into 'work_dir'...
done.

Change to the new directory, checkout the branch 06-documentation, and list the branches

BASH

$ cd work_dir
$ git checkout 06-documentation
$ git branch -a

OUTPUT

* 06-documentation
  remotes/origin/01-command-line
  remotes/origin/02-first-makefile
  remotes/origin/03-no-hands
  remotes/origin/04-testing
  remotes/origin/05-standards
  remotes/origin/06-documentation
  remotes/origin/HEAD -> origin/06-documentation
  remotes/origin/main

Today, we will start with the branch 01-command-line and step through the numbered branches arriving at the branch 06-documentation which uses all of the tool we will introduce.

First let’s peek at the final results

BASH

$ make

View the result with (depending on your OS)

BASH

$ evince build/report.pdf

BASH

$ make docs

View the result with (depending on your OS)

BASH

$ firefox build/docs/html/index.html

Let’s take a couple of minutes to skim the document.

Now clean up and start with the 01-command-line branch

BASH

$ make clean
$ git switch -c my-01-branch remotes/origin/01-command-line
$ ls

The output of the last command should be

OUTPUT

README  abyss.txt  analyze.py  isles.txt  local.bib  report.tex

Now build the document pdf by following the instructions in the README file:

First run the analysis script and display the first ten lines of the results in abyss.dat

BASH

$ python analyze.py abyss.txt abyss.dat abyss.pdf
$ python analyze.py isles.txt isles.dat isles.pdf
$ head -10 abyss.dat

Then use an editor to copy results from the screen into report.tex

BASH

$ emacs report.tex

Now run latexmk to build the pdf

BASH

$ latexmk -pdf report.tex

If you are not sure you did it right, you can start over by running:

BASH

$ rm *.dat *.pdf *.aux *.bbl *.blg *.fdb_latexmk *.fls

or if you have clobbered some source files:

$ rm * ; git checkout .
Key Points
  • We hope to improve on executing commands from a README

Content from First Makefile


Last updated on 2026-10-07 | Edit this page

Estimated time: 40 minutes

Overview

Questions

  • How do I write a simple Makefile?

Objectives

  • Introduce key parts of a Makefile, rules, targets, dependencies and actions.
  • Run Make from the shell.
  • Edit a Makefile
  • Explain when and why to mark targets as .PHONY.

Save your work from the previous episode and checkout the branch for this episode:

BASH

$ git add report.tex
$ git commit -m "my work"
$ git switch -c my-02-branch remotes/origin/02-first-makefile
$ git ls-files

which should yield

OUTPUT

.gitignore
Makefile
README
abyss.txt
analyze.py
isles.txt
local.bib
report.tex

Look at the file called Makefile with

BASH

$ cat Makefile

OUTPUT

SHELL=bash

# First target, report.pdf, is the default.

report.pdf : local.bib report.tex abyss.pdf isles.pdf results.tex abyss.head
	latexmk -pdf report.tex

# Using a new feature of gnu-make as of version 4.3 that permits making
# several targets with one rule.  See
# https://www.gnu.org/software/make/manual/make.html#Multiple-Targets

abyss.dat abyss.pdf abyss.two &: analyze.py abyss.txt
	python analyze.py abyss.txt abyss.dat abyss.pdf |tee abyss.two

isles.dat isles.pdf isles.two &: analyze.py isles.txt
	python analyze.py isles.txt isles.dat isles.pdf |tee isles.two

results.tex: abyss.two isles.two
	@echo Create $@ with an editor from $^ ; exit 1

abyss.head: abyss.dat
	@echo Write a rule in the Makefile to create $@ from $^ ; exit 1

This is a build file, which for Make is called a Makefile - a file executed by Make. All of the indentation in this Makefile consists of TABs. It is unfortunate that while the difference between TABs and multiple SPACEs is invisible, the difference is significant for the program make.

If we try to run make we get

BASH

$ make

OUTPUT

python analyze.py abyss.txt abyss.dat abyss.pdf |tee abyss.two
top_two=[4044, 2807]
python analyze.py isles.txt isles.dat isles.pdf |tee isles.two
top_two=[3822, 2460]
Create results.tex with an editor from abyss.two isles.two
make: *** [Makefile:19: results.tex] Error 1

Let us talk about what happened.

If make is invoked without arguments, it uses the default name Makefile for instructions. In Makefile, since the first target is report.pdf, make will try to build that. Now we will talk about each of the elements of the rule

report.pdf : local.bib report.tex abyss.pdf isles.pdf results.tex abyss.head
	latexmk -pdf report.tex
  • report.pdf is the target, the file to be created, or built.
  • A colon, :, separates targets from dependencies.
  • latexmk -pdf report.tex is the action that will be executed once the dependencies exist.
  • report.tex and local.bib are the first two dependencies, files that are needed by the action to update the target. Targets can have zero or more dependencies. Those two dependencies already existed and were listed by git ls-files.
  • abyss.pdf and isles.pdf are the third and fourth dependencies. Since they did not exist, make invoked the second and third rules in the Makefile to build them using analyze.py.
  • results.tex, the fifth dependency did not exist. So make invoked the fourth rule. The action for the fourth rule (on line 19 of the Makefile) has a return value of 1 which indicates an error. So when it is invoked make terminates and reports the error.
  • Together, the target, dependencies, and actions form a rule.

The dependencies results.tex and abyss.head don’t exist, and the rules in the Makefile to build them return errors.

Callout

Dependencies

The order of rebuilding dependencies is arbitrary. You should not assume that they will be built in the order in which they are listed.

Dependencies must form a directed acyclic graph. A target cannot depend on a dependency which itself, or one of its dependencies, depends on that target.

Let’s create results.tex from the results of running analyze.py with a text editor and create abyss.head with

BASH

$ head -n 10 abyss.dat |awk '{print $1, $2}' > abyss.head

Checking the files we find:

BASH

$ cat abyss.head

OUTPUT

the 4044
and 2807
of 1907
a 1594
to 1515
in 1221
i 974
was 695
it 680
for 675

BASH

$ cat result.tex

OUTPUT

Book & First & Second & Ratio\\ \hline
abyss & 4044 & 2807 & 1.44 \\
isles & 3822 & 2460 & 1.55

Finally we can build and view the pdf document with

BASH

$ make
$ evince report.pdf

In the next episode, we will restructure the Makefile and the python code to automate building result.tex. But before that, we will make some easy changes.

First, put the command we used to build abyss.head into the Makefile. The modified rule should be

abyss.head: abyss.dat
	head -n 10 abyss.dat |awk '{print $$1, $$2}' > abyss.head

Make maps the double $$ to a single $ before executing the action. We can verify that with

BASH

$ rm abyss.head
$ make abyss.head

OUTPUT

head -n 10 abyss.dat |awk '{print $1, $2}' > abyss.head

Next add the following block to Makefile

clean:
	rm -f abyss.head *.pdf *.dat *.two *.aux *.bbl *.blg *.fdb_latexmk \
*.fls *.log
	touch clean

Now try the following sequence

BASH

$ make clean
$ make
$ touch results.tex
$ make

Try again

BASH

$ make clean

OUTPUT

make: 'clean' is up to date.

Change the block to

.PHONY: clean
clean:
	rm -f abyss.head *.pdf *.dat *.two *.aux *.bbl *.blg *.fdb_latexmk *.fls *.log
	touch clean

and test with

BASH

$ make clean
$ make clean
$ make clean

Finally change the block to

.PHONY: clean
clean:
	rm -f abyss.head *.pdf *.dat *.two *.aux *.bbl *.blg *.fdb_latexmk *.fls *.log
Callout

“Up to Date” Versus “Nothing to be Done”

If we ask Make to build a file that already exists and is up to date, then Make informs us that:

BASH

$ make isles.dat

OUTPUT

make: `isles.dat' is up to date.

If we ask Make to build a file that exists but for which there is no rule in our Makefile, then we get message like:

BASH

$ make analyze.py

OUTPUT

make: Nothing to be done for `analyze.py'.

up to date means that the Makefile has a rule with one or more actions whose target is the name of a file (or directory) and the file is up to date.

Nothing to be done means that the file exists but either :

  • the Makefile has no rule for it, or
  • the Makefile has a rule for it, but that rule has no actions

Finally, if we ask Make to build a file that doesn’t exist and for which there is no rule in our Makefile we get a self explanatory message like this

BASH

$ make foo

OUTPUT

make: *** No rule to make target 'foo'.  Stop.
Callout

Makefiles as Documentation

By explicitly recording the inputs to and outputs from steps in our analysis and the dependencies between files, Makefiles act as a type of documentation, reducing the number of things we have to remember.

Callout

Makefiles Do Not Have to be Called Makefile

We don’t have to call our Makefile Makefile. However, if we call it something else we need to tell Make where to find it. This we can do using -f flag. For example, if our Makefile is named MyOtherMakefile:

BASH

$ make -f MyOtherMakefile

Sometimes, the suffix .mk will be used to identify Makefiles that are not called Makefile e.g. install.mk, Rules.mk etc.

When it is asked to build a target, Make checks the ‘last modification time’ of both the target and its dependencies. If any dependency has been updated since the target, then the actions are re-run to update the target. Using this approach, Make knows to only rebuild the files that, either directly or indirectly, depend on the file that changed. This is called an incremental build.

Key Points
  • Use # for comments in Makefiles.
  • Write rules as target: dependencies.
  • Specify update actions in a tab-indented block under the rule.
  • Use .PHONY to mark targets that don’t correspond to files.

Content from No Hands


Last updated on 2026-10-07 | Edit this page

Estimated time: 15 minutes

Overview

Questions

  • How can we make a research project reproducible and comprehensible

Objectives

  • For reproducibility, entirely automate building the final document
  • Organize directory structure of the project
  • Break up monolithic code into understandable blocks
  • Separate plotting from calculations

Use the following commands if you want to save your work from the previous episode

BASH

$ git add Makefile results.tex
$ git commit -m "my work"

Now use the following to get files for this episode

BASH

$ git switch -c my-03-branch remotes/origin/03-no-hands
$ git ls-files

which should yield

Makefile
abyss.txt
analyze.py
isles.txt
local.bib
report.tex
src/analysis/testzipf.py

Next type

BASH

$ make

That should produce several screens of output and build the document as the file build/report.pdf. The last bit of output should be something like

OUTPUT

Output written on build/report.pdf (3 pages, 135727 bytes).
Transcript written on build/report.log.
Latexmk: Getting log file 'build/report.log'
Latexmk: Examining 'build/report.fls'
Latexmk: Examining 'build/report.log'
Latexmk: Found input bbl file 'build/report.bbl'
Latexmk: Log file says output to 'build/report.pdf'
Latexmk: Found bibliography file(s):
  ./local.bib
Latexmk: All targets (build/report.pdf) are up-to-date

Now look at the Makefile

BASH

$ cat Makefile

OUTPUT

SHELL=bash  # This forces the actions below to run in bash shells

export PYTHONPATH := src
export TEXINPUTS := src/TeX//:
export BIBINPUTS := src/TeX//:
export BSTINPUTS := src/TeX//:

build/report.pdf : local.bib report.tex build/abyss.pdf build/isles.pdf build/results.tex build/abyss.head
	mkdir -p $(@D)
	latexmk --outdir=build -pdf report.tex

# A new feature of gnu-make as of version 4.3 supports making several
# targets with one rule.  See
# https://www.gnu.org/software/make/manual/make.html#Multiple-Targets

build/abyss.dat build/abyss.pdf &: analyze.py abyss.txt
	mkdir -p $(@D)
	python analyze.py abyss.txt build/abyss.dat build/abyss.pdf

build/isles.dat build/isles.pdf &: analyze.py isles.txt
	mkdir -p $(@D)
	python analyze.py isles.txt build/isles.dat build/isles.pdf

build/results.tex: src/analysis/testzipf.py build/abyss.dat build/isles.dat
	mkdir -p $(@D)
	python $^ --latex > $@

build/abyss.head : build/abyss.dat
	mkdir -p $(@D)
	head -n 10 $< |awk '{print $$1, $$2}' > $@

.PHONY : clean
clean :
	rm -rf build *.pdf *.dat
Challenge

Explain the Makefile

Write notes explaining each of the following:

  1. $(@D) in the action mkdir -p $(@D)
  2. &: in the line build/abyss.dat build/abyss.pdf &: analyze.py abyss.txt
  3. $^ in the action python $^ --latex > $@
  4. $@ in the action python $^ --latex > $@
  1. In an action, $(@D) is the directory of the target
  2. &: is a new feature of gnu make that supports multiple targets in a rule
  3. $^ is an automatic variable that represents the names of all the dependencies in a rule
  4. $@ is an automatic variable that represents the names of the target of a rule

Tasks


The key to making the build process entirely automatic, was extracting a bit of code from analyze.py and putting it in the new file src/analysis/testzipf.py.

In the following tasks we will extract other chunks of code from analyze.py and put them in locations with names that help identify what the chunks do.

Extract the word counting functionality from analyze.py

We will take the following code from analyze.py, modify it appropraitely, and put it in src/analysis/countwords.py

PYTHON

	with open(input_path, encoding='utf-8', mode='r') as input_fd:
        lines = input_fd.read().splitlines()

    counts = {}
    for line in lines:
        for purge in DELIMITERS:
            line = line.replace(purge, " ")
        words = line.split()
        for word in words:
            word = word.lower().strip()
            if word in counts:
                counts[word] += 1
            else:
                counts[word] = 1
    sorted_counts = sorted(list(counts.items()),
           key=lambda key_value: key_value[1],
           reverse=True)
    stripped = []
    for (word, count) in sorted_counts:
        if len(word) >= min_length:
            stripped.append((word, count))

    total = 0
    for count in stripped:
        total += count[1]
    percentage_counts = [(word, count, (float(count) / total) * 100.0)
              for (word, count) in stripped]

    top_two = [count for (_, count, _) in percentage_counts[0:2]]

    with open(dat_path, encoding='utf-8', mode='w') as output:
        for _tuple in percentage_counts:
            output.write(f"{' '.join(str(item) for item in _tuple)}\n")

    print(f'{top_two=}')

We will use the boilerplate in src/analysis/countwords.py

BASH

$ cat src/analysis/countwords.py

OUTPUT

"""countwords.py:  Count how often each word occurs in a file.

"""
import sys
import argparse

DELIMITERS = ". , ; : ? $ @ ^ < > # % ` ! * - = ( ) [ ] { } / \" '".split()


def word_count(input_path, output_path, min_length=1):
    """Calculate word frequencies in a file.

    Args:
        input_path: Path to book, eg, 'books/abyss.txt'
        output_path: Path to result, eg, 'build/abyss.dat'

    EG, Analyze 'books/abyss.txt' to produce 'build/abyss.dat' with
    first three lines:

    the 4044 6.354494028912634
    and 2807 4.410747957259585
    of 1907 2.9965430546825895

    """
    ###############Put relevant code here#########################

def main(argv=None):
    '''Parses command line and calls function to count words in a file

    '''

    if argv is None:  # Usual case
        argv = sys.argv[1:]

    # Parse command line
    parser = argparse.ArgumentParser(description='Count words in a file')
    parser.add_argument('input_path', help='Path to input, eg, books/abyss.txt')
    parser.add_argument('output_path',
                        nargs='?',
                        help='Path to result, eg, build/abyss.dat')
    parser.add_argument('--min_length',
                        type=int,
                        default=1,
                        help='Drop counts less than "min_length"')
    args = parser.parse_args(argv)

    word_count(args.input_path, args.output_path, args.min_length)

    return 0


if __name__ == "__main__":
    sys.exit(main())

After modifying src/analysis/countwords.py, we put the following in the Makefile.

build/abyss.dat &: src/analysis/countwords.py abyss.txt
	mkdir -p $(@D)
	python $^ $@	

Now we test the code with

BASH

$ make build/abyss.dat

OUTPUT

mkdir -p build
python src/analysis/countwords.py abyss.txt build/abyss.dat
top_two=[4044, 2807]

Check the result

BASH

$ head build/abyss.dat

Extract the plot functionality from analyze.py

We will take the following code from analyze.py, modify it appropraitely, and put it in src/plotscripts/plotcounts.py

PYTHON

    limited_counts = percentage_counts[0:limit]
    word_data = [word for (word, _, _) in limited_counts]
    count_data = [count for (_, count, _) in limited_counts]
    position = np.arange(len(word_data))
    width = 1.0
    fig = plt.figure()
    ax = fig.add_subplot(1, 1, 1)
    ax.set_xticks(position)
    ax.set_xticklabels(word_data)
    plt.bar(position, count_data, width, color='b')
    plt.title("Word Counts")
    ax.set_ylabel("Counts")
    ax.set_xlabel("Word")

Here is the boilerplate in src/plotscripts/plotcounts.py

PYTHON

""" plotcounts.py code to plot word counts for the Zipf project.
"""

import sys
import argparse
import numpy as np
import matplotlib.pyplot as plt

def load_word_counts(filename: str) -> list:
    """Read (word, count, percentage) tuples from a text file.

    Args:
        filename: Path of file to read

    Returns:
        A list of tuples (word:str, count:int, percentage:float)

    Lines starting with # are ignored.

    """
    counts = []
    with open(filename, encoding='utf-8', mode="r") as input_fd:
        for line in input_fd:
            if not line.startswith("#"):
                fields = line.split()
                counts.append((fields[0], int(fields[1]), float(fields[2])))
    return counts

def plot_word_counts(counts, limit=10):
    """
    Given a list of (word, count, percentage) tuples, plot the counts as a
    histogram. Only the first limit tuples are plotted.
    """

    ###################Need chunk of code from analyze.py here##################


def main(argv=None):
    '''Parses command line and calls functions to make specified plot

    '''

    if argv is None:  # Usual case
        argv = sys.argv[1:]

    # Parse command line
    parser = argparse.ArgumentParser(
        description='Make plots for documents or to view')
    parser.add_argument('--limit',
                        type=int,
                        default=10,
                        help='Limit plot the "limit" most frequent words')
    parser.add_argument('--show',
                        action='store_true',
                        help='Display result on screen')
    parser.add_argument('input_path', help='Path to input, eg, books/abyss.txt')
    parser.add_argument('output_path',
                        nargs='?',
                        help='Path to result, eg, build/abyss.pdf')
    args = parser.parse_args(argv)

    counts = load_word_counts(args.input_path)

    # Create matplotlib fig object
    fig = plot_word_counts(counts, args.limit)

    # Display if specfied
    if args.show or args.output_path == 'show':
        plt.show()
        return 0

    fig.savefig(args.output_path)

    return 0


if __name__ == "__main__":
    sys.exit(main())

After modifying src/plotscripts/plotcounts.py, we put the following in the Makefile.

build/abyss.pdf &: src/plotscripts/plotcounts.py abyss.dat
	mkdir -p $(@D)
	python $^ $@	

Now we test the code with

BASH

$ make build/abyss.pdf

OUTPUT

mkdir -p build
python src/plotscripts/plotcounts.py build/abyss.dat build/abyss.pdf
Callout

D.R.Y. (Don’t Repeat Yourself)

Wilson et al. say:

Anything that is repeated in two or more places is more difficult to maintain. Every time a change or correction is made, multiple locations must be updated, which increases the chance of errors and inconsistencies. […] applies to both data and code. […] this maxim holds that every piece of data must have a single authoritative representation in the system.

Now, we have the following block in the Makefile to build a dat file and a pdf for abyss

build/abyss.dat : src/analysis/countwords.py abyss.txt
	mkdir -p $(@D)
	python $^ $@

build/abyss.pdf : src/plotscripts/plotcounts.py build/abyss.dat
	mkdir -p $(@D)
	python $^ $@

we could copy that block and replace abyss with isles, but that would violate the DRY principle. Instead, let’s replace that block with

build/%.dat : src/analysis/countwords.py %.txt
	mkdir -p $(@D)
	python $^ $@

build/%.pdf : src/plotscripts/plotcounts.py build/%.dat
	mkdir -p $(@D)
	python $^ $@

Now test with

BASH

make
Challenge

Explain the use of % in the Makefile

Write a short note

Google AI is correct in saying

In a Makefile, the % character acts as a wildcard operator used to create pattern rules and perform text substitution. It allows you to write a single generic rule that applies to multiple files sharing the same naming convention, significantly reducing redundancy.

Move report.tex and local.bib to src/TeX

First move the files with

BASH

$ mkdir src/TeX
$ mv report.tex local.bib src/TeX

Then modify the rule for build/report.pdf in the Makefile and test with

BASH

$ rm build/report.*
$ make

Move the abyss.txt and isles.txt to the directory books/

First move the files with

BASH

$ mkdir books
$ mv *.txt books

Then modify the rule for build/.dat* in the Makefile and test with

BASH

$ rm build/report.*
$ make

Clean up and Modify the target clean in the Makefile

List the files in the project root directory with

BASH

$ ls

OUTPUT

Makefile  analyze.py  books  build  src

Observe that analyze.py is dead code, and after removing it all of the derived files are in the build directory. So the clean target can become

.PHONY : clean
clean :
	rm -rf build

After editing the Makefile, test with

BASH

$ make clean
$ ls -R
$ make

The following figure shows a graph of the dependencies embodied within our Makefile, involved in building the document:

Dependencies represented within the Makefile
Key Points
  • While the automatic variables in a Makefile look mysterious, they support the DRY principle. Here we have used:
    • $@ which refers to the target of the current rule.
    • $^ which refers to the dependencies of the current rule.
    • $< which refers to the first dependency of the current rule.

Content from Testing


Last updated on 2026-10-07 | Edit this page

Estimated time: 15 minutes

Overview

Questions

  • Why have unit tests?

Objectives

  • Self documenting Makefile
  • Use pytest for unit tests
  • Use coverage to check test coverage

Use the following commands if you want to save your work from the previous episode

BASH

$ make clean
$ git status
$ git add .
$ git status
$ git commit -m "my work"

Now use the following to get files for this episode

BASH

$ git switch -c my-04-branch remotes/origin/04-testing
$ git ls-files

which should yield

OUTPUT

Makefile
books/abyss.txt
books/isles.txt
src/TeX/local.bib
src/TeX/report.tex
src/analysis/countwords.py
src/analysis/testzipf.py
src/plotscripts/plotcounts.py

This episode is primarily about testing, but first we will make the Makefile self documenting. Try

BASH

$ make help

OUTPUT

 help                 : Holy mackerel!  What is this?

BASH

$ tail Makefile

OUTPUT

	head -n 10 $< |awk '{print $$1, $$2}' > $@

.PHONY : clean
clean :
	rm -rf build

## help                 : Holy mackerel!  What is this?
.PHONY : help
help : Makefile
	@sed -n 's/^## / /p' $<
Challenge

Self Documenting Makefiles

Explain how the help block at the end of the Makefile works and how you could use it to document the targets build/report.pdf, build/results.tex, build/abyss.head, clean, and help.

The sed command simply prints each line in the Makefile that starts with “##”. Now we will edit the file to document each of those targets.

So far we have tested each incremental change of our code with something like

BASH

$ make clean
$ make

While that takes less than three seconds to redo all the calculations and build the final document, our toy research project is a surrogate for projects with calculations and other steps that take hours if not days. For such a project one can use unit tests to test each chunk of code quickly in isolation.

Callout

Unit Tests

Google AI says:

Unit tests are automated pieces of code written to verify that a small, isolated part of a software application, usually a single function or method, behaves as expected.

Pytest


We will use the utility pytest to run unit tests, and we will put the testing code in the directory tests/. Here is the block required in the Makefile:

## test                : Run pytest on tests/
.PHONY : test
test :
	pytest tests/

And here is the code for tests/test_countwords.py

PYTHON

"""test_countwords.py

debug with
$py.test --pdb tests/test_countwords.py
"""
import pytest
import pathlib

import analysis.countwords


@pytest.fixture(scope="function")
def shared_count_file(tmp_path_factory):
    # 1. Create a temporary directory using the factory
    temp_dir = tmp_path_factory.mktemp("shared_data")

    # 2. Create the file inside that directory
    count_file = temp_dir / "word_counts"
    count_file.write_text('{"status": "ready"}')

    return count_file


def test_word_count(shared_count_file):
    """Check reading books/abyss.txt

    """
    in_path = 'books/abyss.txt'
    analysis.countwords.word_count(in_path, shared_count_file)


def test_main(shared_count_file: pathlib.PosixPath):
    in_path = 'books/abyss.txt'
    analysis.countwords.main([in_path, f'{shared_count_file}'])

After making those changes to the project we can test testing with

BASH

$ make test
pytest tests/
================================ test session starts ================================
platform linux -- Python 3.12.12, pytest-8.3.5, pluggy-1.5.0
rootdir: /mnt/precious/home/andy_nix/projects/SIAMDS27/learner_setup
plugins: anyio-4.9.0, cov-6.1.0, dependency-0.6.0
collected 2 items

tests/test_countwords.py ..                                                   [100%]

================================= 2 passed in 0.10s =================================

Next we introduce the utility coverage that reports on the sections of code that are covered by a collection of unit tests.

Callout

coverage.py

Google AI says:

Coverage.py is the standard, authoritative code coverage measurement tool for Python. […] it monitors your program during execution to track which lines of code have been run and which have not. It is primarily used to gauge the effectiveness and completeness of your test suites.

Here is a block for the Makefile

## coverage            : make test coverage report in build/htmlcov/
.PHONY : coverage
coverage :
	rm -rf .coverage build/htmlcov
	coverage run --source tests,src/analysis,src/plotscripts -m pytest tests
	coverage html -d build/htmlcov
	@echo To view: "firefox build/htmlcov/index.html & "

Let’s test that with

BASH

$ make coverage

OUTPUT

rm -rf .coverage build/htmlcovcoverage run --source tests,src/analysis,src/plotscripts -m pytest tests
================================ test session starts ================================
platform linux -- Python 3.12.12, pytest-8.3.5, pluggy-1.5.0
rootdir: /mnt/precious/home/andy_nix/projects/SIAMDS27/learner_setup
plugins: anyio-4.9.0, cov-6.1.0, dependency-0.6.0
collected 2 items

tests/test_countwords.py ..                                                   [100%]

================================= 2 passed in 0.16s =================================
coverage html -d build/htmlcov
Wrote HTML report to build/htmlcov/index.html
To view: firefox build/htmlcov/index.html &

After looking at the result in a browser, we see that we should write tests for testzipf.py and plotcounts.py. We have written those and included them in the git branch 05-standards for the next episode.

Key Points
  • Use unit tests for finding bugs quickly, and checking changes.
  • For Python code pytest works
  • Use a coverage tool for finding code that the testing suit missed
  • For Python code coverage works

Content from Coding Standards


Last updated on 2026-10-07 | Edit this page

Estimated time: 15 minutes

Overview

Questions

  • What tools can help achieve the objectives?

Objectives

  • Enforce uniform style in the code
  • Partially compensate for duck-typing in Python code

Use the following commands if you want to save your work from the previous episode

BASH

$ git add Makefile tests/test_countwords.py
$ git commit -m "my work"

Now use the following to get files for this episode

BASH

$ git switch -c my-05-branch remotes/origin/05-standards
$ git ls-files

which should yield

OUTPUT

Makefile
analyze.py
books/abyss.txt
books/isles.txt
src/TeX/local.bib
src/TeX/report.tex
src/analysis/countwords.py
src/analysis/testzipf.py
src/plotscripts/plotcounts.py
tests/test_countwords.py
tests/test_plotcounts.py
tests/test_testzipf.py

Put this block in the Makefile:

## pylintrc            : Fetch standard from google
pylintrc:
	wget https://google.github.io/styleguide/pylintrc

## lint                : Run pylint
.PHONY : lint
lint : pylintrc
	pylint --rcfile pylintrc src

And run pylint with

BASH

$ make lint

Well, let’s focus on one file

BASH

$ pylint --rcfile pylintrc src/analysis/countwords.py

OUTPUT

************* Module analysis.countwords
src/analysis/countwords.py:32:0: W1405: Quote delimiter " is inconsistent with the rest of the file (inconsistent-quotes)
src/analysis/countwords.py:89:0: W1405: Quote delimiter " is inconsistent with the rest of the file (inconsistent-quotes)
src/analysis/countwords.py:59:12: C0103: Variable name "_tuple" doesn't conform to '^[a-z][a-z0-9_]*$' pattern (invalid-name)

After fixing those let’s try again

BASH

$ make lint

Let’s fix those and try again

BASH

$ make lint

OUTPUT

pylint --rcfile pylintrc src

Check consistency of code style


Callout

YAPF Yet Another Python Formatter

Google AI says

YAPF is an open-source Python code formatter developed by Google. Unlike “opinionated” formatters like Black that enforce a single rigid standard, YAPF uses an algorithm inspired by clang-format to search for the “best” layout matching your specific configuration rules. It works by calculating the “lowest-cost” formatting decision to generate code that looks like an experienced human wrote it.

Let’s use the following block in the Makefile to run yapf

## yapf                : Force google format on all python code
.PHONY : yapf
yapf :
	yapf -i --recursive --style "google" src

In order to see what yapf does let’s first stage the files it might change

BASH

$ git add src
$ git status

Since that looks good, we run yapf

BASH

$ make yapf

OUTPUT

yapf -i --recursive --style "google" src

and check on the changes.

Check Python type hints


By default Python code does not check the types of data. At run time, the interpreter executes procedures that are appropriate for the type of the data, eg,

PYTHON

>>> a = '1'
>>> b = '2'
>>> a+b
'12'
>>> a=1
>>> b=2
>>> a+b
3
>>> 

This is called duck typing.

Before running Python code, one can use tools such as mypy to look for consistency in types. To support that, Python supports type hints. EG,

BASH

$ cat > foo.py
a : str = '1'
b : str = '2'
print(a*b)
$ python foo.py 
Traceback (most recent call last):
  File "/mnt/precious/home/andy_nix/projects/SIAMDS27/learner_setup/foo.py", line 3, in <module>
    print(a*b)
          ~^~
TypeError: can't multiply sequence by non-int of type 'str'
ds27$ mypy --strict foo.py
foo.py:3: error: Unsupported operand types for * ("str" and "str")  [operator]
Found 1 error in 1 file (checked 1 source file)
$ rm foo.py

Notice that mypy found the problem before running the code.

Callout

MYPY

Google AI says

Mypy is an optional static type checker for Python that analyzes your source code to detect bugs and type mismatches before your program ever runs. Because Python is traditionally a dynamically typed language, typos or invalid variable types normally crash your application at runtime. Mypy acts like a powerful linter, scanning your type annotations (type hints) to ensure type consistency across your codebase.

Here is the Makefile block that we will use to invoke mypy

## check-types         : Checks type hints
.PHONY : check-types
check-types:
	export MYPYPATH=$$PYTHONPATH; mypy --strict src/analysis

And here is what mypy has to say about our code

BASH

$ make check-types

OUTPUT

export MYPYPATH=$PYTHONPATH; mypy --strict src/analysis
src/analysis/testzipf.py:9: error: Missing type parameters for generic type "list"  [type-arg]
src/analysis/testzipf.py:30: error: Missing type parameters for generic type "list"  [type-arg]
src/analysis/testzipf.py:45: error: Function is missing a type annotation  [no-untyped-def]
src/analysis/testzipf.py:84: error: Call to untyped function "main" in typed context  [no-untyped-call]
src/analysis/countwords.py:10: error: Function is missing a type annotation  [no-untyped-def]
src/analysis/countwords.py:29: error: Need type annotation for "counts" (hint: "counts: dict[<type>, <type>] = ...")  [var-annotated]
src/analysis/countwords.py:65: error: Function is missing a type annotation  [no-untyped-def]
src/analysis/countwords.py:85: error: Call to untyped function "word_count" in typed context  [no-untyped-call]
src/analysis/countwords.py:91: error: Call to untyped function "main" in typed context  [no-untyped-call]
Found 9 errors in 2 files (checked 2 source files)
make: *** [Makefile:49: check-types] Error 1

Let’s address those issues now.

Key Points
  • We use yapf to coerce Python code into standard use of white space
  • We use lint to detect deviations of our Python code from standards specified in pylintrc
  • We use mypy to check our use of Python type hints

Content from Documentation


Last updated on 2026-10-07 | Edit this page

Estimated time: 15 minutes

Overview

Questions

  • How can I use text in Python files as a source of documentation?

Objectives

  • Setup Sphinx
  • Build documentation using text in the Python files

Use the following commands if you want to save your work from the previous episode

BASH

$ git add Makefile pylintrc src/analysis/copuntwords.py src/analysis/testzipf.py src/plotscripts/plotcounts.py
$ git commit -m "my work"

Now use the following to get files for this episode from the branch 06-documentation

BASH

$ git switch -c my-06-branch remotes/origin/06-documentation
$ git ls-files

which should yield

OUTPUT

.gitignore
Makefile
books/abyss.txt
books/isles.txt
docs/source/_templates/custom-module-template.rst
docs/source/autodoc.rst
docs/source/conf.py
docs/source/index.rst
pylintrc
src/TeX/local.bib
src/TeX/report.tex
src/analysis/countwords.py
src/analysis/testzipf.py
src/plotscripts/plotcounts.py
tests/test_countwords.py
tests/test_plotcounts.py
tests/test_testzipf.py

In this episode, we will walk through setting up Sphinx for documentation from scratch, but first let’s again peek at the reuslt with the following:

BASH

$ make doc

The tail of the result is

OUTPUT

writing output... [100%] index
generating indices... genindex py-modindex done
highlighting module code... [100%] plotscripts.plotcounts
writing additional pages... search done
dumping search index in English (code: en)... done
dumping object inventory... done
build succeeded, 3 warnings.

The HTML pages are in build/docs/html.

and we can view the result with

BASH

firefox build/docs/html/index.html &

Navigate to plotscripts:plotcounts and note that the page has information from running python src/plotscripts/plotcounts -h, the definition line of each function, and the docstring of each function.

Start from scratch


Now we will remove the built files, move the doc directory aside and start from scratch

BASH

$ make clean
$ mv docs safe_docs
$ sphinx-quickstart docs

OUTPUT

Welcome to the Sphinx 8.2.3 quickstart utility.

Please enter values for the following settings (just press Enter to
accept a default value, if one is given in brackets).

Selected root path: .

You have two options for placing the build directory for Sphinx output.
Either, you use a directory "_build" within the root path, or you separate
"source" and "build" directories within the root path.
> Separate source and build directories (y/n) [n]:

Here are how to answer the questions:

BASH

> Separate source and build directories (y/n) [n]: y
> Project name: Zipf
> Author name(s): Bat Masterson
> Project release []:
> Project language [en]: 

OUTPUT

Creating file /ROOT-PATH/source/conf.py.
Creating file /ROOT-PATH/source/index.rst.
Creating file /ROOT-PATH/Makefile.
Creating file /ROOT-PATH/make.bat.

Finished: An initial directory structure has been created.

You should now populate your master file /path/source/index.rst and create other documentation
source files. Use the Makefile to build the docs, like so:
   make builder
where "builder" is one of the supported builders, e.g. html, latex or linkcheck.

Let’s see what sphinx-quickstart installed and compare it to safe_docs

BASH

$ ls docs safe_docs

OUTPUT

docs:
Makefile  build  make.bat  source

safe_docs:
source

Our Makefile in the root directory has the following block:

## docs                : Run Sphinx to create documentation
.PHONY : docs
docs : build/docs/html/index.html

build/docs/html/index.html: docs/source/conf.py docs/source/index.rst
        sphinx-build -M html "docs/source" "build/docs"

That block replaces what we need in docs/Makefile and it directs the output from docs/build to build/docs. So we remove the extraneous files in docs/ and look at docs/source

BASH

$ rm -rf docs/Makefile docs/build docs/make.bat
$ ls -R docs/source

OUTPUT

docs/source/:
_static  _templates  conf.py  index.rst

docs/source/_static:

docs/source/_templates:

Now comparing to safe_docs/

BASH

$ ls -R safe_docs

we find that the relevant files are::

  • source/_templates/custom-module-template.rst
  • source/autodoc.rst
  • source/conf.py
  • index.rst

Now look at the differences

BASH

$ sdiff -s safe_docs/source/index.rst docs/source/index.rst 

OUTPUT

   sphinx-quickstart on Sat Oct  3 20:14:28 2026.	      |	   sphinx-quickstart on Mon Oct  5 21:37:05 2026.
   Trees of docstrings <autodoc>			      <

So, we need a line in index.rst that invokes autodoc

Next check conf.py

BASH

$ sdiff -s safe_docs/source/index.rst docs/source/index.rst

The significant part of the output is

OUTPUT


							      >	extensions = []
							      >
extensions = '''sphinx_jinja sphinx.ext.autodoc sphinx.ext.vi <
    sphinx.ext.mathjax sphinx.ext.intersphinx sphinx.ext.cove <
    sphinx.ext.doctest sphinx.ext.autosummary		      <
    sphinx.ext.autosectionlabel sphinx.ext.napoleon	      <
    sphinx_argparse_cli'''.split()			      <
autosummary_generate = True				      <
templates_path = ['_templates']				      <

Before we put that block in docs/source/index.rst, we will look at the other two relevant files to get a feel for what’s going on

BASH

$ cat safe_docs/source/autodoc.rst

OUTPUT

Autodoc
=======

We use the autodoc feature of sphinx to extract documentation from
docstrings in Python scripts.  The autosummary feature of sphinx
recursively descends directory trees extracting documentation.  The
table below provides access to the roots of the extracted
documentation trees and indirectly to individual docstrings.

.. autosummary::
   :toctree: _autosummary
   :template: custom-module-template.rst
   :recursive:

   analysis
   plotscripts

Elsewhere in the documentation we can link to nodes
in trees.  For example, here is a link to the docstring of the code
that :doc:`makes for countwords.py
<_autosummary/analysis.countwords>`.

We can see the result of this block by searching for “autodoc” in the browser. The interesting bit is the autosummary block. Entering “sphinx autosummary” into Google brings up a helpful AI Overview and the documentation. Let’s look at the documentation. The “toctree” and “recursive” options are sort of clear, but we’ll need to look at custom-module-template.rst to understand the “template” option.

BASH

$ less safe_docs/source/_templates/custom-module-template.rst

Yuck! If we ever understood all of that, it was a long time ago. However, we do want to point out a key feature that we find by searching for “cli”. That extension captures the output of eg, “python plotcounts.py -h” and formats it.

Fixing up docs/source

So here’s the plan::

  1. Simply copy custom-module-template.rst and autodoc.rst from safe_docs to docs
  2. Edit conf.py and index.rst
  3. Build the docs

BASH

$ make docs
$ firefox build/docs/html/index.html
Key Points
  • Sphinx can move information from Python souce files to documentation
  • That supports the DRY principal
  • It also supports having good information in Python source files