Summary and Setup

Welcome to our mini-tutorial!

We are going to show you how to use a selection of Software Tools for Reproducible Research. In particular we will show you some of the tools that we used for writing the second edition of <80><9c>Hidden Markov Models and Dynamical Systems<80><9d> which SIAM will publish this year. In the documentation for the source code we have instructions for a variety of operating systems and environments EG, from Nix, one can build the book with the following sequence of commands:

BASH

$ git clone https://gitlab.com/fraserphysics/hmmds
$ cd hmmds
$ nix-shell --pure --keep NIX_PATH nix_files/shell.nix
$ make book

Then, after about ten hours of computation, one may view the book with

BASH

$ evince build/TeX/book/main.pdf

Most of the source code for the book lies in 176 python modules. Since we don<80><99>t have the time to talk about all that code now, we will use the same set of tools to build a smaller project.

The Challenge of Reproducibility is Recognized

Here is a list of papers (some suggested by Claude AI) that have suggestions for better reproducibility:

  1. G. Wilson et al.Best Practices for Scientific Computing in PLOS Biology 2014

  2. T.G. Kolda Taming the Chaos of Computational Experiments in SIAM NEWS 2025

  3. G.K. Sandve et al.Ten Simple Rules for Reproducible Computational Research in PLOS Computational Biology 2013

  4. G. Wilson et al.Good Enough Practices in Scientific Computing in PLOS Computational Biology 2017

  5. R.D. Peng Reproducible Research in Computational Science. Science 2011

  6. L.A. Barba Terminologies for Reproducible Research. 2018

  7. D.L. Donoho et al.Reproducible Research in Computational Harmonic Analysis. Computing in Science & Engineering 2009

  8. R.J. LeVeque, I.M. Mitchell, & V. Stodden, V. Reproducible Research for Scientific Computing: Tools and Strategies for Changing the Culture Computing in Science & Engineering 2012

Best Practices


Here is the summary from (1. Wilson et al.) above with the relevant tools that we use below each item.

Write programs for people, not computers. - A program should not require its readers to hold more than a handful of facts in memory at once. - Make names consistent, distinctive, and meaningful. - Make code style and formatting consistent.

yapf, mypy, and pylint

Let the computer do the work. - Make the computer repeat tasks. - Save recent commands in a file for re-use. - Use a build tool to automate workflows.

make, and sphinx

Make incremental changes. - Work in small steps with frequent feedback and course correction. - Use a version control system. - Put everything that has been created manually in version control.

git

Don<80><99>t repeat yourself (or others). - Every piece of data must have a single authoritative representation in the system. - Modularize code rather than copying and pasting. - Re-use code instead of rewriting it.

Plan for mistakes. - Add assertions to programs to check their operation. - Use an off-the-shelf unit testing library. - Turn bugs into test cases. - Use a symbolic debugger.

pytest, and pdb

Optimize software only after it works correctly. - Use a profiler to identify bottlenecks. - Write code in the highest-level language possible.

cProfile, Python, Cython, cython.parallel, OpenMP

Document design and purpose, not mechanics. - Document interfaces and reasons, not implementations. - Refactor code in preference to explaining how it works. - Embed the documentation for a piece of software in that software.

Sphinx, and sed for make

Collaborate. - Use pre-merge code reviews. - Use pair programming when bringing someone new up to speed and when tackling particularly tricky problems. - Use an issue tracking tool.

GitLab issues

Our Tools


Here are the tools we will use or discuss today

make
We use GNU make to specify the whole sequence of operations from raw data to a pdf document. Makefiles record all parameter settings such as seeds for random number generators. We use make to separate source code from derived data.
Python
In 2001, we started using python because the code is easy to read. Since then it has become popular, and in some fields a standard. To separate stages of processing we write temporary pickle files. In particular we segregate all plotting into scripts that use matplotlib.
Git
We use git in our work generally. Today, we will ask you to use git branches to keep exercises using various tools separate.
yapf
(Yet Another Python Formatter) is an open-source code formatter that we use to automatically coerce code to follow Google<80><99>s Python style.
mypy
Mypy is an optional static type checker for Python.
pylint
Analyzes and scores python code against a user defined style guide, pylintrc.
sphinx
Sphinx is a documentation generator or a tool that translates a set of plain text source files into various output formats. We use the autodoc extension that can import the modules we are documenting, and pull in documentation from docstrings in a semi-automatic way.
pytest
The pytest framework makes it easy to write small, readable tests, and can scale to support complex functional testing for applications and libraries.
pdb
pdb is Python<80><99>s built-in interactive source code debugger.
cProfile
cProfile is Python<80><99>s built-in, recommended tool for deterministic profiling. We use it to find where long running programs are spending the most time.
Cython
Cython is an optimizing static compiler for both the Python programming language and the extended Cython programming language (based on Pyrex). It makes writing C extensions for Python easy enough to be plausible.
cython.parallel
Cython supports native parallelism through the cython.parallel module. It currently supports OpenMP. We use it ease critical bottlenecks.
sed
Sed is a command line stream editor that we use to extract documentation from makefiles in response to <80><9c>make help<80><9d>.
LaTeX
By using make to govern typesetting, we can ensure that reported computational results are products of the most recent version of our source code
Nix
Nix is a package manager and build system that parses reproducible build instructions specified in the Nix Expression Language. Nix stores results of builds in unique addresses specified by a hash of the complete dependency trees. By working in environments specified by such builds, we control the dependency tree and ensure reproducibility.
Prerequisite

Prerequisites

We will assume that you are familiar with the following tools (in order of familiarity)

bash
It<80><99>s probably OK if you use a different shell, but you should be comfortable with command line tools.
A text editor of your choice
Software Carpentry has instructors use Nano to avoid frightening the learners. We will use Emacs or Vim so that we can cover more material.
Python
Most of the code we will discuss is written in Python. If you not comfortable with Python, you should not sign up for this mini-tutorial.
git
We use git to keep code up to date for each learner as we progress from one section to the next throughout the mini-tutorial. While we will provide you with the necessary git commands, it will make more sense to you if you have used the tool before.
LaTeX
We use LaTeX to format results. As with git, it will make more sense to you if you have used the tool before.
Prerequisite

Setup

In order to follow this lesson, you will need to download some files. Please follow instructions on the setup page.

Files


You need to clone a git repository to get files to follow this lesson. We<80><99>ve made a seperate branch in the repository for each episode so that as we transition from one episode to the next you can save your work on the old branch and check-out the next branch to get files that are identical to the ones the instructor is using. Here are the steps:

  1. Open a Bash shell window.

  2. (Optional) Navigate to a convenient directory.

  3. Use the following git command to fetch the necessary files:

    $ git clone https://github.com/fraserphysics/setup_make_novice.git make-lesson

    {: .language-bash}

  4. Change into the make-lesson directory:

    $ cd make-lesson

    {: .language-bash}

  5. Verify that you have a branch for each episode:

    $ git branch -a

    {: .language-bash}

    * 01-intro
    remotes/origin/01-intro
    remotes/origin/02-makefiles
    remotes/origin/03-variables
    remotes/origin/04-dependencies
    remotes/origin/05-patterns
    remotes/origin/06-variables
    remotes/origin/07-functions
    remotes/origin/08-self-doc
    remotes/origin/09-latex
    remotes/origin/HEAD -> origin/01-intro

    {: .output}

Software


You also need to have the following software installed on your computer to follow this lesson:

GNU Make

Linux

Make is a standard tool on most Linux systems and should already be available. Check if you already have Make installed by typing make -v into a terminal.

One exception is Debian, and you should install Make from the terminal using sudo apt-get install make.

OSX

You will need to have Xcode installed (download from the Apple website). Check if you already have Make installed by typing make -v into a terminal.

Windows

Use the Software Carpentry Windows installer.

Python

Python3, Numpy and Matplotlib are required. They can be installed separately, but the easiest approach is to install Anaconda which includes all of the necessary python software.

LaTeX (Optional)

The last episode uses the LaTeX document preparation system. If you find it difficult to install, simply listen to the discussion for that episode without doing any of the exercises.