In notebooks/, you'll find the main.py file, which is the main script for running the analysis. It loads the data from HF, so no need to get additional data files. For collaboration, please create your own main_pascale.py(for example) and run your own experiments.
The canon_tagging.py file contains functions for aggregating the canon data (pushing to HF).