Project Folders And Reruns#
A PYSTILT project is one folder. It holds your settings, your receptors, and every output. With the folder alone you can reopen, extend, or rerun the work.
What’s in a project folder#
my_project/
config.yaml # your settings: meteorology, variants, run options
receptors.csv # your receptors: where and when to release particles
simulations/
variants.yaml # the settings each variant ran with (written by PYSTILT)
by-id/
<receptor id>/ # one folder per receptor
<variant>/ # one folder per simulation
stilt.log # run log, the first place to look when a run fails
<receptor id>_traj.parquet # particle paths
<receptor id>_foot.nc # the footprint, when the variant has a grid
<receptor id>_foot.empty # instead of .nc when no particle reaches the grid
met/, CONTROL, SETUP.CFG ... # HYSPLIT inputs, kept for debugging
A variant declared with from: has no particle file of its own. It uses
the particles of the variant it comes from, so its folder holds only the
footprint and a log.
config.yaml and receptors.csv are yours to edit. PYSTILT never
rewrites a config.yaml it loaded from the folder, and it only appends
new receptors to receptors.csv. If you pass settings to
Model in Python, they replace config.yaml when the model
runs. Receptors you pass are appended to receptors.csv, or start it if
there is none.
simulations/variants.yaml belongs to PYSTILT. It holds the full
settings of every variant that has run, which is how PYSTILT notices a
changed setting (see Configuration). Don’t edit it.
A Slurm run also creates chunks/ and slurm/ folders with the job
scripts and logs (see On An HPC Cluster (Slurm)).
Simulation IDs#
A simulation is one receptor run under one variant. Its ID is the receptor
ID and the variant name joined by a slash. This is also its folder under
simulations/by-id:
{YYYYMMDDHHMM}_{location}/{variant}
202307151800_-111.848_40.766_10/hrrr
For a point receptor, the location is the longitude, latitude, and
altitude. A column receptor has X in place of the altitude. A
multipoint receptor uses multi_ and a short hash of its points, which
does not change if you reorder them.
A project runs every receptor under every variant, so 100 receptors and
three variants make 300 simulations. With no variants in
config.yaml, there is one variant per meteorology source (see
Configuration).
Opening a project again#
config.yaml and receptors.csv are in the folder once the project
has run (or been registered). After that, the folder is all you need:
import stilt
model = stilt.Model(project="./my_project")
model.status() # one row per simulation, with a "complete" column
From the command line:
stilt status ./my_project
To add receptors to an existing project, pass them in and run:
model = stilt.Model(project="./my_project", receptors=new_receptors)
model.run()
New receptors are appended to receptors.csv in the file’s own columns.
Receptors already in the file are left as they are.
Reruns skip finished work#
Before running, PYSTILT checks which simulations are finished and runs only the rest. A simulation is finished when all of its outputs exist:
the particle file, unless the variant is declared with
from:;the footprint file or
.emptymarker, if the variant has a grid.
If the particle file is missing, HYSPLIT runs again. The footprint is then
remade from the new particles, and so are the footprints of any from:
variants that use them.
So after an interruption, a failed Slurm task, or adding a variant to
config.yaml, run the project again. Only what is missing will run. A
new variant runs for every receptor, and nothing else is touched. PYSTILT
refuses to run a variant whose settings changed after it ran. Remove its
outputs first (see Configuration).
To see what is not finished yet:
model.simulations.incomplete().keys() # (receptor, variant) ids
model.simulations.status() # a table of every simulation
To rerun one variant, delete its outputs with stilt rm --variant NAME
or model.remove(NAME). To rerun one simulation, call sim.delete()
on it. To run everything again, pass skip_existing=False to
model.run(), or --no-skip to stilt run.
Storing a project in the cloud#
A project can also live in an s3:// or gs:// bucket. This needs the
cloud extra. HYSPLIT still has to run on a local disk. By default
PYSTILT uses a temporary folder for this. Set compute_root to use your
own scratch space. Each simulation’s outputs are uploaded to the bucket
when it finishes.
model = stilt.Model(
project="gs://my-bucket/wbb_july_case",
compute_root="/scratch/me/pystilt",
)
Outputs read back from a bucket are cached on local disk. Set
PYSTILT_CACHE_DIR to choose where. Otherwise a temporary folder is used.
For more on how this works, see How PYSTILT Tracks Finished Work.