Coming From STILT-R#

PYSTILT does the same science as STILT-R, and its footprints match STILT-R’s cell by cell. The one exception is forward runs with hnf_plume on, where PYSTILT fixes how the plume grows (see STILT-R parity).

What changes is where the settings live. Instead of editing variables in run_stilt.r, you write them in config.yaml or pass them to stilt.Model. Receptors go in receptors.csv.

The workflow side by side#

Step

STILT-R

PYSTILT

Start a project

Rscript -e "uataq::stilt_init('my_project')"

stilt init my_project

Set options

edit variables in r/run_stilt.r

edit config.yaml

Define receptors

build a receptors data frame in run_stilt.r

receptors.csv, or receptor objects in Python

Run

Rscript r/run_stilt.r

stilt run my_project

Outputs

out/by-id/<id>/, _traj.rds and _foot.nc

simulations/by-id/<id>/<variant>/, _traj.parquet and _foot.nc

run_stilt.r settings in PYSTILT#

run_stilt.r

config.yaml

met_path

mets: <name>: directory

met_file_format

mets: <name>: file_format (same strftime codes)

met_file_tres

mets: <name>: file_tres

n_met_min

mets: <name>: n_min

met_subgrid_enable, met_subgrid_buffer, met_subgrid_levels

mets: <name>: subgrid_enable, subgrid_buffer, subgrid_levels. PYSTILT also needs subgrid_bounds.

xmn, xmx, ymn, ymx

grid: xmin, xmax, ymin, ymax

xres, yres, projection

grid: xres, yres, projection

smooth_factor, time_integrate

same names, top level

n_hours, numpar, hnf_plume, rm_dat, timeout, varsiwant

same names, top level

HYSPLIT settings (capemin, delt, kmix0, tlfrac, veght, …)

same names, top level

wind and mixing-depth error settings (siguverr, tluverr, sigzierr, …)

same names, usually on a variant of their own (Transport Error)

n_nodes and n_cores with Slurm

execution: backend: slurm, with n_workers array tasks of cpus_per_task CPUs each (On An HPC Cluster (Slurm))

slurm_options

other keys under execution, passed to sbatch

n_cores without Slurm

execution: n_workers

before_footprint

transforms (Particle Weighting (Transforms))

before_trajec

no equivalent

run_trajec, run_foot

not needed. stilt run runs whatever outputs are missing, and stilt run --no-skip reruns everything. To make a second footprint from the same particles, add a from: variant. To redo one variant, delete its outputs with stilt rm --variant NAME (Project Folders And Reruns).

simulation_id (run a subset)

a smaller receptors.csv, or stilt.execution.run_receptors() with the receptor IDs you want

ziscale as one list per receptor

ziscale per variant, as one factor or one factor per hour. It is the same for every receptor. Per-receptor factors are not supported yet.

stilt_wd, output_wd

the project folder

lib.loc

not needed

Other differences#

A project runs its receptors under named variants. A variant is a meteorology source plus any settings that differ from the defaults (see Configuration). One project can run the same receptors with several meteorology products, a range of ziscale values, or a second footprint grid. Each variant gets its own folder.

The transport error run is a variant too. STILT-R runs HYSPLIT a second time with the wind-error settings and stores both sets of particles in one _traj.rds. In PYSTILT that second run is its own variant, with its own folder and log, and it can repeat several times (Transport Error).

stilt run skips simulations whose outputs already exist. After a failure, run it again and only the missing simulations run. Adding a variant runs only the new variant.

PYSTILT will not change the settings of a variant that has already run. STILT-R deletes out/by-id and starts over whenever run_trajec = TRUE. PYSTILT records the settings each variant ran with and stops with an error if they change under the same name. Give the changed settings a new variant name, or run stilt rm --variant NAME to start that variant over (see Configuration).

Receptors#

receptors.csv holds the same information as STILT-R’s receptors data frame. STILT-R’s column names long, lati, and zagl work as they are. Rename run_time to time:

time,long,lati,zagl
2015-07-05 00:00:00,-111.8472,40.7665,21

STILT-R makes a column receptor from a row with several zagl values. In receptors.csv, a column or multipoint receptor is several rows that share an r_idx value. See Receptors.

Outputs#

Trajectories are Parquet files instead of R .rds files. You can open them in Python with pandas or stilt.Trajectories, and in R with arrow::read_parquet. Footprints are NetCDF as before, with dimensions (time, lat, lon).

A point or column receptor’s ID has the same form as STILT-R’s simulation ID. Each variant is a folder below it, for example 201507050000_-111.8472_40.7665_21/hrrr.

Moving a project over#

  1. Copy your run_stilt.r settings into config.yaml using the table above. Start with one meteorology source and no variants.

  2. Write your receptors to receptors.csv.

  3. Run a few receptors that you already have STILT-R results for, and compare the footprints.

  4. Run everything.