Reading Retrieval Products#
A column retrieval comes as a product file, such as a TROPOMI orbit, an
OCO-2 Lite file, a TCCON site file, or a day of EM27/SUN spectra from GGG.
The readers in stilt.observations turn each of these into a table
with one row per sounding. The columns are the same for every instrument,
so the rest of the workflow (Satellite And Column Observations) is too.
from stilt.observations import read_tropomi_ch4
df = read_tropomi_ch4(
"S5P_OFFL_L2__CH4____20231019T192749_20231019T210918_31176_03_020500_20231021T113213.nc",
lon_range=(-113.5, -110.5), lat_range=(39.5, 42.0),
)
df = df[df.good]
The readers#
Reader |
Files |
Tested on |
|---|---|---|
Operational Sentinel-5P L2 CH4 orbits ( |
A 2023 operational orbit (processor 2.5) and a 2018 blended file. |
|
OCO-2 and OCO-3 L2 Lite XCO2 files ( |
A synthetic file with the documented v11 layout. No real granule has been read yet. |
|
TCCON GGG2020 public site files from CaltechDATA
( |
The Indianapolis GGG2020.R0 file. |
|
GGG2020 |
Real EM27/SUN days. The committed test file is synthetic. |
|
GGG2020 netCDF files: the |
The public layout on a real TCCON file. The private layout only on a synthetic file. |
The satellite readers take lon_range and lat_range to keep only the
pixels inside a box. Use them on a whole orbit. The GGG readers take a
species (xco2, xch4, xco, …), and the netCDF readers also
take a time_range for multi-year site files. The readers do not filter
anything else. Quality flags become the good column, and you decide
what to keep.
EM27/SUN#
GGG (through EGI) writes one .oof file per instrument-day and, from the
same run, one *.private.nc. The .oof has the soundings but no
averaging kernel or prior, so read_ggg_oof()
leaves out the ak, ak_pressure and pressure_levels columns. Take
the kernels from the private file with
read_ggg_netcdf(), or from a site kernel table
keyed by solar zenith angle. The Slant Columns guide shows both.
To read a campaign as one table:
from pathlib import Path
import pandas as pd
from stilt.observations import read_ggg_oof
days = sorted(Path("EM27_oof/ha").glob("ha*.vav.ada.aia.oof"))
df = pd.concat([read_ggg_oof(p, "xch4") for p in days], ignore_index=True)
df = df[df.good]
There is no reader yet for EM27/SUN data processed with PROFFAST (the COCCON pipeline).
The columns#
Vertical arrays run from the surface upward. Pressures are in hPa and
altitudes in metres above mean sea level. Azimuths are degrees clockwise
from north, pointing from the ground toward the instrument or the sun. This
is the convention slant_points() uses.
Every reader returns these columns:
Column |
Meaning |
|---|---|
|
A string that identifies the sounding within the product. |
|
UTC time of the sounding, timezone-naive. |
|
Pixel centre or station location, degrees. |
|
Surface (or station) altitude, m MSL. |
|
Surface pressure the retrieval used, hPa. |
|
The column-average dry-air mole fraction and its one-sigma error. |
|
The gas ( |
|
The product’s recommended quality screen, as a boolean. |
All readers except read_ggg_oof() also return:
Column |
Meaning |
|---|---|
|
The averaging kernel and the pressures it is given at, one array per
row. For a layer product (TROPOMI) the pressures are the layer
midpoints. Pass both to
|
|
The retrieval’s own pressure grid, for
|
Some columns appear only when the product has them:
Column |
Meaning |
|---|---|
|
The viewing geometry for a slant path. For a satellite these are the sensor angles toward the satellite. For a ground-based spectrometer they are the solar angles. The blended TROPOMI files have neither. |
|
The solar angles. In GGG files they equal |
|
The heights of |
|
The prior profile as a mole fraction in |
|
The prior column value (OCO-2). |
|
Pixel corners, one array per row, for
|
Other columns keep the product’s own names, for example qa_value,
quality_flag, operation_mode, xch4_uncorrected, flag,
zmin, and the other x<gas> columns of a GGG file. A private GGG file
adds ak_extrapolated. It is true where the spectrum’s slant xgas fell
outside GGG’s kernel table.
From a file to receptors#
Here is the recipe from Satellite And Column Observations with a real reader. Each TROPOMI sounding becomes a slant receptor on the retrieval’s own levels, up to 3 km above the surface:
import stilt
from stilt.observations import read_tropomi_ch4, slant_points
from stilt.transforms import averaging_kernel_table
model = stilt.Model(project="./xch4")
df = read_tropomi_ch4(path, lon_range=(-113.5, -110.5), lat_range=(39.5, 42.0))
df = df[df.good]
receptors = [
stilt.Receptor.from_points(
r.time,
slant_points(
r.longitude, r.latitude,
r.altitude_levels[r.altitude_levels < r.surface_altitude + 3000],
zenith=r.zenith, azimuth=r.azimuth,
),
altitude_ref="msl",
)
for r in df.itertuples()
]
model.register(receptors=receptors)
averaging_kernel_table(receptors, levels=df.ak_pressure, values=df.ak).to_parquet(
model.project.directory / "kernels.parquet"
)
model.run()
The footprint config then lists two transforms:
transforms:
- kind: averaging_kernel
table: kernels.parquet
coordinate: pres
- kind: pressure_weighting
A few products need a change to this recipe:
For a nadir column, use a
ColumnReceptorfromsurface_altitudein place of the slant points.OCO-2 files give pressures but no heights. Turn
pressure_levelsinto altitudes withpressure_altitudes().A TCCON prior grid starts at sea level, below the station. Drop the levels below
surface_altitudeand putsurface_altitudefirst, so the path starts at the instrument.
Adding an instrument#
Copy the reader module closest to your product from
stilt/observations/readers/. Use tropomi.py for a swath product and
tccon.py or ggg.py for a ground station. The reader holds all of the
product’s conventions, and its output must have the columns above. Keep it a
plain function that returns the table.
Add a small slice of a real file under tests/data/products, and a test
that checks the columns, the units, and that the vertical arrays start at
the surface. tests/data/products/make_samples.py shows how the existing
slices were cut. Readers for PROFFAST EM27/SUN output, MethaneAIR, and
MethaneSAT are welcome. The maintainers have no files to write them
against.