Reading Retrieval Products#

A column retrieval comes as a product file, such as a TROPOMI orbit, an OCO-2 Lite file, a TCCON site file, or a day of EM27/SUN spectra from GGG. The readers in stilt.observations turn each of these into a table with one row per sounding. The columns are the same for every instrument, so the rest of the workflow (Satellite And Column Observations) is too.

from stilt.observations import read_tropomi_ch4

df = read_tropomi_ch4(
    "S5P_OFFL_L2__CH4____20231019T192749_20231019T210918_31176_03_020500_20231021T113213.nc",
    lon_range=(-113.5, -110.5), lat_range=(39.5, 42.0),
)
df = df[df.good]

The readers#

Reader

Files

Tested on

read_tropomi_ch4()

Operational Sentinel-5P L2 CH4 orbits (S5P_*_L2__CH4___, v2.x) and the TROPOMI+GOSAT blended files (S5P_BLND_L2__CH4___)

A 2023 operational orbit (processor 2.5) and a 2018 blended file.

read_oco2()

OCO-2 and OCO-3 L2 Lite XCO2 files (oco2_LtCO2_*.nc4, oco3_LtCO2_*.nc4, Lite FP v10/v11)

A synthetic file with the documented v11 layout. No real granule has been read yet.

read_tccon()

TCCON GGG2020 public site files from CaltechDATA (*.public.nc, *.public.qc.nc)

The Indianapolis GGG2020.R0 file.

read_ggg_oof()

GGG2020 .oof files (*.vav.ada.aia.oof), one instrument-day each. This is how EGI delivers EM27/SUN retrievals.

Real EM27/SUN days. The committed test file is synthetic.

read_ggg_netcdf()

GGG2020 netCDF files: the *.private.nc a GGG run writes, and the public files (read_tccon calls this reader)

The public layout on a real TCCON file. The private layout only on a synthetic file.

The satellite readers take lon_range and lat_range to keep only the pixels inside a box. Use them on a whole orbit. The GGG readers take a species (xco2, xch4, xco, …), and the netCDF readers also take a time_range for multi-year site files. The readers do not filter anything else. Quality flags become the good column, and you decide what to keep.

EM27/SUN#

GGG (through EGI) writes one .oof file per instrument-day and, from the same run, one *.private.nc. The .oof has the soundings but no averaging kernel or prior, so read_ggg_oof() leaves out the ak, ak_pressure and pressure_levels columns. Take the kernels from the private file with read_ggg_netcdf(), or from a site kernel table keyed by solar zenith angle. The Slant Columns guide shows both.

To read a campaign as one table:

from pathlib import Path
import pandas as pd
from stilt.observations import read_ggg_oof

days = sorted(Path("EM27_oof/ha").glob("ha*.vav.ada.aia.oof"))
df = pd.concat([read_ggg_oof(p, "xch4") for p in days], ignore_index=True)
df = df[df.good]

There is no reader yet for EM27/SUN data processed with PROFFAST (the COCCON pipeline).

The columns#

Vertical arrays run from the surface upward. Pressures are in hPa and altitudes in metres above mean sea level. Azimuths are degrees clockwise from north, pointing from the ground toward the instrument or the sun. This is the convention slant_points() uses.

Every reader returns these columns:

Column

Meaning

sounding_id

A string that identifies the sounding within the product.

time

UTC time of the sounding, timezone-naive.

longitude, latitude

Pixel centre or station location, degrees.

surface_altitude

Surface (or station) altitude, m MSL.

surface_pressure

Surface pressure the retrieval used, hPa.

value, uncertainty

The column-average dry-air mole fraction and its one-sigma error.

species, units

The gas (xch4, xco2, …) and the units of value (ppb or ppm).

good

The product’s recommended quality screen, as a boolean.

All readers except read_ggg_oof() also return:

Column

Meaning

ak_pressure, ak

The averaging kernel and the pressures it is given at, one array per row. For a layer product (TROPOMI) the pressures are the layer midpoints. Pass both to averaging_kernel_table() and set coordinate: pres on the averaging_kernel transform.

pressure_levels

The retrieval’s own pressure grid, for pressure_altitudes().

Some columns appear only when the product has them:

Column

Meaning

zenith, azimuth

The viewing geometry for a slant path. For a satellite these are the sensor angles toward the satellite. For a ground-based spectrometer they are the solar angles. The blended TROPOMI files have neither.

solar_zenith, solar_azimuth

The solar angles. In GGG files they equal zenith and azimuth.

altitude_levels

The heights of pressure_levels, m MSL, in operational TROPOMI and GGG netCDF files. Use these for a slant path instead of converting pressures.

apriori

The prior profile as a mole fraction in units. TROPOMI and OCO-2 give it on the kernel’s levels, GGG on pressure_levels.

apriori_column

The prior column value (OCO-2).

longitude_bounds, latitude_bounds

Pixel corners, one array per row, for jitter_points() (TROPOMI, OCO-2).

Other columns keep the product’s own names, for example qa_value, quality_flag, operation_mode, xch4_uncorrected, flag, zmin, and the other x<gas> columns of a GGG file. A private GGG file adds ak_extrapolated. It is true where the spectrum’s slant xgas fell outside GGG’s kernel table.

From a file to receptors#

Here is the recipe from Satellite And Column Observations with a real reader. Each TROPOMI sounding becomes a slant receptor on the retrieval’s own levels, up to 3 km above the surface:

import stilt
from stilt.observations import read_tropomi_ch4, slant_points
from stilt.transforms import averaging_kernel_table

model = stilt.Model(project="./xch4")

df = read_tropomi_ch4(path, lon_range=(-113.5, -110.5), lat_range=(39.5, 42.0))
df = df[df.good]

receptors = [
    stilt.Receptor.from_points(
        r.time,
        slant_points(
            r.longitude, r.latitude,
            r.altitude_levels[r.altitude_levels < r.surface_altitude + 3000],
            zenith=r.zenith, azimuth=r.azimuth,
        ),
        altitude_ref="msl",
    )
    for r in df.itertuples()
]
model.register(receptors=receptors)

averaging_kernel_table(receptors, levels=df.ak_pressure, values=df.ak).to_parquet(
    model.project.directory / "kernels.parquet"
)
model.run()

The footprint config then lists two transforms:

transforms:
  - kind: averaging_kernel
    table: kernels.parquet
    coordinate: pres
  - kind: pressure_weighting

A few products need a change to this recipe:

  • For a nadir column, use a ColumnReceptor from surface_altitude in place of the slant points.

  • OCO-2 files give pressures but no heights. Turn pressure_levels into altitudes with pressure_altitudes().

  • A TCCON prior grid starts at sea level, below the station. Drop the levels below surface_altitude and put surface_altitude first, so the path starts at the instrument.

Adding an instrument#

Copy the reader module closest to your product from stilt/observations/readers/. Use tropomi.py for a swath product and tccon.py or ggg.py for a ground station. The reader holds all of the product’s conventions, and its output must have the columns above. Keep it a plain function that returns the table.

Add a small slice of a real file under tests/data/products, and a test that checks the columns, the units, and that the vertical arrays start at the surface. tests/data/products/make_samples.py shows how the existing slices were cut. Readers for PROFFAST EM27/SUN output, MethaneAIR, and MethaneSAT are welcome. The maintainers have no files to write them against.