Pollutants#

How UATAQ names concentration columns, and how to pick them out.

Concentration columns follow {pollutant}[d|channel]_{units}[_cal|_raw]: O3_ppb, CO2d_ppm_cal (dry mole fraction, calibrated), CH4d_ppm_raw, PM2.5_ugm3, BC6_ngm3 (aethalometer channel 6).

import uataq
from uataq.pollutants import concentration_columns, parse_column

obs = uataq.get_obs("WBB", "CO2", time_range="2024-01")
obs[concentration_columns("CO2", obs.columns)]  # CO2d_ppm_cal, not diagnostics
parse_column("CH4d_ppm_raw")  # pollutant='CH4', units='ppm', dry=True, calibration='raw'

_cal vs _raw: in lin final data _cal is the calibrated value and _raw the uncalibrated one, except for TRX01’s lgr_ugga_manual_cal (from 2023-11-18), which leaves _cal empty and writes the LGR-software calibrated value to _raw.

Labels, units and expected ranges for plotting are instrument-independent and live in lair.pollutants (get_pollutant("CH4").label()).

How UATAQ names pollutant columns.

UATAQ concentration columns follow {pollutant}[d|channel]_{units}[_cal|_raw]: O3_ppb, CO2d_ppm_cal (dry mole fraction, calibrated), CH4d_ppm_raw (uncalibrated), PM2.5_ugm3, BC6_ngm3 (aethalometer channel 6). This module picks those columns out of a DataFrame and parses their names.

Display metadata (long names, LaTeX labels, expected ranges) is instrument-independent and lives in lair.pollutants; uataq does not import lair.

uataq.pollutants.UNITS: tuple[str, ...] = ('ppm', 'ppb', 'ugm3', 'ngm3')#

Units a UATAQ concentration column can carry.

uataq.pollutants.POLLUTANTS: tuple[str, ...] = ('BC', 'CH4', 'CO', 'CO2', 'NO', 'NO2', 'NOx', 'O3', 'PM1', 'PM10', 'PM2.5', 'PM4')#

Every pollutant some configured instrument model measures, in declared case.

uataq.pollutants.concentration_columns(pollutant, columns, raw=False)[source]#

Pick the columns holding a pollutant’s measured concentration.

Matches {pollutant}_{unit} with a concentration unit, e.g. O3_ppb, CO2d_ppm (dry mole fraction), CH4d_ppm_cal (calibrated), PM2.5_ugm3 or BC6_ngm3 (an aethalometer channel). Instrument diagnostics that share the prefix, like O3_Meas_mV or NO2_Slope, spreads like O3_ppb_std, and other pollutants sharing the prefix (NO vs NO2_ppb) are not matches.

Parameters:
  • pollutant (str) – The pollutant, matched case-insensitively.

  • columns (Iterable[str]) – Column names to search.

  • raw (bool) – Also match the uncalibrated ..._raw columns that sit beside ..._cal in lin final data. Default False.

Returns:

The matching columns, in their original order.

Return type:

list[str]

class uataq.pollutants.ConcentrationColumn(name, pollutant, units, dry=False, calibration=None, channel=None)[source]#

A parsed concentration column name.

Parameters:
  • name (str) – The column name, e.g. 'CO2d_ppm_cal'.

  • pollutant (str) – The pollutant in declared case (see POLLUTANTS), e.g. 'CO2'.

  • units (str) – One of UNITS.

  • dry (bool) – Dry mole fraction (the d after the pollutant).

  • calibration (str | None) – 'cal', 'raw', or None when the name carries no suffix.

  • channel (int | None) – Aethalometer wavelength channel, for black carbon.

name: str#
pollutant: str#
units: str#
dry: bool = False#
calibration: str | None = None#
channel: int | None = None#
__init__(name, pollutant, units, dry=False, calibration=None, channel=None)#
uataq.pollutants.parse_column(column)[source]#

Parse a concentration column name.

Parameters:

column (str) – A column name, e.g. 'CH4d_ppm_raw' or 'BC6_ngm3'.

Returns:

The parts of the name, or None if it is not a concentration column of a known pollutant ('Time_UTC', 'O3_Meas_mV', …).

Return type:

ConcentrationColumn | None