Instruments#

UATAQ instruments as classes.

Each instrument class is a subclass of the Instrument abstract base class and implements methods for reading and parsing data files.

The Instrument class provides a common interface for all instrument classes and defines abstract methods that must be implemented by each subclass.

uataq.instruments.GroupSelection = str | collections.abc.Mapping[str, str] | None#

a name, a per-instrument mapping, or None for automatic selection. See Instrument.resolve_group().

Type:

How a caller picks a research group

uataq.instruments.ReadPlan#

which group to read each portion of a time range from. See Instrument.plan_reads().

Type:

A read plan

alias of list[tuple[str, TimeRange]]

class uataq.instruments.Instrument(SID, name, loggers, config)[source]#

Bases: object

Abstract base class for instrument objects.

model#

Model of the instrument.

Type:

str

SID#

Site ID where the instrument is installed.

Type:

str

name#

Name of the instrument.

Type:

str

groups#

Research groups that operate the instrument.

Type:

list[str]

group_dates#

When each group’s archive holds this instrument’s data, for the groups whose archive covers less than the whole installation (config group_dates). A group not listed covers the whole installation.

Type:

dict[str, TimeRange]

loggers#

Loggers used by the research groups to record data.

Type:

set[str]

config#

Configuration settings for the instrument.

Type:

dict

get_files(group: str, lvl: str) → list[str][source]#

Get list of file paths for a given level.

read_data(group: str, lvl: str, time_range: TimeRange, num_processes: int, file_pattern: str) → pd.DataFrame[source]#

Read and parse group data files for the given level and time range using multiple processes.

model: str#
__init__(SID, name, loggers, config)[source]#

Initialize the Instrument object.

Parameters:
  • SID (str) – Site ID where the instrument is installed.

  • name (str) – Name of the instrument.

  • loggers (dict) – Dictionary of loggers used by different research groups.

  • config (dict) – Configuration settings for the instrument.

resolve_group(group=None)[source]#

Pick which research group’s data to read this instrument from.

Parameters:

group (str | Mapping[str, str] | None) – A group name, used as given; a mapping of instrument name to group name, from which this instrument’s entry is used (instruments the mapping does not name fall back to automatic selection); or None to select automatically.

Returns:

The group name.

Return type:

str

Raises:

InvalidGroupError – If no registered groupspace operates this instrument.

Notes

Automatic selection reads the configured operators of this instrument: the default group when it is one of them, otherwise the sole operator, otherwise the first configured. It is a configuration lookup, not a search of the archive – it does not check whether that group actually holds data for a given time range, and it ignores group_dates. plan_reads() is the time-aware version that uataq.sites.Site.read_data() uses.

plan_reads(group=None, time_range=None)[source]#

Plan which group to read each part of a time range from.

Parameters:
  • group (str | Mapping[str, str] | None) – As for resolve_group(). A name, or a mapping entry naming this instrument, reads the whole range from that group.

  • time_range (TimeRange | TimeRangeTypes) – The requested time range. Default None is the whole installation.

Returns:

(group, portion) pairs in time order. The portions are half-open, do not overlap, and lie within the requested range clipped to active_range. Usually a single pair.

Return type:

list[tuple[str, TimeRange]]

Raises:
  • InactiveInstrumentError – If the requested range misses the installation entirely.

  • InvalidGroupError – If no registered groupspace operates this instrument.

  • ReaderError – If no group’s archive covers any of the requested range.

Notes

With no group named, a range that crosses a group_dates boundary is split there. Each portion goes to the most preferred group whose window covers it – the default group, then the configured order, as in resolve_group() – and portions no group covers are skipped. An instrument without group_dates gets [(resolve_group(group), clipped range)], as before.

get_highest_lvl(group)[source]#

Get the highest data level for the instrument.

Parameters:

group (str) – The research group whose data to retrieve.

Returns:

The highest data level.

Return type:

str

get_files(group, lvl)[source]#

Get list of file paths for a given level.

Parameters:
  • group (str) – The research group whose data to retrieve.

  • lvl (str) – The level of the data to retrieve.

Returns:

A list of file paths.

Return type:

list[str]

get_datafiles(group, lvl, time_range, pattern=None)[source]#

Get data files for the given level and time range from the groupspace.

Parameters:
  • group (str) – The research group whose data to retrieve.

  • lvl (str) – The level of the data to retrieve.

  • time_range (TimeRange | TimeRangeTypes) – The time range of the data to retrieve.

  • pattern (str) – A string pattern to filter the file paths.

Returns:

A list of data files.

Return type:

list[DataFile]

property active_range: TimeRange#

When this instrument was installed at the site and when it was removed.

An instrument still installed has no stop. Built from the site configuration’s installation_date / removal_date.

clip_to_active(time_range)[source]#

Narrow a requested time range to when this instrument was installed.

Parameters:

time_range (TimeRange | TimeRangeTypes) – The requested time range.

Returns:

The requested range intersected with active_range. Never wider than what was asked for.

Return type:

TimeRange

Raises:

InactiveInstrumentError – If the requested range does not overlap the active range at all. Both are half-open, so a range that only touches it – stopping at the installation date, or starting at the removal date – does not overlap it.

Notes

Without this, a request that reaches past a swap reads the replacement instrument’s files as though they were this one’s: research groups reuse a file name across an instrument change (the horel group calls both MetOne models esampler), so the file name cannot distinguish them, but the installation and removal dates can.

standardize_data(group, data)[source]#

Standardize the data across research groups.

Rename columns, convert units, map values, etc. as needed.

Parameters:
  • group (str) – The research group whose data to standardize.

  • data (pandas.DataFrame) – The data to standardize.

Returns:

The standardized data.

Return type:

pandas.DataFrame

read_data(group, lvl=None, time_range=None, num_processes=1, file_pattern=None)[source]#

Read and parse data files for the given level and time range.

Uses multiple processes if specified.

Parameters:
  • group (str) – The research group whose data to read.

  • lvl (str) – The level of the data to read.

  • time_range (TimeRange | TimeRangeTypes) – The time range to read data. Default is None which reads all available data.

  • num_processes (int | 'max') – The number of processes to use for parallelization.

  • file_pattern (str) – A string pattern to filter the file paths.

Returns:

A concatenated DataFrame containing the parsed data from files.

Return type:

pandas.DataFrame

uataq.instruments.configure_instrument(SID, name, config, loggers=None)[source]#

Configure an instrument object based on the given configuration settings.

Parameters:
  • SID (str) – Site ID where the instrument is installed.

  • name (str) – Name of the instrument.

  • config (dict) – Configuration settings for the instrument.

  • loggers (dict, optional) – Dictionary of loggers used by different research groups.

Returns:

An instrument object configured with the given settings.

Return type:

Instrument

Raises:
  • ValueError – If the instrument model is not found in the catalog.

  • ValueError – If no loggers are found for the instrument at the site.

class uataq.instruments.InstrumentEnsemble(SID, configs, loggers=None)[source]#

Bases: object

Container for an ensemble of instruments at a site.

SID#

Site ID of the ensemble.

Type:

str

configs#

Dictionary of configuration settings for each instrument.

Type:

dict[str, dict]

names#

List of instrument names in the ensemble.

Type:

list[str]

loggers#

Set of loggers used by the research groups.

Type:

set[str]

groups#

Set of research groups that operate the instruments.

Type:

set[str]

pollutants#

Set of pollutants measured by the instruments.

Type:

set[str]

__init__(SID, configs, loggers=None)[source]#

Initialize the InstrumentEnsemble object.

Parameters:
  • SID (str) – Site ID of the ensemble.

  • configs (dict[instrument, config]) – Dictionary of configuration settings for each instrument.

  • loggers (dict[group, logger], optional) – Dictionary of loggers used by different research groups.

class uataq.instruments.SensorMixin[source]#

Bases: object

Mixin for instrument objects that measure a pollutant.

pollutants(tuple)#
Type:

Tuple of pollutants measured by the instrument.

pollutants: tuple[str, ...]#
class uataq.instruments.BB_205(SID, name, loggers, config)[source]#

Bases: Instrument, SensorMixin

2B Technologies Model 205 ozone monitor (UV absorption).

model: str = '2b_205'#
pollutants: tuple[str, ...] = ('O3',)#
class uataq.instruments.BB_405(SID, name, loggers, config)[source]#

Bases: Instrument, SensorMixin

2B Technologies Model 405 nm NO/NO2/NOx monitor.

model: str = '2b_405'#
pollutants: tuple[str, ...] = ('NO', 'NO2', 'NOx')#
class uataq.instruments.CR1000(SID, name, loggers, config)[source]#

Bases: Instrument

Campbell Scientific CR1000 datalogger.

Not a sensor itself: it records housekeeping such as battery voltage and enclosure temperature.

model: str = 'cr1000'#
class uataq.instruments.GPS(SID, name, loggers, config)[source]#

Bases: Instrument

GPS receiver providing position, and speed and course where logged.

Recorded speed is given as Speed_m_s (lin’s NMEA knots are converted; horel logs m/s). Receivers logging only GPGGA sentences record neither speed nor course; both are then estimated from the positions (uataq.gps.estimate_speed_course()) and flagged in Speed_Estimated / Course_Estimated.

model: str = 'gps'#
estimate_window: int = 1#

Samples on each side of the centered difference used to estimate speed and course from positions. See uataq.gps.estimate_speed_course().

estimate_max_gap: str = '60s'#

Longest interval between consecutive fixes that an estimate may span.

read_data(group, lvl=None, time_range=None, num_processes=1, file_pattern=None, estimate_motion=True)[source]#

Read GPS data, with speed in m/s and course in degrees.

Extends Instrument.read_data(). Recorded speed is Speed_m_s (lin’s files hold NMEA knots, converted here; horel’s are already m/s), and course is Course_deg.

horel data is indexed by the CR1000 logger’s clock, which runs ahead of GPS time by a drifting 1-20 s. Where the receiver’s time of day was logged (Instrument_Time, horel’s raw GTIM), the true time of each fix is added as GPS_Time_UTC. lin’s Time_UTC is already GPS time.

Receivers logging only GPGGA sentences record neither, in which case both are estimated from the positions (see uataq.gps.estimate_speed_course()) and the boolean columns Speed_Estimated / Course_Estimated mark every value that came from positions rather than from the receiver. Recorded values are never overwritten.

Parameters:

estimate_motion (bool) – Fill missing speed and course from the positions. Default True.

Return type:

DataFrame Only reachable through the instrument object – uataq.read_data() does not forward it.

:param See Instrument.read_data() for the other parameters.:

classmethod estimate_motion(data)[source]#

Fill missing Speed_m_s / Course_deg from the positions.

Adds Speed_Estimated and Course_Estimated, which are True exactly where the value was derived from positions rather than recorded by the receiver. Returns data unchanged if it has no positions or no DatetimeIndex to difference against.

Return type:

DataFrame

class uataq.instruments.LGR_NO2(SID, name, loggers, config)[source]#

Bases: Instrument, SensorMixin

Los Gatos Research NO2 analyzer (cavity-enhanced absorption).

model: str = 'lgr_no2'#
pollutants: tuple[str, ...] = ('NO2',)#
class uataq.instruments.LGR_UGGA(SID, name, loggers, config)[source]#

Bases: Instrument, SensorMixin

Los Gatos Research Ultraportable Greenhouse Gas Analyzer.

Measures CO2 and CH4 by off-axis ICOS.

model: str = 'lgr_ugga'#
pollutants: tuple[str, ...] = ('CO2', 'CH4')#
class uataq.instruments.Licor_6262(SID, name, loggers, config)[source]#

Bases: Instrument, SensorMixin

LI-COR LI-6262 infrared CO2/H2O gas analyzer.

model: str = 'licor_6262'#
pollutants: tuple[str, ...] = ('CO2',)#
class uataq.instruments.Licor_7000(SID, name, loggers, config)[source]#

Bases: Licor_6262

LI-COR LI-7000 infrared CO2/H2O gas analyzer.

Parsed like the LI-6262.

model: str = 'licor_7000'#
class uataq.instruments.Magee_AE33(SID, name, loggers, config)[source]#

Bases: Instrument, SensorMixin

Magee Scientific AE33 aethalometer, measuring black carbon.

model: str = 'magee_ae33'#
pollutants: tuple[str, ...] = ('BC',)#
class uataq.instruments.MetOne_ES405(SID, name, loggers, config)[source]#

Bases: Instrument, SensorMixin

Met One E-Sampler ES-405, reporting PM1, PM2.5, PM4 and PM10.

Replaced the ES-642 at several sites. The horel group names both models esampler on disk, so only the configured installation and removal dates separate them – see Instrument.clip_to_active().

model: str = 'metone_es405'#
pollutants: tuple[str, ...] = ('PM1', 'PM2.5', 'PM4', 'PM10')#
class uataq.instruments.MetOne_ES642(SID, name, loggers, config)[source]#

Bases: Instrument, SensorMixin

Met One E-Sampler ES-642, reporting PM2.5 only.

Superseded by the ES-405 at several sites; see MetOne_ES405.

model: str = 'metone_es642'#
pollutants: tuple[str, ...] = ('PM2.5',)#
class uataq.instruments.Teledyne_T200(SID, name, loggers, config)[source]#

Bases: Instrument, SensorMixin

Teledyne API T200 chemiluminescence NO/NO2/NOx analyzer.

model: str = 'teledyne_t200'#
pollutants: tuple[str, ...] = ('NO', 'NO2', 'NOx')#
class uataq.instruments.Teledyne_T300(SID, name, loggers, config)[source]#

Bases: Instrument, SensorMixin

Teledyne API T300 gas-filter-correlation CO analyzer.

model: str = 'teledyne_t300'#
pollutants: tuple[str, ...] = ('CO',)#
class uataq.instruments.Teledyne_T400(SID, name, loggers, config)[source]#

Bases: Instrument, SensorMixin

Teledyne API T400 UV-absorption ozone analyzer.

model: str = 'teledyne_t400'#
pollutants: tuple[str, ...] = ('O3',)#
class uataq.instruments.Teledyne_T500u(SID, name, loggers, config)[source]#

Bases: Instrument, SensorMixin

Teledyne API T500U CAPS NO2 analyzer.

model: str = 'teledyne_t500u'#
pollutants: tuple[str, ...] = ('NO2',)#
class uataq.instruments.Teom_1400ab(SID, name, loggers, config)[source]#

Bases: Instrument, SensorMixin

Thermo/R&P TEOM 1400ab tapered-element oscillating microbalance.

Measures PM2.5 mass.

model: str = 'teom_1400ab'#
pollutants: tuple[str, ...] = ('PM2.5',)#
uataq.instruments.catalog: dict[str, type[Instrument]] = {'2b_205': <class 'uataq.instruments.BB_205'>, '2b_405': <class 'uataq.instruments.BB_405'>, 'cr1000': <class 'uataq.instruments.CR1000'>, 'gps': <class 'uataq.instruments.GPS'>, 'lgr_no2': <class 'uataq.instruments.LGR_NO2'>, 'lgr_ugga': <class 'uataq.instruments.LGR_UGGA'>, 'licor_6262': <class 'uataq.instruments.Licor_6262'>, 'licor_7000': <class 'uataq.instruments.Licor_7000'>, 'magee_ae33': <class 'uataq.instruments.Magee_AE33'>, 'metone_es405': <class 'uataq.instruments.MetOne_ES405'>, 'metone_es642': <class 'uataq.instruments.MetOne_ES642'>, 'teledyne_t200': <class 'uataq.instruments.Teledyne_T200'>, 'teledyne_t300': <class 'uataq.instruments.Teledyne_T300'>, 'teledyne_t400': <class 'uataq.instruments.Teledyne_T400'>, 'teledyne_t500u': <class 'uataq.instruments.Teledyne_T500u'>, 'teom_1400ab': <class 'uataq.instruments.Teom_1400ab'>}#

Instrument catalog