Quick Start#
Laboratory#
It all starts in the lab…
The laboratory object is a singleton instance of the Laboratory
class which is initialized with the UATAQ configuration file.
The configuration file is a JSON file which specifies UATAQ site characteristics
including name, location, status, research groups collecting data, and installed instruments.
The Laboratory object contains the following attributes:
sites : A list of site identifiers.
instruments : A list of instrument names.
Research Sites#
The Site object is the primary interface for accessing data from a UATAQ site.
Each site has a unique site identifier (SID) that corresponds to a key in the configuration file.
The lab is responsible for constructing Site objects from the configuration file,
including building the InstrumentEnsemble for each site.
The InstrumentEnsemble is a container object that hold different
Instrument objects which provide the linkage between a Site and the data files.
The Site object contains the following information as attributes:
SID : The site identifier.
config : A dictionary containing configuration information for the site from the config file.
instruments : An instance of the InstrumentEnsemble class representing the instruments at the site.
groups : The research groups that collect data at the site.
loggers : The loggers used by research groups that record data at a site.
pollutants : The pollutants measured at the site.
There are two primary methods for reading data from a site:
Reading Instrument Data - Data for each instrument at a site is read individually and stored in a dictionary with the instrument name as the key.
Getting Observations - Finalized observations from all instruments at a site are aggregated into a single dataframe.
Site.read_data()andSite.get_obs()have been wrapped inuataq.read_data()anduataq.get_obs()respectively for convenience with an added SID parameter.
Reading Instrument Data#
Using a Site object we can read the data from each instrument
at the site for a specified processing lvl and time range:
The data is returned as a dictionary of pandas dataframes, one for each instrument. The dataframes are indexed by time and have columns for each variable:
Getting Observations#
Or we can only get the finalized observations for a site which aggregates the instruments into a single dataframe:
Finalized observations only include data which has passed QAQC (QAQC_Flag >= 0)
and that are measurements of the ambient atmosphere (ID == -10).
The observations dataframe is indexed by time and aggregates pollutants into a single dataframe.
Two formats are available: wide or long.
The wide format has columns for each pollutant and
the long format has a pollutant column with the pollutant name
and a value column with the measurement value.
Mobile Sites & Observations#
Included as part of UATAQ is the TRAX/eBus project, which collects data from mobile sites.
The MobileSite object is a subclass of the Site object.
The laboratory determines whether to build a Site or
MobileSite object based on the is_mobile attribute in the configuration file.
Mobile sites provide the same functionality as fixed sites, but merge location
data with observations when using the get_obs() method
and return a geodataframe.
Each instrument’s rows are located with the GPS of the research group they were
read from, since that GPS was stamped by the same logger clock (see
locate()). With no group given, a TRAX read can take
methane from lin and ozone from horel, each joined to its own group’s GPS. The
GPS_Group column says which GPS located each row. It also says which clock
the row’s Time_UTC is on: lin rows are on GPS time, horel rows on the horel
logger’s clock, which runs a few seconds ahead.
Or in the long format: