Sites#
Classes
|
A class representing a site where atmospheric measurements are taken. |
|
A class representing a mobile site where atmospheric measurements are taken. |
- class uataq.sites.Site(SID, config, instruments)[source]#
A class representing a site where atmospheric measurements are taken.
- instruments#
An instance of the InstrumentEnsemble class representing the instruments at the site.
- Type:
- read_data(instruments='all', lvl=None, time_range=None, num_processes=1, file_pattern=None)[source]#
Read data for each instrument for specified level.
- read_obs(pollutants='all', format='wide', time_range=None, num_processes=1)#
Read observations for each pollutant, combining instruments by pollutants.
- get_recent_obs(recent=dt.timedelta(days=10), lvl='qaqc')[source]#
Get recent observations from site instruments.
- __init__(SID, config, instruments)[source]#
Initialize a Site object with the given site ID.
- Parameters:
SID (str) – The site identifier.
config (dict) –
A dictionary containing configuration information for the site:
{ name: str, is_active: bool, is_mobile: bool, latitude: float, longitude: float, zagl: float, loggers: dict, instruments: { instrument: { loggers: dict installation_date: str, removal_date: str, } } }
instruments (InstrumentEnsemble) – An instance of the InstrumentEnsemble class representing the instruments at the site.
- read_data(instruments='all', group=None, lvl=None, time_range=None, num_processes=1, file_pattern=None)[source]#
Read data for the specified instruments and level.
- Parameters:
instruments (str or list of str or 'all') – The instrument(s) to read data from. If ‘all’, read data from all instruments. Default is ‘all’.
group (str | Mapping[str, str] | None) – The research group to read data from. A name applies to every instrument; a mapping of instrument name to group name sets it per instrument. Default None selects each instrument’s group automatically, by time when the groups’ archives cover different periods: a range crossing a
group_datesboundary is read from each group in turn and concatenated (seeplan_reads()).lvl (str, optional) – The data level to read. Default is None which reads the highest level available (per group, when the read is split across groups).
time_range (TimeRange | TimeRangeTypes) – The time range to read data. Default is None which reads all available data.
num_processes (int or 'max') – The number of processes to use for reading data. Default is 1.
file_pattern (str, optional) – The file pattern to use for filtering files. Default is None.
- Returns:
A dictionary containing the data for each instrument.
- Return type:
- Raises:
ReaderError – If no data is found for the specified instruments.
- get_obs(pollutants='all', format='wide', group=None, time_range=None, num_processes=1, **kwargs)[source]#
Get observations for each pollutant, combining instruments by pollutants.
- Parameters:
pollutants (str or list of str, optional) – pollutants to read. If ‘all’, read all pollutants. Default is ‘all’.
format (str, optional) – Format of the data to return. Default is ‘wide’.
group (str | Mapping[str, str] | None) – The research group to read data from. A name applies to every instrument; a mapping of instrument name to group name sets it per instrument. Default None selects each instrument’s group automatically (see
resolve_group()).time_range (TimeRange | TimeRangeTypes) – The time range to read data. Default is None which reads all available data.
num_processes (int, optional) – Number of processes to use for reading data. Default is 1.
- Returns:
A dictionary of dataframes, one for each level of data read, or a single dataframe if only one level was read. The keys of the dictionary are the names of the levels (‘calibrated’, ‘qaqc’, ‘raw’), and the values are the corresponding dataframes. If only one level was read, the method returns the corresponding dataframe directly.
- Return type:
Union[Dict[str, pandas.DataFrame], pandas.DataFrame]
- get_recent_obs(recent=datetime.timedelta(days=10), pollutants='all', format='wide', group=None)[source]#
Get recent observations from site instruments.
- Parameters:
recent (str or datetime.timedelta, optional) – Time range to get recent observations. Default is 10 days.
pollutants (str or list of str, optional) – Pollutants to read. If ‘all’, read all pollutants. Default is ‘all’.
format (str, optional) – Format of the data to return. Default is ‘wide’.
group (str, optional) – Research group to read data from. Defaults to None which uses the default group.
- Returns:
A dataframe containing recent observations from site instruments.
- Return type:
- class uataq.sites.MobileSite(SID, config, instruments)[source]#
A class representing a mobile site where atmospheric measurements are taken.
- Parameters:
- static merge_gps(obs, gps, on=None, obs_on=None, gps_on=None)[source]#
Merge observation data with location data from GPS.
- Parameters:
(pd.DataFrame) (gps)
(pd.DataFrame)
(str (gps_on)
optional) (The column name in the GPS data to merge on. If not specified, it will use the value of 'on'.)
(str
optional)
(str
optional)
- Returns:
pd.DataFrame
- Return type:
The merged data with added location information.
- get_obs(pollutants='all', format='wide', group=None, time_range=None, num_processes=1, include_gps=True, **kwargs)[source]#
Get mobile site observations for each pollutant.
Combines instruments by pollutant and, optionally, merges location data from GPS.
- Parameters:
pollutants (str or list of str, optional) – pollutants to read. If ‘all’, read all pollutants. Default is ‘all’.
format (str, optional) – Format of the data to return. Default is ‘wide’.
group (str | Mapping[str, str] | None) – The research group to read data from. A name applies to every instrument; a mapping of instrument name to group name sets it per instrument. Default None selects each instrument’s group automatically (see
resolve_group()).time_range (TimeRange | TimeRangeTypes, optional) – Time range to read data. Default is None.
num_processes (int, optional) – Number of processes to use for reading data. Default is 1.
include_gps (bool, optional) – Whether to include GPS data in the returned dataframe. Default is True.
- Returns:
A dataframe containing mobile site observations for each pollutant with location data merged (a GeoDataFrame when
include_gps).- Return type:
Notes
Each instrument’s rows are located with the GPS logged on the same clock, which is the GPS of the group they were read from (see
locate()). With no group named, one call can read TRAX methane from lin and ozone from horel, so the result can mix both groups’ GPS columns.
- locate(frames, group=None, time_range=None, lvl='final', num_processes=1)[source]#
Merge GPS locations onto instrument data, joining each row on its own clock.
- Parameters:
frames (Mapping[str, pandas.DataFrame]) – Data per instrument name, indexed by
Time_UTC, as read withgroupandtime_range(e.g. fromread_data()).group (str | Mapping[str, str] | None) – The group selection the frames were read with. It is used to replay each instrument’s read plan, so it must be the same one.
time_range (TimeRange | TimeRangeTypes) – The time range the frames were read with.
lvl (str or None, optional) – The GPS data level to read. Default ‘final’; None reads the highest level available.
num_processes (int or 'max', optional) – Number of processes to use for reading GPS data. Default is 1.
- Returns:
The rows that found a location, indexed by
Time_UTC(EPSG:4326).GPS_Groupnames the group whose GPS located each row, and so the clock itsTime_UTCis on (lin: GPS time; horel: CR1000).- Return type:
- Raises:
ReaderError – If no GPS data could be read for any of the rows.
ValueError – If rows would be located with another group’s GPS.
Notes
A row is joined to the GPS of the group it was read from, since the two share a logger clock: lin instruments and lin’s GPS are stamped by the Pi (joined on
Pi_Time), horel’s by the CR1000 (joined onTime_UTC; the instrument and GPS values share one record). The clocks disagree: on TRX01 the CR1000 ran 1-20 s ahead of GPS time on dates sampled 2019-2026, so joining horel rows to lin’s GPS on the Pi clock placed them that many seconds along the track, and dropped rows with no lin GPS record at their second (uataq#42).One group’s rows are never located with another group’s GPS. A GPS group the caller names explicitly (a mapping entry for
gps) that differs from the group some rows were read from raisesValueError, as does a group that logs no GPS at this site.