arlmet.concat_by_time#

arlmet.concat_by_time(directory, dest_dir, *, freq='1D', pattern='*', start=None, end=None, template='{time:%Y%m%d}_arl', sort=True)[source]#

Group every ARL file in a directory by valid time and concatenate each group.

Each input is assigned to a time bin from its valid times — read from the file’s index records, not parsed from its name — floored to freq. All files in a bin are concatenated into one output file. This is the batch form of concat(): e.g. turning a directory of 6-hourly HRRR files into one file per day. Files are never split, so every input must fall entirely within one bin.

Parameters:
  • directory (path-like) – Directory to scan for input ARL files (non-recursive).

  • dest_dir (path-like) – Directory to write the concatenated files into. Created if missing. Should differ from directory.

  • freq (str, default "1D") – Fixed-frequency pandas offset alias giving the size of each output chunk: "1D" = one file per day, "6h" = one per six hours, etc. Each input must fit entirely within one bin: a file whose first and last valid times floor to different bins raises ValueError.

  • pattern (str, default "*") – Glob (relative to directory) selecting input files. Scope it to ARL files; every match must be a readable ARL file.

  • start (str or pandas.Timestamp, optional) – Inclusive bounds on each file’s first valid time; files outside them are skipped. Either may be omitted to leave that side open.

  • end (str or pandas.Timestamp, optional) – Inclusive bounds on each file’s first valid time; files outside them are skipped. Either may be omitted to leave that side open.

  • template (str, default "{time:%Y%m%d}_arl") – str.format template for output filenames, given the bin start time as time (a pandas.Timestamp), e.g. "{time:%Y%m%d}_hrrr". It must encode enough resolution to keep bins distinct at freq; two bins that format to the same filename raise ValueError.

  • sort (bool, default True) – Passed through to concat() for each group.

Returns:

The written output paths, one per non-empty time bin, in time order.

Return type:

list[pathlib.Path]

Raises:

ValueError – If pattern matches no files, a matched file cannot be read as ARL, a file’s valid times straddle a freq bin boundary, or template formats two bins to the same filename. concat()’s grid/axis and duplicate-time checks also apply within each group. All checks run before any output is written.

Examples

Turn a directory of 6-hourly HRRR files into one file per day:

>>> import arlmet
>>> arlmet.concat_by_time(
...     "hrrr/",
...     "daily/",
...     freq="1D",
...     pattern="*_hrrr",
...     template="{time:%Y%m%d}_hrrr",
... )