Downloading Archived Meteorology#

arl-met includes archive classes for downloading ARL meteorology files from the NOAA ARL public archives.

Install the archive dependencies first:

pip install "arlmet[archives]"

Choose an archive#

Each class knows the filename and archive layout for one product, and is registered under a short name.

Available archives#

Name

Class

Product

Typical coverage

"hrrr"

arlmet.archives.HRRRArchive

HRRR 3 km

CONUS, 6-hour files, June 2019–present

"hrrr.v1"

arlmet.archives.HRRRv1Archive

HRRR 3 km, version 1

CONUS, 6-hour files, June 2015–2019

"nam12"

arlmet.archives.NAMArchive

NAM 12 km

North America, daily files

"nams"

arlmet.archives.NAMSArchive

NAM hybrid sigma-pressure

CONUS, Alaska, or Hawaii (domain=), daily files

"gdas1"

arlmet.archives.GDASArchive

GDAS 1 degree

Global, weekly files

"gdas0p5"

arlmet.archives.GDAS0p5Archive

GDAS 0.5 degree

Global, daily files, 2007–2019

"gfs0p25"

arlmet.archives.GFSArchive

GFS 0.25 degree

Global, daily files

"narr"

arlmet.archives.NARRArchive

North American Regional Reanalysis 32 km

North America, monthly files, 1979–2019

"reanalysis"

arlmet.archives.ReanalysisArchive

NCEP/NCAR Reanalysis 2.5 degree

Global, monthly files

Choose an archive by name#

arlmet.archives.ARCHIVES maps each name to its class, and arlmet.archives.get_archive() builds an archive from its name plus any options, which is convenient when the product comes from a config file:

from arlmet.archives import ARCHIVES, get_archive

sorted(ARCHIVES)  # ['gdas0p5', 'gdas1', 'gfs0p25', 'hrrr', ...]
archive = get_archive("nams", domain="ak")
files = archive.fetch("2024-07-18", "2024-07-19", dest_dir="./met/")

An unknown name raises ValueError listing the available ones. A subclass of arlmet.archives.Archive that sets name is registered automatically when it is defined, so your own archives work with get_archive() too.

Download files for a time range#

from arlmet.archives import HRRRArchive

archive = HRRRArchive()
files = archive.fetch(
    "2024-07-18 00:00",
    "2024-07-19 00:00",
    dest_dir="./met",
)

fetch() returns the local paths in chronological order. Duplicate archive keys are removed automatically when the requested time range spans multiple hours within the same ARL file.

Crop on download#

For large global products, pass bbox= so the downloaded file is cropped before it is cached locally.

from arlmet.archives import GFSArchive

archive = GFSArchive()
files = archive.fetch(
    "2024-07-18 00:00",
    "2024-07-19 00:00",
    dest_dir="./met",
    bbox=(-114.0, 39.0, -110.0, 42.0),
)

This uses arlmet.extract_subset() internally after the raw file is downloaded.

Pass levels= to keep only some vertical levels, counted from 0 at the surface. It works with or without bbox.

files = archive.fetch(
    "2024-07-18 00:00",
    "2024-07-19 00:00",
    dest_dir="./met",
    bbox=(-114.0, 39.0, -110.0, 42.0),
    levels=range(20),   # the lowest 20 levels
)

The crop is part of each cached file’s name (for example ...crop_-114.00_39.00_-110.00_42.00.levels_0-19), so files cropped differently are cached separately. Bbox values are written with two decimals unless they have more, which are kept in full (...crop_-111.925_...).

The full, uncropped file is downloaded into dest_dir (not the system temp directory) and deleted once the crop is written, so dest_dir needs room for one full file at a time (about 3 GB for HRRR).

Choose a mirror#

Each archive is served from three mirrors. The default, "s3", is usually the fastest choice.

files = archive.fetch(
    "2024-07-18",
    "2024-07-19",
    dest_dir="./met",
    mirror="ftp",
)

The mirrors are:

  • "s3": NOAA public S3 bucket via s3fs

  • "ftp": NOAA ARL FTP archive

  • "http": READY web archive

Caching and overwrite behavior#

Downloaded files are reused if a matching local file already exists. Pass overwrite=True to force a fresh download.

Files are written to a hidden .<name>.<random>.partial file first and renamed into place only once complete, so an interrupted fetch never leaves a truncated file that a later call would reuse. A .partial file left behind by a killed process is never used and can be deleted.

Requesting a time range that begins before an archive’s start_date (the start of its archive) raises ValueError.

files = archive.fetch(
    "2024-07-18",
    "2024-07-19",
    dest_dir="./met",
    overwrite=True,
)