Skip to content

Development

./scripts/setup.sh          # set up the environment
uv run pytest               # tests
uv run ruff check . && uv run ruff format --check .
uv run mypy
uv run --env-file .env_hcho --env-file .env.local cdk synth   # review infrastructure before deploying
uv run --env-file .env_hcho --env-file .env.local cdk deploy

The Processor class in virtualizarr_processor/processor.py is the sole (non-polymorphic) implementation. The template's synthetic reference implementation lives on as tests/stub_processor.py and still exercises the generic fork/merge mechanics.

Exploration

exploration/ holds standalone PEP 723 scripts used to characterize the source data and to build test stores. Run them with uv run exploration/<script>.py; each declares its own dependencies. All take --collection {hcho,no2} (default hcho) plus --concept-id for any other collection, and need Earthdata Login credentials in ~/.netrc.

  • tempo_dataset_info.py — CMR/UMM-C collection report: extents, granule count, distribution info, most recent granule.
  • inspect_granule_metadata.py — HDF5 structure dump of granules: chunk layouts, codecs, fill values, attributes, and a cross-granule comparison of what varies.
  • combine_twenty_spread_virtual.py — virtualizes N granules spread across the collection's temporal extent, combines them in an in-memory Icechunk store, and reads data back to prove the path end to end.
  • build_titiler_test_store.py — small local store (12 recent granules) for titiler-multidim smoke tests.
  • build_s3_test_store.py — realistic S3-hosted store (100 recent granules, credential-less virtual chunk container); run on in-region compute.

The production inventory builder (scripts/build_backfill_inventory.py, described in Backfill inventory) lives in scripts/ with the other production tooling; it follows the same PEP 723 + --collection conventions.

titiler-multidim smoke test takeaways (2026-08-06)

A 12-granule store per collection was served through titiler-multidim's feat/http-virtual-chunk-auth branch; all checked endpoints (/variables, /info, /tiles, /point) passed for both collections. What carries forward:

  • Flat-at-root layout is required: titiler-xarray does not walk nested groups or resolve group-inherited coordinates. The production stores use it.
  • Clients must always select a time step (sel=time=...&sel_method=nearest). A multi-time variable reaching the renderer fails with "Source data must be 1 band".
  • Tile latency is dominated by per-request virtual-chunk fetches and scales with the area a tile covers: z2 tiles took 3–7 s, z4/z6 tiles 1.5–2.7 s over HTTPS. A dataset/session cache is advisable in production, and Lambda's 1024 file-descriptor limit is a real constraint under tile bursts.
  • Per-scan coverage is inherently partial (single east-west scans, daylight-only retrieval, occasional short rapid-scan slices). Portal time sliders and "latest available" defaults need to account for it.
  • Map clients should set noWrap/maxBounds; tiles crossing ±180 hit an antimeridian error upstream in rio-tiler.
  • CMR publication behavior shaped the forward-processing design: publication order routinely diverges from scan order (the ~43% figure quoted under Forward processing came from an August 2026 14-day window, with the historical archive back-filling at ~1,000 granules/week), republication is rare (~0.3%) and short-window, and median production lag is ~3 h.