Runbook: mirror a store to Source Cooperative¶
Use this to publish (or refresh) a collection's Icechunk store on Source
Cooperative. scripts/mirror_to_source_coop.py streams every object
under the store prefix to
s3://us-west-2.opendata.source.coop/pangeo/tempo-virtual-icechunk/,
building a zip of them on the way that it uploads beside them
(details in Publishing to Source Cooperative).
The stores, from the tracked env files:
| collection | env file | source | destination (under pangeo/tempo-virtual-icechunk/) |
|---|---|---|---|
| NO2 | .env_no2 |
s3://airquality-data-store-develop/tempo/no2/v04/ |
tempo/no2/v04/ and tempo/no2/v04.zip |
| HCHO | .env_hcho |
s3://airquality-data-store-develop/tempo/hcho/v04/ |
tempo/hcho/v04/ and tempo/hcho/v04.zip |
Run it from a terminal on the VEDA JupyterHub. The hub is in us-west-2
with the store bucket, so reads stay in-region; writes go through Source
Coop's proxy at https://data.source.coop, which stores them in its
us-west-2 bucket. Both sides
use short-lived credentials scoped to this job: the source side an SSO
permission set that can only read the store, the destination side the
keys Source Coop issued.
The script is not idempotent and never deletes: every run copies everything again and overwrites what is there. That is fine; it just costs time.
Step 0 — once: a read-only permission set for the store¶
Ask an Identity Center administrator for a permission set in the
account that owns airquality-data-store-develop (say
TempoStoreReader) with this inline policy and a session duration of
12 hours, assigned to whoever runs the mirror. It is exactly the read
half of what the stack grants its own Lambdas
(cdk/stack_constructs/grants.py): listing and reading under
tempo/, which covers both collections and nothing else in the
account.
{
"Version": "2012-10-17",
"Statement": [
{
"Effect": "Allow",
"Action": "s3:GetObject",
"Resource": "arn:aws:s3:::airquality-data-store-develop/tempo/*"
},
{
"Effect": "Allow",
"Action": "s3:ListBucket",
"Resource": "arn:aws:s3:::airquality-data-store-develop",
"Condition": {"StringLike": {"s3:prefix": "tempo/*"}}
}
]
}
Without it, the fallback is to log in with a broader permission set and
assume a role carrying the same policy (aws sts assume-role --role-arn
... --policy file://scoped.json), exporting the three keys it returns.
Those are capped at one hour by role chaining, and the script does not
refresh them, so the whole copy of objects (everything but the final zip
upload) has to finish inside the hour. Prefer the permission set.
Step 1 — set up on the hub¶
In a hub terminal, with the Python (Pangeo) image:
git clone <this repository> && cd tempo-virtual-zarr-pipeline
uv sync # curl -LsSf https://astral.sh/uv/install.sh | sh if missing
aws --version # must be v2 for SSO login
Check free space in the home volume: only the zip lands on disk, so the copy needs room for about the store's size. Store size, after Step 2's login:
aws s3 ls s3://airquality-data-store-develop/tempo/no2/v04/ --recursive --summarize | tail -2
df -h ~
Step 2 — credentials¶
Source. Log in with the scoped permission set. The hub pod already
carries a role of its own in the environment; AWS_PROFILE takes
precedence over it in boto3's credential chain, so the script and the
CLI both use the SSO session.
aws configure sso --profile tempo-reader # SSO start URL, the store's account, TempoStoreReader, us-west-2
export AWS_PROFILE=tempo-reader
aws sts get-caller-identity # ...assumed-role/AWSReservedSSO_TempoStoreReader_.../<you>
aws s3 ls s3://airquality-data-store-develop/tempo/no2/v04/ | head -3 # repo, config.yaml, chunks/
The device-code login prints a URL and a code; open it in your laptop's browser. The session lasts the permission set's duration and the CLI refreshes credentials within it on its own.
Destination. Put the temporary keys Source Coop issued for the
repository in the gitignored .env.local (template: .env.local.sample).
They are Source Coop keys, not AWS ones, valid only at its proxy, which
is where the script sends writes; all three are required:
SOURCE_COOP_ACCESS_KEY_ID=...
SOURCE_COOP_SECRET_ACCESS_KEY=...
SOURCE_COOP_SESSION_TOKEN=...
They are used from the first object on, and last for the zip upload at the end.
Step 3 — run¶
In tmux, so a closed browser tab does not kill the copy (a culled pod
still will; keep the tab open or start early in the day):
tmux new -s mirror
uv run --env-file .env_no2 --env-file .env.local scripts/mirror_to_source_coop.py
uv run --env-file .env_hcho --env-file .env.local scripts/mirror_to_source_coop.py
The env file supplies ICECHUNK_BUCKET, S3_PREFIX and
ICECHUNK_PREFIX; nothing in it needs editing. Each run prints the
object count when the copy starts, uploading when the zip goes up, and
done: at the end.
Step 4 — verify¶
Count objects on both sides; the destination should have the source's
count plus one (the zip). For NO2 (swap no2 for hcho):
aws s3 ls s3://airquality-data-store-develop/tempo/no2/v04/ --recursive --summarize | tail -2
aws s3 ls --no-sign-request \
s3://us-west-2.opendata.source.coop/pangeo/tempo-virtual-icechunk/tempo/no2/v04/ \
--recursive --summarize | tail -2
aws s3 ls --no-sign-request \
s3://us-west-2.opendata.source.coop/pangeo/tempo-virtual-icechunk/tempo/no2/v04.zip
Then open the public copy as a reader would (metadata only; no Earthdata credentials needed for that):
uv run python -c "
import icechunk, zarr
repo = icechunk.Repository.open(icechunk.s3_storage(
bucket='us-west-2.opendata.source.coop',
prefix='pangeo/tempo-virtual-icechunk/tempo/no2/v04',
region='us-west-2', anonymous=True))
print(zarr.open_group(repo.readonly_session('main').store, mode='r').tree())
"
Verify the Icechunk store¶
The public copy should be the same repository as the source, at the same
version. Compare the snapshot main points at on each side and the
length of its history; a run that stopped partway, or a source that
moved on since the copy, shows up as a mismatch:
uv run --env-file .env_no2 python -c "
import icechunk, os
from virtualizarr_processor.manifest import storage_prefix
p = storage_prefix()
sides = {
'source': icechunk.s3_storage(bucket=os.environ['ICECHUNK_BUCKET'], prefix=p,
region='us-west-2', from_env=True),
'public': icechunk.s3_storage(bucket='us-west-2.opendata.source.coop',
prefix=f'pangeo/tempo-virtual-icechunk/{p}',
region='us-west-2', anonymous=True),
}
seen = set()
for name, storage in sides.items():
repo = icechunk.Repository.open(storage)
main = repo.lookup_branch('main')
n = len(list(repo.ancestry(branch='main')))
print(f'{name}: main={main} snapshots={n}')
seen.add((main, n))
print('OK' if len(seen) == 1 else 'MISMATCH')
"
Then read it the way a user would. check_virtual_containers.py, pointed
at the public copy anonymously, confirms the store declares its virtual
chunk containers, covers every manifest URL with them, and reads a chunk
back through them. That chunk read needs Earthdata credentials in the
environment (EARTHDATA_TOKEN or username/password; see the script's
docstring):
uv run --env-file .env_no2 python scripts/check_virtual_containers.py \
--bucket us-west-2.opendata.source.coop \
--prefix pangeo/tempo-virtual-icechunk/tempo/no2/v04 \
--region us-west-2 --anonymous
Never pass --fix here; the public copy is written only by the mirror.
To fix a container, fix the source store and mirror again.
Finally, look at it. compare_to_gibs.py renders a scan of the public
copy beside the GIBS image Worldview shows for the same scan (the imagery
https://tempo.si.edu/data_for_public.html links to), with a per-pixel
difference panel; its docstring says what a good result looks like. It
reads chunks, so it needs the same Earthdata credentials:
uv run scripts/compare_to_gibs.py --collection no2 # writes gibs-no2-<scan>.png
uv run scripts/compare_to_gibs.py --collection hcho --time 2026-10-01T18:30 # the scan Worldview shows at 18:30
Verify the zip¶
Check the zip from the local copy at stores/<prefix>.zip (it is
what was uploaded). No script checks the zip itself, but both store
checks open a local directory when ICECHUNK_BUCKET is empty and
ICECHUNK_LOCAL_PATH is set. Clear the bucket with env inside
uv run, so the env file cannot set it again:
unzip -tq stores/tempo/no2/v04.zip # every entry's CRC
unzip -l stores/tempo/no2/v04.zip | tail -1 # file count = source object count
unzip -q stores/tempo/no2/v04.zip -d /tmp/v04
uv run --env-file .env_no2 env ICECHUNK_BUCKET= ICECHUNK_LOCAL_PATH=/tmp/v04 \
python scripts/check_virtual_containers.py
uv run --env-file .env_no2 env ICECHUNK_BUCKET= ICECHUNK_LOCAL_PATH=/tmp/v04 \
python scripts/verify_store.py --samples 8
rm -rf /tmp/v04
check_virtual_containers.py confirms the store declares its virtual
chunk containers, covers every manifest URL with them, and reads a chunk
back; verify_store.py compares sampled time steps against CMR and the
source granules. Both read granule bytes, so export EARTHDATA_TOKEN (or
EARTHDATA_USERNAME and EARTHDATA_PASSWORD) first. Without it they fall
back to the hub's own AWS role, which asdc-prod-protected refuses with
AccessDenied. Neither compares the zip's file list
against the source prefix, so the count from unzip -l is the only check
of that.
Step 5 — clean up¶
The home volume persists between sessions, so leave nothing behind:
rm -rf stores/ .env.local
aws sso logout # drops the cached SSO token
The zip holds nothing that is not already in both buckets.