Skip to content

Add Pangeo CMIP6 ArchiveIndex accessor (closes #284) - #287

Open
ahmedmohiduet wants to merge 2 commits into
ACCESS-Community-Hub:developfrom
ahmedmohiduet:pangeo-cmip6-accessor
Open

ahmedmohiduet wants to merge 2 commits into
ACCESS-Community-Hub:developfrom
ahmedmohiduet:pangeo-cmip6-accessor

Conversation

@ahmedmohiduet

Copy link
Copy Markdown

Closes #284 (part of #276).

What this adds

A new third-party package, packages/third_party/pangeo_site_archive (import name site_archive_pangeo), containing PangeoCMIP6, an ArchiveIndex for the Pangeo CMIP6 cloud archive (ARCO Zarr on Google Cloud, anonymous access).

import pyearthtools.data
from site_archive_pangeo import PangeoCMIP6

index = PangeoCMIP6(["tas", "pr"], source_id="CanESM5", experiment_id="historical")
ds = index.retrieve(pyearthtools.data.Petdt("2000-01"))

Also changed: requirements_cicd.txt (editable install of the new package), .github/workflows/python-app.yml (adds it to --cov), and a short section in docs/data.md.

Design choices

  • Plain pandas + xarray instead of intake-esm. The catalogue is a single CSV; pandas.read_csv with usecols is enough, keeps the dependency list to pandas, xarray, zarr, gcsfs, and makes the tests offline (they use a small fixture CSV).
  • Catalogue is a constructor argument (catalog=), with no ROOT_DIRECTORIES. register_archive does a setattr on pyearthtools.data.archive for each registered package, and both the NCI and AWS packages register their own ROOT_DIRECTORIES under the same name, so importing both silently clobbers one. Passing the catalogue explicitly avoids that for this package.
  • filesystem() returns one Zarr store URL per variable and ignores the time. The base retrieve already does data.sel(time=...) after load(), so load() can return the lazy store. Nothing is downloaded until the data is computed.
  • Newest version is used when a variable has several; several grids raise an error asking for grid_label, rather than silently choosing one.
  • Variables are merged with xarray.merge(..., join="exact", compat="minimal") after selecting [[variable]] from each store. join="exact" fails loudly if the grids differ. compat="no_conflicts" raised a MergeError on the scalar height coordinate (2 m for tas, 10 m for uas), and override would silently keep the wrong height.
  • __init__ does no network access (the catalogue is read lazily and once), because register_archive attaches .sample(), which constructs the class from sample_kwargs.
  • exists() is overridden, because the base implementation turns the gs:// URLs into Path objects.
  • data_interval is left unset in this first PR. When it is set, ArchiveIndex.search calls .values() on a list and the blanket except Exception: pass hides the failure. series() therefore needs an explicit interval.

Known limitation

When variables disagree on a scalar coordinate (e.g. height for tas + uas), that coordinate is dropped from the merged dataset. It is kept when they agree (tas alone, or tas + pr). This is covered by a regression test.

Testing

  • 15 offline unit tests, 100% statement and branch coverage of the new package (pytest --cov=site_archive_pangeo --cov-branch). No test touches the network; xarray.open_zarr is monkeypatched.
  • Checked manually against the real archive (CanESM5, historical, r1i1p1f1, Amon) for ["tas"], ["tas","pr"], ["tas","uas"] and ["tas","uas","pr"]. All four load and return time=1, latitude=64, longitude=128 for Petdt("2000-01").
  • pre-commit (end-of-file-fixer, trailing-whitespace, black, ruff) passes.

Questions for reviewers

  1. Are the names OK? I used folder pangeo_site_archive, import site_archive_pangeo, distribution PyEarthTools-archive-Pangeo, and archive name pangeo.
  2. Is a licence header expected on source files in third_party?
  3. Would you like a noci-marked live test, or is the offline suite enough?

I'm happy to make changes or add more tests wherever you think they're needed. Please let me know which cases or behaviours you'd like covered.

Checklist

  • Existing issue (Create ArchiveIndex derived accessor class #284), and I stated I was working on it
  • Docstrings complete, Google style
  • All new code covered by unit tests
  • Documentation updated (docs/data.md)
  • Generative AI was used (see below)

Generative AI declaration

  • Tool and version: Claude Sonnet 5.5 (Anthropic), used through Claude Code / claude.ai.
  • Scope: I used it as an assistant while I worked through the issue. Specifically, it helped (a) explain how the existing accessors (ArchiveIndex, AdvancedTimeIndex, register_archive, the NCI and AWS packages) work, (b) draft pangeo_cmip6.py, the unit tests, the docs/data.md section and this description, and (c) debug the height merge conflict.
  • Human in the loop: I ran the tests, linting, pre-commit and the live checks above myself, I reviewed all of the code, and I can explain each change in review.

Adds packages/third_party/pangeo_site_archive with PangeoCMIP6, an ArchiveIndex that reads CMIP6 variables from the Pangeo ARCO cloud archive, plus offline unit tests, CI install/coverage entries and a docs section. Closes ACCESS-Community-Hub#284.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Unqqtp91DWVC8ro9Y724iN
@ahmedmohiduet

ahmedmohiduet commented Oct 6, 2026 •

Copy link
Copy Markdown
Author

Hi @stevehadd,

A note about this PR in relation to ICCS Hacktoberfest 2026, which I'm taking part in. The event asks participants not to use AI tools for their pull requests. As declared in the description above, I used a generative AI assistant (Claude Sonnet 5.5) substantially to prepare this PR. I had not read the event's rule when I opened it, so I wrote to the organisers.

Marion Weinzierl at ICCS replied that if you are happy to accept this contribution, then it is OK for this one case.

So my question is: are you happy to accept this PR as a contribution for the event? I understand completely if you would rather not. Either way it remains a normal PyEarthTools contribution, and I'm glad to make any changes you'd like, including on the design questions in the description.

Thank you, and sorry for the extra work.

@stevehadd stevehadd left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks, @ahmedmohiduet that PR requests looks, code is clear, good tests. I have tested the accessor and loaded fine, and also run the unit tests, everything passed.
I only added one suggestion about the catalog, which you can ignore. I'll now pass you over to @tennlee for review and approval before merging.

)
self.record_initialisation()

def _load_catalog(self) -> pandas.DataFrame:

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

my only suggestion was about helping people find what is available in this dataset. I wondered if we could have a static function on the class which returns the catalog as a pandas dataframe, so the user can programmtically discover what data is available?

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

got it. Will add static method PangeoCMIP6.catalog() returning catalogue as pandas.DataFrame, so users can browse source_id, experiment_id and variable_id values before constructing index.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

pushed commit adding PangeoCMIP6.catalog().

@stevehadd
stevehadd requested a review from tennlee October 9, 2026 14:41
Returns the Pangeo CMIP6 catalogue as a pandas.DataFrame so users can discover available data. Includes a unit test.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Create ArchiveIndex derived accessor class

2 participants