Skip to content

Read ShapePipe v1 and v2 products side by side - #343

Open
cailmdaley wants to merge 41 commits into
developfrom
feat/v2-catalogue-readers
Open

cailmdaley wants to merge 41 commits into
developfrom
feat/v2-catalogue-readers

Conversation

@cailmdaley

@cailmdaley cailmdaley commented Sep 10, 2026 •

Copy link
Copy Markdown
Collaborator

Closes #340. Closes #342. Closes #361.

Where to look

82 files change, but most are mechanical. The design is in four files; read them first:

  1. src/sp_validation/io.py: the one catalogue reader (file format from contents, lazy or column-restricted).
  2. src/sp_validation/grammar.py: v1 → v2 column names and units, plus column_map.
  3. config/columns/shapepipe_v2.yaml: every column sp_validation reads, with its meaning.
  4. src/sp_validation/galaxy.py: the mask cut and the v2 selection choices (NGMIX_N_EPOCH > 0, mask bits).

Skim:

  • Call sites: catalog_builders.py, rho_tau.py, cosmo_val/*, calibration.py, masks.py and scripts/ now read through the above, and most of their diffs are that routing.
  • Configs: config/calibration/mask_v1.X.* renames mask columns.
  • Tests: test_grammar, test_io, test_grammar_calibration and test_release_regression are the behavioural evidence.

Skip:

  • Deleted files (patch-era scripts, survey.py, compute_theory_cov.py): about 1,700 lines, nothing replaces them.
  • catalog.py's removed readers, which moved to io.py.

What changes

sp_validation now reads both generations of ShapePipe products through one code path written against the v2 grammar:

  • v2: the Snakemake campaign products (final_cat_<campaign>.hdf5, full_starcat_<campaign>.hdf5, per-bit mask columns, no patches).
  • v1: every existing catalogue, v1.3–v1.6, including the fiducial v1.4.6.3.

Only two things know about v1: the reader and the adapter. Everything else — builders, workflow, configs, scripts — speaks the v2 campaign model.

One reader, any container. sp_validation.io is the only place that opens catalogue files, and every call site goes through it, cosmo_val, ρ/τ and the calibration scripts included.

  • open_catalogue opens a file as one lazy table without reading data up front. FITS is memory-mapped; HDF5 columns are read when indexed.
  • read_catalogue reads a restricted set of columns into memory and, given a key_column, records which chunk each row came from (EXPID on the star catalogue).
  • Format from contents: the reader tells FITS from HDF5 by the file's contents, not its extension. It handles a FITS table HDU, a single HDF5 dataset, the comprehensive data+data_ext, and ShapePipe's campaign files of one dataset per tile or exposure, at any group nesting (checked against n_tiles/n_exposures).
  • Memory: the campaign merge streams, so peak memory is the merged array plus one tile.

Column names are configuration. The v2 names are the canonical schema, and config/columns/shapepipe_v2.yaml lists every column sp_validation reads by fixed name, each with its meaning.

  • column_map: any catalogue config (cat_config blocks, calibration params) may carry column_map: {canonical name: name in the file}, with one-* patterns allowed ({"NGMIX_*": "MYFIT_*"}), so a catalogue in another naming convention needs a few lines of YAML.
  • Role keys: the cat_config *_col keys still pick which column plays a role (e.g. w_iv vs w_des).
  • Guard test: test_column_schema checks code and schema against each other in both directions.
  • Docs: the docs build renders the schema as the "Catalogue columns" page.

Adapter (v1). grammar.adapt presents a v1 table in v2 names and units, and reads the generation from the file's columns.

  • It renames E1_PSF_HSM → HSM_G1_PSF, NGMIX_ELL_* → NGMIX_G1/G2_*, and so on.
  • It converts SIGMA_*_HSM = σ → HSM_T_* = 2σ².
  • It corrects one v1 defect: the no-shear reconvolved-PSF size is taken from the 1P column. Remove vestigial FHP/MK NOSHEAR←1P Tpsf substitution (guaranteed no-op) #267 removed this substitution as a no-op, which holds only for v2 products.
  • It applies a config column_map ahead of its v1 rules; a map entry overrides only the rule for its own name.
  • It works lazily over numpy, FITS_rec and h5py. A row selection on h5py reads each dataset once.
  • The map is in docs/ngmix_psf_column_migration.md.

Patches retired (#340). No code branches on, loops over, or names P1–P8. JointCat merges a list of campaign files and adds a campaign column. Patch-only scripts are deleted: prepare_patch_for_spval.sh, plot_rho_stats_patches.py, survey_stats_all.sh, star_match_stats.py, check_tile_IDs_SP_LF.py, combine_results.py and compute_area.py (the last two fed a create_joint_shape_cat.py that no longer exists), and cosmo_val/compute_theory_cov.py. That last one had hard-coded home paths and version keys that are not in cat_config; rho_tau builds the same CovTauTh from cat_config, and the workflow drives that path. merge_psf_cat.py and stats_tile_id_gal_counts.py take campaigns. Released v1 merged catalogues are flat (data / data_ext) and keep their int8 patch column, which passes through untouched.

Masks (#342). v1 and v2 share one mask vocabulary: one boolean column per UNIONS mask bit, named MASK_{bit}_{Label} (MASK_1_Faint_star_halos, MASK_4_Stars, MASK_1024_Maximask, …). These are the bit/label pairs ApplyHspMasks has always written into v1 data_ext, now prefixed.

  • grammar.MASK_LABELS is the single table, and ApplyHspMasks writes through it.
  • The adapter presents v1 {bit}_{Label} and the pre-release MASK_n{bit} of the v2 integration campaigns under the same names.
  • ShapePipe #886 writes the same names (eleven columns; bit 512 has no map).
  • CalibrateCat.read_cat returns one table (data joined with data_ext). Every mask config is a single dat: list, and a config with a dat_ext: section raises an error.
  • galaxy.mask_cut ORs the six reason bits {1, 2, 4, 8, 64, 1024} and drops NaN rows with a warning.
  • v1 configs keep their IMAFLAGS_ISO == 0 cut, whose bits mean something different from the mask columns.

ρ/τ reads catalogues through the adapter and passes tables to shear_psf_leakage (#44 there). cat_config column names were corrected to match their files, and a candide test now checks every entry. Dropping a scalar mask= argument fixes the DES and jackknife ρ/τ paths, which built empty catalogues.

Breaking changes for existing scripts and notebooks:

  • catalog.read_campaign_catalogue, read_star_catalogue, iter_campaign_tiles and read_shape_catalog are replaced by io.open_catalogue / io.read_catalogue; metacal loses its prefix argument (use column_map).
  • read_cat() returns one table, not (dat, dat_ext).
  • get_masks_from_config(config, dat, ...) takes that table.
  • Mask configs name MASK_{bit}_{Label}.
  • extract_info no longer runs on v1 per-patch final_cat files, and v1 per-patch FITS can no longer be re-merged; the merged v1 releases are what the readers open.
  • The nine SExtractor-only columns DR6 catalogue mode lacks (shapepipe #924) are no longer passed through.

Verification.

  • Release regression: a candide test recalibrates rows 200M–201M of the v1.4 comprehensive catalogue and matches released v1.4.6.3 exactly: the same 131,526 objects in the same order, and 17 per-object columns bit-identical. On a separate 5M-row / 240-tile subset, the selection is identical (684,385 objects) and so is every per-object column. The weights and leakage correction are not bit-identical to the release, because the release fed its weight and leakage binning the uncorrected PSF size while this branch uses the corrected size throughout.
  • ρ/τ on SP_v1.4.6.3 through the adapter is identical to the earlier run on a hand-translated PSF file (≤1e-13σ with the same patch seed).
  • v2: exercised on the smk-g7 campaign (64 tiles; see comment).
  • Masks: a table written with each of the three spellings gives identical columns and identical cuts.
  • Generality: a catalogue under made-up column names reads to the same table from FITS, flat HDF5, data+data_ext and chunked HDF5. A metacal catalogue renamed NGMIX_*→MYFIT_* and given a column_map calibrates to bit-identical R and g.
  • Memory: reading the v1.4.6.3 shear and star catalogues through the leakage path peaks at 15.4 GiB RSS, the same as the upstream reader, with c1/c2 bit-identical.
  • Unit tests: the fast suite passes (448), and the slow suite (release regression, additive bias) and the candide config tests pass at b1fbf84.

— Claude (Opus) on behalf of Cail

🤖 Generated with Claude Code

@github-actions

github-actions Bot commented Sep 10, 2026 •

Copy link
Copy Markdown
Contributor

✅ ruff is clean — nothing to fix here.

@cailmdaley

Copy link
Copy Markdown
Collaborator Author

Exercised against the first real v2 products (smk-g7, 64 tiles): galaxy reader 1,851,100 rows / 77 columns in 11 s at 1.2 GB RSS; star reader 53,264 stars over 127 exposures; default mask cut keeps 72.98%; full selection 1,105,851 rows with mean e1 = −0.00020 ± 0.00019, e2 = −0.00015 ± 0.00019. Three follow-up commits from that run: an explicit NGMIX_N_EPOCH > 0 guard (CosmoStat/shapepipe#889 — 1.03% of objects are never fit but carry MCAL_FLAGS = 0 with −10 sentinels; the old != −10 PSF comparison happened to catch them), a NaN-safe mask cut, and n_exposures validation + an EXPID column on the star reader. One decision left open on purpose: config/calibration/mask_v2.0.yaml has no NGMIX_MCAL_FLAGS cut — adding one changes row counts, so it is Cail's call.

— Claude (Fable) on behalf of Cail

cailmdaley and others added 22 commits September 28, 2026 06:17
… key

NGMIX_MCAL_FLAGS carries bit 30 (shapepipe#854's absent-measurement flag)
alongside the native fitter bits. Format I truncates it to int16 and
silently zeroes bit 30 on write, in both the FITS and the HDF5 branch of
write_shape_catalog (the HDF5 path reuses the FITS column's array).

The per-type NGMIX_FLAGS_{1P,1M,2P,2M,NOSHEAR} columns had the same
problem from the other direction: their format entry was keyed as
"FLAGS_{suffix}", which never matches the real column name
"{prefix}_FLAGS_{suffix}", so the writer's float64 default masked the
dead key rather than narrowing anything. Both parameter files now key
and format all six metacal bitmasks the same way, at K, so they round-
trip exactly and survive JointCat's optional memory-reduction pass
(which only downcasts int32/float64, not int64).

NGMIX_MCAL_TYPES_FAIL stays at I: it is a count in [0, 5], not a bitmask.

Adds a round-trip test parametrized over both parameter files, FITS and
HDF5, and reduce_mem on/off, checking 0, bit 30 alone, and bit 30 with a
native bit together.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
ShapePipe sets bit 2**30 in the metacal flags for a missing measurement,
which needs 32 bits. The parameter files now write NGMIX_MCAL_FLAGS and
the five NGMIX_FLAGS_* columns as FITS J (int32).

JointCat.dtype_out with reduce_mem narrowed every int32 column outside a
keep-list to int8, silently wrapping any value above 127. It now only
narrows float64 to float32 (RA/Dec excepted) and leaves integer columns
alone.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Replace read_hdf5_file's hardcoded patches/<name>/<tile-ID> lookup with
find_dataset_group(), which descends from the file root through single
container groups until it reaches the per-unit datasets. This reads the
legacy patches/<campaign>/ layout that ShapePipe still writes as a
compatibility shim, a future flat tiles/ layout, and the exposures/<exp>
layout of full_starcat_<campaign>.hdf5 with the same code.

read_star_catalogue() keeps the FITS path for files ending in .fits.
Requested columns missing from the data now raise a clear KeyError.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
ShapePipe v2 drops IMAFLAGS_ISO for eleven boolean MASK_n* columns.
galaxy.mask_cut() ORs a configurable list of them (default MASK_n4,
MASK_n1, MASK_n2, MASK_n8, MASK_n1024 — stars, star halos, manual galaxy
mask, MaxiMask) and returns the keep mask; a catalogue missing any of the
requested columns raises a KeyError naming them.

classification_galaxy_base takes mask_columns; extract_info.py passes the
params.py mask_columns list and uses it for the star-sample cut too, and
now reads both catalogues through the new readers. Column lists in
params.py and masking.py's SPATIAL_CUTS updated to the MASK_n* names.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
ShapePipe v2 processes a campaign (a tile list); P1-P7 no longer exist.

- catalog_builders.JointCat: get_patches()/get_n_obj() and the per-patch
  FITS merge are replaced by a merge over a list of campaign hdf5 files
  (-i final_cat_A.hdf5+final_cat_B.hdf5), read through
  read_campaign_catalogue. The 'patch' int8 column becomes a 'campaign'
  string column; the hdf5 root attr becomes 'campaigns'.
- survey.get_footprint() deleted: it was a lookup table of P1-P7 (plus W3)
  RA/Dec boundaries, meaningless for a campaign. Its only caller,
  catalog.check_matching, used it as an optional pre-filter via a 'name'
  argument that every caller passed as None; the argument goes too, along
  with the test_survey test that exercised P5.
- merge_psf_cat.py, combine_results.py, stats_tile_id_gal_counts.py,
  compute_area.py: patch vocabulary and v1/v1.5/v1.6 P-name shortcuts
  generalised to an explicit list of campaigns.
- params.py: 'name = "P7"' becomes 'campaign = None'.

Deleted (only ever meaningful for the P1-P7 era):
- scripts/prepare_patch_for_spval.sh: symlinks a v1 per-patch run tree
  (~/psfex/${patch}/output/run_sp_Ms/.../full_starcat-0000000.fits,
  tiles_${patch}.txt) into a working dir; neither the layout nor the file
  names exist in v2.
- scripts/plot_rho_stats_patches.py: globs P* directories and reads
  P*/output/run_sp_Pl/mccd_plots_runner/output/rho_stats_id.fits, one
  curve per patch. No campaign analogue.
- scripts/survey_stats_all.sh: hardcoded 'for patch in P1 ... P7' over v1
  run-directory bookkeeping.
- scripts/star_match_stats.py: sums stats_file.txt over the seven patches.
- scripts/check_tile_IDs_SP_LF.py: compares per-patch ShapePipe tile IDs
  against CFIS3500_THELI_P<n>.list Lensfit files; both sides P-named.

The word 'patch' survives only in catalog.py's comment naming the legacy
hdf5 group, and in the treecorr jackknife sense (npatch/patch_number).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
Seventeen unit tests on tiny synthetic hdf5 fixtures built in a temp dir:
both campaign layouts (legacy patches/<campaign>/<tile-ID> and flat
tiles/<tile-ID>) read identically, param-list restriction, missing-column
and ambiguous-layout errors; the star reader on exposures/<exp> hdf5 and
on FITS; galaxy.mask_cut defaults, explicit list, empty list, missing
column, and a v1 IMAFLAGS_ISO-only catalogue; JointCat.merge_catalogues
across two campaigns of different layouts, plus its error paths.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
- concatenate_datasets preallocates the output and fills it column by
  column, so peak memory is the packed output plus one tile instead of the
  full-width uncut catalogue; the verbose estimate uses the real itemsize.
- validate the requested columns against every tile dataset, not only the
  first, and name the offending dataset in the error.
- read_campaign_catalogue checks the root n_tiles attribute and refuses a
  truncated file; campaign_shape reports row count and dtype from metadata.
- JointCat.merge_catalogues preallocates the merged array from that first
  pass (no more accumulate-then-concatenate, which doubled peak memory),
  promotes each column's dtype across all campaigns so a wider string or
  integer column in a later file is no longer silently truncated, and
  refuses multi-dimensional columns explicitly.
- reduce_mem reduces int32 to int16 (int8 wrapped N_EPOCH/CCD_NB values)
  and every assignment is range-checked, raising instead of wrapping.
- check_matching drops the identity index over d1 that the retired
  footprint prefilter left behind, and extract_info applies mask_cut to
  the matched subset instead of the whole catalogue.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
Add config/calibration/mask_v2.0.yaml: the v1.X.11 cut set with
IMAFLAGS_ISO and the v1 post-processing masks replaced by the boolean
MASK_n<bit> columns (True = masked, hence kind: equal, value: False), the
coverage bits and MASK_n2048 listed but commented out. The v1.X configs are
left untouched: each describes a legacy catalogue that really has
IMAFLAGS_ISO, and rewriting them would break reproducing published versions.

plots.sky_plots no longer hardcodes the v1 label set; it combines whichever
of IMAFLAGS_ISO / MASK_n* / npoint3 / 1024_Maximask the config declared.
The comprehensive-to-minimal demo asks for the v2 mask labels.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
- params.py interpolated the removed 'name' into two paths (NameError on
  import); use 'campaign', give it a real default, and point star_cat_path
  at full_starcat_<campaign>.hdf5 (hdu_star_cat now documented as legacy
  FITS only). params_im_sim.py renames 'name' to 'campaign' likewise.
- combine_results.get_area matches both the campaign and the legacy patch
  wording, and raises on a missing file or unmatched pattern instead of
  returning None / a 1 deg^2 placeholder that silently rescales densities.
- merge_psf_cat writes the campaign *name* as a string column (FITS 'A<n>'),
  matching JointCat; the old 1-based ordinal depended on -p argument order.
- compute_m_bias_image_sims counts tiles via the n_tiles attribute or
  find_dataset_group, not a hardcoded legacy group inside a bare except.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
Tests: assert the exact concatenation instead of sorted values, add a
fixture whose keys are inserted out of order, and cover the truncated
n_tiles file, a column missing from a later tile, dtype promotion across
campaigns in both argument orders, reduce_mem overflow, and a
multi-dimensional column.

Docs: CLAUDE.md, scripts/calibration/README.md, homogenize_cat_extended.py
and the catalog_builders docstrings now say campaign.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
Three defects in the campaign reading path, found in review:

- concatenate_datasets and campaign_shape both took the output dtype from
  the first dataset alone, so a campaign whose tiles differ (S7 next to
  S12 tile IDs, i2 next to i4, f4 next to f8 after a partial
  reprocessing) had the later tiles silently truncated, downcast or
  wrapped. np.concatenate, which this code replaced, promoted. Both now
  build the dtype with group_dtype(), promoting every column across every
  dataset; catalog_builders._promote becomes an alias of the shared
  catalog.promote_dtypes rather than a second copy of it.

- merge_catalogues still held a whole campaign in memory next to the
  preallocated output, the very thing its comment claimed the rewrite
  avoided -- and with one campaign per merge in v2, that is the normal
  case, ~2x the merged catalogue at DR6 scale. It now fills the output
  tile by tile through the new catalog.iter_campaign_tiles(), so peak
  memory is the output plus a single tile.

- write_hdf5_file wrote the merged array twice, create_dataset(data=...)
  followed by an immediate dset[:] = dat_all.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
mask_v2.0.yaml dropped v1.X's '64_r' r-band imaging cut without a
replacement, so a v2 calibration run admitted objects outside the r-band
footprint that v1 rejected -- a silent change of effective area, n(z) and
galaxy-density normalisation. Enable MASK_n64, v1's '64_r' equivalent.

v1's other coverage cut, npoint3 >= 3, came from an external
post-processing catalogue and has no v2 counterpart; say so in the config
rather than leaving its absence unexplained.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
extract_info fancy-indexed the full-width catalogue, dd[ind_star], only
to read ~5 boolean mask columns from it -- a copy of every column for
every matched star (~GB at DR6 scale). Mask first, index the resulting
bool array.

plot_leakage still labelled its curves "all", "P1" ... "P7": a P-named
survivor of the #340 retirement, and a fixed length that silently
mismatched the number of input files. Labels and colours now follow the
input files.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
The v2 branch inferred from v1's '64_r' naming that MASK_n64 flags r-band
imaging coverage. It does not: bit 64 is an undocumented *reason* bit of
the r-band default bitmask, and OR{n1,n2,n4,n8,n64,n1024} reproduces
mask_r, the v1 r-band mask, exactly on the P3 region.

Add MASK_n64 to DEFAULT_MASK_COLUMNS so the default galaxy cut is exactly
that set, documented as reproducing mask_r, and mirror it in the
calibration params, the v2 mask config and the minimal-catalogue demo.

Document n16/n32/n128/n256 as the u/g/i/z coverage flags (no r flag: the
catalogue is r-selected) and n2048 as absent Pan-STARRS z2. The faint vs
bright assignment of n1/n2 is unconfirmed for the Aug-2026 products, so
the labels no longer claim one.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
ShapePipe's make_cat pre-fills the NGMIX_* columns with sentinels
(G1/G2 = -10, T/FLUX = 0) and overwrites them only for objects present
in the ngmix output, so an object ngmix never fit keeps
NGMIX_MCAL_FLAGS == 0 and passes a flag-only cut. In
final_cat_smk-g7.hdf5 that is 18,983 of 1,851,100 objects (1.03%);
admitting them drags mean e1 to -0.096 (std 0.98) from +0.0001.

classification_galaxy_ngmix already rejected all 18,983 via the
NGMIX_G1_PSF_ORIG_NOSHEAR != -10 guard, so the production selection was
never affected -- verified on the real file, which gives 1,105,851 rows
out with and without the new cut. But that protection was incidental:
it is an exact float equality against a sentinel ShapePipe may change,
and the coadd N_EPOCH >= 2 cut in classification_galaxy_base does not
substitute for it (18,750 of the 18,983 have N_EPOCH >= 1). Cut on
NGMIX_N_EPOCH > 0 explicitly so the guarantee is stated, not inferred.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
mask_cut decided on truthiness via astype(bool). ShapePipe writes the
MASK_n* columns as float64 {0, 1} rather than bool (being fixed
upstream), and astype(bool) reads NaN as True, so an incomplete mask
column would have silently deleted sky. Decide on the value instead
(masked iff > 0.5), accept bool, int and float alike, and treat NaN as
"no verdict recorded" -- keep the object, but count and warn, since a
nonzero count means the product is defective. final_cat_smk-g7.hdf5
carries no NaNs and only exact 0.0/1.0, so this is defensive: the real
file gives 1,105,851 rows out before and after.

read_star_catalogue silently accepted a truncated file and threw away
exposure provenance. It now validates the n_exposures root attribute
against the datasets found, as the galaxy reader validates n_tiles
(check_n_tiles generalised to check_n_units), and adds an EXPID column
carrying the exposure number each star came from. The datasets are
named by that number and concatenating them discarded it, leaving no
way to group stars by exposure downstream. Names may be bare
("2086324", as smk-g7 writes them) or carry the CFIS suffix
("2110000p"), so EXPID takes the leading digits. On the real star
catalogue: 53,264 stars over 127 exposures, matching n_exposures.

Also note in group_dtype that ShapePipe writes TILE_ID as f8, so the
string-promotion branch is for a future string-valued TILE_ID, with a
TODO recording that as an open schema decision. No behaviour change.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
mask_v2.0.yaml, applied downstream to the comprehensive catalogue, cut
on neither NGMIX_MCAL_FLAGS nor any epoch column. It rejected the
never-fit objects only through its NGMIX_G1/G2_PSF_ORIG_NOSHEAR != -10
cuts, and only because make_cat happens to fill the PSF columns from
the same -10 literal it uses for the galaxy ellipticities. That is the
same accidental immunity just removed from
classification_galaxy_ngmix, one stage further downstream.

Add NGMIX_N_EPOCH >= 1. The column is already carried into the
comprehensive catalogue via add_cols_pre_cal in params.py. On
final_cat_smk-g7.hdf5 the cut keeps 1,832,117 of 1,851,100 objects,
removing exactly the 18,983 (1.03%) never-fit rows.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QbnPCyzuDNTgkg715pHhar
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VXzqmMYw7Kp8QMtoiVHjVq
sp_validation reads only the v2 column grammar, while every catalogue
on disk (v1.3.x-v1.6.x) is v1, so e.g. rho/tau raised KeyError on the
v1.4.a PSF file. New module sp_validation.grammar is the one place that
knows the difference:

- V1_RULES: the v1->v2 map of the shape-measurement columns as data (HSM
  renames, SIGMA -> T = 2 sigma^2 via cs_util.size.sigma_to_T, 2-vector or
  flattened NGMIX_ELL* split into G1/G2, PSFo/Tpsf/MOM_FAIL renames),
  applied to tables detect_generation calls v1. Mapping is by name, so v1
  values reach the names the code reads.
- MASK_RULES: the healsparse mask bits {b}_{label} (the names
  ApplyHspMasks gave them in the comprehensive HDF5's data_ext) ->
  MASK_n{b}, applied whatever the generation: they are the same bits of
  the same UNIONS bitmask ShapePipe v2 writes as MASK_n{b}
  (MASK_LABELS; 512 = outside the tile's unique region). IMAFLAGS_ISO is
  not mapped: its v1 bits mean different things.
- adapt(table, *tables): tables with nothing to rename come back
  unchanged; otherwise a V2View, a lazy column view over numpy, FITS_rec
  or h5py that joins row-aligned tables (data + data_ext), computes derived
  columns on access, and composes row selections as ranges or selected
  indices, reading only the window of rows they span.
- read_catalogue(path, hdu), materialise, v2_names, read_column_names.

rho_tau.get_rho_tau / get_jackknife_cov / get_theory_cov read each
catalogue once through read_catalogue and hand the table to
shear_psf_leakage (needs its loaded-catalogue support).

Tests: v1 twins adapt to their v2 twins for numpy (vector and flattened
ELL), FITS_rec and h5py; row selection commutes; dtype matches every
column read; data_ext mask names rename with or without a generation and
conflict with their new names; h5py selections read only their window;
the psf_size_error field and full get_rho_tau outputs agree between a v1
PSF catalogue and its v2 twin (and a rename-only sigma is caught).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
cailmdaley and others added 11 commits September 28, 2026 06:17
test_cat_config_columns_exist_on_candide reads each cat_config
catalogue's FITS header through grammar.read_column_names and checks
the psf block's declared columns and the shear block's *_col columns
are among the names the file presents. Known gaps are listed with
reasons and fail the test once healed.

Config fixes it surfaced:
- psf dec_col Dec -> DEC: every PSF file (and ShapePipe v2) names it
  DEC; only FITS_rec's case-insensitive lookup hid the mismatch.
- SP_v1.4.12.3 / SP_v1.4.13.3 psf blocks named v1 columns and the
  retired square_size flag; now the v2 names like every other entry.
- SP_axel_v0.0 / SP_v1.4.5.A shear: declare their lowercase ra/dec;
  SP_v1.4.5.A's PSF shape columns are psf_g1/psf_g2.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
build_catalog indexes each column with ``mask``; a scalar True only
adds an axis treecorr reshapes away (no stars are cut), and a scalar
False selects nothing, so treecorr raises "Input arrays have zero
length". That broke get_rho_tau for DES and get_jackknife_cov for
every catalogue. No flag cut was ever applied, so dropping the
argument leaves the non-DES results unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Every reader on the calibration path now presents its tables in the v2
column grammar, so a ShapePipe v1 product runs through the same code and
the same configs as a v2 one:

- catalog.read_campaign_catalogue, campaign_shape, iter_campaign_tiles and
  the JointCat merge adapt each tile (param_list names v2 columns);
  read_star_catalogue adapts FITS and HDF5 star catalogues.
- CalibrateCat.read_cat returns one table: the comprehensive HDF5's data
  and data_ext joined by grammar.adapt, so mask columns read the same
  whether they sit in data (v2) or data_ext (post-processed v1), and
  galaxy.mask_cut works on either.
- get_masks_from_config takes that one table, and each mask config is one
  `dat` cut list: the v1 configs' dat_ext cuts move into it under their
  MASK_n{b} names (the selections are unchanged; IMAFLAGS_ISO stays in the
  v1 configs). The image-sim overlay and scripts/masking.py follow.
- ApplyHspMasks writes mask bit b as MASK_n{b}, from grammar.MASK_LABELS.
- cosmo_val/compute_theory_cov.py hands CovTauTh loaded, adapted tables.

Tests: a v1 comprehensive HDF5 (v1 data + data_ext with the old mask
names) and its v2 twin give identical mask_cut and per-config mask
selections for every config in config/calibration, and identical metacal
inputs and response; the campaign and star readers present v1 files in v2.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…builders, seeded patches)

Pinned below develop's tip 372980f, whose scipy>=1.18 requirement cannot
resolve against cs_util<0.3.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LsAQRxNg6yDsUX4y2SWxTJ
ShapePipe v1 wrote a wrong NGMIX_Tpsf_NOSHEAR: it differs by ~2% for
nearly every object (v1.4, v1.5, v1.6 comprehensive) from the
reconvolution kernel metacal applied, which NGMIX_Tpsf_{1P,1M,2P,2M}
record (they agree to ~1e-5). The v1.4.6.3 release used the 1P value
in its metacal size cut and wrote it as the no-shear column; #267
dropped that substitution, which is a no-op only on the v2 stack, so
the branch's v1 calibration selected ~3% more objects than the release.

The adapter now presents NGMIX_T_PSF_RECONV_NOSHEAR from NGMIX_Tpsf_1P
when the table has it, falling back to NGMIX_Tpsf_NOSHEAR (a cut
catalogue such as v1.4.6.3's already holds the 1P value there), and
hides the raw column. Rules gain a fallback source; a derived column is
presented where the first of its sources sits. The module docstring
and the migration doc state the rule: name mapping, except this one
documented v1 defect.

Every consumer now sees the 1P value, including the w_des and
PSF-leakage size-ratio binning, where the release used the raw value;
those two columns therefore do not reproduce the release bit for bit.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LsAQRxNg6yDsUX4y2SWxTJ
metacal reads ~40 columns of data[mask]. Over h5py datasets the lazy
view read each column over the whole span of the selection, so a
full-sky v1 calibration re-read the 238 GB comprehensive file once per
column. Selecting rows of a view over h5py Datasets now reads the
selected rows of every column in one pass per dataset (whole rows, in
256 MB blocks over the selection's span, skipping empty blocks), as
indexing the Dataset did on develop, and wraps them in a view over the
in-memory arrays so renames and derived columns still apply. Views
over in-memory tables stay lazy. to_structured reads each dataset once
for all requested fields.

On a cold 10M-row window of v1.4.c (4.8M rows selected):
h5py Dataset[mask] 19.2 s; adapt(data, data_ext)[mask] plus the 40
metacal columns 23.5 s; the previous per-column path took 16.5 s for
its first column alone, i.e. one full re-read per column.

dtype is computed once per view and passed to row selections
(group_dtype asked for it per column, quadratically). view[()] and
view[...] select every row, as for an h5py Dataset; other tuples raise.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LsAQRxNg6yDsUX4y2SWxTJ
A column missing from some tiles (several v1 tiles carry no
SPREAD_MODEL) surfaced as numpy's bare "no field of name" KeyError.
group_dtype now checks every table first and raises a KeyError naming
how many tables lack which columns, and which ones; callers pass the
tables keyed by dataset name so the message names tiles.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LsAQRxNg6yDsUX4y2SWxTJ
Mask configs are one dat: list over the joined data + data_ext table,
but get_masks_from_config and scripts/masking.py read only config["dat"],
so an older config with a dat_ext: list (several sit under v1.4.x and
v1.5.x) silently lost those cuts. masks.catalogue_cuts returns the dat
list and raises on dat_ext, saying to merge it into dat with {b}_{label}
renamed MASK_n{b}; the calibration, footprint and demo scripts all read
cuts through it.

scripts/masking.py's footprint cuts had dropped IMAFLAGS_ISO, which the
v1 configs still cut on and which is spatial (halo, border, Messier,
NGC, spike bits); it is back in SPATIAL_CUTS.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LsAQRxNg6yDsUX4y2SWxTJ
…s do

mask_cut kept objects whose mask column is NaN (with a warning), while
the config-driven cut that calibration runs (kind: equal, value: False)
drops them, so the two selections disagreed on a defective product.
mask_cut now drops them too, still counting and warning. Keeping them
was chosen to stop astype(bool) from dropping NaN rows silently; the
warning covers that, and dropping an object with no mask verdict is
the conservative choice for a shear catalogue.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LsAQRxNg6yDsUX4y2SWxTJ
Drop wording that goes stale ("being fixed upstream", "unconfirmed for
the Aug-2026 products", "legacy" layouts and files) in favour of what
is true now, and ApplyHspMasks.write_hdf5_header's documented but
nonexistent campaigns parameter.

test_configured_paths_exist_on_candide resolved a calibration config's
relative params.input_path against config/calibration; such configs
run as config_mask.yaml in their run directory, where the relative
path names a file, so the guard skips those.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LsAQRxNg6yDsUX4y2SWxTJ
…, slow)

Runs scripts/calibration/calibrate_comprehensive_cat.py with
mask_v1.X.6.yaml on rows 200M..201M of the v1.4.c comprehensive HDF5
and checks the released cut catalogue's rows from that window, matched
by (RA, Dec) and bracketed by rows from outside it: same objects, same
order, and identical per-object columns, including the no-shear
reconvolved-PSF size against the release's NGMIX_Tpsf_NOSHEAR. Globally
calibrated columns (e1, e2, w_des, leakage-corrected) are not compared.
Skipped where the release products are absent.

Passes in ~30 s on n09 (131,526 objects); before the no-shear
reconvolved-PSF correction it fails, selecting 134,936.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LsAQRxNg6yDsUX4y2SWxTJ
@cailmdaley
cailmdaley force-pushed the feat/v2-catalogue-readers branch from 9ec5502 to 330ac8d Compare September 28, 2026 10:24
@cailmdaley cailmdaley changed the title Read ShapePipe v2 campaign products: hdf5 catalogues, MASK_n* cut, patches retired Read ShapePipe v1 and v2 products side by side Sep 28, 2026

@martinkilbinger martinkilbinger left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Very nice, I like the campaign tag superceding the patch IDs. And the adopt mechanism.

Comment thread config/calibration/mask_v1.X.10.yaml Outdated
Comment thread docs/ngmix_psf_column_migration.md Outdated
Comment thread docs/ngmix_psf_column_migration.md Outdated
cailmdaley and others added 6 commits September 29, 2026 16:25
ShapePipe v2 catalogues in DR6 catalogue mode no longer carry MAG_WIN,
MAGERR_WIN, SNR_WIN, FLUX_AUTO, FLUXERR_AUTO, FLUX_APER, FLUXERR_APER,
FWHM_IMAGE, FWHM_WORLD (shapepipe #924). Stop passing them through in
extract_info (params.add_cols) and calibrate_comprehensive_cat, and drop
the SExtractor SNR_WIN curve from the galaxy SNR histogram. The v1.4.6.3
release regression compares the remaining per-object columns.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BeQNTBcjTqu6TPktxysovQ
uv.lock taken from develop; `uv lock` resolves it unchanged (its
shear_psf_leakage pin, develop@649edf4, already contains #44).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The external-mask column for bit b is MASK_{b}_{label}, labels from
grammar.MASK_LABELS (grammar.mask_column), so the column carries both the
bit identity and its meaning. ApplyHspMasks writes these names; every mask
config, script and the galaxy defaults cut on them (galaxy derives its
tuples from mask_column).

The adapter presents the v1 data_ext spelling {b}_{label} and the
pre-release ShapePipe v2 spelling MASK_n{b} under the canonical name;
neither marks a generation. Tests cover both spellings mapping to the same
column and the same cuts, including the fixed-name faint/bright halo check.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Delete scripts/combine_results.py and scripts/compute_area.py (no caller;
  their R.txt/c.txt outputs feed only a create_joint_shape_cat.py that no
  longer exists) and the post_processing.md sections describing them.
- Delete cosmo_val/compute_theory_cov.py: its version keys, home-directory
  paths and base_dir + absolute-subdir join no longer resolve, and
  rho_tau's CovTauTh path, which the workflow drives, does the same job.
- using_the_catalogues.md: the v1.0 example selects on the file's `patch`
  column; v1.4.1+ comprehensive HDF5 (flat data/data_ext) is read with
  grammar.adapt, cut FITS catalogues with grammar.read_catalogue.
- index.rst: per-campaign merging. mask_r agreement stated as verified on
  the P3 sky area. Drop the uncalled ApplyHspMasks.get_label_struct.
  Campaign-reader fixture named CAMPAIGN.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@cailmdaley

Copy link
Copy Markdown
Collaborator Author

this is now ready for review @sachaguer!

@sachaguer sachaguer left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I am confused by the new fits to hdf5 catalog loading. Until this PR, sp_validation was able to read .fits file given some configuration giving the name of the columns. Now my understanding of the PR is that there are two branches to read the ShapePipe output. One that reads the catalog from hdf5 files and one that reads from FITS files.

What is puzzling me is whether the catalog loading is now kind of hard coded on UNIONS and I could not find my answer in the many modifications that were made. For example, in galaxy.py the star columns written by ShapePipe v2 is hardcoded in the script. My worry is if someone wants to run our pipeline and its data and, say, the star catalog is a FITS file or a hdf5 file with a different name convention for the columns. Is it possible to run sp_validation with this updated API.

I would be happy to have clarification on that matter. I was not able to disentangle it myself just going through the many modified files.

cailmdaley and others added 2 commits October 1, 2026 16:37
Container and column naming are now two axes, each known in one place.

sp_validation.io.Catalogue / read_catalogue detect the container from file
contents (h5py.is_hdf5, else FITS) and read FITS table HDUs, a single HDF5
dataset, a comprehensive data + data_ext HDF5, and ShapePipe's HDF5 groups
of per-tile/per-exposure chunks (n_tiles / n_exposures checked; key_column
keeps chunk names, e.g. EXPID). Every catalogue read routes through it:
extract_info (fixing the np.load-on-.fits path), the builders' read_cat
(BaseCat and CalibrateCat merged) and campaign merge, cosmo_val pseudo-Cl,
catalogue characterisation and the shear_psf_leakage-backed results, the
rho/tau loader (temp FITS only when the input is not already plain HDU 1),
glass_mock, calibration and scripts/masking.py.

grammar.adapt takes a column_map {canonical name: name in the file},
explicit or one-'*' patterns, applied lazily ahead of the v1 rules; it
overrides any rule for the same name and the v1 detection runs on the rest.
Catalogue configs carry it (cat_config blocks, calibration params,
params.py galaxy_/star_column_map). The *_col keys still pick roles,
naming columns as presented after the map.

config/columns/shapepipe_v2.yaml lists every column the code reads by
fixed name; test_column_schema keeps code and file in step both ways, and
a Sphinx extension renders it as docs/source/catalogue_columns.md.

Removed: catalog.read_campaign_catalogue, read_star_catalogue,
campaign_shape, iter_campaign_tiles, concatenate_datasets,
STAR_CAT_COLUMNS, read_shape_catalog; grammar.read_catalogue,
read_column_names, requires_adaptation; metacal's prefix argument.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01816Up3mpuhwYq7cnmyRHqv
io.open_catalogue / open_entry return a catalogue as one table without
reading it: FITS memory-mapped, unchunked HDF5 read per column on access.
The cosmo_val leakage subclasses, pseudo-Cl and catalogue
characterisation use them instead of reading every column into memory.
Through the leakage path the v1.4.6.3 shear catalogue (61.4M rows, v1
names, so adapted) peaked at 29.6 GiB RSS, 15.6 GiB of it anonymous; it
now peaks at 15.4 GiB with 0.2 GiB anonymous, as upstream's memmap read
does (the rest is file-backed page cache).

grammar: a column_map entry overrides only the rule for its own canonical
name; the columns it renames still feed the other rules (a map onto
NGMIX_Tpsf_1P no longer sends NGMIX_T_PSF_RECONV_NOSHEAR to the wrong
NGMIX_Tpsf_NOSHEAR). adapt(V2View, column_map=...) applies the map over
the view instead of ignoring it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01816Up3mpuhwYq7cnmyRHqv
@cailmdaley

cailmdaley commented Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator Author

Replying to Sacha's question about hard-coded catalogue loading. That was a fair worry, and it pointed at something real. Column names were still configurable, but the file format was decided separately at each call site, and cosmo_val could only open FITS. Since b1fbf84, both the file format and the column names are configuration:

  • File format: sp_validation.io is the only code that opens catalogues, and everything goes through it. It tells FITS from HDF5 by the file's contents, and reads a FITS table, a single HDF5 dataset, or ShapePipe's campaign files of one dataset per tile or exposure. A star catalogue can be FITS or HDF5 under any name.
  • Column names: the code reads only the names listed in config/columns/shapepipe_v2.yaml, which gives each column's meaning and is rendered as the "Catalogue columns" docs page. A catalogue with other names adds column_map: {canonical name: name in your file} to its config entry, and patterns like {"NGMIX_*": "MYFIT_*"} work. A test fails if the code reads a name the schema doesn't list.
  • Your galaxy.py example: the mask columns there are defaults for the mask_columns parameter, and they can be renamed through the map like everything else.

A test reads one catalogue with made-up column names from each of these layouts and gets identical tables. Another renames a metacal catalogue's columns and recovers bit-identical R and calibrated g through the map. The PR description now opens with a "Where to look" section: the four files that carry the design, which diffs are mechanical, and what can be skipped, so the 82 files are less of a wall.

— Claude (Opus) on behalf of Cail

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

3 participants