Skip to content

[SYSTEMDS-3971] Improve Scuro test coverage - #2615

Open
shieru1214 wants to merge 6 commits into
apache:mainfrom
shieru1214:scuro-coverage
Open

shieru1214 wants to merge 6 commits into
apache:mainfrom
shieru1214:scuro-coverage

Conversation

@shieru1214

@shieru1214 shieru1214 commented Sep 16, 2026 •

Copy link
Copy Markdown

This change adds 13 tests for persistence and results retrieval in unimodal_optimizer.py.

Persistence

  • store_results: It saves the optimizer results and execution statistics in two separate files. The tests check the saved contents, the automatically generated file name, and the fallback when the file name does not end with .pkl.
  • load_results: It loads previously saved results. The test stores one result, clears it from memory, and then checks that it can be loaded again.
  • _count_results_by_modality: It counts the results from the first task of each modality for checkpoint progress. The test checks the returned count for both a modality with results and an empty modality.
  • resume_from_checkpoint: It restores results from a checkpoint. The tests cover both restoring saved results and keeping the current results when no checkpoint is found.

Results Retrieval

  • get_k_best_results: It returns the top k results according to their scores. The tests check the score order, using existing cached data, executing the result DAGs when the cache is empty, and returning empty lists when there are no results.
  • get_dag_by_id: It searches a list of DAGs by dag_id. The tests cover both returning the matching DAG and returning None when the ID is not found.

The five optimizer tests built one modality each and then called the same
helper, which holds all the assertions. They become a table of modality
sets and a small factory, driven by one subTest per set, so the number of
combinations stays the same. The keyword arguments keep the inputs as the
separate tests had them, including the text case that used one sentence
instead of ten.

test_audio_representations built its audio modality inline although
_create_audio_modality already exists and the other two audio tests use
it. It now calls the helper as well.

Assisted-by: AI
The three per-modality window aggregation tests ran the same helper over
data of the same shape and dtype, because window_aggregation dispatches on
the data layout and not on the modality type. They become one test with
the modality and the aggregation as subTest dimensions, which also splits
the inner aggregation loop into separately reported cases.

The 3d and 2d shape tests differed only in the input dimensions and in the
hard coded expected shape. Window aggregation only changes the first axis,
so one expression covers both.

Assisted-by: AI
The four tests ran the same eleven line sequence and differed only in the
operator and in the expected booleans. The booleans are the actual
content, so they move into a table with one subTest per operator.

The commutativity check now compares the measured result with the
operator's own "commutative" attribute instead of a second hand written
copy of it. The concatenation test also compared the wrong operand, so its
chain against n-ary property was never covered.

Assisted-by: AI
Three groups of methods in unimodal_optimizer had no test at all, because
they only run when a flag is set that no test sets: store_results and
load_results need save_all_results, resume_from_checkpoint needs resume,
and add_result is only called by the hyperparameter tuner.

Thirteen tests build the optimizer and the result container directly
instead of running a search, so they finish in milliseconds. They cover
storing and loading the results, the file name fallbacks, the checkpoint
progress count, resuming, reading back the k best results including both
cache states, and looking up a dag by id.

store_cache and load_cache stay untested for now: store_cache writes the
whole cache to one file under result_path while load_cache expects one
file per modality and task in the working directory, so a store followed
by a load raises FileNotFoundError.

Assisted-by: AI
@shieru1214

shieru1214 commented Sep 29, 2026 •

Copy link
Copy Markdown
Author

This change improves the test coverage of hyperparameter_tuner.py. The added tests focus on three parts: search space construction, trial parameter, and tuning result handling.

1. Search space construction

This part converts the parameter ranges declared by operators into search spaces that Optuna can use. _build_param_specs is the entry function and it was already fully covered, so I mainly added tests for three helper functions. For _param_values_to_spec, I tested list, tuple, scalar, other iterable, and invalid inputs.
For _window_input_stats, I tested whether it can create input length information from window_size. This information is later used to remove candidate values that do not fit the input.
For _expand_aggregation_param_specs, I added five cases for flattening nested aggregation parameters, narrowing the domain based on the input length, and inputs that cannot be expanded.

2. Trial parameter

The first part flattens the parameters, while this part puts the values selected by Optuna back into the DAG nodes. The main function is _apply_trial_params_to_node. It first takes the parameters belonging to the current node, and then decides where to place them based on the form of the node. I added tests for two branches: a node with a pushed-down aggregation and a node that is itself an AggregatedRepresentation.
For _apply_pushdown_trial_params, I tested whether aggregation parameters are correctly stored in _pushdown_aggregation, whether other parameters stay at the top level, and whether the original parameters remain unchanged. I also added tests for _materialize_node_params and the two operation checks, including non-class inputs, import failures, leaf nodes, and empty parameters.

3. Tuning result handling

This part tests how tuning results are stored and read. For add_result, I added two cases: storing multimodal results in mm_results, and skipping a tuning result when it is None. setup_mm uses optimize_unimodal to decide whether the result structure should be replaced with an mm_results entry. This flag indicates whether the unimodal representations at the leaf nodes are also optimized. The tests for get_k_best_dags check that best_params are applied to the rebuilt DAG and that the original DAG is not changed. The test for get_k_best_results checks that the rebuilt DAG is executed and the last value is taken from its output dictionary.

@codecov

codecov Bot commented Oct 1, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 75.66%. Comparing base (8c7d556) to head (b3134b4).
⚠️ Report is 19 commits behind head on main.

Additional details and impacted files
@@              Coverage Diff              @@
##               main    #2615       +/-   ##
=============================================
+ Coverage     71.81%   75.66%    +3.84%     
=============================================
  Files          2075      436     -1639     
  Lines        222176    25465   -196711     
  Branches      38264        0    -38264     
=============================================
- Hits         159559    19268   -140291     
+ Misses        51561     6197    -45364     
+ Partials      11056        0    -11056     
Flag Coverage Δ
java ?
python 75.66% <100.00%> (+0.70%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@shieru1214

shieru1214 commented Oct 2, 2026 •

Copy link
Copy Markdown
Author

This change adds tests for multimodal_optimizer.py, which cover three parts: fusion DAG generation, constructor setup, and the optimize loop.

Fusion DAG Generation

Before this change, the three methods that generate the fusion search space had no active tests.

  • _generate_modality_combinations: test_modality_combinations_respect_min_max_modalities uses three modalities and checks the exact subsets for max_modalities values of 2, 3, and 5. This also checks the case where the maximum is larger than the number of available modalities.
  • _generate_representation_combinations: test_representation_combinations_pick_one_per_modality uses one modality with two representations and another with three. It checks all six possible combinations.
  • _generate_fusion_dags: test_fusion_dags_contain_every_selected_representation checks every generated DAG. It checks that the selected modality and representation leaves are present, the number of fusion nodes is correct, each fusion node has two inputs, and the root is a fusion node.
  • test_fusion_dags_use_every_fusion_operator checks that both configured operators, Concatenation and Average, are used. This is tested separately because a DAG can have the correct structure but still miss one operator.

Constructor Setup

The generator tests already called __init__, but it did not directly check the values created during construction. I added three tests for this part.

  • test_min_modalities_clamped_to_two checks that min_modalities=1 is changed to 2. This prevents single-modality DAGs from being treated as fusion DAGs.
  • test_max_modalities_defaults_to_number_of_modalities checks that the default maximum is the number of available modalities.
  • test_k_best_representations_hold_the_cached_data covers _extract_k_best_representations. It checks that every modality is stored by its ID and that the optimizer keeps the cached representation data returned by the unimodal results, instead of the score records.

Optimize Loop

Before this change, none of the tests called optimize. I added three tests for the main logic of this loop.

  • test_stops_at_max_combinations uses three modalities and sets max_combinations=5. It checks that _evaluate_dag is called exactly five times and that five results are stored.
  • test_drops_failed_evaluations makes every second evaluation return None. It checks that six evaluations are attempted but only three results are stored. Failed evaluations still count towards the limit, and the loop continues after them.
  • test_keeps_results_and_budget_per_task uses two tasks with a limit of two. It checks that each task receives two evaluations and that every result is stored under the correct task name.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: In Progress

Development

Successfully merging this pull request may close these issues.

1 participant