[SYSTEMDS-3971] Improve Scuro test coverage - #2615
shieru1214 wants to merge 6 commits into
Conversation
The five optimizer tests built one modality each and then called the same helper, which holds all the assertions. They become a table of modality sets and a small factory, driven by one subTest per set, so the number of combinations stays the same. The keyword arguments keep the inputs as the separate tests had them, including the text case that used one sentence instead of ten. test_audio_representations built its audio modality inline although _create_audio_modality already exists and the other two audio tests use it. It now calls the helper as well. Assisted-by: AI
The three per-modality window aggregation tests ran the same helper over data of the same shape and dtype, because window_aggregation dispatches on the data layout and not on the modality type. They become one test with the modality and the aggregation as subTest dimensions, which also splits the inner aggregation loop into separately reported cases. The 3d and 2d shape tests differed only in the input dimensions and in the hard coded expected shape. Window aggregation only changes the first axis, so one expression covers both. Assisted-by: AI
The four tests ran the same eleven line sequence and differed only in the operator and in the expected booleans. The booleans are the actual content, so they move into a table with one subTest per operator. The commutativity check now compares the measured result with the operator's own "commutative" attribute instead of a second hand written copy of it. The concatenation test also compared the wrong operand, so its chain against n-ary property was never covered. Assisted-by: AI
Three groups of methods in unimodal_optimizer had no test at all, because they only run when a flag is set that no test sets: store_results and load_results need save_all_results, resume_from_checkpoint needs resume, and add_result is only called by the hyperparameter tuner. Thirteen tests build the optimizer and the result container directly instead of running a search, so they finish in milliseconds. They cover storing and loading the results, the file name fallbacks, the checkpoint progress count, resuming, reading back the k best results including both cache states, and looking up a dag by id. store_cache and load_cache stay untested for now: store_cache writes the whole cache to one file under result_path while load_cache expects one file per modality and task in the working directory, so a store followed by a load raises FileNotFoundError. Assisted-by: AI
|
This change improves the test coverage of 1. Search space constructionThis part converts the parameter ranges declared by operators into search spaces that Optuna can use. 2. Trial parameterThe first part flattens the parameters, while this part puts the values selected by Optuna back into the DAG nodes. The main function is 3. Tuning result handlingThis part tests how tuning results are stored and read. For |
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #2615 +/- ##
=============================================
+ Coverage 71.81% 75.66% +3.84%
=============================================
Files 2075 436 -1639
Lines 222176 25465 -196711
Branches 38264 0 -38264
=============================================
- Hits 159559 19268 -140291
+ Misses 51561 6197 -45364
+ Partials 11056 0 -11056
Flags with carried forward coverage won't be shown. Click here to find out more. ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
… optimizer Assisted-by: AI
|
This change adds tests for Fusion DAG GenerationBefore this change, the three methods that generate the fusion search space had no active tests.
Constructor SetupThe generator tests already called
Optimize LoopBefore this change, none of the tests called
|
This change adds 13 tests for persistence and results retrieval in
unimodal_optimizer.py.Persistence
store_results: It saves the optimizer results and execution statistics in two separate files. The tests check the saved contents, the automatically generated file name, and the fallback when the file name does not end with.pkl.load_results: It loads previously saved results. The test stores one result, clears it from memory, and then checks that it can be loaded again._count_results_by_modality: It counts the results from the first task of each modality for checkpoint progress. The test checks the returned count for both a modality with results and an empty modality.resume_from_checkpoint: It restores results from a checkpoint. The tests cover both restoring saved results and keeping the current results when no checkpoint is found.Results Retrieval
get_k_best_results: It returns the topkresults according to their scores. The tests check the score order, using existing cached data, executing the result DAGs when the cache is empty, and returning empty lists when there are no results.get_dag_by_id: It searches a list of DAGs bydag_id. The tests cover both returning the matching DAG and returningNonewhen the ID is not found.