Repository navigation
Conversation
ISplineTransformer integrated each M-spline with a fixed 200-point trapezoid grid, so basis functions narrower than one grid step came out identically zero or ramped over the wrong interval. Each I-spline is now the B-spline antiderivative normalized by its analytic integral, which is exact for any knot spacing. Fixes #53
MSplineTransformer zeroed spans below 1e-6 and clipped the design to +/-1e3. M-spline values scale like 1/span, so these absolute limits made the basis depend on the feature's units: small-scale features were flattened and tiny-range ones became all zero. The clips are removed; only a basis function with degenerate support relative to the knot range is zeroed. Fixes #63
Quantile knots on top-coded, zero-inflated or low-cardinality features repeat and land on the range boundary. The B/M/I base clipped knots into the closed range, and the natural-cubic, cubic-regression and tensor-product paths used them verbatim, so basis functions collapsed to a point: dead or duplicate columns, and all-zero B/M-spline rows at x_max. A shared supplement_interior_knots helper now drops boundary and repeated knots and tops the set up from strictly interior quantile and uniform candidates, keeping every existing knot. Both spline construction paths use it. Fixes #56
…e spacing CARTLocationSelector returned its split candidates in ascending location order and the minimum-spacing filter keeps the first of two close candidates, so a weak split just below the dominant one evicted it. Candidates are now ordered by weighted impurity decrease (as the LightGBM selector already does with gains), and the same ranking drives the over-max trim. Fixes #70
…points on tied data On discrete or heavily tied features the selector's quantile top-up and its quantile fallbacks collapsed onto the tied values and the range boundary: too few locations (narrower than output_dim, breaking CrossFittedTransformer), duplicates, and locations on x_min/x_max that gave PLE a constant column. The top-up and both fallbacks now keep every selector-found location and fill the shortfall with strictly interior, spaced quantile candidates first, then uniform ones. LightGBM's +/-1e-35 split-at-zero sentinel is mapped to the midpoint of the values it separates, and PLE treats any positive bin width as non-degenerate instead of using an absolute 1e-10 tolerance. Fixes #57
…ement
LightGBMLocationSelector always trained the binary objective on the raw
labels. Labels other than {0, 1} looked single-class (silently falling back to
quantiles), multiclass targets only found the class-0 boundary, and string
labels raised. Targets are now label-encoded and trained with the binary or
multiclass objective as appropriate, matching the CART selector.
Fixes #61
…placement The B/M/I, cubic-regression and natural-cubic splines asked their target-aware selector for a fixed 3-15 basis-function window and then reduced the result to the needed count with select_knots, which keeps evenly spaced indices of the sorted list. That positional trim routinely discarded the split where the target changes, and capped adaptive widths at 15 (14 / 12 for the cubic families). The spline adapter now accepts a per-call knot window; the transformers pass the window resolved from output_dim or [min_output_dim, max_output_dim], so the selector's importance ranking keeps the most informative splits, and a short set is only topped up. The adaptive tutorial output is updated accordingly. Fixes #69
…step names valid Preprocessor passed each column label to scikit-learn's ColumnTransformer both as the column selector and inside the step name. ColumnTransformer reads an integer selector as a position, so integer labels that differ from their positions silently swapped columns (or crashed), and labels starting with '_' or containing '__' produced step names scikit-learn rejects. The ColumnTransformer is now fitted and applied on a frame whose labels are strings (a shallow relabelled copy only when needed), and columns are selected by that string label. Steps keep the num_<label> / cat_<label> name whenever scikit-learn accepts it and otherwise get a unique positional name; block keys, feature names and lineage are derived from the column label, and output_dims_ / get_feature_info map back to the input labels. Fixes #55 Fixes #71
Boolean columns are detected as categorical, but the categorical pipeline starts with a SimpleImputer, which rejects the bool dtype, so any frame with a plain bool column failed to fit under the default settings and every preset. Boolean columns are now cast to object (True / False, with a nullable boolean's missing value as NaN) before they reach the ColumnTransformer; the caller's frame is left unchanged. Fixes #54
…presentation With imputation enabled, add_missing_indicator=True and missing_policy='impute_with_indicator' used SimpleImputer(add_indicator=True) as the first pipeline step, so every later step saw the indicator as a second input feature: it was scaled, expanded into basis columns or crossed into polynomial terms, and get_feature_lineage() crashed. The imputer now runs without its indicator and scikit-learn's MissingIndicator (the one SimpleImputer used internally) runs on the raw input in a parallel branch, mirroring the separate_state construction. The indicator stays a raw 0/1 column with its previous name, the representation columns are unchanged, and lineage reads the indicator width from the fitted branch. Fixes #62
… values numerical_method='box-cox' rescales into (1e-3, 1) with a MinMaxScaler fitted on the training data, so any value below the training minimum mapped to <= 0 and PowerTransformer raised at transform, breaking prediction and plain cross-validation. The scaler's output is now floored at 1e-3: such values transform like the training minimum, while every other output, including values above the training maximum, is unchanged. Fixes #59
Preprocessor recorded n_features_in_ at fit but never checked it. Array columns are named feature_0, feature_1, ... by position, and the inner ColumnTransformer selects those names, so an array with an extra leading column was shifted by one (dropping the last feature) and a trailing extra column was ignored, with no error. transform now raises the usual 'X has n features, but ... is expecting m' error for NumPy input whose width differs from n_features_in_. Fixes #72
Preprocessor forwarded input_features to its ColumnTransformer, which was fitted on the synthetic feature_0, feature_1, ... names of array input and so required exactly those names. That broke Pipeline.get_feature_names_out() whenever an upstream step outputs arrays. Following scikit-learn's convention, Preprocessor now records feature_names_in_ only when fitted on string-labelled columns and keeps validating input_features against it. Without it, any n_features_in_ names are accepted: each step's output names are rebuilt from its fitted transformer with the given labels (on shallow copies without their recorded feature names), so they match the names a fit on named columns would produce. Fixes #65
With output_format='auto' the sparse/dense choice was re-made on every transform call from that batch's density, so the same fitted preprocessor returned a dense array for some batches and a CSR matrix for others (often single rows), breaking dense-only estimators at predict time. The choice is now made once at fit from the training output (which the inner ColumnTransformer computes anyway) and stored in output_format_; transform uses it, and output_report_ still reports each call's own density. Fixes #73
…n specs to_spec emitted object arrays with a raw tolist(), so numpy scalars such as np.bool_ in fitted state (e.g. SimpleImputer.statistics_) leaked into the 'JSON-compatible' spec: fingerprint_, reproducibility_report() and json.dumps(to_spec()) raised, and to_spec(path) left a truncated file behind. Builtin dtypes such as Preprocessor(dtype=float) were written as builtins:float, which from_spec always refuses. Object-array elements are now encoded one by one (the format is unchanged, so existing specs still load), builtin dtype specifiers are stored as the numpy dtype they denote, types a spec could never load raise PretabSerializationError at save time, and to_spec serializes to text before opening the target file so a failure never truncates an existing spec. Fixes #66
… root handlers configure_logging returned early whenever the 'pretab' logger had any real handler, so it could not tell PreTab's own stream handler from a host's: the first configuration pinned the level for the rest of the process (a later verbose=2 fit only printed the level-1 summary), and under a host-configured root logger (logging.basicConfig) PreTab still added its own handler, printing every line twice. PreTab now marks the handler it attaches. A host handler on the 'pretab' logger still makes the call a no-op; otherwise the level is always applied and PreTab's handler is reused, not duplicated. When a host handler on an ancestor already receives the records, no handler is attached and one attached earlier is removed. get_feature_info() no longer lowers a DEBUG level set earlier. Fixes #67
CrossFittedTransformer converted X with np.asarray and sliced folds from that array, so a mixed-dtype DataFrame reached the wrapped estimator as an object array without column names. A wrapped Preprocessor then treated every column as categorical (every out-of-fold value of a numeric feature became the 'unseen category' code), and name-based estimators raised. DataFrames are now passed through unchanged, with each fold taken as a row subset by position; other input is still converted to a 2D array. Fixes #58
This was referenced Oct 9, 2026
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes 18 of the 20 open bug reports, one commit per issue. Each commit has its own regression tests, and each of those tests fails on
main.Not included, because they already have open PRs:
Both PRs touch
preprocessor.py,compose/factory.pyandcompose/output.py, so whichever merges second may need a small rebase.Fixes
Splines and placement
supplement_interior_knotshelper drops knots that sit on the range boundary or repeat. It tops the set up with strictly interior quantile and uniform knots. The B/M/I base and the natural-cubic, cubic-regression and tensor-product paths all use it.binaryobjective for two classes andmulticlassotherwise.output_dimor[min_output_dim, max_output_dim]) to the selector, so its importance ranking picks the knots instead of a positional trim. Adaptive widths are no longer capped at 15. The adaptive tutorial output is updated.Preprocessor and composition
ColumnTransformeris fitted on string column labels and selects each column by that label. Step names staynum_<label>/cat_<label>whenever scikit-learn accepts them; otherwise they get a unique positional name. Block keys, feature names and lineage are derived from the column label, so they are unchanged.output_dims_andget_feature_infoare keyed by the original labels.objectbefore theColumnTransformer. A missing value in pandas' nullablebooleandtype becomesNaN.MissingIndicator(the classSimpleImputeruses internally) runs on the raw input in a parallel branch, the same wayseparate_stateis built. The indicator stays a raw 0/1 column with its previous name, and lineage works.MinMaxScaler(clip=True), which would also clip values above the training maximum that work today.transformraises the usual "X has n features, but … is expecting m" error for NumPy input whose width differs from the fitted one.feature_names_in_is recorded only for string-labelled input, following scikit-learn's convention. Without it, anyinput_featuresof the right length rename the inputs. Names are rebuilt from each fitted step, so they match a fit on named columns.output_format="auto"is decided once atfitfrom the training output and stored inoutput_format_.output_report_still reports each call's own density.floatare stored as the numpy dtype they denote. Types a spec could never load raisePretabSerializationErrorat save time.to_spec(path)builds the text before opening the file, so a failed save can no longer truncate an existing spec.logging.basicConfig), PreTab attaches no handler and removes one it attached earlier.get_feature_info()no longer lowers a DEBUG level.CrossFittedTransformerpasses DataFrames to the wrapped estimator unchanged, taking each fold by row position.Behaviour changes to review
add_missing_indicator=Trueorimpute_with_indicator, a basis method now gets one extra column per feature with missing values. Before, it got an expanded copy of the indicator. Representation column names are unchanged.output_format_andfeature_names_in_change the fingerprint of newly fitted preprocessors. Specs written by 1.0.0 still load and keep their previous"auto"behaviour.1and"1", raisePretabDataError.Testing
maingives 1438 passed in the same environment (Python 3.12, numpy 2.5.3, pandas 2.3.3, scikit-learn 1.9.1, scipy 1.18.1, lightgbm 4.7.0).scripts/quickstart.pypasses.ruff checkandruff format --checkare clean.test_preprocessor.pyandtest_edge_cases.py, also appear onmainin this environment.main).output_dim6 and 7 (198/200 and 123/200 without the Target-aware spline placement trims selector knots by position, discarding the most informative splits; adaptive width capped at 15 #69 change).docs/tutorials/adaptive_resolution.mdanddocs/core_concepts/outputs_and_inspection.md.Follow-ups (not in this PR)
placement_strategy="quantile"for the feature maps (RBF, ReLU, sigmoid, tanh) still repeats centers on tied data. On a zero-inflated feature,output_dim=6gives only 4 distinct columns. No open issue covers this yet.infer_objects()on object ndarrays into_dataframe, is not done. It would change type detection for object-array input.Closes #53, closes #54, closes #55, closes #56, closes #57, closes #58, closes #59, closes #61, closes #62, closes #63, closes #65, closes #66, closes #67, closes #69, closes #70, closes #71, closes #72, closes #73