Common Errors & Gotchas
Purpose
A troubleshooting reference for the mistakes and errors that come up most often when working with scikit-learn — many are conceptual (data leakage) rather than syntax errors, and are far more damaging because the code runs fine but produces misleadingly optimistic results.
Data Leakage — The #1 Silent Killer
# BAD — scaler sees the ENTIRE dataset, including test data, before splitting
scaler = StandardScaler()
X_scaled = scaler.fit_transform(X) # leaks test-set statistics into training
X_train, X_test, y_train, y_test = train_test_split(X_scaled, y)
# GOOD — split FIRST, fit scaler ONLY on training data
X_train, X_test, y_train, y_test = train_test_split(X, y)
scaler = StandardScaler()
X_train_scaled = scaler.fit_transform(X_train)
X_test_scaled = scaler.transform(X_test) # transform only, never fit againWhy this matters
Fitting any preprocessing step (scaler, imputer, feature selector, encoder) on data that includes the test set means the test set is no longer a true “unseen data” evaluation — reported performance will be optimistically biased and won’t reflect real-world generalization.
The fix: always use a Pipeline
Wrapping preprocessing + model in a Pipeline and passing the WHOLE pipeline to
train_test_split/cross_val_score/GridSearchCVmakes leakage structurally impossible — each fold refits preprocessing only on that fold’s training portion.
NotFittedError
model = LogisticRegression()
model.predict(X_test) # NotFittedError — .fit() was never calledAlways call
.fit()before.predict()/.transform(). If using a saved model, confirmjoblib.load()actually returned a fitted object (check for attributes ending in_, e.g.model.coef_).
Shape Mismatches — 1D vs 2D
model.fit(df["feature"], y) # ValueError — Series is 1D, sklearn expects 2D X
model.fit(df[["feature"]], y) # correct — double brackets keep it a DataFrame (2D)
model.fit(X, y_2d_array) # warning/error if y has shape (n,1) instead of (n,)
model.fit(X, y.values.ravel()) # flatten to 1D if neededCategorical Data Not Encoded
model.fit(X, y) # ValueError: could not convert string to floatMost sklearn estimators require purely numeric input. Encode categorical columns first — see 02-Preprocessing-Scaling for
OneHotEncoder/OrdinalEncoder, ideally within a ColumnTransformer.
Forgetting handle_unknown="ignore" on OneHotEncoder
ohe = OneHotEncoder() # default: errors on unseen categories at transform time
ohe = OneHotEncoder(handle_unknown="ignore") # unseen categories -> all-zero row instead of crashingIf production/test data contains a category not seen during training (e.g. a new city name), the default
OneHotEncoderraises an error at.transform()time —handle_unknown="ignore"avoids this by encoding unseen categories as all zeros.
Mismatched Feature Order/Names Between Train and Predict
model.fit(X_train[["age", "income"]], y_train)
model.predict(X_test[["income", "age"]]) # WRONG ORDER — silently produces garbage predictions in older sklearnModern sklearn (≥1.0) checks feature names
If
X_trainwas a DataFrame, sklearn storesfeature_names_in_and will raise a warning/error if.predict()receives differently-ordered or differently-named columns — but always double-check when working with raw NumPy arrays, where no such check exists.
Class Imbalance Ignored
model = LogisticRegression() # default: treats all classes equally, biased toward majority class
model = LogisticRegression(class_weight="balanced") # auto-adjusts for imbalanceEvaluate with precision/recall/F1 (see 11-Model-Evaluation-Metrics), not just accuracy, on imbalanced datasets.
ConvergenceWarning
# LogisticRegression, MLPClassifier, etc. — didn't converge within max_iter
model = LogisticRegression(max_iter=1000) # increase from the default (100)Also consider scaling features first — unscaled data often causes slow/failed convergence in gradient-based solvers.
Refitting the Vectorizer/Scaler on Test Data
tfidf.fit_transform(X_test_text) # WRONG — creates a DIFFERENT vocabulary than training
tfidf.transform(X_test_text) # CORRECT — reuses the training-fitted vocabularyCross-Validation Score Looks “Too Good”
Suspiciously high CV scores usually mean leakage
Common causes: preprocessing fit before splitting, duplicate rows across train/test, a feature that’s a proxy for or directly derived from the target (e.g. accidentally including a post-outcome column), or time-series data split randomly instead of chronologically (use
TimeSeriesSplit).
Comparing Floats / Random State Confusion
model1 = RandomForestClassifier(random_state=42)
model2 = RandomForestClassifier(random_state=42)
# Same random_state + same data + same params = reproducible identical results — useful for debugging