Estimator API Basics

Definition

Every scikit-learn model/transformer follows the same Estimator API β€” a consistent contract of methods (fit, predict, transform, fit_transform, score) that makes any estimator swappable inside a Pipeline without changing surrounding code.


The Core Method Contract

MethodUsed byPurpose
.fit(X, y)all estimatorslearn parameters from training data
.predict(X)supervised modelsgenerate predictions on new data
.predict_proba(X)classifiersclass probability estimates
.transform(X)transformersapply a learned transformation
.fit_transform(X)transformersfit + transform in one call (more efficient for some transformers)
.score(X, y)supervised modelsdefault evaluation metric (RΒ² for regressors, accuracy for classifiers)
.get_params() / .set_params()all estimatorsinspect/modify hyperparameters
from sklearn.linear_model import LogisticRegression
 
model = LogisticRegression()      # 1. Instantiate with hyperparameters
model.fit(X_train, y_train)         # 2. Fit β€” learn from training data
predictions = model.predict(X_test)   # 3. Predict on new data
probabilities = model.predict_proba(X_test)  # class probabilities
accuracy = model.score(X_test, y_test)  # quick built-in evaluation

Estimator Types

graph TD
    A[Estimator] --> B[Predictor]
    A --> C[Transformer]
    B --> B1["Classifier
.predict / .predict_proba"]
    B --> B2["Regressor
.predict"]
    C --> C1["Preprocessor
.transform / .fit_transform"]

X and y Conventions

X.shape      # (n_samples, n_features) β€” always 2D, even with 1 feature
y.shape        # (n_samples,) β€” 1D for single-target problems
 
X = df[["feature1", "feature2"]]     # DataFrame or NumPy array both work
y = df["target"]                       # Series or 1D array

X with a single feature must still be 2D

df[["feature1"]] (double brackets, DataFrame) works; df["feature1"] (single brackets, Series) raises a shape error when passed directly to .fit(). Reshape a 1D array explicitly with .reshape(-1, 1) if needed.


Inspecting Fitted Attributes (trailing underscore convention)

model.coef_              # learned coefficients (linear models)
model.intercept_           # learned intercept
model.feature_importances_   # tree-based models
model.classes_                 # class labels seen during fit (classifiers)
model.n_features_in_             # number of features seen during fit

Trailing underscore = "learned from data"

Any attribute ending in _ (e.g. coef_, labels_) only exists after .fit() has been called β€” accessing it beforehand raises NotFittedError.


Hyperparameters vs Learned Parameters

model = LogisticRegression(C=1.0, penalty="l2", max_iter=1000)  # hyperparameters set at construction
model.get_params()          # dict of current hyperparameters
model.set_params(C=0.5)       # change a hyperparameter (before re-fitting)

random_state β€” Reproducibility

model = RandomForestClassifier(random_state=42)
train_test_split(X, y, random_state=42)

Set random_state everywhere

Any estimator or splitting function with inherent randomness (tree splits, bootstrapping, shuffling) accepts random_state β€” always set it for reproducible experiments and debugging.