Ensemble Learning in Machine Learning: A Practical Guide

Multiple machine learning model paths converging into one prediction.
Multiple machine learning model paths converging into one prediction.

Ensemble learning combines predictions from multiple models into one result. The right method can reduce variance, capture patterns a single model misses, or exploit complementary model strengths. The gain depends on model diversity, sound validation, and an acceptable cost at inference time.

What is ensemble learning?

An ensemble in machine learning is a predictive system that combines two or more models, called base learners or base estimators, to produce one prediction. The combiner may be a majority vote, an average, a weighted sum, or another model trained to interpret the base predictions.

Consider three binary classifiers. If each has an independent error probability p, majority voting fails only when at least two classifiers fail:

P(ensemble error) = 3p²(1 - p) + p³ = p²(3 - 2p)

At p = 0.1, the ensemble error is 0.028. That result comes with strict assumptions. The classifiers must have equal error rates, and their errors must be independent. Real models trained on the same data rarely satisfy either condition. When all three fail on the same examples, voting repeats the error instead of correcting it.

The main ensemble methods differ in how they create useful variation and combine predictions.

MethodTraining patternCombination ruleGood starting point when
Voting or averagingModels train separatelyVote, mean, median, or weighted meanSeveral competent models make different errors
BaggingOne algorithm trains on resampled data, often in parallelVote or averageAn unstable model overfits small changes in the data
BoostingLearners train in sequenceWeighted additive predictionA simple learner underfits important structure
StackingDifferent models produce out-of-fold predictionsA meta-model learns the combinationModel families have complementary strengths

Random forests are a canonical bagging-style ensemble. Gradient-boosted trees, including XGBoost, are boosting ensembles. A soft-voting classifier and a stack can combine models from different algorithm families.

Why ensembles work

An ensemble helps when its members are individually useful and make different errors. Adding more copies of the same failure mode adds cost without adding much information.

For a simple averaging ensemble with M estimators, equal prediction variance σ², and average pairwise error correlation ρ, the variance of the mean is:

Var(mean prediction) = σ² [ρ + (1 - ρ) / M]

The second term shrinks as the ensemble grows. The correlated part, ρσ², remains. This explains why the 500th near-identical tree may add little, while a smaller set of genuinely different models can still improve a prediction. It also explains why diversity alone is insufficient. Random, inaccurate models can be diverse and still produce a poor ensemble.

Correlated models repeat one error while diverse models can combine into a correct result

Ensembles can improve a model in three related ways:

  • Reduce variance. Averaging unstable predictors makes the final prediction less sensitive to one training sample. Bagging deep decision trees is the standard example.
  • Reduce bias. Boosting adds learners that capture structure missed by the current ensemble, which can turn many simple rules into a flexible predictor.
  • Expand representation. A learned combination can represent relationships that no base learner captures alone. Stacking uses this idea across different model families.

These statistical, computational, and representational reasons have shaped ensemble research since the early work summarized in Dietterich's ensemble taxonomy.

Voting and averaging

Voting and averaging are the simplest ways to combine models because the base learners can be trained independently.

For classification, hard voting returns the class predicted by the most models. Soft voting averages predicted class probabilities and selects the class with the largest average probability. A weighted vote gives more influence to selected models.

Soft voting retains more information than hard labels, but its probabilities need care. A model that emits extreme, poorly calibrated probabilities can dominate a mean even when its ranking accuracy is good. Measure log loss or Brier score alongside a discrimination metric such as area under the receiver operating characteristic curve (ROC AUC). If decisions depend on the probability itself, inspect calibration and use a separate calibration split or cross-validation. Our confidence calibration guide explains the underlying problem in more detail.

For regression, an arithmetic mean is the usual starting point. A median or trimmed mean can be more robust when one model occasionally produces extreme predictions. Choose the aggregation rule using out-of-sample results and the actual business loss, rather than the training score.

Bagging and random forests

Bagging, short for bootstrap aggregating, trains copies of an estimator on bootstrap samples. Each sample draws n rows with replacement from a training set of n rows. One bootstrap sample contains about 63.2% unique training rows on average, leaving about 36.8% out of bag.

The original bagging paper showed why this procedure is especially useful for unstable learners. A deep decision tree can change substantially when a few training rows change. Training many such trees on different samples and averaging them reduces that variance. A stable learner may gain less from the same treatment, while still paying the training and inference cost.

A random forest adds feature randomness. At each split, a tree considers a random subset of features. This decorrelates the trees so that averaging has more room to help. Breiman's random forest paper formalized the relationship between individual tree strength and correlation. Our random forest guide goes deeper into the algorithm and its main controls.

Out-of-bag predictions also provide a convenient internal performance estimate. Each row can be predicted by trees that did not train on it. This is useful for iteration, but a final untouched test set is still valuable when model choice, thresholds, and hyperparameters have all been influenced by the same training data.

Bagging fits when you have a high-variance learner, enough compute for many members, and an inference budget that can absorb them. Its members can usually train in parallel.

Boosting and XGBoost

Boosting builds an additive ensemble in sequence. Each new learner focuses on weaknesses in the current combined prediction.

AdaBoost increases the influence of misclassified training examples, then combines learners with performance-based weights. Gradient boosting takes a broader approach: each stage fits the direction that reduces a chosen differentiable loss. With squared error, that direction resembles the current residual. With classification losses, the update occurs in the loss function's gradient space.

This sequential correction can reduce bias and fit complex nonlinear structure with shallow trees. It also creates tighter coupling between members. Training is less parallel than bagging, and excessive depth, learning rate, or iteration count can overfit. Noisy labels and outliers may attract repeated attention, depending on the loss.

XGBoost is an ensemble algorithm. It is an implementation of gradient-boosted decision trees with a regularized objective and engineering for sparse data and efficient tree construction, described in the XGBoost paper. Other gradient-boosted tree implementations make different choices around tree growth, categorical features, histogram construction, and distributed training. Treat the library and the ensemble method as separate decisions.

For boosting, tune the learning rate and number of iterations together. Use early stopping on a validation set that respects the deployment split, and keep the final test data untouched.

Stacking without data leakage

Stacking, or stacked generalization, trains a second-level model called a meta-learner to combine base predictions. A typical classification stack might combine logistic regression, a random forest, and a gradient-boosted tree, then use logistic regression as the meta-learner.

The critical implementation detail is out-of-fold prediction:

  1. Divide the training set into K folds.
  2. For each base model and fold, train on the other K - 1 folds and predict the held-out fold.
  3. Join those held-out predictions into one out-of-fold column per base model.
  4. Train the meta-learner on those columns.
  5. Refit each base model on all available training data for inference.

Training the meta-learner on predictions from models that already saw the same rows leaks target information. The meta-learner sees unrealistically clean signals and often fails on new data. Wolpert's stacked generalization paper frames stacking as a way to learn and correct the generalization biases of the base learners.

Use grouped, time-series, or stratified folds when the data requires them. The stacking folds must follow the same leakage boundaries as the final evaluation.

A compact ensemble learning example in Python

This scikit-learn example compares three base classifiers with a soft-voting ensemble on the breast cancer dataset. Scaling stays inside the logistic-regression pipeline, so each cross-validation fold fits preprocessing only on its training portion.

Python
from sklearn.datasets import load_breast_cancer from sklearn.ensemble import ( HistGradientBoostingClassifier, RandomForestClassifier, VotingClassifier, ) from sklearn.linear_model import LogisticRegression from sklearn.model_selection import RepeatedStratifiedKFold, cross_validate from sklearn.pipeline import make_pipeline from sklearn.preprocessing import StandardScaler X, y = load_breast_cancer(return_X_y=True) models = [ ("logistic", make_pipeline( StandardScaler(), LogisticRegression(max_iter=2000), )), ("forest", RandomForestClassifier( n_estimators=300, random_state=42, n_jobs=-1, )), ("boosting", HistGradientBoostingClassifier(random_state=42)), ] ensemble = VotingClassifier(estimators=models, voting="soft") cv = RepeatedStratifiedKFold( n_splits=5, n_repeats=3, random_state=42, ) for name, model in models + [("soft_vote", ensemble)]: scores = cross_validate( model, X, y, cv=cv, scoring=("roc_auc", "neg_log_loss"), n_jobs=-1, ) print( name, "ROC AUC:", scores["test_roc_auc"].mean(), "log loss:", -scores["test_neg_log_loss"].mean(), )

The ensemble is a candidate, not an automatic winner. Keep it only if the improvement is consistent across folds, meaningful for the chosen metric, and worth its extra inference cost. If you tune models or weights using cross-validation results, estimate final performance on a separate test set or use nested cross-validation.

The official scikit-learn ensemble guide covers current APIs for voting, bagging, forests, boosting, and stacking.

How to choose an ensemble method

Start from the failure pattern of a strong single-model baseline.

  1. Choose the evaluation split first. Use time-based splits for forecasting, group-based splits when entities repeat, and stratification when class balance matters. A random split cannot rescue a deployment mismatch.
  2. Diagnose bias, variance, and error overlap. Large train-to-validation gaps point toward variance. Consistently weak train and validation scores point toward bias or missing signal. Compare which rows each candidate gets wrong, rather than comparing one aggregate score alone.
  3. Match the method to the problem. Try bagging for an unstable learner, boosting for a learner that needs additive corrections, simple averaging for competent models with different errors, and stacking when a learned combiner has enough data to generalize.
  4. Validate the whole pipeline. Fit imputation, encoding, scaling, feature selection, calibration, and the ensemble inside each training fold. Generate stacking features out of fold.
  5. Benchmark the operating cost. Measure training time, model size, median and tail inference latency, memory, throughput, and failure handling. Include every base model and the combiner.
  6. Compare against the strongest single model. The ensemble must earn its extra code, compute, monitoring, and incident surface.

For classification, check class-level performance, probability calibration, and threshold behavior. For regression, inspect residuals across important slices and compare performance under the loss used by the application.

Advantages and limitations of ensemble learning

The main advantages are better generalization, lower sensitivity to one training sample, and the ability to combine complementary representations. Ensembles can also provide a smoother path from a solid baseline to higher predictive performance without designing a new model family.

The limitations are equally practical:

  • Training and inference use more compute, memory, and storage.
  • More components create more versioning, monitoring, and rollback work.
  • Explanation becomes harder because the prediction passes through several models.
  • Stacking and blending can leak data when prediction features are generated incorrectly.
  • Averaged probabilities can still be poorly calibrated.
  • Shared blind spots remain when every member learns from the same biased data or missing features.

A single model is often the better production choice when it meets the target, responds faster, is easier to explain, and is cheaper to operate. It is also safer when the available dataset is too small to estimate ensemble weights or meta-model behavior reliably.

Start with one credible baseline. Add the simplest ensemble that targets a measured failure mode. Keep it only when held-out performance and production constraints both improve.

Share

Subscribe

Sign up to our e-mail list to get the best of the Dasha blog sent directly to your inbox.

We use cookies for functional and analytical purposes. Please refer to our Privacy Policy for details.