Wednesday, September 30, 2026
HomeData ScienceAutoML and Hyperparameter Optimisation – Optuna, TPOT and Bayesian Search

AutoML and Hyperparameter Optimisation – Optuna, TPOT and Bayesian Search

Table of Content

📋 KEY INSIGHTS

  • Hyperparameter optimisation is one of the highest-leverage activities in the ML workflow — the difference between a poorly-tuned and well-tuned XGBoost model can exceed the difference between model families.
  • Grid Search exhaustively tries all combinations (expensive but reproducible). Random Search samples randomly (surprisingly effective). Bayesian optimisation (Optuna, Hyperopt) learns from past trials to focus on promising regions.
  • Optuna uses Tree-structured Parzen Estimators (TPE) by default — a Bayesian method that models P(params | good trial) and samples parameters that are likely to yield low loss.
  • Pruning eliminates unpromising trials early (after seeing the first 20% of training) — it dramatically reduces the total compute budget without sacrificing final model quality.
  • AutoML systems (TPOT, Auto-sklearn, H2O AutoML) go beyond hyperparameter tuning to also search over pipeline structure — which preprocessing, which model family, which ensembling.
  • Optuna studies are persistent (saved to SQLite or PostgreSQL) — you can stop an optimisation run and resume it later, or distribute it across multiple machines.

Hyperparameter tuning is where many data science projects are lost or won. The choice of learning rate, regularisation strength, tree depth, or number of attention heads can shift model performance by 10–20% — more than switching between model families. Yet manual tuning is time-consuming, irreproducible, and leaves the best configurations unexplored. Optuna, released by Preferred Networks in 2019, has become the standard tool for hyperparameter optimisation in the Python ML ecosystem. This guide covers the full Optuna workflow — from basic objective functions to distributed search, pruning, and integration with XGBoost, PyTorch, and scikit-learn — plus an introduction to full AutoML pipelines with TPOT and H2O.

Grid Search tries every combination in a predefined grid — if you have 5 values for each of 6 hyperparameters, that is 5^6 = 15,625 training runs. It scales exponentially with the number of parameters and is completely uninformed — it does not use the results of past trials to inform future ones. Random Search (Bergstra & Bengio, 2012) samples from the search space randomly, which sounds naïve but is surprisingly effective — because most hyperparameter spaces have low effective dimensionality (only a few parameters really matter), random search hits the good regions nearly as often as grid search while exploring far more of the space.

Bayesian Optimisation improves on both by treating hyperparameter optimisation as a sequential decision problem. After each trial, it updates a probabilistic surrogate model of the objective function (the loss landscape) and uses it to choose the next configuration most likely to improve on the best result found so far. Optuna’s TPE (Tree-structured Parzen Estimator) models P(hyperparams | good trial) and P(hyperparams | bad trial) separately, then samples configurations with a high ratio — focusing search on regions similar to past good trials.

import optuna
import numpy as np
from sklearn.datasets import make_classification
from sklearn.model_selection import cross_val_score, StratifiedKFold
from sklearn.ensemble import RandomForestClassifier
from xgboost import XGBClassifier
import warnings
warnings.filterwarnings('ignore')

X, y = make_classification(n_samples=5000, n_features=20, n_informative=10,
                            random_state=42)
cv = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)

# --- Basic Optuna objective for XGBoost ---
def xgb_objective(trial):
    params = {
        'n_estimators':       trial.suggest_int('n_estimators', 100, 800, step=50),
        'max_depth':          trial.suggest_int('max_depth', 3, 10),
        'learning_rate':      trial.suggest_float('learning_rate', 1e-4, 0.3, log=True),
        'subsample':          trial.suggest_float('subsample', 0.5, 1.0),
        'colsample_bytree':   trial.suggest_float('colsample_bytree', 0.4, 1.0),
        'reg_alpha':          trial.suggest_float('reg_alpha', 1e-6, 10.0, log=True),
        'reg_lambda':         trial.suggest_float('reg_lambda', 1e-6, 10.0, log=True),
        'min_child_weight':   trial.suggest_int('min_child_weight', 1, 10),
        'gamma':              trial.suggest_float('gamma', 0.0, 5.0),
        'eval_metric':        'logloss',
        'use_label_encoder':  False,
        'random_state':       42,
        'n_jobs':             -1,
    }
    model = XGBClassifier(**params, verbosity=0)
    scores = cross_val_score(model, X, y, cv=cv, scoring='roc_auc', n_jobs=-1)
    return scores.mean()

# Create and optimise study
optuna.logging.set_verbosity(optuna.logging.WARNING)
study = optuna.create_study(
    direction  = 'maximize',   # we are maximising ROC-AUC
    sampler    = optuna.samplers.TPESampler(seed=42),
    pruner     = optuna.pruners.MedianPruner(n_startup_trials=5, n_warmup_steps=3),
)
study.optimize(xgb_objective, n_trials=50, show_progress_bar=True)

print('
Best trial:')
print('  ROC-AUC:', round(study.best_value, 4))
print('  Params: ', study.best_params)
#Study cutout signage
Photo by Artem Beliaikin on Unsplash

Pruning eliminates trials that are unlikely to improve on the current best, based on intermediate values reported during training. For iterative algorithms (XGBoost, neural networks trained epoch-by-epoch), you can report the validation loss after each round and let Optuna stop bad trials early — dramatically reducing total compute with minimal sacrifice in final result quality.

# --- Pruning with XGBoost callbacks ---
import xgboost as xgb

def xgb_objective_with_pruning(trial):
    params = {
        'max_depth':        trial.suggest_int('max_depth', 3, 10),
        'learning_rate':    trial.suggest_float('learning_rate', 1e-3, 0.3, log=True),
        'subsample':        trial.suggest_float('subsample', 0.5, 1.0),
        'colsample_bytree': trial.suggest_float('colsample_bytree', 0.4, 1.0),
        'reg_alpha':        trial.suggest_float('reg_alpha', 1e-6, 1.0, log=True),
        'eval_metric':      'auc',
        'seed':             42,
    }
    from sklearn.model_selection import train_test_split
    X_tr, X_val, y_tr, y_val = train_test_split(X, y, test_size=0.2, random_state=42)
    dtrain = xgb.DMatrix(X_tr, label=y_tr)
    dval   = xgb.DMatrix(X_val, label=y_val)
    evals_result = {}

    pruning_callback = optuna.integration.XGBoostPruningCallback(trial, 'validation-auc')

    booster = xgb.train(
        params,
        dtrain,
        num_boost_round = 500,
        evals           = [(dval, 'validation')],
        early_stopping_rounds = 20,
        callbacks       = [pruning_callback],
        verbose_eval    = False,
        evals_result    = evals_result,
    )
    best_auc = max(evals_result['validation']['auc'])
    return best_auc

study_pruned = optuna.create_study(
    direction = 'maximize',
    sampler   = optuna.samplers.TPESampler(seed=42),
    pruner    = optuna.pruners.MedianPruner(n_startup_trials=5),
)
study_pruned.optimize(xgb_objective_with_pruning, n_trials=30)
print('Best pruned AUC:', round(study_pruned.best_value, 4))

# --- Persistent study: save to SQLite ---
# storage = 'sqlite:///optuna_studies.db'
# study_persistent = optuna.create_study(
#     study_name='xgb_optimisation', direction='maximize',
#     storage=storage, load_if_exists=True   # resume if it already exists
# )
# study_persistent.optimize(xgb_objective, n_trials=100)

# --- Parallel search on multiple machines ---
# Each machine runs:
# study = optuna.load_study(study_name='xgb_optimisation', storage=storage)
# study.optimize(xgb_objective, n_trials=25)   # all machines share the same study

Full AutoML — TPOT and H2O AutoML

Full AutoML goes beyond hyperparameter tuning to also select the model family and preprocessing pipeline. TPOT (Tree-based Pipeline Optimisation Tool) uses genetic algorithms to search over the space of sklearn pipelines — finding the best combination of scalers, feature selectors, and classifiers. H2O AutoML trains and stacks multiple algorithms (GBM, XGBoost, Deep Learning, Random Forest, GLM) and selects the best via cross-validation leaderboard. These tools are excellent for rapid baseline development and for teams without dedicated ML engineering resources.

# --- TPOT: genetic algorithm pipeline search ---
# pip install tpot
# from tpot import TPOTClassifier
# tpot = TPOTClassifier(
#     generations=10, population_size=50, cv=5,
#     random_state=42, verbosity=2, scoring='roc_auc',
#     max_time_mins=30,    # stop after 30 minutes
#     n_jobs=-1,
# )
# tpot.fit(X_train, y_train)
# print('TPOT best pipeline:', tpot.fitted_pipeline_)
# tpot.export('best_pipeline.py')   # exports a reproducible sklearn pipeline script

# --- H2O AutoML ---
# import h2o
# from h2o.automl import H2OAutoML
# h2o.init()
# train_h2o = h2o.H2OFrame(pd.DataFrame(np.c_[X, y], columns=...))
# aml = H2OAutoML(max_models=20, seed=42, max_runtime_secs=300)
# aml.train(x=feature_cols, y='target', training_frame=train_h2o)
# print(aml.leaderboard.head(10))

# --- Optuna: multi-objective optimisation (minimise loss AND training time) ---
def multi_objective(trial):
    n_estimators  = trial.suggest_int('n_estimators', 50, 500, step=50)
    max_depth     = trial.suggest_int('max_depth', 3, 10)
    model = RandomForestClassifier(n_estimators=n_estimators,
                                   max_depth=max_depth, random_state=42, n_jobs=-1)
    import time
    t0     = time.time()
    scores = cross_val_score(model, X, y, cv=3, scoring='roc_auc')
    elapsed = time.time() - t0
    return 1 - scores.mean(), elapsed   # minimise (1-AUC) and training time

study_mo = optuna.create_study(
    directions=['minimize', 'minimize'],   # Pareto front of accuracy vs speed
    sampler=optuna.samplers.NSGAIISampler(seed=42)
)
study_mo.optimize(multi_objective, n_trials=30, show_progress_bar=False)
pareto = study_mo.best_trials
print('Pareto-optimal trials:', len(pareto))
for t in pareto[:3]:
    print('  1-AUC =', round(t.values[0], 4), '  Time =', round(t.values[1], 2), 's',
          ' | Params:', t.params)
MethodSearch StrategyBest ForReproducible?
Grid SearchExhaustive gridSmall search spaces (<4 params)Yes
Random SearchRandom samplingLarge spaces, few important paramsYes (with seed)
Optuna TPEBayesian (TPE)General purpose, 5–20 paramsYes (with seed)
Optuna + PruningBayesian + early stoppingIterative models (XGB, neural nets)Yes
TPOTGenetic algorithmsPipeline structure searchPartial
H2O AutoMLEnsemble + stackingRapid baseline, non-specialistsYes

✦ SUMMARIZE THIS ARTICLE WITH AI

The gradient boosting models (XGBoost, LightGBM) whose hyperparameters Optuna optimises are explained in our Gradient Boosting Deep Dive. The broader model evaluation and cross-validation methodology behind the objective function is in our Model Evaluation guide. For neural network hyperparameter search (learning rate, architecture, optimiser) using Optuna with PyTorch, our Neural Network Architectures guide covers the setup. MLOps patterns for packaging and deploying the best trial as a production model are in our MLOps Interview Q&A.

Leave feedback about this

  • Rating

Durgesh Kekare
Durgesh Kekarehttps://www.dataexpertise.in
Durgesh Kekare is a data science educator and founder of DataExpertise.in. With expertise in Python, machine learning, and analytics, he helps 10,000+ learners break into data careers.

Latest Posts

List of Categories