Forecasting with XGBoost, LightGBM and other gradient boosting models¶
Gradient boosting models have gained popularity in the machine learning community due to their ability to achieve excellent results in a wide range of use cases, including both regression and classification. Although these models have traditionally been less common in forecasting, they have proven to be highly effective in this domain: for example, most of the top-performing solutions of the M5 forecasting competition were based on LightGBM. Some of the key advantages of using gradient boosting models for forecasting include:
The ease with which exogenous variables, in addition to autoregressive variables, can be incorporated into the model.
The ability to capture non-linear relationships between variables.
High scalability, which enables the models to handle large volumes of data.
There are several implementations of gradient boosting in Python, four of the most popular being XGBoost, LightGBM, scikit-learn HistGradientBoostingRegressor and CatBoost. All of these libraries follow the scikit-learn API, making them compatible with skforecast. In practice, the four implementations reach a similar accuracy on most problems, and the gain obtained from good features (in particular, exogenous variables) is usually larger than the difference between libraries. The choice is often driven by training speed, GPU support and the handling of categorical features.
✏️ Note
All of the gradient boosting libraries mentioned above (XGBoost, LightGBM, HistGradientBoostingRegressor and CatBoost) can handle categorical features natively. Since version 0.22.0, the machine learning forecasters of skforecast take care of the configuration: with the default categorical_features='auto', the non-numeric exogenous columns are detected, encoded internally and passed to the estimator as categorical, so no estimator-specific arguments (such as enable_categorical in XGBoost or cat_features in CatBoost) are needed. Detailed information can be found in the categorical features guide and in the following use cases:
- Forecasting time series with gradient boosting: Skforecast, XGBoost, LightGBM, Scikit-learn and CatBoost
- Forecasting time series with XGBoost
- Forecasting time series with LightGBM
- Forecasting energy demand with machine learning
XGBoost, LightGBM and CatBoost can benefit from GPU acceleration, which can significantly speed up the training process. To learn more about how to enable GPU acceleration, please refer to the following link: Skforecast GPU acceleration.
💡 Tip
Tree-based models, including decision trees, random forests and gradient boosting machines (GBMs), have limitations when it comes to extrapolation, i.e., making predictions or estimates beyond the range of observed data. This limitation becomes particularly critical when forecasting time series data with a trend. Because these models cannot predict values beyond the observed range during training, their predicted values will deviate from the underlying trend.
Several strategies have been proposed to address this challenge, with one of the most common approaches being the use of differentiation. The differentiation process involves computing the differences between consecutive observations in the time series. Instead of directly modeling the absolute values, the focus shifts to modeling the change between consecutive observations. Skforecast includes the differentiation parameter in ForecasterRecursive and the other machine learning forecasters to indicate that a differentiation process must be applied before training the model. It is worth noting that the entire differentiation process is automated and its effects are seamlessly reversed during the prediction phase. This ensures that the resulting forecast values are in the original scale of the time series data.
For more details on this topic, see the time series differentiation user guide and Modelling time series trend with tree-based models.
Libraries and data¶
# Libraries
# ==============================================================================
from xgboost import XGBRegressor
from lightgbm import LGBMRegressor
from sklearn.metrics import mean_absolute_error
from skforecast.datasets import fetch_dataset
from skforecast.preprocessing import RollingFeatures
from skforecast.recursive import ForecasterRecursive
The h2o_exog dataset, downloaded with fetch_dataset, contains the monthly expenditure on corticosteroid drugs of the Australian health system (column y) together with two simulated exogenous variables (exog_1 and exog_2). The index is a DatetimeIndex that already has a monthly frequency (MS, month start); skforecast requires a frequency when the index is a DatetimeIndex. The last 36 months are kept as the test set and the rest of the series is used for training.
# Download data
# ==============================================================================
data = fetch_dataset('h2o_exog')
╭─────────────────────────────────── h2o_exog ────────────────────────────────────╮ │ Description: │ │ Monthly expenditure ($AUD) on corticosteroid drugs that the Australian health │ │ system had between 1991 and 2008. Two additional variables (exog_1, exog_2) are │ │ simulated. │ │ │ │ Source: │ │ Hyndman R (2023). fpp3: Data for Forecasting: Principles and Practice (3rd │ │ Edition). http://pkg.robjhyndman.com/fpp3package/, │ │ https://github.com/robjhyndman/fpp3package, http://OTexts.com/fpp3. │ │ │ │ URL: │ │ https://raw.githubusercontent.com/skforecast/skforecast- │ │ datasets/main/data/h2o_exog.csv │ │ │ │ Shape: 195 rows x 3 columns │ ╰─────────────────────────────────────────────────────────────────────────────────╯
# Data preprocessing and train-test split
# ==============================================================================
data.index.name = 'date'
exog_features = ['exog_1', 'exog_2']
n_test = 36
data_train = data.iloc[:-n_test, :]
data_test = data.iloc[-n_test:, :]
print(
f"Train dates : {data_train.index.min().date()} --- "
f"{data_train.index.max().date()} (n={len(data_train)})"
)
print(
f"Test dates : {data_test.index.min().date()} --- "
f"{data_test.index.max().date()} (n={len(data_test)})"
)
Train dates : 1992-04-01 --- 2005-06-01 (n=159) Test dates : 2005-07-01 --- 2008-06-01 (n=36)
⚠ Warning
Time series must be split chronologically: the test set always contains the most recent observations. Random splits, such as train_test_split with shuffling, place future observations in the training set and lead to overly optimistic results.
In this guide, the values of exog_1 and exog_2 for the test period are passed to predict() as if they were known in advance. This is only valid for variables whose future values are available at prediction time (for example, calendar features or planned promotions). Otherwise, it is a form of lookahead bias and forecasts or estimates of those variables must be used instead.
LightGBM forecaster¶
A ForecasterRecursive is created with an LGBMRegressor, the last 8 values of the series as lags, a 7-month rolling mean built with RollingFeatures and the two exogenous variables of the dataset. The window size of the forecaster (the number of past observations needed to create the predictors) is 8, the maximum between the largest lag (8) and the size of the rolling window (7). After training, the forecaster predicts the 36 months of the test set (only the first 10 predictions are displayed). The predictions are recursive: each predicted value is used as a lag to predict the next step. Finally, the importance of each predictor is obtained with get_feature_importances().
# Create and fit forecaster
# ==============================================================================
forecaster_lgbm = ForecasterRecursive(
estimator = LGBMRegressor(random_state=123, verbose=-1),
lags = 8,
window_features = RollingFeatures(stats=['mean'], window_sizes=[7]),
)
forecaster_lgbm.fit(y=data_train['y'], exog=data_train[exog_features])
forecaster_lgbm
ForecasterRecursive
General Information
- Estimator: LGBMRegressor
- Lags: [1 2 3 4 5 6 7 8]
- Window features: ['roll_mean_7']
- Calendar features: None
- Window size: 8
- Series name: y
- Exogenous included: True
- Categorical features: auto
- Weight function included: False
- Differentiation order: None
- Drop NaN from series: False
- Creation date: 2026-10-08 19:49:12
- Last fit date: 2026-10-08 19:49:12
- Skforecast version: 0.26.0
- Python version: 3.14.3
- Forecaster id: None
Exogenous Variables
exog_1, exog_2
Data Transformations
- Transformer for y: None
- Transformer for exog: None
Training Information
- Training range: [Timestamp('1992-04-01 00:00:00'), Timestamp('2005-06-01 00:00:00')]
- Training index type: DatetimeIndex
- Training index frequency: MS
Estimator Parameters
-
{'boosting_type': 'gbdt', 'class_weight': None, 'colsample_bytree': 1.0, 'importance_type': 'split', 'learning_rate': 0.1, 'max_depth': -1, 'min_child_samples': 20, 'min_child_weight': 0.001, 'min_split_gain': 0.0, 'n_estimators': 100, 'n_jobs': None, 'num_leaves': 31, 'objective': None, 'random_state': 123, 'reg_alpha': 0.0, 'reg_lambda': 0.0, 'subsample': 1.0, 'subsample_for_bin': 200000, 'subsample_freq': 0, 'verbose': -1}
Fit Kwargs
-
{}
# Predict
# ==============================================================================
predictions_lgbm = forecaster_lgbm.predict(
steps = n_test,
exog = data_test[exog_features]
)
predictions_lgbm.head(10)
2005-07-01 0.941524 2005-08-01 0.927756 2005-09-01 1.065236 2005-10-01 1.088430 2005-11-01 1.083675 2005-12-01 1.168836 2006-01-01 0.967025 2006-02-01 0.751336 2006-03-01 0.816116 2006-04-01 0.801357 Freq: MS, Name: pred, dtype: float64
# Feature importances
# ==============================================================================
forecaster_lgbm.get_feature_importances()
| feature | importance | |
|---|---|---|
| 10 | exog_2 | 121 |
| 1 | lag_2 | 94 |
| 0 | lag_1 | 59 |
| 9 | exog_1 | 43 |
| 5 | lag_6 | 42 |
| 3 | lag_4 | 41 |
| 4 | lag_5 | 37 |
| 7 | lag_8 | 25 |
| 6 | lag_7 | 22 |
| 2 | lag_3 | 14 |
| 8 | roll_mean_7 | 11 |
XGBoost forecaster¶
The same configuration is used with an XGBRegressor. Only the estimator changes: lags, window features and exogenous variables are defined in the same way. The same applies to HistGradientBoostingRegressor (scikit-learn) or CatBoostRegressor: only the estimator passed to the forecaster needs to be replaced.
# Create and fit forecaster
# ==============================================================================
forecaster_xgb = ForecasterRecursive(
estimator = XGBRegressor(random_state=123),
lags = 8,
window_features = RollingFeatures(stats=['mean'], window_sizes=[7]),
)
forecaster_xgb.fit(y=data_train['y'], exog=data_train[exog_features])
forecaster_xgb
ForecasterRecursive
General Information
- Estimator: XGBRegressor
- Lags: [1 2 3 4 5 6 7 8]
- Window features: ['roll_mean_7']
- Calendar features: None
- Window size: 8
- Series name: y
- Exogenous included: True
- Categorical features: auto
- Weight function included: False
- Differentiation order: None
- Drop NaN from series: False
- Creation date: 2026-10-08 19:49:12
- Last fit date: 2026-10-08 19:49:12
- Skforecast version: 0.26.0
- Python version: 3.14.3
- Forecaster id: None
Exogenous Variables
exog_1, exog_2
Data Transformations
- Transformer for y: None
- Transformer for exog: None
Training Information
- Training range: [Timestamp('1992-04-01 00:00:00'), Timestamp('2005-06-01 00:00:00')]
- Training index type: DatetimeIndex
- Training index frequency: MS
Estimator Parameters
-
{'objective': 'reg:squarederror', 'base_score': None, 'booster': None, 'callbacks': None, 'colsample_bylevel': None, 'colsample_bynode': None, 'colsample_bytree': None, 'device': None, 'early_stopping_rounds': None, 'enable_categorical': False, 'eval_metric': None, 'feature_types': None, 'feature_weights': None, 'gamma': None, 'grow_policy': None, 'importance_type': None, 'interaction_constraints': None, 'learning_rate': None, 'max_bin': None, 'max_cat_threshold': None, 'max_cat_to_onehot': None, 'max_delta_step': None, 'max_depth': None, 'max_leaves': None, 'min_child_weight': None, 'missing': nan, 'monotone_constraints': None, 'multi_strategy': None, 'n_estimators': None, 'n_jobs': None, 'num_parallel_tree': None, 'random_state': 123, 'reg_alpha': None, 'reg_lambda': None, 'sampling_method': None, 'scale_pos_weight': None, 'subsample': None, 'tree_method': None, 'validate_parameters': None, 'verbosity': None}
Fit Kwargs
-
{}
# Predict
# ==============================================================================
predictions_xgb = forecaster_xgb.predict(
steps = n_test,
exog = data_test[exog_features]
)
predictions_xgb.head(10)
2005-07-01 0.881743 2005-08-01 0.985714 2005-09-01 1.070262 2005-10-01 1.099444 2005-11-01 1.116030 2005-12-01 1.206317 2006-01-01 0.977022 2006-02-01 0.679524 2006-03-01 0.740902 2006-04-01 0.742273 Freq: MS, Name: pred, dtype: float64
# Feature importances
# ==============================================================================
forecaster_xgb.get_feature_importances()
| feature | importance | |
|---|---|---|
| 0 | lag_1 | 0.299734 |
| 10 | exog_2 | 0.255852 |
| 1 | lag_2 | 0.105777 |
| 6 | lag_7 | 0.100981 |
| 4 | lag_5 | 0.083807 |
| 7 | lag_8 | 0.063430 |
| 9 | exog_1 | 0.055471 |
| 3 | lag_4 | 0.017550 |
| 5 | lag_6 | 0.010896 |
| 2 | lag_3 | 0.004139 |
| 8 | roll_mean_7 | 0.002362 |
The two importance tables are not directly comparable. By default, LightGBM reports the number of times each feature is used in a split (importance_type='split'), hence the integer values, whereas XGBoost reports the average gain of the splits that use each feature, normalized so that the importances add up to 1. Even so, both models agree on the main points: exog_2 and the most recent lags (lag_1 and lag_2) are among the most important predictors, while the rolling mean roll_mean_7 is the least important one.
These importances should be interpreted with caution. They are computed from a single model trained on a short series: 159 observations, of which the first 8 are only used to create the lags of the first training row, leaving 151 training rows. Therefore, small changes in the data or in the hyperparameters can change the ranking. In addition, split counts and gain-based importances tend to favor features with many distinct values, such as continuous variables, and they do not indicate the direction of the effect. For more reliable explanations, such as SHAP values, see the explainability user guide.
Model comparison¶
The predictions of both models are compared with the observed values of the test set using the mean absolute error (MAE), which is expressed in the same units as the series.
# Test error
# ==============================================================================
mae_lgbm = mean_absolute_error(data_test['y'], predictions_lgbm)
mae_xgb = mean_absolute_error(data_test['y'], predictions_xgb)
print(f"Mean of y in the test set : {data_test['y'].mean():.4f}")
print(f"Test MAE LightGBM : {mae_lgbm:.4f}")
print(f"Test MAE XGBoost : {mae_xgb:.4f}")
Mean of y in the test set : 0.9364 Test MAE LightGBM : 0.0925 Test MAE XGBoost : 0.0774
With the default hyperparameters, XGBoost obtains a lower test error than LightGBM (MAE of 0.0774 versus 0.0925). Since the average value of the series in the test set is about 0.94, this corresponds to an average error of roughly 8% and 10% of the level of the series, respectively.
This result should not be read as a general ranking of the two libraries, nor does it contradict the statement in the introduction that they reach a similar accuracy on most problems. Both models use the default hyperparameters of their library, and these defaults differ (for example, in tree depth, number of leaves or minimum number of observations per leaf), so part of the gap reflects those defaults rather than the algorithms themselves. In addition, the comparison is based on a single train-test split with only 36 observations. A more reliable evaluation repeats the process over several folds with backtesting_forecaster and tunes the hyperparameters and lags of each model, as explained in the hyperparameter tuning and lags selection user guide.