Feature selection¶
Feature selection is the process of selecting a subset of relevant features (variables, predictors) for use in model construction. Feature selection techniques are used for several reasons: to simplify models to make them easier to interpret, to reduce training time, to avoid the curse of dimensionality, to improve generalization by reducing overfitting (formally, variance reduction), and others.
Skforecast is compatible with the feature selection methods implemented in scikit-learn and feature-engine libraries. There are several methods for feature selection, but the most common are:
Feature selection based on threshold (SelectFromModel)
SelectFromModel can be used along with any estimator that has a coef_ or feature_importances_ attribute after fitting. Features are considered unimportant and removed if the corresponding coef_ or feature_importances_ values are below the given threshold parameter. In addition to specifying the threshold numerically, there are built-in heuristics for finding a threshold using a string argument. Available heuristics are 'mean', 'median' and float multiples of these, such as '0.1*mean'.
This method is very fast compared to the others because it does not require any additional model training. However, it does not evaluate the impact of feature removal on the model. It is often used for an initial selection before applying another more computationally expensive feature selection method.
Recursive feature elimination
Given an external estimator that assigns weights to features (e.g., the coefficients of a linear model), the goal of recursive feature elimination (RFE) is to select features by recursively considering smaller and smaller sets of features. First, the estimator is trained on the initial set of features, and the importance of each feature is obtained either by a specific attribute (such as coef_, feature_importances_) or by a callable. Then, the least important features are pruned from the current set of features. This procedure is repeated recursively on the pruned set until the desired number of features to select is eventually reached. RFECV performs RFE in a cross-validation loop to find the optimal number of features.
Sequential Feature Selection
Sequential Feature Selection (SFS) can be either forward or backward, with the direction parameter controlling whether forward or backward SFS is used.
Forward-SFS is a greedy procedure that iteratively finds the best new feature to add to the set of selected features. It starts with zero features and finds the one that maximizes a cross-validated score when an estimator is trained on that single feature. Once this first feature is selected, the procedure is repeated, adding one new feature to the set of selected features. The procedure stops when the desired number of selected features is reached, as determined by the
n_features_to_selectparameter.Backward-SFS follows the same idea, but works in the opposite direction. Instead of starting with no features and greedily adding features, it starts with all features and greedily removes features from the set.
In general, forward and backward selection do not produce equivalent results. Also, one can be much faster than the other depending on the requested number of selected features: if we have 10 features and ask for 7 selected features, forward selection would need to perform 7 iterations while backward selection would only need to perform 3.
SFS does not require the underlying model to expose a coef_ or feature_importances_ attribute. However, it may be slower compared to the other approaches, considering that more models have to be evaluated. For example in backward selection, the iteration going from features to features using k-fold cross-validation requires fitting models to be evaluated.
Minimum Redundancy Maximum Relevance (MRMR)
Minimum Redundancy Maximum Relevance (MRMR) is a filter-based feature selection method that aims to identify a subset of features that are both highly relevant to the target variable and minimally redundant with respect to each other. Relevance is typically measured using mutual information between each feature and the target, while redundancy is assessed via the mutual information between pairs of features. By optimizing both criteria, MRMR helps reduce overfitting and improve model interpretability, especially in high-dimensional settings. The MRMR class from feature-engine can be used to implement this method.
Choosing a method
The following table summarizes the main characteristics of each method to help choose the most appropriate one.
| Method | Type | Relative cost | Needs coef_ or feature_importances_ |
When to use it |
|---|---|---|---|---|
SelectFromModel |
Embedded | Very low (one model fit) | Yes | Quick first pass to discard clearly irrelevant features, or as the first step of a combined strategy. |
RFECV |
Wrapper | Medium (one fit per elimination step and fold) | Yes | General purpose choice when the number of features to keep is unknown. |
SequentialFeatureSelector |
Wrapper | High (one fit per candidate feature, step and fold) | No | Estimators without feature importances, or when a small number of features is wanted. |
MRMR |
Filter | Low with mutual information methods such as method='MIQ' (no model fit) |
No | Many redundant features (for example, highly correlated lags or engineered features), or to obtain a model-agnostic ranking. |
💡 Tip
Feature selection is a powerful tool for improving the performance of machine learning models. However, it is computationally expensive and can be time-consuming. Since the goal is to find the best subset of features, not the best model, it is not necessary to use the entire data set or a highly complex model. Instead, it is recommended to use a small subset of the data and a simple model. Once the best subset of features has been identified, the model can then be trained using the entire dataset and a more complex configuration.
For example, in this use case, the model is an LGBMRegressor with 900 trees and a maximum depth of 7. However, to find the best subset of features, only 100 trees (50 for the more expensive SequentialFeatureSelector) and a maximum depth of 5 are used.
✏️ Note
For autoregressive forecasters, the natural precursor to feature selection is identifying a sensible set of candidate lags from the data itself. The Autocorrelation analysis and lag selection guide shows how to read the ACF and PACF to propose those candidates. The methods described below are then used to prune the redundant ones and, optionally, to select among the exogenous features.
Feature selection with skforecast¶
The select_features and select_features_multiseries functions select the best subset of features (autoregressive, exogenous and calendar) using any selector that implements the fit and get_feature_names_out methods, such as those in sklearn.feature_selection or MRMR from feature-engine. The main parameters are:
forecaster:ForecasterRecursiveorForecasterDirect(select_features),ForecasterRecursiveMultiSeriesorForecasterDirectMultiVariate(select_features_multiseries). For direct forecasters, only the features of the first step are used.selector: feature selector. It is fitted in place, so it can be inspected after the call.y(orseries) andexog: training data.select_only: group(s) of features evaluated by the selector ('autoreg','exog','calendar', a list of them, orNonefor all). The other groups are kept unchanged.force_inclusion: features that are always kept, as a list of names or a regular expression.subsampleandrandom_state: proportion of rows used for the selection and the seed of the random sampling. Rows are sampled without replacement and keep their time order, so inselect_featuresaTimeSeriesSplitcan be passed as thecvof the selector. Even so, the scores computed inside the selector are not an estimate of the multi-step forecasting error: validate the selection with backtesting.
They return four lists: selected_lags, selected_window_features, selected_exog and selected_calendar_features. See the API reference for all the details.
Libraries and data¶
# Libraries
# ==============================================================================
import warnings
import numpy as np
import pandas as pd
from lightgbm import LGBMRegressor
from sklearn.feature_selection import RFECV
from sklearn.feature_selection import SequentialFeatureSelector
from sklearn.feature_selection import SelectFromModel
from sklearn.model_selection import ShuffleSplit
from feature_engine.selection import MRMR
from skforecast.datasets import fetch_dataset
from skforecast.preprocessing import RollingFeatures
from skforecast.preprocessing import CalendarFeatures
from skforecast.recursive import ForecasterRecursive
from skforecast.recursive import ForecasterRecursiveMultiSeries
from skforecast.model_selection import TimeSeriesFold
from skforecast.model_selection import backtesting_forecaster
from skforecast.feature_selection import select_features
from skforecast.feature_selection import select_features_multiseries
# Warnings
# ==============================================================================
# Some scikit-learn selectors fit the estimator with feature names and score it
# on numpy arrays, which makes LightGBM warn about missing feature names.
warnings.filterwarnings(
"ignore",
message="X does not have valid feature names, but .* was fitted with feature names"
)
# MRMR ('MIQ') divides relevance by redundancy, which warns when the redundancy
# of a feature is zero.
warnings.filterwarnings(
"ignore",
category=RuntimeWarning,
module="feature_engine.selection.mrmr"
)
# Download data
# ==============================================================================
data = fetch_dataset(name="bike_sharing_extended_features")
data.head(3)
╭──────────────────────── bike_sharing_extended_features ─────────────────────────╮ │ Description: │ │ Hourly usage of the bike share system in the city of Washington D.C. during the │ │ years 2011 and 2012. In addition to the number of users per hour, the dataset │ │ was enriched by introducing supplementary features. Addition includes calendar- │ │ based variables (day of the week, hour of the day, month, etc.), indicators for │ │ sunlight, incorporation of rolling temperature averages, and the creation of │ │ polynomial features generated from variable pairs. All cyclic variables are │ │ encoded using sine and cosine functions to ensure accurate representation. │ │ │ │ Source: │ │ Fanaee-T,Hadi. (2013). Bike Sharing Dataset. UCI Machine Learning Repository. │ │ https://doi.org/10.24432/C5W894. │ │ │ │ URL: │ │ https://raw.githubusercontent.com/skforecast/skforecast- │ │ datasets/main/data/bike_sharing_extended_features.csv │ │ │ │ Shape: 17352 rows x 90 columns │ ╰─────────────────────────────────────────────────────────────────────────────────╯
| users | weather | month_sin | month_cos | week_of_year_sin | week_of_year_cos | week_day_sin | week_day_cos | hour_day_sin | hour_day_cos | ... | temp_roll_mean_1_day | temp_roll_mean_7_day | temp_roll_max_1_day | temp_roll_min_1_day | temp_roll_max_7_day | temp_roll_min_7_day | holiday_previous_day | holiday_next_day | temp | holiday | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| date_time | |||||||||||||||||||||
| 2011-01-08 00:00:00 | 25.0 | mist | 0.5 | 0.866025 | 0.120537 | 0.992709 | -0.781832 | 0.62349 | 0.258819 | 0.965926 | ... | 8.063334 | 10.127976 | 9.02 | 6.56 | 18.86 | 4.92 | 0.0 | 0.0 | 7.38 | 0.0 |
| 2011-01-08 01:00:00 | 16.0 | mist | 0.5 | 0.866025 | 0.120537 | 0.992709 | -0.781832 | 0.62349 | 0.500000 | 0.866025 | ... | 8.029166 | 10.113334 | 9.02 | 6.56 | 18.86 | 4.92 | 0.0 | 0.0 | 7.38 | 0.0 |
| 2011-01-08 02:00:00 | 16.0 | mist | 0.5 | 0.866025 | 0.120537 | 0.992709 | -0.781832 | 0.62349 | 0.707107 | 0.707107 | ... | 7.995000 | 10.103572 | 9.02 | 6.56 | 18.86 | 4.92 | 0.0 | 0.0 | 7.38 | 0.0 |
3 rows × 90 columns
# Data selection (reduce data size to speed up the example)
# ==============================================================================
data = data.drop(columns="weather")
data = data.loc["2012-01-01 00:00:00":]
# Split data into train and test
# ==============================================================================
# Feature selection is performed only with the training data. The test data is
# kept aside to evaluate, with backtesting, the forecaster trained with the
# selected features.
end_train = "2012-11-30 23:00:00"
data_train = data.loc[:end_train, :].copy()
data_test = data.loc[end_train:, :].iloc[1:].copy()
print(f"Dates train : {data_train.index.min()} --- {data_train.index.max()} (n={len(data_train)})")
print(f"Dates test : {data_test.index.min()} --- {data_test.index.max()} (n={len(data_test)})")
Dates train : 2012-01-01 00:00:00 --- 2012-11-30 23:00:00 (n=8040) Dates test : 2012-12-01 00:00:00 --- 2012-12-30 23:00:00 (n=720)
Create forecaster¶
A forecasting model is created to predict the number of users using the last 48 values (last two days) as lags, window features, and the exogenous features available in the dataset.
# Create forecaster
# ==============================================================================
window_features = RollingFeatures(
stats = ['mean', 'mean', 'sum'],
window_sizes = [24, 48, 24]
)
forecaster = ForecasterRecursive(
estimator = LGBMRegressor(
n_estimators = 900,
random_state = 15926,
max_depth = 7,
verbose = -1
),
lags = 48,
window_features = window_features
)
Feature selection with Recursive Feature Elimination (RFECV)¶
Selection of autoregressive and exogenous features¶
By default (select_only=None), the select_features function evaluates all the features of the forecaster at the same time: autoregressive (lags and window features), exogenous and, if any, calendar features.
RFECV removes one feature at each iteration (step=1) and uses cross-validation to find the best number of features. The argument min_features_to_select sets the minimum number of features that RFECV can end up with; it is a lower bound, not the final number. Since the forecaster has 139 candidate features, a lower bound of 25 prevents the selector from keeping too few features and reduces the number of iterations.
# Feature selection (autoregressive and exog) with scikit-learn RFECV
# ==============================================================================
estimator = LGBMRegressor(n_estimators=100, max_depth=5, random_state=15926, verbose=-1)
selector = RFECV(estimator=estimator, step=1, cv=3, min_features_to_select=25)
selected_lags, selected_window_features, selected_exog, selected_calendar_features = select_features(
forecaster = forecaster,
selector = selector,
y = data_train["users"],
exog = data_train.drop(columns="users"),
select_only = None,
force_inclusion = None,
subsample = 0.5,
random_state = 123,
verbose = True,
)
Feature selection (RFECV)
-------------------------
Total number of records available: 7992
Total number of records used for feature selection: 3996
Number of features available: 139
Lags (n=48)
Window features (n=3)
Exog (n=88)
Calendar (n=0)
Number of features selected: 40
Lags (n=25) : [1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 13, 14, 16, 21, 23, 24, 25, 29, 34, 36, 41, 42, 46, 48]
Window features (n=1) : ['roll_mean_48']
Exog (n=14) : ['hour_day_sin', 'hour_day_cos', 'poly_week_of_year_sin__hour_day_sin', 'poly_week_of_year_sin__sunset_hour_sin', 'poly_week_of_year_cos__week_day_sin', 'poly_week_of_year_cos__week_day_cos', 'poly_week_of_year_cos__hour_day_sin', 'poly_week_day_sin__hour_day_sin', 'poly_week_day_sin__hour_day_cos', 'poly_week_day_cos__hour_day_cos', 'poly_hour_day_sin__hour_day_cos', 'temp_roll_mean_1_day', 'temp_roll_mean_7_day', 'temp']
Calendar (n=0) : []
Then, a new forecaster is created with the selected features. Selected lags and exogenous variables can be used directly, while window features require a new RollingFeatures object with only the selected ones. In this case, only roll_mean_48 has been kept, so the new RollingFeatures object computes a single rolling mean of 48 hours.
# Create a forecaster with the selected features
# ==============================================================================
forecaster_selected = ForecasterRecursive(
estimator = LGBMRegressor(
n_estimators = 900,
random_state = 15926,
max_depth = 7,
verbose = -1
),
lags = selected_lags,
window_features = RollingFeatures(stats=['mean'], window_sizes=[48])
)
Evaluate the selected features¶
The selection is useful only if the forecaster with the selected features predicts at least as well as the one with all the features. This is checked with backtesting on the test set (December 2012), which has not been used during the selection.
# Backtesting: all features vs selected features
# ==============================================================================
cv = TimeSeriesFold(
steps=24, initial_train_size=len(data_train), refit=False
)
metric_all, _ = backtesting_forecaster(
forecaster = forecaster,
y = data["users"],
exog = data.drop(columns="users"),
cv = cv,
metric = "mean_absolute_error",
show_progress = False
)
metric_selected, _ = backtesting_forecaster(
forecaster = forecaster_selected,
y = data["users"],
exog = data[selected_exog],
cv = cv,
metric = "mean_absolute_error",
show_progress = False
)
print(f"MAE all features : {metric_all.iloc[0, 0]:.3f}")
print(f"MAE selected features : {metric_selected.iloc[0, 0]:.3f}")
MAE all features : 44.874 MAE selected features : 44.273
With less than a third of the features (40 out of 139), the forecaster achieves a slightly lower error (44.27 vs 44.87). From here on, this forecaster is used in the rest of the examples, except the next one.
# Forecaster used in the next examples
# ==============================================================================
forecaster_all = forecaster # all candidate features, used in the next example
forecaster = forecaster_selected
Selection on a subset of features¶
With select_only, the selector evaluates only some groups of features. The other groups are kept unchanged and returned as they are:
'autoreg': lags and window features.'exog': exogenous features.'calendar': calendar features (see Selection of calendar features).A list with any combination of the above.
The autoregressive features of the selected forecaster have already been chosen in the previous section, so the 'autoreg' example starts from the forecaster with all the candidate features (forecaster_all). Since fewer features are evaluated in the following examples, min_features_to_select is lowered to 10.
# Feature selection (only autoregressive) with scikit-learn RFECV
# ==============================================================================
estimator = LGBMRegressor(n_estimators=100, max_depth=5, random_state=15926, verbose=-1)
selector = RFECV(estimator=estimator, step=1, cv=3, min_features_to_select=10)
selected_lags, selected_window_features, selected_exog, selected_calendar_features = select_features(
forecaster = forecaster_all,
selector = selector,
y = data_train["users"],
exog = data_train.drop(columns="users"),
select_only = 'autoreg',
subsample = 0.5,
random_state = 123,
verbose = True,
)
Feature selection (RFECV)
-------------------------
Total number of records available: 7992
Total number of records used for feature selection: 3996
Number of features available: 139
Lags (n=48)
Window features (n=3)
Exog (n=88)
Calendar (n=0)
Number of features selected: 130
Lags (n=41) : [1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 31, 33, 34, 35, 36, 39, 41, 42, 43, 44, 46, 47, 48]
Window features (n=1) : ['roll_mean_48']
Exog (n=88) : ['month_sin', 'month_cos', 'week_of_year_sin', 'week_of_year_cos', 'week_day_sin', 'week_day_cos', 'hour_day_sin', 'hour_day_cos', 'sunrise_hour_sin', 'sunrise_hour_cos', 'sunset_hour_sin', 'sunset_hour_cos', 'poly_month_sin__month_cos', 'poly_month_sin__week_of_year_sin', 'poly_month_sin__week_of_year_cos', 'poly_month_sin__week_day_sin', 'poly_month_sin__week_day_cos', 'poly_month_sin__hour_day_sin', 'poly_month_sin__hour_day_cos', 'poly_month_sin__sunrise_hour_sin', 'poly_month_sin__sunrise_hour_cos', 'poly_month_sin__sunset_hour_sin', 'poly_month_sin__sunset_hour_cos', 'poly_month_cos__week_of_year_sin', 'poly_month_cos__week_of_year_cos', 'poly_month_cos__week_day_sin', 'poly_month_cos__week_day_cos', 'poly_month_cos__hour_day_sin', 'poly_month_cos__hour_day_cos', 'poly_month_cos__sunrise_hour_sin', 'poly_month_cos__sunrise_hour_cos', 'poly_month_cos__sunset_hour_sin', 'poly_month_cos__sunset_hour_cos', 'poly_week_of_year_sin__week_of_year_cos', 'poly_week_of_year_sin__week_day_sin', 'poly_week_of_year_sin__week_day_cos', 'poly_week_of_year_sin__hour_day_sin', 'poly_week_of_year_sin__hour_day_cos', 'poly_week_of_year_sin__sunrise_hour_sin', 'poly_week_of_year_sin__sunrise_hour_cos', 'poly_week_of_year_sin__sunset_hour_sin', 'poly_week_of_year_sin__sunset_hour_cos', 'poly_week_of_year_cos__week_day_sin', 'poly_week_of_year_cos__week_day_cos', 'poly_week_of_year_cos__hour_day_sin', 'poly_week_of_year_cos__hour_day_cos', 'poly_week_of_year_cos__sunrise_hour_sin', 'poly_week_of_year_cos__sunrise_hour_cos', 'poly_week_of_year_cos__sunset_hour_sin', 'poly_week_of_year_cos__sunset_hour_cos', 'poly_week_day_sin__week_day_cos', 'poly_week_day_sin__hour_day_sin', 'poly_week_day_sin__hour_day_cos', 'poly_week_day_sin__sunrise_hour_sin', 'poly_week_day_sin__sunrise_hour_cos', 'poly_week_day_sin__sunset_hour_sin', 'poly_week_day_sin__sunset_hour_cos', 'poly_week_day_cos__hour_day_sin', 'poly_week_day_cos__hour_day_cos', 'poly_week_day_cos__sunrise_hour_sin', 'poly_week_day_cos__sunrise_hour_cos', 'poly_week_day_cos__sunset_hour_sin', 'poly_week_day_cos__sunset_hour_cos', 'poly_hour_day_sin__hour_day_cos', 'poly_hour_day_sin__sunrise_hour_sin', 'poly_hour_day_sin__sunrise_hour_cos', 'poly_hour_day_sin__sunset_hour_sin', 'poly_hour_day_sin__sunset_hour_cos', 'poly_hour_day_cos__sunrise_hour_sin', 'poly_hour_day_cos__sunrise_hour_cos', 'poly_hour_day_cos__sunset_hour_sin', 'poly_hour_day_cos__sunset_hour_cos', 'poly_sunrise_hour_sin__sunrise_hour_cos', 'poly_sunrise_hour_sin__sunset_hour_sin', 'poly_sunrise_hour_sin__sunset_hour_cos', 'poly_sunrise_hour_cos__sunset_hour_sin', 'poly_sunrise_hour_cos__sunset_hour_cos', 'poly_sunset_hour_sin__sunset_hour_cos', 'temp_roll_mean_1_day', 'temp_roll_mean_7_day', 'temp_roll_max_1_day', 'temp_roll_min_1_day', 'temp_roll_max_7_day', 'temp_roll_min_7_day', 'holiday_previous_day', 'holiday_next_day', 'temp', 'holiday']
Calendar (n=0) : []
# Feature selection (only exog) with scikit-learn RFECV
# ==============================================================================
estimator = LGBMRegressor(n_estimators=100, max_depth=5, random_state=15926, verbose=-1)
selector = RFECV(estimator=estimator, step=1, cv=3, min_features_to_select=10)
selected_lags, selected_window_features, selected_exog, selected_calendar_features = select_features(
forecaster = forecaster,
selector = selector,
y = data_train["users"],
exog = data_train.drop(columns="users"),
select_only = 'exog',
subsample = 0.5,
random_state = 123,
verbose = True,
)
Feature selection (RFECV)
-------------------------
Total number of records available: 7992
Total number of records used for feature selection: 3996
Number of features available: 114
Lags (n=25)
Window features (n=1)
Exog (n=88)
Calendar (n=0)
Number of features selected: 40
Lags (n=25) : [1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 13, 14, 16, 21, 23, 24, 25, 29, 34, 36, 41, 42, 46, 48]
Window features (n=1) : ['roll_mean_48']
Exog (n=14) : ['hour_day_sin', 'hour_day_cos', 'poly_month_sin__week_of_year_cos', 'poly_week_of_year_sin__week_day_sin', 'poly_week_of_year_sin__week_day_cos', 'poly_week_of_year_sin__hour_day_sin', 'poly_week_day_sin__hour_day_sin', 'poly_week_day_sin__hour_day_cos', 'poly_week_day_cos__hour_day_sin', 'poly_week_day_cos__hour_day_cos', 'poly_hour_day_sin__hour_day_cos', 'temp_roll_mean_1_day', 'temp_roll_mean_7_day', 'temp']
Calendar (n=0) : []
Force selection of specific features¶
The force_inclusion argument keeps certain features regardless of the decision of the selector. It accepts a list of names or a regular expression.
A common case is cyclical encoding: the sine and cosine of a feature are only meaningful together, but the selector evaluates them independently and can keep one without the other. In the previous example, RFECV discarded both month_sin and month_cos. To keep the yearly seasonality anyway, the regular expression "^month_" forces both components.
# Feature selection (only exog) with scikit-learn RFECV forcing the inclusion
# of both components of the cyclical month
# ==============================================================================
estimator = LGBMRegressor(n_estimators=100, max_depth=5, random_state=15926, verbose=-1)
selector = RFECV(estimator=estimator, step=1, cv=3, min_features_to_select=10)
selected_lags, selected_window_features, selected_exog, selected_calendar_features = select_features(
forecaster = forecaster,
selector = selector,
y = data_train["users"],
exog = data_train.drop(columns="users"),
select_only = 'exog',
force_inclusion = "^month_",
subsample = 0.5,
random_state = 123,
verbose = True,
)
Feature selection (RFECV)
-------------------------
Total number of records available: 7992
Total number of records used for feature selection: 3996
Number of features available: 114
Lags (n=25)
Window features (n=1)
Exog (n=88)
Calendar (n=0)
Number of features selected: 42
Lags (n=25) : [1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 13, 14, 16, 21, 23, 24, 25, 29, 34, 36, 41, 42, 46, 48]
Window features (n=1) : ['roll_mean_48']
Exog (n=16) : ['month_sin', 'month_cos', 'hour_day_sin', 'hour_day_cos', 'poly_month_sin__week_of_year_cos', 'poly_week_of_year_sin__week_day_sin', 'poly_week_of_year_sin__week_day_cos', 'poly_week_of_year_sin__hour_day_sin', 'poly_week_day_sin__hour_day_sin', 'poly_week_day_sin__hour_day_cos', 'poly_week_day_cos__hour_day_sin', 'poly_week_day_cos__hour_day_cos', 'poly_hour_day_sin__hour_day_cos', 'temp_roll_mean_1_day', 'temp_roll_mean_7_day', 'temp']
Calendar (n=0) : []
# Check the decision of RFECV and the final selection
# ==============================================================================
print("month_cos selected by RFECV :", "month_cos" in selector.get_feature_names_out())
print("month_cos in selected_exog :", "month_cos" in selected_exog)
month_cos selected by RFECV : False month_cos in selected_exog : True
Feature selection with Sequential Feature Selection (SFS)¶
Sequential Feature Selection is a robust method for selecting features, but it is computationally expensive: at each iteration, one model is trained for every remaining candidate feature (and every cross-validation fold). Several strategies can be combined to reduce its computational cost:
Use a single validation split to evaluate each candidate model instead of the default 5-fold cross-validation.
Use a lighter estimator. In this example, 50 trees are used instead of 100.
Reduce
subsampleandn_features_to_select, since the number of models to train grows with both.
# Feature selection (only exog) with scikit-learn SequentialFeatureSelector
# ==============================================================================
estimator = LGBMRegressor(n_estimators=50, max_depth=5, random_state=15926, verbose=-1)
selector = SequentialFeatureSelector(
estimator = estimator,
n_features_to_select = 15,
direction = "forward",
cv = ShuffleSplit(n_splits=1, test_size=0.3, random_state=951),
scoring = "neg_mean_absolute_error",
)
selected_lags, selected_window_features, selected_exog, selected_calendar_features = select_features(
forecaster = forecaster,
selector = selector,
y = data_train["users"],
exog = data_train.drop(columns="users"),
select_only = 'exog',
subsample = 0.2,
random_state = 123,
verbose = True,
)
Feature selection (SequentialFeatureSelector)
---------------------------------------------
Total number of records available: 7992
Total number of records used for feature selection: 1598
Number of features available: 114
Lags (n=25)
Window features (n=1)
Exog (n=88)
Calendar (n=0)
Number of features selected: 41
Lags (n=25) : [1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 13, 14, 16, 21, 23, 24, 25, 29, 34, 36, 41, 42, 46, 48]
Window features (n=1) : ['roll_mean_48']
Exog (n=15) : ['hour_day_sin', 'hour_day_cos', 'sunset_hour_cos', 'poly_month_cos__week_of_year_cos', 'poly_week_of_year_cos__hour_day_cos', 'poly_week_day_sin__hour_day_sin', 'poly_week_day_sin__hour_day_cos', 'poly_week_day_cos__hour_day_cos', 'poly_hour_day_sin__hour_day_cos', 'poly_hour_day_sin__sunset_hour_cos', 'poly_sunrise_hour_sin__sunset_hour_sin', 'poly_sunrise_hour_sin__sunset_hour_cos', 'poly_sunrise_hour_cos__sunset_hour_cos', 'poly_sunset_hour_sin__sunset_hour_cos', 'holiday']
Calendar (n=0) : []
Feature selection with Minimum Redundancy Maximum Relevance (MRMR)¶
Minimum Redundancy Maximum Relevance (MRMR) is a filter-based feature selection method that is model-agnostic and fast to compute, making it suitable for high-dimensional datasets. However, it relies on statistical criteria like mutual information rather than evaluating model performance directly, so it may be less tailored to a specific estimator compared to wrapper methods.
# Feature selection with feature-engine MRMR
# ==============================================================================
# `random_state` makes the mutual information estimates, and therefore the
# selected features, reproducible.
selector = MRMR(method="MIQ", max_features=25, regression=True, random_state=123)
selected_lags, selected_window_features, selected_exog, selected_calendar_features = select_features(
forecaster = forecaster,
selector = selector,
y = data_train["users"],
exog = data_train.drop(columns="users"),
select_only = None,
subsample = 0.5,
random_state = 123,
verbose = True,
)
Feature selection (MRMR)
------------------------
Total number of records available: 7992
Total number of records used for feature selection: 3996
Number of features available: 114
Lags (n=25)
Window features (n=1)
Exog (n=88)
Calendar (n=0)
Number of features selected: 25
Lags (n=17) : [1, 2, 3, 4, 5, 7, 9, 10, 13, 14, 21, 23, 24, 25, 34, 46, 48]
Window features (n=0) : []
Exog (n=8) : ['hour_day_sin', 'hour_day_cos', 'poly_week_of_year_sin__week_of_year_cos', 'poly_week_day_sin__sunrise_hour_cos', 'poly_week_day_cos__hour_day_sin', 'poly_hour_day_sin__sunset_hour_sin', 'holiday_next_day', 'temp']
Calendar (n=0) : []
Combination of feature selection methods¶
Combining feature selection methods can help speed up the process. An effective approach is to first use SelectFromModel to eliminate the less important features, and then use SequentialFeatureSelector to determine the best subset of features from this reduced list. This two-step method often improves efficiency by focusing on the most important features: in this example, SequentialFeatureSelector only has to evaluate half of the candidate features.
# Feature selection (autoregressive and exog) with SelectFromModel + SequentialFeatureSelector
# ==============================================================================
estimator = LGBMRegressor(n_estimators=100, max_depth=5, random_state=15926, verbose=-1)
# Step 1: Select the 50% most important features with SelectFromModel
# `max_features` accepts a callable that receives the matrix of candidate features
selector_1 = SelectFromModel(
estimator = estimator,
max_features = lambda X: int(X.shape[1] * 0.5),
threshold = -np.inf
)
selected_lags_1, selected_window_features_1, selected_exog_1, selected_calendar_features_1 = select_features(
forecaster = forecaster,
selector = selector_1,
y = data_train["users"],
exog = data_train.drop(columns="users"),
select_only = None,
subsample = 0.2,
random_state = 123,
verbose = True,
)
Feature selection (SelectFromModel)
-----------------------------------
Total number of records available: 7992
Total number of records used for feature selection: 1598
Number of features available: 114
Lags (n=25)
Window features (n=1)
Exog (n=88)
Calendar (n=0)
Number of features selected: 57
Lags (n=25) : [1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 13, 14, 16, 21, 23, 24, 25, 29, 34, 36, 41, 42, 46, 48]
Window features (n=1) : ['roll_mean_48']
Exog (n=31) : ['hour_day_sin', 'hour_day_cos', 'poly_month_sin__week_of_year_cos', 'poly_month_sin__week_day_cos', 'poly_month_sin__hour_day_cos', 'poly_month_cos__week_day_cos', 'poly_month_cos__hour_day_sin', 'poly_week_of_year_sin__week_day_sin', 'poly_week_of_year_sin__week_day_cos', 'poly_week_of_year_sin__hour_day_sin', 'poly_week_of_year_sin__hour_day_cos', 'poly_week_of_year_sin__sunrise_hour_cos', 'poly_week_of_year_sin__sunset_hour_sin', 'poly_week_of_year_cos__week_day_sin', 'poly_week_of_year_cos__hour_day_sin', 'poly_week_of_year_cos__hour_day_cos', 'poly_week_day_sin__hour_day_sin', 'poly_week_day_sin__hour_day_cos', 'poly_week_day_cos__hour_day_sin', 'poly_week_day_cos__hour_day_cos', 'poly_week_day_cos__sunrise_hour_cos', 'poly_hour_day_sin__hour_day_cos', 'poly_hour_day_sin__sunrise_hour_sin', 'poly_hour_day_cos__sunrise_hour_cos', 'poly_hour_day_cos__sunset_hour_sin', 'poly_hour_day_cos__sunset_hour_cos', 'temp_roll_mean_1_day', 'temp_roll_mean_7_day', 'temp_roll_max_1_day', 'temp_roll_max_7_day', 'temp']
Calendar (n=0) : []
# Step 2: Select the 15 best features with SequentialFeatureSelector
# ==============================================================================
# Only the features kept in step 1 are evaluated. Step 1 kept all the lags and
# the window feature, so the forecaster is unchanged and only the exogenous
# variables are reduced.
estimator = LGBMRegressor(n_estimators=50, max_depth=5, random_state=15926, verbose=-1)
selector_2 = SequentialFeatureSelector(
estimator = estimator,
n_features_to_select = 15,
direction = "forward",
cv = ShuffleSplit(n_splits=1, test_size=0.3, random_state=951),
scoring = "neg_mean_absolute_error",
)
selected_lags, selected_window_features, selected_exog, selected_calendar_features = select_features(
forecaster = forecaster,
selector = selector_2,
y = data_train["users"],
exog = data_train[selected_exog_1],
select_only = None,
subsample = 0.2,
random_state = 123,
verbose = True,
)
Feature selection (SequentialFeatureSelector)
---------------------------------------------
Total number of records available: 7992
Total number of records used for feature selection: 1598
Number of features available: 57
Lags (n=25)
Window features (n=1)
Exog (n=31)
Calendar (n=0)
Number of features selected: 15
Lags (n=7) : [1, 3, 6, 9, 16, 21, 42]
Window features (n=0) : []
Exog (n=8) : ['hour_day_sin', 'hour_day_cos', 'poly_month_cos__week_day_cos', 'poly_week_of_year_cos__hour_day_sin', 'poly_week_day_sin__hour_day_cos', 'poly_hour_day_sin__hour_day_cos', 'poly_hour_day_cos__sunrise_hour_cos', 'poly_hour_day_cos__sunset_hour_sin']
Calendar (n=0) : []
Selection of calendar features¶
Forecasters can create calendar features from the datetime index with the calendar_features argument (see the Calendar features guide). They form a third group that can be evaluated alone (select_only='calendar') or together with the others. The selector evaluates the encoded columns (month_sin, month_cos), but selected_calendar_features returns the source features (month), which is what CalendarFeatures expects.
Since the dataset already contains precomputed calendar columns, only the temperature and holiday variables are used as exogenous variables in this example.
# Forecaster with calendar features
# ==============================================================================
calendar_features = CalendarFeatures(
features = ['month', 'quarter', 'week', 'day_of_week', 'day_of_month', 'weekend', 'hour'],
encoding = 'cyclical'
)
forecaster_calendar = ForecasterRecursive(
estimator = LGBMRegressor(
n_estimators = 900,
random_state = 15926,
max_depth = 7,
verbose = -1
),
lags = 24,
calendar_features = calendar_features
)
# Exogenous variables: temperature and holidays
exog_features = [
col for col in data_train.columns
if col.startswith(("temp", "holiday"))
]
exog_features
['temp_roll_mean_1_day', 'temp_roll_mean_7_day', 'temp_roll_max_1_day', 'temp_roll_min_1_day', 'temp_roll_max_7_day', 'temp_roll_min_7_day', 'holiday_previous_day', 'holiday_next_day', 'temp', 'holiday']
All groups of features are evaluated together (select_only=None).
# Feature selection (autoregressive, exog and calendar) with scikit-learn RFECV
# ==============================================================================
estimator = LGBMRegressor(n_estimators=100, max_depth=5, random_state=15926, verbose=-1)
selector = RFECV(estimator=estimator, step=1, cv=3, min_features_to_select=10)
selected_lags, selected_window_features, selected_exog, selected_calendar_features = select_features(
forecaster = forecaster_calendar,
selector = selector,
y = data_train["users"],
exog = data_train[exog_features],
select_only = None,
subsample = 0.5,
random_state = 123,
verbose = True,
)
Feature selection (RFECV)
-------------------------
Total number of records available: 8016
Total number of records used for feature selection: 4008
Number of features available: 47
Lags (n=24)
Window features (n=0)
Exog (n=10)
Calendar (n=13)
Number of features selected: 15
Lags (n=11) : [1, 2, 3, 4, 6, 9, 10, 14, 22, 23, 24]
Window features (n=0) : []
Exog (n=2) : ['temp_roll_mean_1_day', 'temp']
Calendar (n=2) : ['day_of_week', 'hour']
Only day_of_week and hour are kept, the calendar features that describe the weekly and daily patterns of the number of users. Finally, the forecaster is trained with a new CalendarFeatures object that includes only the selected calendar features.
# Train forecaster with selected features (including calendar features)
# ==============================================================================
calendar_features = CalendarFeatures(
features = selected_calendar_features,
encoding = 'cyclical'
)
forecaster_calendar = ForecasterRecursive(
estimator = LGBMRegressor(
n_estimators = 900,
random_state = 15926,
max_depth = 7,
verbose = -1
),
lags = selected_lags,
calendar_features = calendar_features
)
forecaster_calendar.fit(y=data_train["users"], exog=data_train[selected_exog])
Feature Selection in Global Forecasting Models¶
As with single series, feature selection can be applied to global forecasting models with the select_features_multiseries function. It works like select_features, with series (a dict of pandas Series is the recommended format) instead of y. For ForecasterDirectMultiVariate, selected_lags is a dict with the selected lags of each series.
In this example, the standard calendar features are created by the forecaster with CalendarFeatures, while other calendar indicators (start and end of month, quarter and year) are passed as exogenous variables.
# Data
# ==============================================================================
data = fetch_dataset(name="items_sales")
╭─────────────────────── items_sales ───────────────────────╮ │ Description: │ │ Simulated time series for the sales of 3 different items. │ │ │ │ Source: │ │ Simulated data. │ │ │ │ URL: │ │ https://raw.githubusercontent.com/skforecast/skforecast- │ │ datasets/main/data/simulated_items_sales.csv │ │ │ │ Shape: 1097 rows x 3 columns │ ╰───────────────────────────────────────────────────────────╯
# Exogenous variables: calendar indicators not available in CalendarFeatures
# ==============================================================================
data["is_month_start"] = data.index.is_month_start.astype(int)
data["is_month_end"] = data.index.is_month_end.astype(int)
data["is_quarter_start"] = data.index.is_quarter_start.astype(int)
data["is_quarter_end"] = data.index.is_quarter_end.astype(int)
data["is_year_start"] = data.index.is_year_start.astype(int)
data["is_year_end"] = data.index.is_year_end.astype(int)
# Training data used for feature selection
# ==============================================================================
data_train = data.loc[:"2014-06-30", :]
data_train.head()
| item_1 | item_2 | item_3 | is_month_start | is_month_end | is_quarter_start | is_quarter_end | is_year_start | is_year_end | |
|---|---|---|---|---|---|---|---|---|---|
| date | |||||||||
| 2012-01-01 | 8.253175 | 21.047727 | 19.429739 | 1 | 0 | 1 | 0 | 1 | 0 |
| 2012-01-02 | 22.777826 | 26.578125 | 28.009863 | 0 | 0 | 0 | 0 | 0 | 0 |
| 2012-01-03 | 27.549099 | 31.751042 | 32.078922 | 0 | 0 | 0 | 0 | 0 | 0 |
| 2012-01-04 | 25.895533 | 24.567708 | 27.252276 | 0 | 0 | 0 | 0 | 0 | 0 |
| 2012-01-05 | 21.379238 | 18.191667 | 20.357737 | 0 | 0 | 0 | 0 | 0 | 0 |
# Create forecaster
# ==============================================================================
calendar_features = CalendarFeatures(
features = ['month', 'quarter', 'week', 'day_of_week', 'day_of_month'],
encoding = 'cyclical'
)
forecaster = ForecasterRecursiveMultiSeries(
estimator = LGBMRegressor(n_estimators=900, random_state=15926, max_depth=7, verbose=-1),
lags = 24,
window_features = RollingFeatures(stats=['mean', 'mean', 'mean'], window_sizes=[24, 48, 72]),
calendar_features = calendar_features
)
# Feature selection (autoregressive, exog and calendar) with scikit-learn RFECV
# ==============================================================================
series_columns = ["item_1", "item_2", "item_3"]
exog_columns = [col for col in data.columns if col not in series_columns]
series_dict_train = {col: data_train[col] for col in series_columns}
estimator = LGBMRegressor(n_estimators=100, max_depth=5, random_state=15926, verbose=-1)
selector = RFECV(estimator=estimator, step=1, cv=3, min_features_to_select=10)
selected_lags, selected_window_features, selected_exog, selected_calendar_features = select_features_multiseries(
forecaster = forecaster,
selector = selector,
series = series_dict_train,
exog = data_train[exog_columns],
select_only = None,
subsample = 0.5,
random_state = 123,
verbose = True,
)
Feature selection (RFECV)
-------------------------
Total number of records available: 2520
Total number of records used for feature selection: 1260
Number of features available: 43
Lags (n=24)
Window features (n=3)
Exog (n=6)
Calendar (n=10)
Number of features selected: 32
Lags (n=24) : [1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24]
Window features (n=3) : ['roll_mean_24', 'roll_mean_48', 'roll_mean_72']
Exog (n=0) : []
Calendar (n=5) : ['month', 'quarter', 'week', 'day_of_week', 'day_of_month']
All the lags and window features are kept, together with five calendar features. None of the exogenous variables has been selected. The global forecasting model is then trained with the selected features.
# Train forecaster with selected features
# ==============================================================================
forecaster = ForecasterRecursiveMultiSeries(
estimator = LGBMRegressor(n_estimators=900, random_state=15926, max_depth=7, verbose=-1),
lags = selected_lags,
window_features = RollingFeatures(stats=['mean', 'mean', 'mean'], window_sizes=[24, 48, 72]),
calendar_features = CalendarFeatures(features=selected_calendar_features, encoding='cyclical')
)
forecaster.fit(series=series_dict_train)
forecaster
ForecasterRecursiveMultiSeries
General Information
- Estimator: LGBMRegressor
- Lags: [ 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24]
- Window features: ['roll_mean_24', 'roll_mean_48', 'roll_mean_72']
- Calendar features: ['month', 'quarter', 'week', 'day_of_week', 'day_of_month']
- Window size: 72
- Series encoding: ordinal
- Exogenous included: False
- Categorical features: auto
- Weight function included: False
- Series weights: None
- Differentiation order: None
- Drop NaN from series: False
- Creation date: 2026-10-08 20:08:37
- Last fit date: 2026-10-08 20:08:39
- Skforecast version: 0.26.0
- Python version: 3.14.3
- Forecaster id: None
Exogenous Variables
None
Data Transformations
- Transformer for series: None
- Transformer for exog: None
Training Information
- Series names (levels): item_1, item_2, item_3
- Training range: 'item_1': ['2012-01-01', '2014-06-30'], 'item_2': ['2012-01-01', '2014-06-30'], 'item_3': ['2012-01-01', '2014-06-30']
- Training index type: DatetimeIndex
- Training index frequency: D
Estimator Parameters
-
{'boosting_type': 'gbdt', 'class_weight': None, 'colsample_bytree': 1.0, 'importance_type': 'split', 'learning_rate': 0.1, 'max_depth': 7, 'min_child_samples': 20, 'min_child_weight': 0.001, 'min_split_gain': 0.0, 'n_estimators': 900, 'n_jobs': None, 'num_leaves': 31, 'objective': None, 'random_state': 15926, 'reg_alpha': 0.0, 'reg_lambda': 0.0, 'subsample': 1.0, 'subsample_for_bin': 200000, 'subsample_freq': 0, 'verbose': -1}
Fit Kwargs
-
{}
Key takeaways¶
Perform the selection only with training data, using a simple estimator and a subsample, and confirm the result with backtesting.
Selected lags and exogenous variables can be used directly. Window and calendar features must be recreated with new
RollingFeaturesandCalendarFeaturesobjects.SelectFromModelandMRMRare fast,RFECVis a good general purpose choice, andSequentialFeatureSelectoris the most expensive. Combining a fast method with an expensive one reduces the total cost.
Related guides: Autocorrelation analysis and lag selection, Calendar features, Backtesting and the API reference.