Categorical features¶
Categorical features (variables that take a discrete set of values, such as weather conditions or holiday status) provide valuable signals in time series forecasting. Before most machine learning models can use them, however, they must be converted into numerical representations. This encoding step must be learned exclusively from training data to avoid information leakage.
Since version 0.22.0, all skforecast machine learning forecasters (for example, ForecasterRecursive) provide a built-in categorical_features parameter that automatically handles encoding and natively configures gradient boosting estimators (XGBoost, LightGBM, CatBoost and HistGradientBoosting), requiring no manual encoder pipelines or estimator-specific parameters. This is the recommended approach for most use cases and is covered first in this guide.
This document also covers three manual encoding techniques based on scikit-learn encoders: one-hot encoding and ordinal encoding, which are applied via transformer_exog, and target encoding, which must be applied outside the forecaster because TargetEncoder needs the target variable to be fitted. These approaches are useful when working with non-gradient-boosting estimators or when fine-grained control over encoding behavior is needed.
For a comprehensive walkthrough of categorical features in gradient boosting forecasting models, see Forecasting time series with gradient boosting: Skforecast, XGBoost, LightGBM, Scikit-learn and CatBoost.
Choosing an encoding strategy
The table below summarizes when to use each approach.
| Method | How to use it | Estimator | Trade-offs |
|---|---|---|---|
| Built-in | categorical_features='auto' or a list |
Any (native in LightGBM, XGBoost, HistGradientBoosting, CatBoost) | Simplest; native support configured automatically |
| One-hot encoding | transformer_exog |
Any | No ordinal assumption; adds one column per category |
| Ordinal encoding | transformer_exog |
Any, best with trees | Single column; arbitrary order unless categories is set |
| Target encoding | Before the forecaster | Any | Uses the target signal, good for high cardinality; fit on training data only to avoid leakage |
General recommendation:
For gradient boosting models, use
categorical_features='auto', which requires the least code and lets the model use native categorical splits.For linear models or non-gradient-boosting trees, use one-hot or ordinal encoding via
transformer_exog.For high-cardinality features (e.g., hundreds of category levels), target encoding, applied manually outside the forecaster, is the most compact option.
✏️ Note
All of the transformations described in this document are compatible with all forecasters. However, it is important to ensure that they are learned only from the training data to avoid information leakage, and that the same fitted transformation is applied during prediction. To reduce the likelihood of errors, it is advisable to include the transformation within the forecaster object (via transformer_exog or categorical_features), so that it is handled internally during both fit() and predict().
Libraries and data¶
The dataset used in this user guide, loaded with fetch_dataset, consists of information on the number of users of a bicycle rental service, in addition to weather variables and holiday data. Two of the variables in the dataset, holiday and weather, are categorical.
# Libraries
# ==============================================================================
import numpy as np
import pandas as pd
from sklearn.preprocessing import (
OneHotEncoder, OrdinalEncoder, TargetEncoder, StandardScaler
)
from sklearn.compose import make_column_transformer, make_column_selector
from sklearn.model_selection import KFold
from sklearn.ensemble import HistGradientBoostingRegressor
from lightgbm import LGBMRegressor
from xgboost import XGBRegressor
from catboost import CatBoostRegressor
from skforecast.datasets import fetch_dataset
from skforecast.recursive import ForecasterRecursive
# Downloading data
# ==============================================================================
data = fetch_dataset(name='bike_sharing', raw=True)
╭───────────────────────────────── bike_sharing ──────────────────────────────────╮ │ Description: │ │ Hourly usage of the bike share system in the city of Washington D.C. during the │ │ years 2011 and 2012. In addition to the number of users per hour, information │ │ about weather conditions and holidays is available. │ │ │ │ Source: │ │ Fanaee-T,Hadi. (2013). Bike Sharing Dataset. UCI Machine Learning Repository. │ │ https://doi.org/10.24432/C5W894. │ │ │ │ URL: │ │ https://raw.githubusercontent.com/skforecast/skforecast- │ │ datasets/main/data/bike_sharing_dataset_clean.csv │ │ │ │ Shape: 17544 rows x 12 columns │ ╰─────────────────────────────────────────────────────────────────────────────────╯
# Preprocess data
# ==============================================================================
data['date_time'] = pd.to_datetime(data['date_time'], format='%Y-%m-%d %H:%M:%S')
data = data.set_index('date_time')
data = data.sort_index()
data = data.asfreq('h')
data = data[['holiday', 'weather', 'temp', 'hum', 'users']]
data['holiday'] = data['holiday'].astype(int)
data[['holiday', 'weather']] = data[['holiday', 'weather']].astype(str)
print(data.dtypes)
data.head(3)
holiday object weather object temp float64 hum float64 users float64 dtype: object
| holiday | weather | temp | hum | users | |
|---|---|---|---|---|---|
| date_time | |||||
| 2011-01-01 00:00:00 | 0 | clear | 9.84 | 81.0 | 16.0 |
| 2011-01-01 01:00:00 | 0 | clear | 9.02 | 80.0 | 40.0 |
| 2011-01-01 02:00:00 | 0 | clear | 9.02 | 80.0 | 32.0 |
Only a subset of the available data is used to keep execution time short. The training period covers June and July 2012 and the test period covers the first half of August 2012. The training period contains the three weather categories (clear, mist and rain) and one public holiday (4 July, 24 hourly observations with holiday = '1'). The test period also contains the three weather categories, but no holidays.
# Split train-test
# ==============================================================================
start_train = '2012-06-01 00:00:00'
end_train = '2012-07-31 23:59:00'
end_test = '2012-08-15 23:59:00'
data_train = data.loc[start_train:end_train, :]
data_test = data.loc[end_train:end_test, :]
print(
f"Dates train : {data_train.index.min()} --- {data_train.index.max()}"
f" (n={len(data_train)})"
)
print(
f"Dates test : {data_test.index.min()} --- {data_test.index.max()}"
f" (n={len(data_test)})"
)
Dates train : 2012-06-01 00:00:00 --- 2012-07-31 23:00:00 (n=1464) Dates test : 2012-08-01 00:00:00 --- 2012-08-15 23:00:00 (n=360)
Built-in categorical features handling¶
New in version 0.22.0
The categorical_features parameter allows the forecaster to handle categorical exogenous variables internally. This removes the need to build encoder pipelines in transformer_exog or configure estimator-specific parameters manually.
categorical_features |
Behavior |
|---|---|
'auto' (default) |
Any exogenous column with a non-numeric, non-boolean dtype (after applying transformer_exog) is treated as categorical. |
list |
Only the listed column names are treated as categorical. Numeric columns can also be included in the list. |
None |
No internal encoding is applied. |
For the most popular gradient boosting frameworks (LightGBM, XGBoost, HistGradientBoostingRegressor, CatBoost), the forecaster also automatically configures the estimator's native categorical support, so no fit_kwargs or estimator-level parameters are needed.
✏️ Note
Internally, the forecaster applies OrdinalEncoder to the categorical features, converting them to float codes (0.0, 1.0, 2.0, ...). This avoids errors when fitting the internal estimator. For estimators that natively support categorical features (LightGBM, XGBoost, CatBoost and HistGradientBoosting), the forecaster also tells the estimator which columns are categorical (see the diagram below), so they are treated as categorical rather than numeric. For all other estimators, the float codes are passed as-is and treated as continuous numeric features.
┌─────────────────────────────────────────────────────────┐
│ User provides exog │
│ (with categorical string columns) │
└───────────────────────┬─────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────┐
│ transformer_exog applied │
│ (if configured by the user) │
└───────────────────────┬─────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────┐
│ categorical_features detection │
│ │
│ 'auto' → detect non-numeric, non-boolean columns │
│ list → use explicit column names │
│ None → skip (user manages encoding) │
└───────────────────────┬─────────────────────────────────┘
│
┌───────────┴───────────┐
▼ ▼
Categorical cols found No categorical cols
│ │
▼ ▼
┌───────────────────────┐ ┌───────────────────────┐
│ OrdinalEncoder │ │ Pass data as-is to │
│ (internal) │ │ estimator │
│ │ └───────────────────────┘
│ string → float codes │
│ 'a' → 0.0 │
│ 'b' → 1.0 │
│ 'c' → 2.0 │
└───────────┬───────────┘
│
▼
┌─────────────────────────────────────────────────────────┐
│ Configure estimator native categorical support │
│ │
│ LightGBM → fit(categorical_feature=[indices]) │
│ CatBoost → fit(cat_features=[indices]), codes as int │
│ XGBoost → set_params(feature_types, enable_cat) │
│ HistGBR → set_params(categorical_features) │
│ Other → no-op (float codes treated as numeric) │
└───────────────────────┬─────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────┐
│ estimator.fit(X_train, y_train) │
│ (float-encoded cats + native cat config) │
└─────────────────────────────────────────────────────────┘
⚠ Warning
The internal OrdinalEncoder is configured with unknown_value=np.nan. This means that any category encountered during predict() that was not seen during fit() will be encoded as NaN. Whether this causes an error depends on the estimator: NaN-tolerant models (LightGBM, CatBoost, XGBoost, and scikit-learn tree-based models such as HistGradientBoosting, RandomForest, ExtraTrees and DecisionTree) will handle it gracefully, but other estimators will raise an error.
LightGBM¶
# Forecaster with LightGBM and categorical feature handling
# ==============================================================================
exog_features = ['holiday', 'weather', 'temp', 'hum']
forecaster = ForecasterRecursive(
estimator = LGBMRegressor(random_state=123, verbose=-1),
lags = 5,
categorical_features = 'auto' # Detects any non-numeric column
)
forecaster.fit(
y = data_train['users'],
exog = data_train[exog_features]
)
forecaster
ForecasterRecursive
General Information
- Estimator: LGBMRegressor
- Lags: [1 2 3 4 5]
- Window features: None
- Calendar features: None
- Window size: 5
- Series name: users
- Exogenous included: True
- Categorical features: auto
- Weight function included: False
- Differentiation order: None
- Drop NaN from series: False
- Creation date: 2026-10-08 19:47:08
- Last fit date: 2026-10-08 19:47:09
- Skforecast version: 0.26.0
- Python version: 3.14.3
- Forecaster id: None
Exogenous Variables
holidayCAT, weatherCAT, temp, hum
Data Transformations
- Transformer for y: None
- Transformer for exog: None
Training Information
- Training range: [Timestamp('2012-06-01 00:00:00'), Timestamp('2012-07-31 23:00:00')]
- Training index type: DatetimeIndex
- Training index frequency: h
Estimator Parameters
-
{'boosting_type': 'gbdt', 'class_weight': None, 'colsample_bytree': 1.0, 'importance_type': 'split', 'learning_rate': 0.1, 'max_depth': -1, 'min_child_samples': 20, 'min_child_weight': 0.001, 'min_split_gain': 0.0, 'n_estimators': 100, 'n_jobs': None, 'num_leaves': 31, 'objective': None, 'random_state': 123, 'reg_alpha': 0.0, 'reg_lambda': 0.0, 'subsample': 1.0, 'subsample_for_bin': 200000, 'subsample_freq': 0, 'verbose': -1}
Fit Kwargs
-
{}
# Categorical features detected and the encoding applied
# ==============================================================================
print('Detected categorical features:', forecaster.categorical_features_names_in_)
print('Encoding mapping:')
for feature, cats in zip(
forecaster.categorical_features_names_in_,
forecaster.categorical_encoder.categories_
):
mapping = {cat: float(i) for i, cat in enumerate(cats)}
print(f" '{feature}': {mapping}")
Detected categorical features: ['holiday', 'weather']
Encoding mapping:
'holiday': {'0': 0.0, '1': 1.0}
'weather': {'clear': 0.0, 'mist': 1.0, 'rain': 2.0}
✏️ Note
Use the create_train_X_y() method to inspect the exact feature matrix that the forecaster passes to the estimator during training. This is a useful debugging tool to verify that the categorical encoding (and any transformer) is applied correctly and that the final column names and dtypes match your expectations.
As the training matrix below shows, categorical features are encoded as floats. This is required because skforecast stores the training matrix internally as a single numeric numpy array, so every column must be numeric. However, if the estimator supports native categorical features, the forecaster automatically configures it to treat these columns as categorical, so the model can still leverage the categorical nature of the data.
# Training matrices
# ==============================================================================
X_train, y_train = forecaster.create_train_X_y(
y = data_train['users'],
exog = data_train[exog_features]
)
print(X_train.dtypes)
X_train.head(3)
lag_1 float64 lag_2 float64 lag_3 float64 lag_4 float64 lag_5 float64 holiday float64 weather float64 temp float64 hum float64 dtype: object
| lag_1 | lag_2 | lag_3 | lag_4 | lag_5 | holiday | weather | temp | hum | |
|---|---|---|---|---|---|---|---|---|---|
| date_time | |||||||||
| 2012-06-01 05:00:00 | 6.0 | 4.0 | 16.0 | 34.0 | 86.0 | 0.0 | 1.0 | 25.42 | 78.0 |
| 2012-06-01 06:00:00 | 45.0 | 6.0 | 4.0 | 16.0 | 34.0 | 0.0 | 2.0 | 25.42 | 88.0 |
| 2012-06-01 07:00:00 | 141.0 | 45.0 | 6.0 | 4.0 | 16.0 | 0.0 | 1.0 | 26.24 | 89.0 |
# Predictions
# ==============================================================================
forecaster.predict(steps=3, exog=data_test[exog_features])
2012-08-01 00:00:00 90.197049 2012-08-01 01:00:00 43.383265 2012-08-01 02:00:00 22.098846 Freq: h, Name: pred, dtype: float64
The categorical_features parameter also accepts an explicit list of column names. It is also possible to include numeric columns in the explicit list, forcing the model to treat them as categorical even though they have an integer or float dtype (an example is shown in the Ordinal encoding section). In this example, the list selects the same columns that 'auto' detects, so both forecasters are equivalent.
# Explicit list of categorical column names
# ==============================================================================
forecaster = ForecasterRecursive(
estimator = LGBMRegressor(random_state=123, verbose=-1),
lags = 5,
categorical_features = ['holiday', 'weather'] # Explicitly specified
)
forecaster.fit(
y = data_train['users'],
exog = data_train[exog_features]
)
print('Categorical features:', forecaster.categorical_features_names_in_)
Categorical features: ['holiday', 'weather']
XGBoost¶
For XGBoost, the forecaster sets the estimator parameters feature_types ('c' for categorical columns, 'q' for numeric ones) and enable_categorical=True.
# Forecaster with XGBoost and categorical feature handling
# ==============================================================================
forecaster = ForecasterRecursive(
estimator = XGBRegressor(random_state=123),
lags = 5,
categorical_features = 'auto'
)
forecaster.fit(
y = data_train['users'],
exog = data_train[exog_features]
)
print('Detected categorical features:', forecaster.categorical_features_names_in_)
print('feature_types :', forecaster.estimator.get_params()['feature_types'])
print('enable_categorical :', forecaster.estimator.get_params()['enable_categorical'])
Detected categorical features: ['holiday', 'weather'] feature_types : ['q', 'q', 'q', 'q', 'q', 'c', 'c', 'q', 'q'] enable_categorical : True
HistGradientBoostingRegressor¶
For HistGradientBoostingRegressor, the forecaster sets the estimator parameter categorical_features to the positions (column indices) of the categorical columns in the training matrix.
# Forecaster with HistGradientBoostingRegressor and categorical feature handling
# ==============================================================================
forecaster = ForecasterRecursive(
estimator = HistGradientBoostingRegressor(random_state=123),
lags = 5,
categorical_features = 'auto'
)
forecaster.fit(
y = data_train['users'],
exog = data_train[exog_features]
)
print('Detected categorical features:', forecaster.categorical_features_names_in_)
print(
'categorical_features (column indices):',
forecaster.estimator.get_params()['categorical_features']
)
Detected categorical features: ['holiday', 'weather'] categorical_features (column indices): [5, 6]
CatBoost¶
For CatBoost, the column indices are passed to fit() through the cat_features argument. Since CatBoost requires categorical columns to be integers or strings, the float codes are cast to integers, and unknown categories (encoded as NaN) are mapped to -1, which CatBoost treats as a new category.
# Forecaster with CatBoostRegressor and categorical feature handling
# ==============================================================================
forecaster = ForecasterRecursive(
estimator = CatBoostRegressor(
random_state = 123,
verbose = 0,
allow_writing_files = False
),
lags = 5,
categorical_features = 'auto'
)
forecaster.fit(
y = data_train['users'],
exog = data_train[exog_features]
)
print('Detected categorical features:', forecaster.categorical_features_names_in_)
Detected categorical features: ['holiday', 'weather']
Combining categorical_features and transformer_exog¶
Arguments categorical_features and transformer_exog can be used together. The former focuses on handling categorical features internally, while the latter allows for custom transformations (including encoding) to be applied to the exogenous data inside the forecaster, before the categorical features are detected. This means that users can apply any desired transformations to their exogenous features using transformer_exog, and then rely on categorical_features to automatically detect and handle any categorical columns in the transformed data.
For example, it is possible to apply a standard scaler to some of the numerical columns in the exogenous data using transformer_exog, and then use categorical_features='auto' to let the forecaster automatically handle any categorical columns. Scaling is not needed for tree-based models such as LightGBM, since their splits do not depend on the scale of the features; it is used here only to illustrate the mechanism.
✏️ Note
Because transformer_exog is applied before categorical_features detection, ensure that the transformer does not convert categorical columns to numeric, otherwise categorical_features='auto' will no longer detect them.
# Example of how to use transformer_exog together with categorical_features
# ==============================================================================
# Scale 'temp' and 'hum' columns, while leaving 'holiday' and 'weather' unchanged.
# Then, use categorical_features='auto' to let the forecaster automatically handle
# the categorical columns.
transformer_exog = make_column_transformer(
(StandardScaler(), ['temp', 'hum']),
remainder='passthrough',
verbose_feature_names_out=False
).set_output(transform='pandas')
forecaster = ForecasterRecursive(
estimator = LGBMRegressor(random_state=123, verbose=-1),
lags = 5,
transformer_exog = transformer_exog,
categorical_features = 'auto' # Detects any non-numeric column
)
forecaster.fit(
y = data_train['users'],
exog = data_train[exog_features]
)
# Training matrices with transformer_exog applied and categorical features detected
# ==============================================================================
X_train, y_train = forecaster.create_train_X_y(
y = data_train['users'],
exog = data_train[exog_features]
)
X_train.head(3)
| lag_1 | lag_2 | lag_3 | lag_4 | lag_5 | temp | hum | holiday | weather | |
|---|---|---|---|---|---|---|---|---|---|
| date_time | |||||||||
| 2012-06-01 05:00:00 | 6.0 | 4.0 | 16.0 | 34.0 | 86.0 | -0.913070 | 1.076344 | 0.0 | 1.0 |
| 2012-06-01 06:00:00 | 45.0 | 6.0 | 4.0 | 16.0 | 34.0 | -0.913070 | 1.621649 | 0.0 | 2.0 |
| 2012-06-01 07:00:00 | 141.0 | 45.0 | 6.0 | 4.0 | 16.0 | -0.719565 | 1.676179 | 0.0 | 1.0 |
Note that the column order of the training matrix has changed: temp and hum now appear before holiday and weather. ColumnTransformer places the output of each transformer first, in the order in which they are defined, and appends the remainder='passthrough' columns at the end. The categorical columns are still detected and encoded as float codes, and the forecaster uses their new positions when configuring the estimator, so they are still treated as categorical.
Manual encoding with scikit-learn encoders¶
The built-in categorical_features parameter covers the most common use case: gradient boosting estimators with native categorical support. For other scenarios (linear models, SVMs, non-gradient-boosting tree ensembles) or when precise control over encoding behavior is required (custom category orderings, infrequent-category grouping or target encoding), scikit-learn encoders can be used instead. One-hot and ordinal encoders are passed to the forecaster inside a ColumnTransformer via transformer_exog, whereas target encoding must be applied outside the forecaster (see the warning in the Target encoding section).
The following sections demonstrate three encoding strategies: one-hot encoding, ordinal encoding and target encoding.
✏️ Note
transformer_exog is applied before categorical_features detection. If transformer_exog already converts all categorical columns to numeric, categorical_features = 'auto' (the default) will find nothing to encode and will have no effect. Setting categorical_features = None explicitly is not required in this case, but can be used to make the intent clearer.
One-hot encoding¶
One-hot encoding, also known as one-of-K encoding (and closely related to dummy encoding, which drops one of the resulting columns), consists of replacing the categorical variable with a set of binary variables that take the value 0 or 1 to indicate whether a particular category is present in an observation. For example, suppose a dataset contains a categorical variable called "color" with the possible values of "red," "blue," and "green". Using one-hot encoding, this variable is converted into three binary variables such as color_red, color_blue, and color_green, where each variable takes a value of 0 or 1 depending on the category.
The OneHotEncoder class in scikit-learn can be used to transform any categorical feature with n possible values into n new binary features, where one of them takes the value 1, and all the others take the value 0. The OneHotEncoder can be configured to handle certain corner cases, including unknown categories, missing values, and infrequent categories.
When
handle_unknown='ignore', categories not seen during fit are encoded as all zeros. Ifdropis also set, an unknown category becomes indistinguishable from the dropped one. Additionally, if a feature contains bothnp.nanandNone, they are considered separate categories.It supports the aggregation of infrequent categories into a single output for each feature. The parameters to enable the aggregation of infrequent categories are
min_frequencyandmax_categories. By settinghandle_unknownto'infrequent_if_exist', unknown categories are considered infrequent.To avoid collinearity between features, it is possible to drop one of the categories per feature using the
dropargument. This is especially important when using linear models.
ColumnTransformer objects in scikit-learn provide a powerful way to define transformations and apply them to specific features. By encapsulating the OneHotEncoder in a ColumnTransformer object, it can be passed to a forecaster using the transformer_exog argument.
One-hot encoding is particularly well-suited for linear models, support vector machines, and neural networks. For gradient boosting models with native categorical support, the built-in categorical_features parameter is generally the preferred approach. A LightGBM model is used in the following example only to allow a direct comparison with the results of the built-in approach.
# ColumnTransformer with one-hot encoding
# ==============================================================================
# A ColumnTransformer is used to transform categorical features (non-numeric)
# using one-hot encoding. Numeric features are left untouched. For binary
# features, only one column is created.
one_hot_encoder = make_column_transformer(
(
OneHotEncoder(sparse_output=False, drop='if_binary'),
make_column_selector(dtype_exclude=np.number)
),
remainder='passthrough',
verbose_feature_names_out=False,
).set_output(transform='pandas')
# Create and fit forecaster with a transformer for exogenous features
# ==============================================================================
forecaster = ForecasterRecursive(
estimator = LGBMRegressor(random_state=123, verbose=-1),
lags = 5,
transformer_exog = one_hot_encoder
)
forecaster.fit(
y = data_train['users'],
exog = data_train[exog_features]
)
forecaster
ForecasterRecursive
General Information
- Estimator: LGBMRegressor
- Lags: [1 2 3 4 5]
- Window features: None
- Calendar features: None
- Window size: 5
- Series name: users
- Exogenous included: True
- Categorical features: auto
- Weight function included: False
- Differentiation order: None
- Drop NaN from series: False
- Creation date: 2026-10-08 19:47:11
- Last fit date: 2026-10-08 19:47:11
- Skforecast version: 0.26.0
- Python version: 3.14.3
- Forecaster id: None
Exogenous Variables
holiday, weather, temp, hum
Data Transformations
- Transformer for y: None
- Transformer for exog: ColumnTransformer(remainder='passthrough',
transformers=[('onehotencoder',
OneHotEncoder(drop='if_binary',
sparse_output=False),
)], verbose_feature_names_out=False)
Training Information
- Training range: [Timestamp('2012-06-01 00:00:00'), Timestamp('2012-07-31 23:00:00')]
- Training index type: DatetimeIndex
- Training index frequency: h
Estimator Parameters
-
{'boosting_type': 'gbdt', 'class_weight': None, 'colsample_bytree': 1.0, 'importance_type': 'split', 'learning_rate': 0.1, 'max_depth': -1, 'min_child_samples': 20, 'min_child_weight': 0.001, 'min_split_gain': 0.0, 'n_estimators': 100, 'n_jobs': None, 'num_leaves': 31, 'objective': None, 'random_state': 123, 'reg_alpha': 0.0, 'reg_lambda': 0.0, 'subsample': 1.0, 'subsample_for_bin': 200000, 'subsample_freq': 0, 'verbose': -1}
Fit Kwargs
-
{}
Once the forecaster has been trained, the transformer can be inspected by accessing the transformer_exog attribute.
# Access to the transformer used for exogenous features
# ==============================================================================
print(forecaster.transformer_exog.get_feature_names_out())
forecaster.transformer_exog
['holiday_1' 'weather_clear' 'weather_mist' 'weather_rain' 'temp' 'hum']
ColumnTransformer(remainder='passthrough',
transformers=[('onehotencoder',
OneHotEncoder(drop='if_binary',
sparse_output=False),
<sklearn.compose._column_transformer.make_column_selector object at 0x165c59550>)],
verbose_feature_names_out=False)In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook. On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
Parameters
Fitted attributes
['holiday', 'weather']
Parameters
Fitted attributes
4 features
| holiday_1 |
| weather_clear |
| weather_mist |
| weather_rain |
['temp', 'hum']
2 features
| temp |
| hum |
6 features
| holiday_1 |
| weather_clear |
| weather_mist |
| weather_rain |
| temp |
| hum |
# Predictions
# ==============================================================================
forecaster.predict(steps=3, exog=data_test[exog_features])
2012-08-01 00:00:00 90.197049 2012-08-01 01:00:00 43.383265 2012-08-01 02:00:00 22.098846 Freq: h, Name: pred, dtype: float64
These predictions are identical to the ones obtained with the built-in categorical_features handling (90.197, 43.383 and 22.099). This is expected: when a categorical feature has at most max_cat_to_onehot levels (4 by default), LightGBM splits it with a one-vs-rest strategy, which is exactly what a tree does with one-hot encoded columns. Both holiday (2 levels) and weather (3 levels) are below this limit. With higher-cardinality features, the native handling can group several categories in a single split, and the two approaches would no longer match.
# Create training matrices
# ==============================================================================
X_train, y_train = forecaster.create_train_X_y(
y = data_train['users'],
exog = data_train[exog_features]
)
print(X_train.dtypes)
X_train.head(3)
lag_1 float64 lag_2 float64 lag_3 float64 lag_4 float64 lag_5 float64 holiday_1 float64 weather_clear float64 weather_mist float64 weather_rain float64 temp float64 hum float64 dtype: object
| lag_1 | lag_2 | lag_3 | lag_4 | lag_5 | holiday_1 | weather_clear | weather_mist | weather_rain | temp | hum | |
|---|---|---|---|---|---|---|---|---|---|---|---|
| date_time | |||||||||||
| 2012-06-01 05:00:00 | 6.0 | 4.0 | 16.0 | 34.0 | 86.0 | 0.0 | 0.0 | 1.0 | 0.0 | 25.42 | 78.0 |
| 2012-06-01 06:00:00 | 45.0 | 6.0 | 4.0 | 16.0 | 34.0 | 0.0 | 0.0 | 0.0 | 1.0 | 25.42 | 88.0 |
| 2012-06-01 07:00:00 | 141.0 | 45.0 | 6.0 | 4.0 | 16.0 | 0.0 | 0.0 | 1.0 | 0.0 | 26.24 | 89.0 |
Ordinal encoding¶
Ordinal encoding converts categorical variables into integers by assigning each category a unique numeric code. Unlike one-hot encoding, it produces a single column per feature, which makes it memory-efficient. The encoding is most appropriate when categories have a meaningful natural order (e.g., low < medium < high), or when used with tree-based models, which can isolate individual categories through successive splits and are therefore less sensitive to the arbitrary order of the codes than linear models.
The scikit-learn library provides the OrdinalEncoder class, which assigns integers from 0 to n_categories-1. By default, the order is determined alphabetically. It is important to note that this ordering is arbitrary for unordered categories; users should specify explicit categories if a meaningful order exists. The class also exposes the encoded_missing_value parameter to handle missing values, and handle_unknown='use_encoded_value' to gracefully deal with unseen categories at prediction time.
Using OrdinalEncoder via transformer_exog gives full control over encoding parameters: for example, setting unknown_value=-1 (instead of the NaN used by the built-in categorical_features encoder) or specifying explicit category orderings. The dtype of the encoder has no practical effect, since the training matrix is converted to float64 (see the output below).
# ColumnTransformer with ordinal encoding
# ==============================================================================
# A ColumnTransformer is used to transform categorical (not numerical) features
# using ordinal encoding. Numeric features are left untouched. Missing values
# are coded as -1. If a new category is found in the test set, it is encoded
# as -1.
ordinal_encoder = make_column_transformer(
(
OrdinalEncoder(
handle_unknown='use_encoded_value',
unknown_value=-1,
encoded_missing_value=-1
),
make_column_selector(dtype_exclude=np.number)
),
remainder='passthrough',
verbose_feature_names_out=False,
).set_output(transform='pandas')
# Create and fit a forecaster with a transformer for exogenous features
# ==============================================================================
forecaster = ForecasterRecursive(
estimator = LGBMRegressor(random_state=123, verbose=-1),
lags = 5,
transformer_exog = ordinal_encoder
)
forecaster.fit(
y = data_train['users'],
exog = data_train[exog_features]
)
# Create training matrices
# ==============================================================================
X_train, y_train = forecaster.create_train_X_y(
y = data_train['users'],
exog = data_train[exog_features]
)
print(X_train.dtypes)
X_train.head(3)
lag_1 float64 lag_2 float64 lag_3 float64 lag_4 float64 lag_5 float64 holiday float64 weather float64 temp float64 hum float64 dtype: object
| lag_1 | lag_2 | lag_3 | lag_4 | lag_5 | holiday | weather | temp | hum | |
|---|---|---|---|---|---|---|---|---|---|
| date_time | |||||||||
| 2012-06-01 05:00:00 | 6.0 | 4.0 | 16.0 | 34.0 | 86.0 | 0.0 | 1.0 | 25.42 | 78.0 |
| 2012-06-01 06:00:00 | 45.0 | 6.0 | 4.0 | 16.0 | 34.0 | 0.0 | 2.0 | 25.42 | 88.0 |
| 2012-06-01 07:00:00 | 141.0 | 45.0 | 6.0 | 4.0 | 16.0 | 0.0 | 1.0 | 26.24 | 89.0 |
# Predictions
# ==============================================================================
forecaster.predict(steps=3, exog=data_test[exog_features])
2012-08-01 00:00:00 85.516374 2012-08-01 01:00:00 40.843863 2012-08-01 02:00:00 19.843919 Freq: h, Name: pred, dtype: float64
Unlike one-hot encoding, ordinal encoding gives different predictions from the built-in approach (85.516 instead of 90.197 for the first step). LGBMRegressor is not told that the encoded columns are categorical, so it treats the codes as numbers and can only make threshold splits such as weather <= 1.5. With the codes clear = 0, mist = 1 and rain = 2, a single split can separate {clear} from {mist, rain}, or {clear, mist} from {rain}, but never {mist} from {clear, rain}; that requires two consecutive splits. Since holiday has only two levels, its ordinal code is equivalent to a binary column, so the whole difference comes from weather. In this dataset, the alphabetical order happens to match the severity of the weather, but this is a coincidence: with other category names, the order imposed by the encoder would be arbitrary.
The ordinal-encoded columns are numeric, so categorical_features='auto' does not detect them. However, the explicit list accepts numeric columns: declaring weather as categorical makes LightGBM use its native categorical splits again. As the output below shows, this recovers the predictions of the built-in approach (90.197, 43.383 and 22.099) and confirms that the difference comes from how weather is split.
# Ordinal encoding with weather declared as categorical
# ==============================================================================
forecaster = ForecasterRecursive(
estimator = LGBMRegressor(random_state=123, verbose=-1),
lags = 5,
transformer_exog = ordinal_encoder,
categorical_features = ['weather'] # Ordinal codes as categorical
)
forecaster.fit(
y = data_train['users'],
exog = data_train[exog_features]
)
forecaster.predict(steps=3, exog=data_test[exog_features])
2012-08-01 00:00:00 90.197049 2012-08-01 01:00:00 43.383265 2012-08-01 02:00:00 22.098846 Freq: h, Name: pred, dtype: float64
Target encoding¶
Target encoding is a technique that encodes categorical variables based on the relationship between the categories and the target variable. Each category is encoded based on a shrinkage estimate of the average target values for observations belonging to that category. The encoding scheme mixes the global target mean with the target mean conditioned on the value of the category.
For example, suppose a categorical variable "City" with categories "New York," "Los Angeles," and "Chicago," and a target variable "Salary." One can calculate the mean salary for each city based on the training data, and use these mean values to encode the categories.
This encoding scheme is useful for categorical features with high cardinality, where one-hot encoding would inflate the feature space, making it more expensive for a downstream model to process. A classic example of high-cardinality categorical variables is location data, such as zip codes or regions.
The TargetEncoder class is available in scikit-learn (since version 1.3). TargetEncoder considers missing values, such as np.nan or None, as another category and encodes them like any other category. Categories that are not seen during fit are encoded with the target mean, i.e. target_mean_. A more detailed description of target encoding can be found in the scikit-learn user guide.
⚠ Warning
TargetEncoder differs from the other transformers in scikit-learn in that it requires not only the features to be transformed but also the response variable (target), which in forecasting is the time series itself. Currently, the only transformers allowed inside a forecaster are those that do not require the target variable to be fitted. Therefore, to use target encoding, transformations must be applied outside the forecaster.
# ColumnTransformer with target encoding
# ==============================================================================
# A ColumnTransformer is used to transform categorical (not numerical) features
# using target encoding. Numeric features are left untouched. TargetEncoder
# considers missing values, such as np.nan or None, as another category and
# encodes them like any other category. Categories that are not seen during fit
# are encoded with the target mean.
# Passing a splitter to `cv` requires scikit-learn >= 1.9. In previous versions,
# use `cv=5, shuffle=True, random_state=123` instead.
target_encoder = make_column_transformer(
(
TargetEncoder(
categories = 'auto',
target_type = 'continuous',
smooth = 'auto',
cv = KFold(n_splits=5, shuffle=True, random_state=123),
),
make_column_selector(dtype_exclude=np.number)
),
remainder='passthrough',
verbose_feature_names_out=False,
).set_output(transform='pandas')
# Transform the exogenous features using the transformer outside the forecaster
# ==============================================================================
exog_transformed = target_encoder.fit_transform(
X = data_train[exog_features],
y = data_train['users']
)
exog_transformed.head(3)
| holiday | weather | temp | hum | |
|---|---|---|---|---|
| date_time | ||||
| 2012-06-01 00:00:00 | 271.605660 | 284.348564 | 27.06 | 50.0 |
| 2012-06-01 01:00:00 | 278.826822 | 294.781049 | 26.24 | 53.0 |
| 2012-06-01 02:00:00 | 278.826822 | 294.781049 | 26.24 | 57.0 |
# Transform test exog using the already-fitted encoder
# ==============================================================================
# Use .transform() (not .fit_transform()) so that the encoding learned on the
# training data is reused and no target values are needed.
exog_test_transformed = target_encoder.transform(data_test[exog_features])
exog_test_transformed.head(3)
| holiday | weather | temp | hum | |
|---|---|---|---|---|
| date_time | ||||
| 2012-08-01 00:00:00 | 277.107302 | 292.259998 | 27.88 | 79.0 |
| 2012-08-01 01:00:00 | 277.107302 | 292.259998 | 27.06 | 83.0 |
| 2012-08-01 02:00:00 | 277.107302 | 292.259998 | 26.24 | 83.0 |
# Create and fit forecaster with pre-transformed exogenous features
# ==============================================================================
forecaster = ForecasterRecursive(
estimator = LGBMRegressor(random_state=123, verbose=-1),
lags = 5
)
forecaster.fit(
y = data_train['users'],
exog = exog_transformed # pre-transformed training exog (all numeric)
)
# Predictions
# ==============================================================================
forecaster.predict(steps=3, exog=exog_test_transformed)
2012-08-01 00:00:00 86.332010 2012-08-01 01:00:00 42.424949 2012-08-01 02:00:00 19.994938 Freq: h, Name: pred, dtype: float64