Weighted time series forecasting¶
In many real-world scenarios, historical data is available for forecasting, but not all of it is reliable. For example, IoT sensors capture raw data from the physical world, but they are often prone to failure, malfunction, and attrition due to harsh deployment environments, leading to unusual or erroneous readings. Similarly, factories may shut down for maintenance, repair, or overhaul, resulting in gaps in the data. The COVID-19 pandemic has also affected population behavior, impacting many time series such as production, sales, and transportation.
The presence of unreliable or unrepresentative values in the data history can be a problem, since the model may learn patterns that will not be repeated in the future. For most forecasting algorithms, removing that part of the data is not an option because they require the time series to be complete. An alternative solution is to reduce the weight of the affected observations during model training. This document shows with an example how skforecast makes it easy to apply this strategy, and how to check whether it actually improves the predictions.
✏️ Note
The example that follows demonstrates how a portion of the time series can be excluded from model training by assigning it a weight of zero. However, the use of weights extends beyond the inclusion or exclusion of observations and can also balance the degree of influence that each observation has on the forecasting model. For instance, an observation assigned a weight of 10 will have ten times more impact on the model training than an observation assigned a weight of 1.
⚠ Warning
In most gradient boosting implementations, such as LightGBM, XGBoost, and CatBoost, samples with zero weight are typically excluded when calculating gradients and Hessians. However, these samples are still taken into account when constructing the feature histograms, which can result in a model that differs from one trained without zero-weighted samples. For more information on this issue, please refer to this GitHub issue.
For linear models such as the Ridge regressor used in this document, a weight of zero is equivalent to removing the observation from the training set.
Libraries and data¶
# Libraries
# ==============================================================================
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
from sklearn.linear_model import Ridge
from skforecast.recursive import ForecasterRecursive
from skforecast.preprocessing import CalendarFeatures
from skforecast.model_selection import TimeSeriesFold, backtesting_forecaster
from skforecast.plot import set_dark_theme
# Data download
# ==============================================================================
url = (
'https://raw.githubusercontent.com/skforecast/skforecast-datasets/refs/heads/'
'main/data/energy_production_shutdown.csv'
)
data = pd.read_csv(url)
# Data preprocessing
# ==============================================================================
data['date'] = pd.to_datetime(data['date'], format='%Y-%m-%d')
data = data.set_index('date')
data = data.sort_index()
data = data.asfreq('D')
data.head()
| production | |
|---|---|
| date | |
| 2012-01-01 | 375.1 |
| 2012-01-02 | 474.5 |
| 2012-01-03 | 573.9 |
| 2012-01-04 | 539.5 |
| 2012-01-05 | 445.4 |
The dataset contains the daily energy production of a power plant from 2012-01-01 to 2014-12-30. The data from 2012 and 2013 is used to train the model (731 days) and the data from 2014 as the test set (364 days, since the dataset ends on 2014-12-30). The plant was shut down between 2012-06-01 and 2012-09-30, a period highlighted in the plot below, during which the production dropped to around 300, compared with around 465 during the rest of 2012 and 2013.
# Split data into train-test
# ==============================================================================
end_train = '2013-12-31 23:59:00'
data_train = data.loc[: end_train, :]
data_test = data.loc[end_train:, :]
print(
f"Dates train : {data_train.index.min()} --- {data_train.index.max()} "
f"(n={len(data_train)})"
)
print(
f"Dates test : {data_test.index.min()} --- {data_test.index.max()} "
f"(n={len(data_test)})"
)
Dates train : 2012-01-01 00:00:00 --- 2013-12-31 00:00:00 (n=731) Dates test : 2014-01-01 00:00:00 --- 2014-12-30 00:00:00 (n=364)
# Time series plot
# ==============================================================================
set_dark_theme()
fig, ax = plt.subplots(figsize=(7, 3))
data_train['production'].plot(ax=ax, label='train', linewidth=1)
data_test['production'].plot(ax=ax, label='test', linewidth=1)
ax.axvspan(
pd.to_datetime('2012-06-01'),
pd.to_datetime('2012-09-30'),
label='Shutdown',
color='gray',
alpha=0.3
)
ax.set_title('Energy production')
ax.set_xlabel('')
ax.legend();
Include the whole time series¶
First, a ForecasterRecursive is created with a Ridge regressor. Its predictors are the values of the previous 21 days (lags=21) and two calendar features, the month and the day of the week, created with CalendarFeatures and one-hot encoded. The forecaster generates these features from the index of the series, both when it is trained and when it predicts. All observations have the same weight during training, including those of the shutdown period. This model is used as the baseline.
# Create a recursive multi-step forecaster (ForecasterRecursive)
# ==============================================================================
calendar_features = CalendarFeatures(
features = ['month', 'day_of_week'],
encoding = 'onehot'
)
forecaster_baseline = ForecasterRecursive(
estimator = Ridge(),
lags = 21,
calendar_features = calendar_features
)
Once the model is created, a backtesting process is run with backtesting_forecaster and TimeSeriesFold to simulate the behavior of the forecaster if it had predicted the test set in batches of 10 days. With refit=False, the model is trained only once, using the training set, and then predicts the test set in 37 consecutive folds (the last one contains only 4 days, since incomplete folds are allowed by default with allow_incomplete_fold=True). The mean absolute error (MAE) is used as the metric. The same cv object is used later to evaluate the model with weights, so both results are comparable.
# Backtesting: predict batches of 10 days
# ==============================================================================
cv = TimeSeriesFold(
steps = 10,
initial_train_size = len(data_train),
refit = False
)
metric_baseline, predictions_baseline = backtesting_forecaster(
forecaster = forecaster_baseline,
y = data['production'],
cv = cv,
metric = 'mean_absolute_error',
verbose = False
)
metric_baseline
| mean_absolute_error | |
|---|---|
| 0 | 31.781508 |
Exclude part of the time series¶
Between 2012-06-01 and 2012-09-30, the plant underwent a shutdown. This period is especially harmful for a model that uses the month as a predictor: the effect of June, July, August, and September is learned from two years of data, one of which (2012) has an abnormally low production. Since the test set contains these months again, their predictions are biased downward.
To reduce the impact of these dates on the model, a custom function is created and passed to the forecaster through the weight_func argument. Each time the forecaster is fitted, this function receives the index of the training matrix (the date of the target value of each row) and must return one weight per row.
The function below assigns a weight of 0 to any date that falls within the shutdown period or up to 21 days later, and a weight of 1 to all other dates. The extra 21 days are needed because the forecaster uses lags=21: the rows of the training matrix whose target is dated between 2012-10-01 and 2012-10-21 still use values from the shutdown as predictors, so they are excluded as well. Observations assigned a weight of 0 have no influence on the model training.
✏️ Note
The weights returned by weight_func must be non-negative, must not contain NaN values, and cannot all be zero. If any of these conditions is not met, a ValueError is raised when the forecaster is fitted. The weights of a training set can be inspected with the create_sample_weights method.
⚠ Warning
If the estimator does not accept sample_weight in its fit method, the weight_func argument is ignored and an IgnoredArgumentWarning is issued when the forecaster is created. The model is then trained as if no weights had been provided.
✏️ Note
Weights are only used when the model is trained. The observations with a weight of 0 are still part of the time series, so they are used as lags whenever they fall within the last window of data needed to create the predictors. Therefore, weighting does not help if the anomalous period is located right before the dates to be predicted.
# Custom function to create weights
# ==============================================================================
def custom_weights(index):
"""
Return a weight of 0 for dates between 2012-06-01 and 2012-10-21
(shutdown plus the 21 days affected through the lags) and 1 otherwise.
"""
weights = np.where((index >= '2012-06-01') & (index <= '2012-10-21'), 0, 1)
return weights
A new ForecasterRecursive is created with the same configuration as the baseline, this time passing custom_weights to the weight_func argument. The function is called automatically every time the forecaster is fitted: only once in this backtesting, since refit=False, and on every refit if refit=True.
# Create a recursive multi-step forecaster with weights
# ==============================================================================
forecaster_weighted = ForecasterRecursive(
estimator = Ridge(),
lags = 21,
calendar_features = calendar_features,
weight_func = custom_weights
)
The source code of the function is stored in the source_code_weight_func attribute of the forecaster. This makes it possible to check how the weights were calculated, for example after loading a forecaster saved in a previous session.
# Source code of the weight function stored in the forecaster
# ==============================================================================
print(forecaster_weighted.source_code_weight_func)
def custom_weights(index):
"""
Return a weight of 0 for dates between 2012-06-01 and 2012-10-21
(shutdown plus the 21 days affected through the lags) and 1 otherwise.
"""
weights = np.where((index >= '2012-06-01') & (index <= '2012-10-21'), 0, 1)
return weights
Before running the backtesting, the weights assigned to each observation are inspected with the create_sample_weights method. It takes the training matrix created with create_train_X_y, which has 710 rows: the 731 training observations minus the first 21, which are only used as lags. Of these rows, 143 have a weight of 0 and 567 a weight of 1.
# Weights assigned to each row of the training matrix
# ==============================================================================
X_train, y_train = forecaster_weighted.create_train_X_y(y=data_train['production'])
sample_weights = pd.Series(
forecaster_weighted.create_sample_weights(X_train=X_train),
index = X_train.index,
name = 'weight'
)
print(sample_weights.value_counts())
fig, ax = plt.subplots(figsize=(7, 2))
sample_weights.plot(ax=ax, linewidth=1)
ax.set_title('Weights of the training observations')
ax.set_xlabel('');
weight 1 567 0 143 Name: count, dtype: int64
The backtesting is repeated with the same cv object, so the only difference between both models is the weighting of the training observations.
# Backtesting: predict batches of 10 days
# ==============================================================================
metric_weighted, predictions_weighted = backtesting_forecaster(
forecaster = forecaster_weighted,
y = data['production'],
cv = cv,
metric = 'mean_absolute_error',
verbose = False
)
metric_weighted
| mean_absolute_error | |
|---|---|
| 0 | 26.123422 |
# Comparison of backtesting metrics
# ==============================================================================
mae_baseline = metric_baseline.loc[0, 'mean_absolute_error']
mae_weighted = metric_weighted.loc[0, 'mean_absolute_error']
results = pd.DataFrame(
{'mean_absolute_error': [mae_baseline, mae_weighted]},
index = ['Without weights', 'With weights']
)
results['improvement (%)'] = 100 * (1 - results['mean_absolute_error'] / mae_baseline)
results
| mean_absolute_error | improvement (%) | |
|---|---|---|
| Without weights | 31.781508 | 0.000000 |
| With weights | 26.123422 | 17.803077 |
# Backtesting predictions of both models
# ==============================================================================
fig, ax = plt.subplots(figsize=(7, 3))
data_test['production'].plot(ax=ax, label='test', linewidth=1)
predictions_baseline['pred'].plot(ax=ax, label='without weights', linewidth=1)
predictions_weighted['pred'].plot(ax=ax, label='with weights', linewidth=1)
ax.set_title('Energy production: backtesting predictions')
ax.set_xlabel('')
ax.legend();
Giving a weight of 0 to the shutdown period and the 21 days that follow it (excluding them from the model training) reduces the mean absolute error of the backtesting from 31.8 to 26.1, an improvement of 17.8%.
Conclusions¶
The
weight_funcargument makes it possible to reduce or remove the influence of part of the time series during training, without breaking its continuity. The function receives the index of the training matrix and returns one weight per row.When a period is excluded, the rows of the training matrix that use it as lags must be excluded as well (in this document, the following 21 days).
Weights only act during training, and only if the estimator accepts
sample_weight. For linear models, a weight of 0 is equivalent to removing the observation; for gradient boosting models it is not exactly the same.The benefit depends on the model and on how the anomalous period relates to the dates to be predicted. In this example it is clear because the month is a predictor and the months of the shutdown appear again in the test set. The effect of the weights should therefore be verified with backtesting in each case rather than taken for granted.
In this example, the weights were used only to include or exclude observations, but any non-negative value is allowed: a weight between 0 and 1 reduces the influence of a period without removing it completely, and a weight greater than 1 increases it.
For another example, in which the anomalous periods are the COVID-19 lockdown and a snowstorm, see Mitigating the Impact of Covid on Forecasting Models. When part of the data is missing rather than unreliable, see Forecasting time series with missing values.