Aggregate a time series¶
Turn a long series into a few typical periods. Every other guide here is a variation on this call.
In [1]:
Copied!
import pandas as pd
import plotly.io as pio
import tsam
pio.renderers.default = "notebook_connected"
raw = pd.read_csv("../data/testdata.csv", index_col=0, parse_dates=True)
data = raw.loc["2010-01-01":"2010-02-11"] # six weeks of hourly data
UNITS = {"GHI": "W/m²", "T": "°C", "Wind": "m/s", "Load": "MW"}
data.head()
import pandas as pd
import plotly.io as pio
import tsam
pio.renderers.default = "notebook_connected"
raw = pd.read_csv("../data/testdata.csv", index_col=0, parse_dates=True)
data = raw.loc["2010-01-01":"2010-02-11"] # six weeks of hourly data
UNITS = {"GHI": "W/m²", "T": "°C", "Wind": "m/s", "Load": "MW"}
data.head()
Out[1]:
| GHI | T | Wind | Load | |
|---|---|---|---|---|
| 2010-01-01 00:30:00 | 0 | -2.8 | 8.6 | 364.541326 |
| 2010-01-01 01:30:00 | 0 | -3.3 | 9.7 | 357.416844 |
| 2010-01-01 02:30:00 | 0 | -3.2 | 9.8 | 350.191306 |
| 2010-01-01 03:30:00 | 0 | -3.2 | 9.4 | 345.161449 |
| 2010-01-01 04:30:00 | 0 | -3.2 | 10.0 | 340.678216 |
Aggregate¶
Two parameters carry the decision:
n_clusters— how many typical periods to keep.period_duration— the length of one:"1D","1W", or a number of hours.
Add temporal_resolution (e.g. "15min") if your index is irregular or its step cannot be
inferred.
In [2]:
Copied!
result = tsam.aggregate(data, n_clusters=8, period_duration="1D")
print("reduced", len(data), "hours to", result.n_clusters, "typical days")
result = tsam.aggregate(data, n_clusters=8, period_duration="1D")
print("reduced", len(data), "hours to", result.n_clusters, "typical days")
reduced 1008 hours to 8 typical days
Read the outputs¶
Three pieces are what a downstream model consumes:
cluster_representatives— the typical period profiles.cluster_counts— how many real periods each stands for (its weight).accuracy— RMSE / MAE per column.
In [3]:
Copied!
print("counts:", result.cluster_counts)
print("\nper-column RMSE:")
print(result.accuracy.rmse.round(3).to_string())
result.cluster_representatives.head()
print("counts:", result.cluster_counts)
print("\nper-column RMSE:")
print(result.accuracy.rmse.round(3).to_string())
result.cluster_representatives.head()
counts: {0: 7.0, 1: 3.0, 2: 10.0, 3: 4.0, 4: 8.0, 5: 1.0, 6: 7.0, 7: 2.0}
per-column RMSE:
GHI 0.065
T 0.105
Wind 0.155
Load 0.091
Out[3]:
| GHI | T | Wind | Load | ||
|---|---|---|---|---|---|
| timestep | |||||
| 0 | 0 | 0.0 | -0.950653 | 3.955787 | 407.677352 |
| 1 | 0.0 | -1.460029 | 3.955787 | 403.959016 | |
| 2 | 0.0 | -1.154403 | 3.955787 | 410.950870 | |
| 3 | 0.0 | -1.561904 | 3.955787 | 415.455196 | |
| 4 | 0.0 | -1.460029 | 4.944733 | 431.688908 |
In [4]:
Copied!
result.plot.compare(
columns=["Load"],
time_slice=slice("2010-01-11", "2010-01-17"),
color="source",
units=UNITS,
title="One week: original vs. reconstructed Load",
)
result.plot.compare(
columns=["Load"],
time_slice=slice("2010-01-11", "2010-01-17"),
color="source",
units=UNITS,
title="One week: original vs. reconstructed Load",
)
From here¶
| To… | Go to |
|---|---|
| make it smaller within each period | Segmentation |
| hit a target size | How small can you go? |
| change how periods are grouped | Clustering methods |
| change what each typical period keeps | Representations |
| keep the peak day exactly | Extreme periods |
| know what it will cost to run | How long will this take? |
| hand the result to a model | Optimization workflow |
Unsure which lever you need? Choosing a method walks all four on one dataset.