Extreme periods¶
Clustering finds typical behavior, so a once-a-year peak gets blended into an average and lost.
ExtremeConfig forces chosen extremes to be kept exactly, alongside the normal clustering.
In [1]:
Copied!
import pandas as pd
import plotly.io as pio
import tsam
pio.renderers.default = "notebook_connected"
raw = pd.read_csv("../data/testdata.csv", index_col=0, parse_dates=True)
data = raw.loc["2010-01-01":"2010-02-11"] # six weeks of hourly data
import pandas as pd
import plotly.io as pio
import tsam
pio.renderers.default = "notebook_connected"
raw = pd.read_csv("../data/testdata.csv", index_col=0, parse_dates=True)
data = raw.loc["2010-01-01":"2010-02-11"] # six weeks of hourly data
The peak gets averaged away¶
In [2]:
Copied!
base = tsam.aggregate(data, n_clusters=6, period_duration="1D")
print(f"original peak Load: {data['Load'].max():.1f}")
print(f"typical-period peak: {base.cluster_representatives['Load'].max():.1f}")
base = tsam.aggregate(data, n_clusters=6, period_duration="1D")
print(f"original peak Load: {data['Load'].max():.1f}")
print(f"typical-period peak: {base.cluster_representatives['Load'].max():.1f}")
original peak Load: 636.5 typical-period peak: 601.5
Keep it exactly¶
Target the period holding the single highest/lowest value (max_value/min_value) or the
highest/lowest period total (max_period/min_period):
In [3]:
Copied!
from tsam import ExtremeConfig
kept = tsam.aggregate(
data,
n_clusters=6,
period_duration="1D",
extremes=ExtremeConfig(method="new_cluster", max_value=["Load"]),
)
print(f"clusters: {base.n_clusters} -> {kept.n_clusters}")
print(f"typical-period peak now: {kept.cluster_representatives['Load'].max():.1f}")
from tsam import ExtremeConfig
kept = tsam.aggregate(
data,
n_clusters=6,
period_duration="1D",
extremes=ExtremeConfig(method="new_cluster", max_value=["Load"]),
)
print(f"clusters: {base.n_clusters} -> {kept.n_clusters}")
print(f"typical-period peak now: {kept.cluster_representatives['Load'].max():.1f}")
clusters: 6 -> 7 typical-period peak now: 636.5
/home/docs/checkouts/readthedocs.org/user_builds/tsam/checkouts/latest/src/tsam/pipeline/orchestrator.py:211: UserWarning: At least one maximal value of the aggregated time series exceeds the maximal value the input time series for: {'T': 1.7763568394002505e-15, 'Load': 1.1368683772161603e-13}. To silence the warning set the 'numerical_tolerance' to a higher value.
warnings.warn(
How the extreme enters: method¶
new_cluster— add it as its own cluster, kept exactly, never blended. Recommended.append— add it as an extra typical period.replace— substitute the most similar existing cluster (cluster count unchanged).
In [4]:
Copied!
for method in ["new_cluster", "append", "replace"]:
r = tsam.aggregate(
data,
n_clusters=6,
period_duration="1D",
extremes=ExtremeConfig(method=method, max_value=["Load"]),
)
print(
f"{method:12s} -> {r.n_clusters} clusters, "
f"peak {r.cluster_representatives['Load'].max():.1f}"
)
for method in ["new_cluster", "append", "replace"]:
r = tsam.aggregate(
data,
n_clusters=6,
period_duration="1D",
extremes=ExtremeConfig(method=method, max_value=["Load"]),
)
print(
f"{method:12s} -> {r.n_clusters} clusters, "
f"peak {r.cluster_representatives['Load'].max():.1f}"
)
new_cluster -> 7 clusters, peak 636.5 append -> 7 clusters, peak 636.5 replace -> 6 clusters, peak 636.5
/home/docs/checkouts/readthedocs.org/user_builds/tsam/checkouts/latest/src/tsam/pipeline/orchestrator.py:211: UserWarning: At least one maximal value of the aggregated time series exceeds the maximal value the input time series for: {'T': 1.7763568394002505e-15, 'Load': 1.1368683772161603e-13}. To silence the warning set the 'numerical_tolerance' to a higher value.
warnings.warn(
/home/docs/checkouts/readthedocs.org/user_builds/tsam/checkouts/latest/src/tsam/pipeline/orchestrator.py:211: UserWarning: At least one maximal value of the aggregated time series exceeds the maximal value the input time series for: {'T': 1.7763568394002505e-15, 'Load': 1.1368683772161603e-13}. To silence the warning set the 'numerical_tolerance' to a higher value.
warnings.warn(
/home/docs/checkouts/readthedocs.org/user_builds/tsam/checkouts/latest/src/tsam/pipeline/orchestrator.py:211: UserWarning: At least one maximal value of the aggregated time series exceeds the maximal value the input time series for: {'T': 1.7763568394002505e-15, 'Load': 1.1368683772161603e-13}. To silence the warning set the 'numerical_tolerance' to a higher value.
warnings.warn(
Several at once¶
Request as many as you need. A period that is extreme on more than one criterion is added once.
In [5]:
Copied!
multi = tsam.aggregate(
data,
n_clusters=6,
period_duration="1D",
extremes=ExtremeConfig(method="new_cluster", max_value=["Load"], min_value=["T"]),
)
print(f"clusters: {base.n_clusters} -> {multi.n_clusters}")
kept.plot.cluster_counts()
multi = tsam.aggregate(
data,
n_clusters=6,
period_duration="1D",
extremes=ExtremeConfig(method="new_cluster", max_value=["Load"], min_value=["T"]),
)
print(f"clusters: {base.n_clusters} -> {multi.n_clusters}")
kept.plot.cluster_counts()
/home/docs/checkouts/readthedocs.org/user_builds/tsam/checkouts/latest/src/tsam/pipeline/orchestrator.py:211: UserWarning: At least one maximal value of the aggregated time series exceeds the maximal value the input time series for: {'T': 1.7763568394002505e-15, 'Load': 1.1368683772161603e-13}. To silence the warning set the 'numerical_tolerance' to a higher value.
warnings.warn(
clusters: 6 -> 8
Extreme days stand in for just themselves, so they carry a small occurrence count next to the
typical clusters. They are also excluded from rescaling by design, so unlike a maxoid
representation they do not fight preserve_column_means.
- Choosing a method — the three different ways to preserve a peak, and why they are not interchangeable.
- Representations —
maxoid/MinMaxMean(...), the lighter touch. - Rescaling — how mean preservation works.