Pipeline Guide¶
When you call tsam.aggregate() it builds a configuration and
hands off to run_pipeline(), which runs the aggregation in four phases:
- Prepare data — normalize the input and reshape it into clustering candidates.
- Cluster — group the periods and pick a representative for each.
- Refine — the optional adjustments: extremes, rescaling, segmentation.
- Build the result — denormalize, rebuild the series, pack the result object.
Phases 2 and 3 are split along what is mandatory and what is not: phase 2 is the one step every aggregation takes, and everything a caller can switch on or off sits in phase 3.
This page is the stable conceptual map of that flow. Each phase below names the stage functions it runs and links to their full reference — the precise signatures, parameters, and behavior live in the auto-generated API reference, which tracks the code. The public input and output types live in the API reference — Configuration and Results.
Overview¶
The diagram below traces the user-facing data flow on the left — a time series
and a Config go into aggregate(), which returns the
clustered data — through the four-phase run_pipeline() down the center, with
the milestone dataclass passed between phases. The right column lists the
clustering, representation, and segmentation options that Phases 2 and 3 draw on.
Relation to Hoffmann et al. (2020)
Phases 1–3 implement the three feature-based merging steps from the time-series-aggregation review by Hoffmann et al. (2020): Preprocessing and Normalization (Phase 1), Algorithms, Distance Metrics, Representation (Phase 2), and Rescaling (Phase 3 step 4b). Phase 4 then reconstructs and packages the result. See Methodological positioning for how tsam sits in the review's overall taxonomy.
Entry points¶
Two ways in, both ending in run_pipeline():
tsam.aggregate() — the primary API. It validates inputs,
derives n_timesteps_per_period from period_duration / temporal_resolution,
and calls run_pipeline(). See its API reference for all
parameters.
ClusteringResult.apply() — reuse a fitted
clustering on new data, skipping clustering in favor of the stored assignments:
Phase 1 — Prepare data¶
Turns the raw input into the candidate matrix the clustering stage consumes
(steps 1–2, plus optional 2a / 2b). Orchestrated by
prepare_data.
- Normalize — scale every column to
[0, 1]so no column dominates the distance. →normalize - Unstack to periods — reshape the flat series into a
(period x timestep-feature)matrix. →unstack_to_periods - 2a · Apply weights (optional,
weights) — bake per-column weights into a copy of the candidates so they influence clustering distance only. - 2b · Add period-sum features (optional,
include_period_sums) — append per-period column sums as extra distance-only features. →add_period_sum_features
Milestone → PreparedData — normalized
data, period profiles, the candidate matrix, and the weight vector.
Phase 2 — Cluster¶
Groups the periods and picks a representative for each (step 3). This is the
one phase every aggregation runs in full — nothing in it is optional.
Orchestrated by
cluster_candidates.
- Cluster centers — group periods and pick a representative for each. →
cluster_periods(the duration-curve and transfer variants handleuse_duration_curvesandClusteringResult.apply()). Any period-sum features are trimmed back off the representatives afterwards, which stay in weighted space for the extreme detection that follows.
Milestone → ClusterAssignment —
representatives, cluster order, and center indices.
Phase 3 — Refine¶
Applies everything that can still change which periods are represented or
what they contain (step 4, plus optional 4a / 4b / 4c). Every stage here
is switchable; with no extremes, no rescaling and no segmentation the phase only
unweights and counts. Orchestrated by
refine_representatives.
- 4a · Add extremes (optional,
ExtremeConfig) — inject extreme-value periods so peaks and troughs survive. →add_extreme_periods - Unweight · count — divide the weights back out, count cluster occurrences, and correct the padded last period's weight.
- 4b · Rescale (optional,
preserve_column_means) — scale non-extreme centers so their occurrence-weighted means match the original totals. →rescale_representatives - 4c · Segment (optional,
SegmentConfig) — merge adjacent timesteps within each period into fewer segments. →segment_typical_periods
Milestone →
RefinedRepresentatives — the
final typical periods, still normalized, plus occurrence counts and the
extreme/rescale/segmentation metadata.
Phase 4 — Build the result¶
Expresses the refined representatives in the user's units and packs them up
(steps 5–7). Nothing here changes the aggregation any more. Orchestrated by
build_result.
- Denormalize — convert the representatives back to the user's units. →
denormalize - Reconstruct + accuracy — expand the typical periods back to a
full-length series and score it. →
reconstruct,compute_accuracy - Assemble — build the serializable, transferable
ClusteringResultand pack it with the typical periods, counts, reconstruction, and metadata into the result thattsam.aggregate()returns as anAggregationResult.
Milestone → PipelineResult — the
internal result that tsam.aggregate() wraps as an AggregationResult.
Reference¶
Full signatures and options live in the API reference — Configuration, Results, and Pipeline internals (the phase and stage functions above link straight into it). The source-tree module map is below.
Public surface
| Module | Responsibility |
|---|---|
api.py |
aggregate() — the entry point: builds a PipelineConfig, runs the pipeline, wraps the output as an AggregationResult. |
config.py |
Config dataclasses (ClusterConfig, SegmentConfig, ExtremeConfig, Distribution, MinMaxMean) plus the transfer object ClusteringResult. |
result.py |
AggregationResult, AccuracyMetrics. |
tuning.py |
Sweep configurations and rank by accuracy (loop aggregate()). |
plot.py |
Plotly-based visualization (lazy import). |
options.py |
Global numerical options and tolerances. |
Pipeline internals
| Module | Responsibility |
|---|---|
pipeline/orchestrator.py |
run_pipeline() plus the four phase functions and the glue with no dedicated stage module. |
pipeline/normalize.py |
Scale columns to [0, 1] and invert it (normalize / denormalize). |
pipeline/periods.py |
Reshape the flat series into a (period, timestep) matrix; optional period-sum features. |
pipeline/clustering.py |
Config-aware clustering stage: adapts ClusterConfig (plus the duration-curve and transfer variants) onto algorithms/clustering. |
pipeline/extremes.py |
Inject extreme-value periods into the cluster set. |
pipeline/rescale.py |
Adjust representatives so column means match the original. |
pipeline/segmentation.py |
Merge adjacent timesteps within a typical period. |
pipeline/accuracy.py |
Reconstruct the full series and compute accuracy metrics. |
pipeline/types.py |
Internal dataclasses: PipelineConfig, the phase milestones, PipelineResult. |
algorithms/clustering.py · algorithms/representations.py |
Clustering dispatch — to scikit-learn or the algorithms/ k-medoids/k-maxoids solvers — and representative computation (shared by clustering and segmentation). |
algorithms/k_medoids_exact.py · algorithms/k_maxoids.py |
k-medoids (MILP) / k-maxoids solvers. |
algorithms/duration_representation.py |
Duration-curve representation (for distribution). |
algorithms/segmentation.py |
Constrained agglomerative segmentation. |
weights.py · exceptions.py |
Weight validation; custom warnings. |