Skip to content

Pipeline Guide

When you call tsam.aggregate() it builds a configuration and hands off to run_pipeline(), which runs the aggregation in four phases:

  1. Prepare data — normalize the input and reshape it into clustering candidates.
  2. Cluster — group the periods and pick a representative for each.
  3. Refine — the optional adjustments: extremes, rescaling, segmentation.
  4. Build the result — denormalize, rebuild the series, pack the result object.

Phases 2 and 3 are split along what is mandatory and what is not: phase 2 is the one step every aggregation takes, and everything a caller can switch on or off sits in phase 3.

This page is the stable conceptual map of that flow. Each phase below names the stage functions it runs and links to their full reference — the precise signatures, parameters, and behavior live in the auto-generated API reference, which tracks the code. The public input and output types live in the API reference — Configuration and Results.


Overview

The diagram below traces the user-facing data flow on the left — a time series and a Config go into aggregate(), which returns the clustered data — through the four-phase run_pipeline() down the center, with the milestone dataclass passed between phases. The right column lists the clustering, representation, and segmentation options that Phases 2 and 3 draw on.

Pipeline data flow

Relation to Hoffmann et al. (2020)

Phases 1–3 implement the three feature-based merging steps from the time-series-aggregation review by Hoffmann et al. (2020): Preprocessing and Normalization (Phase 1), Algorithms, Distance Metrics, Representation (Phase 2), and Rescaling (Phase 3 step 4b). Phase 4 then reconstructs and packages the result. See Methodological positioning for how tsam sits in the review's overall taxonomy.


Entry points

Two ways in, both ending in run_pipeline():

tsam.aggregate() — the primary API. It validates inputs, derives n_timesteps_per_period from period_duration / temporal_resolution, and calls run_pipeline(). See its API reference for all parameters.

result = tsam.aggregate(df, n_clusters=8, period_duration=24)

ClusteringResult.apply() — reuse a fitted clustering on new data, skipping clustering in favor of the stored assignments:

result1 = tsam.aggregate(df_wind, n_clusters=8)
result2 = result1.clustering.apply(df_all)

Phase 1 — Prepare data

Turns the raw input into the candidate matrix the clustering stage consumes (steps 1–2, plus optional 2a / 2b). Orchestrated by prepare_data.

  1. Normalize — scale every column to [0, 1] so no column dominates the distance. → normalize
  2. Unstack to periods — reshape the flat series into a (period x timestep-feature) matrix. → unstack_to_periods
  3. 2a · Apply weights (optional, weights) — bake per-column weights into a copy of the candidates so they influence clustering distance only.
  4. 2b · Add period-sum features (optional, include_period_sums) — append per-period column sums as extra distance-only features. → add_period_sum_features

Milestone → PreparedData — normalized data, period profiles, the candidate matrix, and the weight vector.

Phase 2 — Cluster

Groups the periods and picks a representative for each (step 3). This is the one phase every aggregation runs in full — nothing in it is optional. Orchestrated by cluster_candidates.

  1. Cluster centers — group periods and pick a representative for each. → cluster_periods (the duration-curve and transfer variants handle use_duration_curves and ClusteringResult.apply()). Any period-sum features are trimmed back off the representatives afterwards, which stay in weighted space for the extreme detection that follows.

Milestone → ClusterAssignment — representatives, cluster order, and center indices.

Phase 3 — Refine

Applies everything that can still change which periods are represented or what they contain (step 4, plus optional 4a / 4b / 4c). Every stage here is switchable; with no extremes, no rescaling and no segmentation the phase only unweights and counts. Orchestrated by refine_representatives.

  • 4a · Add extremes (optional, ExtremeConfig) — inject extreme-value periods so peaks and troughs survive. → add_extreme_periods
  • Unweight · count — divide the weights back out, count cluster occurrences, and correct the padded last period's weight.
  • 4b · Rescale (optional, preserve_column_means) — scale non-extreme centers so their occurrence-weighted means match the original totals. → rescale_representatives
  • 4c · Segment (optional, SegmentConfig) — merge adjacent timesteps within each period into fewer segments. → segment_typical_periods

Milestone → RefinedRepresentatives — the final typical periods, still normalized, plus occurrence counts and the extreme/rescale/segmentation metadata.

Phase 4 — Build the result

Expresses the refined representatives in the user's units and packs them up (steps 5–7). Nothing here changes the aggregation any more. Orchestrated by build_result.

  1. Denormalize — convert the representatives back to the user's units. → denormalize
  2. Reconstruct + accuracy — expand the typical periods back to a full-length series and score it. → reconstruct, compute_accuracy
  3. Assemble — build the serializable, transferable ClusteringResult and pack it with the typical periods, counts, reconstruction, and metadata into the result that tsam.aggregate() returns as an AggregationResult.

Milestone → PipelineResult — the internal result that tsam.aggregate() wraps as an AggregationResult.


Reference

Full signatures and options live in the API reference — Configuration, Results, and Pipeline internals (the phase and stage functions above link straight into it). The source-tree module map is below.

Public surface
Module Responsibility
api.py aggregate() — the entry point: builds a PipelineConfig, runs the pipeline, wraps the output as an AggregationResult.
config.py Config dataclasses (ClusterConfig, SegmentConfig, ExtremeConfig, Distribution, MinMaxMean) plus the transfer object ClusteringResult.
result.py AggregationResult, AccuracyMetrics.
tuning.py Sweep configurations and rank by accuracy (loop aggregate()).
plot.py Plotly-based visualization (lazy import).
options.py Global numerical options and tolerances.
Pipeline internals
Module Responsibility
pipeline/orchestrator.py run_pipeline() plus the four phase functions and the glue with no dedicated stage module.
pipeline/normalize.py Scale columns to [0, 1] and invert it (normalize / denormalize).
pipeline/periods.py Reshape the flat series into a (period, timestep) matrix; optional period-sum features.
pipeline/clustering.py Config-aware clustering stage: adapts ClusterConfig (plus the duration-curve and transfer variants) onto algorithms/clustering.
pipeline/extremes.py Inject extreme-value periods into the cluster set.
pipeline/rescale.py Adjust representatives so column means match the original.
pipeline/segmentation.py Merge adjacent timesteps within a typical period.
pipeline/accuracy.py Reconstruct the full series and compute accuracy metrics.
pipeline/types.py Internal dataclasses: PipelineConfig, the phase milestones, PipelineResult.
algorithms/clustering.py · algorithms/representations.py Clustering dispatch — to scikit-learn or the algorithms/ k-medoids/k-maxoids solvers — and representative computation (shared by clustering and segmentation).
algorithms/k_medoids_exact.py · algorithms/k_maxoids.py k-medoids (MILP) / k-maxoids solvers.
algorithms/duration_representation.py Duration-curve representation (for distribution).
algorithms/segmentation.py Constrained agglomerative segmentation.
weights.py · exceptions.py Weight validation; custom warnings.