Skip to content

API Reference

ETHOS.TSAM's public API. Most users only need aggregate — the single entry point — plus the configuration objects and the result it returns. The pipeline internals are here for completeness.

Topic Contents
Configuration ClusterConfig, SegmentConfig, ExtremeConfig, Distribution, MinMaxMean
Results AggregationResult, AccuracyMetrics, ConcurrencyMetrics, ClusteringResult
Pipeline internals run_pipeline, the four phases, and the stage functions
Tuning Hyperparameter sweeps over aggregation settings
Utilities Options, weights, plotting, low-level aggregation helpers

Aggregation

tsam.api

New simplified API for tsam aggregation.

Functions:

Name Description
aggregate

Aggregate time series data into typical periods.

unstack_to_periods

Reshape time series data into period structure for visualization.

aggregate

aggregate(
    data: DataFrame,
    n_clusters: int,
    *,
    period_duration: int | float | str = 24,
    temporal_resolution: float | str | None = None,
    cluster: ClusterConfig | None = None,
    segments: SegmentConfig | None = None,
    extremes: ExtremeConfig | None = None,
    weights: dict[str, float] | None = None,
    preserve_column_means: bool = True,
    rescale_exclude_columns: list[str] | None = None,
    round_decimals: int | None = None,
    numerical_tolerance: float = 1e-13,
) -> AggregationResult

Aggregate time series data into typical periods.

This function reduces a time series dataset to a smaller set of representative "typical periods" using clustering algorithms.

Parameters:

Name Type Description Default
data DataFrame

Input time series data with a datetime index. Each column represents a different variable (e.g., solar, wind, demand). The index should be a DatetimeIndex with regular intervals.

required
n_clusters int

Number of clusters (typical periods) to create. Higher values mean more accuracy but less data reduction. Typical range: 4-20 for energy system models.

required
period_duration int | float | str

Length of each period. Accepts an int/float in hours (e.g., 24 for daily, 168 for weekly) or a pandas Timedelta string (e.g., '24h', '1d', '1w').

24
temporal_resolution float | str | None

Time resolution of input data. Accepts a float in hours (e.g., 1.0 for hourly, 0.25 for 15-minute) or a pandas Timedelta string (e.g., '1h', '15min', '30min'). If not provided, inferred from the datetime index.

None
cluster ClusterConfig | None

Clustering configuration. If not provided, uses defaults: method "hierarchical" and representation "medoid".

None
segments SegmentConfig | None

Segmentation configuration for reducing temporal resolution within periods. If not provided, no segmentation is applied.

None
extremes ExtremeConfig | None

Configuration for preserving extreme periods. If not provided, no extreme period handling is applied.

None
weights dict[str, float] | None

Per-column weights that influence all pipeline stages (clustering, segmentation, representation, rescaling). Higher weight means more influence on distance calculations, e.g. {"demand": 2.0, "solar": 1.0}.

None
preserve_column_means bool

Rescale typical periods so each column's weighted mean matches the original data's mean. Ensures total energy/load is preserved when weights represent occurrence counts.

True
rescale_exclude_columns list[str] | None

Column names to exclude from rescaling when preserve_column_means is True. Useful for binary/indicator columns (0/1 values) that should not be rescaled. If None, all columns are rescaled.

None
round_decimals int | None

Round output values to this many decimal places. If not provided, no rounding is applied.

None
numerical_tolerance float

Tolerance for numerical precision issues. Controls when warnings are raised for aggregated values exceeding the original time series bounds. Increase this value to silence warnings caused by floating-point precision errors.

1e-13

Returns:

Type Description
AggregationResult

An object containing the aggregated periods

AggregationResult

(cluster_representatives), the cluster assignment of each original

AggregationResult

period, the occurrence count per cluster (cluster_counts), accuracy

AggregationResult

metrics (RMSE, MAE), and a to_dict() method.

Raises:

Type Description
ValueError

If input data is invalid or parameters are inconsistent.

TypeError

If parameter types are incorrect.

Examples:

Basic usage with defaults:

>>> import tsam
>>> result = tsam.aggregate(df, n_clusters=8)
>>> typical = result.cluster_representatives

With custom clustering:

>>> from tsam import aggregate, ClusterConfig
>>> result = aggregate(
...     df,
...     n_clusters=8,
...     cluster=ClusterConfig(method="kmeans", representation="mean"),
... )

With segmentation (reduce to 12 timesteps per period):

>>> from tsam import aggregate, SegmentConfig
>>> result = aggregate(
...     df,
...     n_clusters=8,
...     segments=SegmentConfig(n_segments=12),
... )

Preserving peak demand periods:

>>> from tsam import aggregate, ExtremeConfig
>>> result = aggregate(
...     df,
...     n_clusters=8,
...     extremes=ExtremeConfig(max_value=["demand"]),
... )

Transferring assignments to new data:

>>> result1 = aggregate(df_wind, n_clusters=8)
>>> result2 = result1.clustering.apply(df_all)
Note

See ClusterConfig, SegmentConfig, ExtremeConfig, and AggregationResult for related configuration and result objects.

Source code in src/tsam/api.py
def aggregate(
    data: pd.DataFrame,
    n_clusters: int,
    *,
    period_duration: int | float | str = 24,
    temporal_resolution: float | str | None = None,
    cluster: ClusterConfig | None = None,
    segments: SegmentConfig | None = None,
    extremes: ExtremeConfig | None = None,
    weights: dict[str, float] | None = None,
    preserve_column_means: bool = True,
    rescale_exclude_columns: list[str] | None = None,
    round_decimals: int | None = None,
    numerical_tolerance: float = 1e-13,
) -> AggregationResult:
    """Aggregate time series data into typical periods.

    This function reduces a time series dataset to a smaller set of
    representative "typical periods" using clustering algorithms.

    Args:
        data: Input time series data with a datetime index. Each column
            represents a different variable (e.g., solar, wind, demand). The
            index should be a DatetimeIndex with regular intervals.
        n_clusters: Number of clusters (typical periods) to create. Higher
            values mean more accuracy but less data reduction. Typical range:
            4-20 for energy system models.
        period_duration: Length of each period. Accepts an int/float in hours
            (e.g., 24 for daily, 168 for weekly) or a pandas Timedelta string
            (e.g., '24h', '1d', '1w').
        temporal_resolution: Time resolution of input data. Accepts a float in
            hours (e.g., 1.0 for hourly, 0.25 for 15-minute) or a pandas
            Timedelta string (e.g., '1h', '15min', '30min'). If not provided,
            inferred from the datetime index.
        cluster: Clustering configuration. If not provided, uses defaults:
            method "hierarchical" and representation "medoid".
        segments: Segmentation configuration for reducing temporal resolution
            within periods. If not provided, no segmentation is applied.
        extremes: Configuration for preserving extreme periods. If not provided,
            no extreme period handling is applied.
        weights: Per-column weights that influence all pipeline stages
            (clustering, segmentation, representation, rescaling). Higher weight
            means more influence on distance calculations, e.g.
            ``{"demand": 2.0, "solar": 1.0}``.
        preserve_column_means: Rescale typical periods so each column's weighted
            mean matches the original data's mean. Ensures total energy/load is
            preserved when weights represent occurrence counts.
        rescale_exclude_columns: Column names to exclude from rescaling when
            ``preserve_column_means`` is True. Useful for binary/indicator
            columns (0/1 values) that should not be rescaled. If None,
            all columns are rescaled.
        round_decimals: Round output values to this many decimal places. If not
            provided, no rounding is applied.
        numerical_tolerance: Tolerance for numerical precision issues. Controls
            when warnings are raised for aggregated values exceeding the original
            time series bounds. Increase this value to silence warnings caused by
            floating-point precision errors.

    Returns:
        An object containing the aggregated periods
        (``cluster_representatives``), the cluster assignment of each original
        period, the occurrence count per cluster (``cluster_counts``), accuracy
        metrics (RMSE, MAE), and a ``to_dict()`` method.

    Raises:
        ValueError: If input data is invalid or parameters are inconsistent.
        TypeError: If parameter types are incorrect.

    Examples:
        Basic usage with defaults:

        >>> import tsam
        >>> result = tsam.aggregate(df, n_clusters=8)
        >>> typical = result.cluster_representatives

        With custom clustering:

        >>> from tsam import aggregate, ClusterConfig
        >>> result = aggregate(
        ...     df,
        ...     n_clusters=8,
        ...     cluster=ClusterConfig(method="kmeans", representation="mean"),
        ... )

        With segmentation (reduce to 12 timesteps per period):

        >>> from tsam import aggregate, SegmentConfig
        >>> result = aggregate(
        ...     df,
        ...     n_clusters=8,
        ...     segments=SegmentConfig(n_segments=12),
        ... )

        Preserving peak demand periods:

        >>> from tsam import aggregate, ExtremeConfig
        >>> result = aggregate(
        ...     df,
        ...     n_clusters=8,
        ...     extremes=ExtremeConfig(max_value=["demand"]),
        ... )

        Transferring assignments to new data:

        >>> result1 = aggregate(df_wind, n_clusters=8)
        >>> result2 = result1.clustering.apply(df_all)

    Note:
        See ``ClusterConfig``, ``SegmentConfig``, ``ExtremeConfig``, and
        ``AggregationResult`` for related configuration and result objects.
    """
    # Validate input
    if not isinstance(data, pd.DataFrame):
        raise TypeError(f"data must be a pandas DataFrame, got {type(data).__name__}")

    if not isinstance(n_clusters, int) or n_clusters < 1:
        raise ValueError(f"n_clusters must be a positive integer, got {n_clusters}")

    # Parse duration parameters to hours
    period_duration = parse_duration_hours(period_duration, "period_duration")
    if period_duration <= 0:
        raise ValueError(f"period_duration must be positive, got {period_duration}")

    temporal_resolution = (
        parse_duration_hours(temporal_resolution, "temporal_resolution")
        if temporal_resolution is not None
        else None
    )
    if temporal_resolution is not None and temporal_resolution <= 0:
        raise ValueError(
            f"temporal_resolution must be positive, got {temporal_resolution}"
        )

    # Apply defaults
    if cluster is None:
        cluster = ClusterConfig()

    # Compute n_timesteps_per_period
    if temporal_resolution is not None:
        resolution = temporal_resolution
    else:
        # Infer resolution from data index
        if isinstance(data.index, pd.DatetimeIndex) and len(data.index) > 1:
            resolution = (data.index[1] - data.index[0]).total_seconds() / 3600
        else:
            resolution = 1.0  # Default to hourly

    timesteps_per_period = period_duration / resolution
    n_timesteps_per_period = round(timesteps_per_period)
    if abs(timesteps_per_period - n_timesteps_per_period) > 1e-9 * max(
        1.0, timesteps_per_period
    ):
        raise ValueError(
            "The combination of period_duration and temporal_resolution "
            "does not result in an integer number of time steps per period"
        )
    n_timesteps_per_period = int(n_timesteps_per_period)

    # Validate segments against data
    if segments is not None:
        if segments.n_segments > n_timesteps_per_period:
            raise ValueError(
                f"n_segments ({segments.n_segments}) cannot exceed "
                f"timesteps per period ({n_timesteps_per_period})"
            )
        seg_rep = segments.representation
        if isinstance(seg_rep, Distribution) and (
            seg_rep.reference_attribute is not None
            or (
                seg_rep.concurrency is not None and seg_rep.concurrency != "independent"
            )
        ):
            raise ValueError(
                "concurrency / reference_attribute are not supported for segment "
                "representations: each segment collapses to a single value per "
                "attribute (one time step), so there is no within-period time "
                "axis left to order. Set them on the cluster representation "
                "instead — the ordering is applied before segmentation."
            )

    # Validate extreme columns exist in data
    if extremes is not None:
        all_extreme_cols = (
            extremes.max_value
            + extremes.min_value
            + extremes.max_period
            + extremes.min_period
        )
        missing = set(all_extreme_cols) - set(data.columns)
        if missing:
            raise ValueError(f"Extreme period columns not found in data: {missing}")

    # Validate and normalize weights (a pipeline input, not part of ClusterConfig)
    validated_weights = validate_weights(data.columns, weights)

    # Build pipeline config
    cfg = PipelineConfig(
        n_clusters=n_clusters,
        n_timesteps_per_period=n_timesteps_per_period,
        cluster=cluster,
        weights=validated_weights,
        extremes=extremes if extremes and extremes.has_extremes() else None,
        segments=segments,
        rescale_cluster_periods=preserve_column_means,
        rescale_exclude_columns=rescale_exclude_columns,
        round_decimals=round_decimals,
        numerical_tolerance=numerical_tolerance,
        temporal_resolution=temporal_resolution,
    )

    result = run_pipeline(data=data, cfg=cfg)

    return _build_aggregation_result(result, is_transferred=False)

unstack_to_periods

unstack_to_periods(
    data: DataFrame, period_duration: int | float | str = 24
) -> pd.DataFrame

Reshape time series data into period structure for visualization.

Transforms a flat time series into a DataFrame with periods as rows and timesteps as a MultiIndex level, suitable for creating heatmaps with plotly.

Parameters:

Name Type Description Default
data DataFrame

Time series data with datetime index.

required
period_duration int | float | str

Length of each period. Accepts an int/float in hours (e.g., 24 for daily, 168 for weekly) or a pandas Timedelta string (e.g., '24h', '1d', '1w').

24

Returns:

Type Description
DataFrame

Reshaped data with shape (n_periods, n_timesteps_per_period) for each

DataFrame

column. Suitable for px.imshow(result["column"].values.T) to create

DataFrame

heatmaps.

Examples:

>>> import tsam
>>> import plotly.express as px
>>>
>>> # Reshape data for heatmap visualization
>>> unstacked = tsam.unstack_to_periods(df, period_duration=24)
>>>
>>> # Create heatmap with plotly
>>> px.imshow(
...     unstacked["Load"].values.T,
...     labels={"x": "Day", "y": "Hour", "color": "Load"},
...     title="Load Heatmap"
... )
Source code in src/tsam/api.py
def unstack_to_periods(
    data: pd.DataFrame,
    period_duration: int | float | str = 24,
) -> pd.DataFrame:
    """Reshape time series data into period structure for visualization.

    Transforms a flat time series into a DataFrame with periods as rows and
    timesteps as a MultiIndex level, suitable for creating heatmaps with plotly.

    Args:
        data: Time series data with datetime index.
        period_duration: Length of each period. Accepts an int/float in hours
            (e.g., 24 for daily, 168 for weekly) or a pandas Timedelta string
            (e.g., '24h', '1d', '1w').

    Returns:
        Reshaped data with shape (n_periods, n_timesteps_per_period) for each
        column. Suitable for ``px.imshow(result["column"].values.T)`` to create
        heatmaps.

    Examples:
        >>> import tsam
        >>> import plotly.express as px
        >>>
        >>> # Reshape data for heatmap visualization
        >>> unstacked = tsam.unstack_to_periods(df, period_duration=24)
        >>>
        >>> # Create heatmap with plotly
        >>> px.imshow(
        ...     unstacked["Load"].values.T,
        ...     labels={"x": "Day", "y": "Hour", "color": "Load"},
        ...     title="Load Heatmap"
        ... )
    """
    period_hours = parse_duration_hours(period_duration, "period_duration")

    # Infer timestep resolution from data index
    timestep_hours = 1.0  # Default to hourly
    if isinstance(data.index, pd.DatetimeIndex) and len(data.index) > 1:
        timestep_hours = (data.index[1] - data.index[0]).total_seconds() / 3600

    # Calculate timesteps per period
    timesteps_per_period = round(period_hours / timestep_hours)
    if timesteps_per_period < 1:
        raise ValueError(
            f"period_duration ({period_hours}h) is smaller than "
            f"data timestep resolution ({timestep_hours}h)"
        )

    from tsam.pipeline.periods import unstack_to_periods as _unstack

    profiles = _unstack(data.copy(), timesteps_per_period)
    return cast("pd.DataFrame", profiles.profiles_dataframe)