API Reference

Package Root

The public loss classes are available from the package root:

from ddtw import dDTW, SDTW, CTC, DTW

Core Loss

class ddtw.ddtw.dDTW(*args: Any, **kwargs: Any)[source]

Bases: Module

Initialize the general dDTW loss function.

See [1] for the graph formulation.

[1] Johannes Zeitler and Meinard Müller. A Unified Perspective on CTC and Soft-DTW Using Differentiable DTW. IEEE Transactions on Audio, Speech and Language Processing, vol. 34, pages 936-951, 2026.

Parameters:
  • cost_function (str or callable, optional) – Local cost function used when X and Y are passed to forward(). Built-in strings are "MSE", "BCE", and "CTC". A callable must return a cost tensor with shape (B, N, M). Default: "MSE".

  • min_function (str, optional) – Recursive minimum or differentiable approximation. Choose among "softmin", "sparsemin", "smoothmin", and "hardmin". Default: "softmin".

  • gamma (float, optional) – Temperature parameter used by differentiable minimum functions. Default: 1.0.

  • step_sizes (list of list of int, optional) – Alignment step sizes [dn, dm]. Each step points from the current cell (n, m) to predecessor (n-dn, m-dm). Default: [[1, 0], [0, 1], [1, 1]].

  • global_step_weights (list of float, optional) – Local cost weights associated with step_sizes. Must contain one scalar per step. Default: [1.0, 1.0, 1.0].

  • normalization (str, optional) – Normalization applied to each batch loss before averaging. Choose among "N", "M", "NM", "ctc", and "none". Default: "N".

  • backend (str, optional) – Backend to use. Choose among "auto", "torch", "cpu_numba", and "cuda_cpp". "auto" tries CUDA first, then Numba CPU, then pure PyTorch. Default: "auto".

  • dtype_float (torch.dtype, optional) – Floating-point dtype for internal tensors. The CUDA C++ backend currently requires torch.float32. Default: torch.float32.

  • cuda_device (str or torch.device, optional) – CUDA device to use, for example "cuda:0". Default: None.

  • store_debug (bool, optional) – If True, retain intermediate matrices on the backend class for inspection after forward/backward. Default: False.

forward(X=None, Y=None, C=None, B_start=None, B_end=None, list_N=None, list_M=None, local_step_weights=None, start_penalty=None, end_penalty=None, num_start_conditions=None, num_end_conditions=None)[source]

Compute the dDTW loss.

Pass either X and Y or a precomputed cost matrix C. If X and Y are provided, self.cost_function computes C.

Parameters:
  • X (torch.Tensor, optional) – First input sequence with shape (B, N, D). Usually the model predictions.

  • Y (torch.Tensor, optional) – Second input sequence with shape (B, M, D). Usually the target or reference sequence.

  • C (torch.Tensor, optional) – Precomputed local cost matrix with shape (B, N, M).

  • B_start (list or torch.Tensor, optional) – Start boundary conditions. For each batch item, stores one or more zero-based [n, m] cells. Tensor form must have shape (B, max_start_conditions, 2). If None, defaults to [[0, 0]] for each batch item.

  • B_end (list or torch.Tensor, optional) – End boundary conditions. For each batch item, stores one or more zero-based [n, m] cells. Tensor form must have shape (B, max_end_conditions, 2). If None, defaults to [[list_N[b] - 1, list_M[b] - 1]].

  • list_N (list or torch.Tensor, optional) – Active lengths along the X/row axis with shape (B,). If None, all batch items use the full padded length N.

  • list_M (list or torch.Tensor, optional) – Active lengths along the Y/column axis with shape (B,). If None, all batch items use the full padded length M.

  • local_step_weights (torch.Tensor, optional) – Cell-wise step weights with shape (B, N, M, S), where S is the number of configured steps. If None, global_step_weights are broadcast to all cells.

  • start_penalty (list or torch.Tensor, optional) – Multiplicative local-cost weights for start boundary conditions. Tensor form must have shape (B, max_start_conditions). If None, defaults to 1 for every start condition.

  • end_penalty (list or torch.Tensor, optional) – Multiplicative local-cost weights for end boundary conditions. Tensor form must have shape (B, max_end_conditions). If None, defaults to 0 for every end condition.

  • num_start_conditions (list or torch.Tensor, optional) – Number of valid start conditions for each batch item when B_start is padded. Shape (B,).

  • num_end_conditions (list or torch.Tensor, optional) – Number of valid end conditions for each batch item when B_end is padded. Shape (B,).

Returns:

Scalar batch-mean dDTW loss after the configured normalization.

Return type:

torch.Tensor

Loss Variants

class ddtw.ddtw_variants.SDTW(*args: Any, **kwargs: Any)[source]

Bases: dDTW

Initialize SDTW loss function, see [1, 2].

[1] Marco Cuturi and Mathieu Blondel. Soft-DTW: A Differentiable Loss Function for Time-Series. In Proceedings of the International Conference on Neural Information Processing Systems (NIPS), vol. 2, pages 2292-2300, 2013.

[2] Johannes Zeitler and Meinard Müller. A Unified Perspective on CTC and Soft-DTW Using Differentiable DTW. IEEE Transactions on Audio, Speech and Language Processing, vol. 34, pages 936-951, 2026.

Parameters:
  • cost_function (str) – Local cost function for pair-wise comparison of sequence elements. Choose among (“MSE”, “BCE”, “CTC”). Default: “MSE”.

  • gamma (float) – Softmin temperature hyperparameter. Default: 1.0

  • step_sizes (list) – Alignment step sizes in [n,m] direction, given as list of tuples. Default: [[1,0], [0,1], [1,1]]

  • global_step_weights (list) – Step weights associated to the step sizes. Default: [1.0, 1.0, 1.0]

  • normalization (str) – Normalization of SDTW cost. Choose among (“M”, “N”, “NM”, “ctc”, “none”). “N”: divide by N. “M”: divide by M. “NM”: divide by (N*M). “ctc”: divide by (M-1)/2. “none”: no normalization. Default: “N”

  • backend (str) – Backend to use. Choose among (“auto”, “torch”, “cpu_numba”, “cuda_cpp”). Default: “auto”

  • dtype_float (torch.dtype) – Number format for internal computations. Default: torch.float32

  • cuda_device (str or torch.device, optional) – CUDA device to use, for example "cuda:0". Default: None.

  • store_debug (bool) – Whether to retain intermediate backend matrices for inspection. Default: False.

forward(X=None, Y=None, C=None, list_N=None, list_M=None)[source]

Compute the SDTW loss.

Parameters:
  • X (torch.tensor [shape=(B, N, D)]) – Input sequence, usually the DNN predictions.

  • Y (torch.tensor [shape=(B, M, D)]) – Input sequence, usually the weak targets.

  • C (torch.tensor [shape=(B, N, M)]) – Pre-computed local cost matrix C.

  • list_N (torch.tensor [shape=(B)]) – Sequence lengths of X (<= N) of the individual batch elements. If None, defaults to [N, N, …, N]

  • list_M (torch.tensor [shape=(B)]) – Sequence lengths of Y (<= M) of the individual batch elements. If None, defaults to [M, M, …, M]

Returns:

Scalar batch-mean SDTW loss.

Return type:

torch.Tensor

class ddtw.ddtw_variants.DTW(*args: Any, **kwargs: Any)[source]

Bases: dDTW

Initialize DTW loss function, see [3].

[3]. Meinard Müller. Fundamentals of Music Processing - Using Python and Jupyter Notebooks. Springer Verlag, 2nd edition, 2021.

Parameters:
  • cost_function (str) – Local cost function for pair-wise comparison of sequence elements. Choose among (“MSE”, “BCE”, “CTC”). Default: “MSE”.

  • step_sizes (list) – Alignment step sizes in [n,m] direction, given as list of tuples. Default: [[1,0], [0,1], [1,1]]

  • global_step_weights (list) – Step weights associated to the step sizes. Default: [1.0, 1.0, 1.0]

  • normalization (str) – Normalization of SDTW cost. Choose among (“M”, “N”, “NM”, “ctc”, “none”). “N”: divide by N. “M”: divide by M. “NM”: divide by (N*M). “ctc”: divide by (M-1)/2. “none”: no normalization. Default: “N”

  • backend (str) – Backend to use. Choose among (“auto”, “torch”, “cpu_numba”, “cuda_cpp”). Default: “auto”

  • dtype_float (torch.dtype) – Number format for internal computations. Default: torch.float32

  • cuda_device (str or torch.device, optional) – CUDA device to use, for example "cuda:0". Default: None.

  • store_debug (bool) – Whether to retain intermediate backend matrices for inspection. Default: False.

forward(X=None, Y=None, C=None, list_N=None, list_M=None)[source]

Compute the DTW loss.

Parameters:
  • X (torch.tensor [shape=(B, N, D)]) – Input sequence, usually the DNN predictions.

  • Y (torch.tensor [shape=(B, M, D)]) – Input sequence, usually the weak targets.

  • C (torch.tensor [shape=(B, N, M)]) – Pre-computed local cost matrix C.

  • list_N (torch.tensor [shape=(B)]) – Sequence lengths of X (<= N) of the individual batch elements. If None, defaults to [N, N, …, N]

  • list_M (torch.tensor [shape=(B)]) – Sequence lengths of Y (<= M) of the individual batch elements. If None, defaults to [M, M, …, M]

Returns:

Scalar batch-mean DTW loss. With a precomputed C, gradients with respect to C mark the selected hard warping path.

Return type:

torch.Tensor

class ddtw.ddtw_variants.smoothDTW(*args: Any, **kwargs: Any)[source]

Bases: dDTW

Initialize smoothDTW loss function, see [4].

[4]. Isma Hadji, K. Derpanis, and A. Jepson. Representation learning via global temporal alignment and cycle-consistency. In IEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 11068-11077, 2021.

Parameters:
  • cost_function (str) – Local cost function for pair-wise comparison of sequence elements. Choose among (“MSE”, “BCE”, “CTC”). Default: “MSE”.

  • gamma (float) – Softmin temperature hyperparameter. Default: 1.0

  • step_sizes (list) – Alignment step sizes in [n,m] direction, given as list of tuples. Default: [[1,0], [0,1], [1,1]]

  • global_step_weights (list) – Step weights associated to the step sizes. Default: [1.0, 1.0, 1.0]

  • normalization (str) – Normalization of SDTW cost. Choose among (“M”, “N”, “NM”, “ctc”, “none”). “N”: divide by N. “M”: divide by M. “NM”: divide by (N*M). “ctc”: divide by (M-1)/2. “none”: no normalization. Default: “N”

  • backend (str) – Backend to use. Choose among (“auto”, “torch”, “cpu_numba”, “cuda_cpp”). Default: “auto”

  • dtype_float (torch.dtype) – Number format for internal computations. Default: torch.float32

  • cuda_device (str or torch.device, optional) – CUDA device to use, for example "cuda:0". Default: None.

  • store_debug (bool) – Whether to retain intermediate backend matrices for inspection. Default: False.

forward(X=None, Y=None, C=None, list_N=None, list_M=None)[source]

Compute the smoothDTW loss.

Parameters:
  • X (torch.tensor [shape=(B, N, D)]) – Input sequence, usually the DNN predictions.

  • Y (torch.tensor [shape=(B, M, D)]) – Input sequence, usually the weak targets.

  • C (torch.tensor [shape=(B, N, M)]) – Pre-computed local cost matrix C.

  • list_N (torch.tensor [shape=(B)]) – Sequence lengths of X (<= N) of the individual batch elements. If None, defaults to [N, N, …, N]

  • list_M (torch.tensor [shape=(B)]) – Sequence lengths of Y (<= M) of the individual batch elements. If None, defaults to [M, M, …, M]

Returns:

Scalar batch-mean smoothDTW loss.

Return type:

torch.Tensor

class ddtw.ddtw_variants.sparseDTW(*args: Any, **kwargs: Any)[source]

Bases: dDTW

Initialize sparseDTW loss function, see [5].

[5] Arthur Mensch and Mathieu Blondel. Differentiable Dynamic Programming for Structured Prediction and Attention. In Proceedings of the International Converence on Machine Learning (ICML), pages 3459-3468, Stockholm, Sweden, 2018.

Parameters:
  • cost_function (str) – Local cost function for pair-wise comparison of sequence elements. Choose among (“MSE”, “BCE”, “CTC”). Default: “MSE”.

  • gamma (float) – Sparsemin temperature hyperparameter. Default: 1.0

  • step_sizes (list) – Alignment step sizes in [n,m] direction, given as list of tuples. Default: [[1,0], [0,1], [1,1]]

  • global_step_weights (list) – Step weights associated to the step sizes. Default: [1.0, 1.0, 1.0]

  • normalization (str) – Normalization of SDTW cost. Choose among (“M”, “N”, “NM”, “ctc”, “none”). “N”: divide by N. “M”: divide by M. “NM”: divide by (N*M). “ctc”: divide by (M-1)/2. “none”: no normalization. Default: “N”

  • backend (str) – Backend to use. Choose among (“auto”, “torch”, “cpu_numba”, “cuda_cpp”). Default: “auto”

  • dtype_float (torch.dtype) – Number format for internal computations. Default: torch.float32

  • cuda_device (str or torch.device, optional) – CUDA device to use, for example "cuda:0". Default: None.

  • store_debug (bool) – Whether to retain intermediate backend matrices for inspection. Default: False.

forward(X=None, Y=None, C=None, list_N=None, list_M=None)[source]

Compute the sparseDTW loss.

Parameters:
  • X (torch.tensor [shape=(B, N, D)]) – Input sequence, usually the DNN predictions.

  • Y (torch.tensor [shape=(B, M, D)]) – Input sequence, usually the weak targets.

  • C (torch.tensor [shape=(B, N, M)]) – Pre-computed local cost matrix C.

  • list_N (torch.tensor [shape=(B)]) – Sequence lengths of X (<= N) of the individual batch elements. If None, defaults to [N, N, …, N]

  • list_M (torch.tensor [shape=(B)]) – Sequence lengths of Y (<= M) of the individual batch elements. If None, defaults to [M, M, …, M]

Returns:

Scalar batch-mean sparseDTW loss.

Return type:

torch.Tensor

class ddtw.ddtw_variants.subSDTW(*args: Any, **kwargs: Any)[source]

Bases: dDTW

Initialize subsequence SDTW loss function, see [7].

[7] Johannes Zeitler and Meinard Müller. Subsequence Soft Dynamic Time Warping. In Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing (CASSP), Barcelona, Spain, 2026.

Parameters:
  • cost_function (str) – Local cost function for pair-wise comparison of sequence elements. Choose among (“MSE”, “BCE”, “CTC”). Default: “MSE”.

  • min_function (str) – Minimum function or approximation thereof. Choose among (“softmin”, “sparsemin”, “smoothmin”, “hardmin”). Default: “softmin”.

  • gamma (float) – Min. function temperature hyperparameter. Default: 1.0

  • step_sizes (list) – Alignment step sizes in [n,m] direction, given as list of tuples. Default: [[1,0], [0,1], [1,1]]

  • global_step_weights (list) – Step weights associated to the step sizes. Default: [1.0, 1.0, 1.0]

  • normalization (str) – Normalization of SDTW cost. Choose among (“M”, “N”, “NM”, “ctc”, “none”). “N”: divide by N. “M”: divide by M. “NM”: divide by (N*M). “ctc”: divide by (M-1)/2. “none”: no normalization. Default: “N”

  • backend (str) – Backend to use. Choose among (“auto”, “torch”, “cpu_numba”, “cuda_cpp”). Default: “auto”

  • dtype_float (torch.dtype) – Number format for internal computations. Default: torch.float32

  • cuda_device (str or torch.device, optional) – CUDA device to use, for example "cuda:0". Default: None.

  • store_debug (bool) – Whether to retain intermediate backend matrices for inspection. Default: False.

  • sub_X (bool) – Whether to allow subsequence starts and ends along the X-direction (row axis). Default: True.

  • sub_Y (bool) – Whether to allow subsequence starts and ends along the Y-direction (column axis). Default: True.

  • compensate_subseq (bool) – Whether to add start/end penalties for skipped prefixes or suffixes. Default: True.

forward(X=None, Y=None, C=None, list_N=None, list_M=None)[source]

Compute the subsequence SDTW loss.

Parameters:
  • X (torch.tensor [shape=(B, N, D)]) – Input sequence, usually the DNN predictions.

  • Y (torch.tensor [shape=(B, M, D)]) – Input sequence, usually the weak targets.

  • C (torch.tensor [shape=(B, N, M)]) – Pre-computed local cost matrix C.

  • list_N (torch.tensor [shape=(B)]) – Sequence lengths of X (<= N) of the individual batch elements. If None, defaults to [N, N, …, N]

  • list_M (torch.tensor [shape=(B)]) – Sequence lengths of Y (<= M) of the individual batch elements. If None, defaults to [M, M, …, M]

Returns:

Scalar batch-mean subsequence SDTW loss.

Return type:

torch.Tensor

class ddtw.ddtw_variants.CTC(*args: Any, **kwargs: Any)[source]

Bases: dDTW

Initialize CTC loss function, parameterized within the dDTW framework, see [8, 2].

[2] Johannes Zeitler and Meinard Müller. A Unified Perspective on CTC and Soft-DTW Using Differentiable DTW. IEEE Transactions on Audio, Speech and Language Processing, vol. 34, pages 936-951, 2026.

[8] Alex Graves et al. Connectionist Temporal Classification: Labelling Unsegmented Sequence Data with Recurrent Neural Networks. In Proceedings of the International Conference on Machine Learning (ICML), pages 369-376, Pittsburgh, Pennsylvania, USA, 2006.

Parameters:
  • blank_penalty_weight (float) – Penalty for alignment of the blank symbol. Default: 1.0 (no penalty).

  • blank_index (int) – Class index of the blank symbol. Default: 0.

  • gamma (float) – Softmin temperature hyperparameter. Default: 1.0

  • global_step_weights (list) – Step weights associated to the CTC step sizes [[1, 0], [1, 1], [1, 2]]. Default: [1.0, 1.0, 1.0].

  • backend (str) – Backend to use. Choose among (“auto”, “torch”, “cpu_numba”, “cuda_cpp”). Default: “auto”

  • dtype_float (torch.dtype) – Number format for internal computations. Default: torch.float32

  • cuda_device (str or torch.device, optional) – CUDA device to use, for example "cuda:0". Default: None.

  • store_debug (bool) – Whether to retain intermediate backend matrices for inspection. Default: False.

Notes

The CTC variant fixes cost_function="CTC", min_function="softmin", step_sizes=[[1, 0], [1, 1], [1, 2]], and normalization="ctc". The normalization divides each batch item by its unexpanded target length before averaging.

forward(X=None, Y=None, list_N=None, list_M=None)[source]

Compute the CTC loss in the dDTW graph formulation.

Parameters:
  • X (torch.tensor [shape=(B, N, D)]) – Input log-probabilities. The CTC local cost is the negative log-probability of the active target or blank state.

  • Y (torch.tensor [shape=(B, M)]) – Integer target-label indices before blank expansion. Labels should use blank_index only for padding beyond list_M.

  • list_N (torch.tensor [shape=(B)]) – Sequence lengths of X (<= N) of the individual batch elements. If None, defaults to [N, N, …, N]

  • list_M (torch.tensor [shape=(B)]) – Target-label lengths before CTC blank expansion. If None, defaults to [M, M, …, M].

Returns:

Scalar batch-mean CTC loss.

Return type:

torch.Tensor

class ddtw.ddtw_variants.partial_matching(*args: Any, **kwargs: Any)[source]

Bases: dDTW

Initialize partial matching loss function, parameterized within the dDTW framework, see [9, 2].

[2] Johannes Zeitler and Meinard Müller. A Unified Perspective on CTC and Soft-DTW Using Differentiable DTW. IEEE Transactions on Audio, Speech and Language Processing, vol. 34, pages 936-951, 2026.

[9] Pavel A. Pevzner. Computational Molecular Biology: An Algorithmic Approach. MIT Press, 2000.

Parameters:
  • cost_function (str) – Local cost function used when X and Y are supplied. Default: “CTC”.

  • min_function (str) – Minimum function or differentiable approximation thereof. Default: “hardmin”.

  • gamma (float) – Temperature parameter for differentiable minimum functions. Default: 1.0.

  • normalization (str) – Normalization of PM cost. Choose among (“M”, “N”, “NM”, “ctc”, “none”). “N”: divide by N. “M”: divide by M. “NM”: divide by (N*M). “ctc”: divide by (M-1)/2. “none”: no normalization. Default: “none”.

  • backend (str) – Backend to use. Choose among (“auto”, “torch”, “cpu_numba”, “cuda_cpp”). Default: “auto”

  • dtype_float (torch.dtype) – Number format for internal computations. Default: torch.float32

  • cuda_device (str or torch.device, optional) – CUDA device to use, for example "cuda:0". Default: None.

  • store_debug (bool) – Whether to retain intermediate backend matrices for inspection. Default: False.

Notes

This variant fixes step_sizes=[[1, 0], [0, 1], [1, 1]] and global_step_weights=[0.0, 0.0, 1.0]. Horizontal and vertical moves therefore do not accumulate local cost; only diagonal matches do.

forward(X=None, Y=None, C=None, list_N=None, list_M=None)[source]

Compute the partial matching loss.

Parameters:
  • X (torch.tensor [shape=(B, N, D)]) – Input sequence, usually the DNN predictions.

  • Y (torch.tensor [shape=(B, M, D)]) – Input sequence, usually the weak targets.

  • C (torch.tensor [shape=(B, N, M)]) – Pre-computed local cost matrix C. To compare against a score-maximizing partial matching reference, pass C=-S for score matrix S.

  • list_N (torch.tensor [shape=(B)]) – Sequence lengths of X (<= N) of the individual batch elements. If None, defaults to [N, N, …, N]

  • list_M (torch.tensor [shape=(B)]) – Sequence lengths of Y (<= M) of the individual batch elements. If None, defaults to [M, M, …, M]

Returns:

Scalar batch-mean partial matching loss.

Return type:

torch.Tensor