API Reference
Package Root
The public loss classes are available from the package root:
from ddtw import dDTW, SDTW, CTC, DTW
Core Loss
- class ddtw.ddtw.dDTW(*args: Any, **kwargs: Any)[source]
Bases:
ModuleInitialize the general dDTW loss function.
See [1] for the graph formulation.
[1] Johannes Zeitler and Meinard Müller. A Unified Perspective on CTC and Soft-DTW Using Differentiable DTW. IEEE Transactions on Audio, Speech and Language Processing, vol. 34, pages 936-951, 2026.
- Parameters:
cost_function (str or callable, optional) – Local cost function used when
XandYare passed toforward(). Built-in strings are"MSE","BCE", and"CTC". A callable must return a cost tensor with shape(B, N, M). Default:"MSE".min_function (str, optional) – Recursive minimum or differentiable approximation. Choose among
"softmin","sparsemin","smoothmin", and"hardmin". Default:"softmin".gamma (float, optional) – Temperature parameter used by differentiable minimum functions. Default:
1.0.step_sizes (list of list of int, optional) – Alignment step sizes
[dn, dm]. Each step points from the current cell(n, m)to predecessor(n-dn, m-dm). Default:[[1, 0], [0, 1], [1, 1]].global_step_weights (list of float, optional) – Local cost weights associated with
step_sizes. Must contain one scalar per step. Default:[1.0, 1.0, 1.0].normalization (str, optional) – Normalization applied to each batch loss before averaging. Choose among
"N","M","NM","ctc", and"none". Default:"N".backend (str, optional) – Backend to use. Choose among
"auto","torch","cpu_numba", and"cuda_cpp"."auto"tries CUDA first, then Numba CPU, then pure PyTorch. Default:"auto".dtype_float (torch.dtype, optional) – Floating-point dtype for internal tensors. The CUDA C++ backend currently requires
torch.float32. Default:torch.float32.cuda_device (str or torch.device, optional) – CUDA device to use, for example
"cuda:0". Default:None.store_debug (bool, optional) – If
True, retain intermediate matrices on the backend class for inspection after forward/backward. Default:False.
- forward(X=None, Y=None, C=None, B_start=None, B_end=None, list_N=None, list_M=None, local_step_weights=None, start_penalty=None, end_penalty=None, num_start_conditions=None, num_end_conditions=None)[source]
Compute the dDTW loss.
Pass either
XandYor a precomputed cost matrixC. IfXandYare provided,self.cost_functioncomputesC.- Parameters:
X (torch.Tensor, optional) – First input sequence with shape
(B, N, D). Usually the model predictions.Y (torch.Tensor, optional) – Second input sequence with shape
(B, M, D). Usually the target or reference sequence.C (torch.Tensor, optional) – Precomputed local cost matrix with shape
(B, N, M).B_start (list or torch.Tensor, optional) – Start boundary conditions. For each batch item, stores one or more zero-based
[n, m]cells. Tensor form must have shape(B, max_start_conditions, 2). IfNone, defaults to[[0, 0]]for each batch item.B_end (list or torch.Tensor, optional) – End boundary conditions. For each batch item, stores one or more zero-based
[n, m]cells. Tensor form must have shape(B, max_end_conditions, 2). IfNone, defaults to[[list_N[b] - 1, list_M[b] - 1]].list_N (list or torch.Tensor, optional) – Active lengths along the
X/row axis with shape(B,). IfNone, all batch items use the full padded lengthN.list_M (list or torch.Tensor, optional) – Active lengths along the
Y/column axis with shape(B,). IfNone, all batch items use the full padded lengthM.local_step_weights (torch.Tensor, optional) – Cell-wise step weights with shape
(B, N, M, S), whereSis the number of configured steps. IfNone,global_step_weightsare broadcast to all cells.start_penalty (list or torch.Tensor, optional) – Multiplicative local-cost weights for start boundary conditions. Tensor form must have shape
(B, max_start_conditions). IfNone, defaults to1for every start condition.end_penalty (list or torch.Tensor, optional) – Multiplicative local-cost weights for end boundary conditions. Tensor form must have shape
(B, max_end_conditions). IfNone, defaults to0for every end condition.num_start_conditions (list or torch.Tensor, optional) – Number of valid start conditions for each batch item when
B_startis padded. Shape(B,).num_end_conditions (list or torch.Tensor, optional) – Number of valid end conditions for each batch item when
B_endis padded. Shape(B,).
- Returns:
Scalar batch-mean dDTW loss after the configured normalization.
- Return type:
torch.Tensor
Loss Variants
- class ddtw.ddtw_variants.SDTW(*args: Any, **kwargs: Any)[source]
Bases:
dDTWInitialize SDTW loss function, see [1, 2].
[1] Marco Cuturi and Mathieu Blondel. Soft-DTW: A Differentiable Loss Function for Time-Series. In Proceedings of the International Conference on Neural Information Processing Systems (NIPS), vol. 2, pages 2292-2300, 2013.
[2] Johannes Zeitler and Meinard Müller. A Unified Perspective on CTC and Soft-DTW Using Differentiable DTW. IEEE Transactions on Audio, Speech and Language Processing, vol. 34, pages 936-951, 2026.
- Parameters:
cost_function (str) – Local cost function for pair-wise comparison of sequence elements. Choose among (“MSE”, “BCE”, “CTC”). Default: “MSE”.
gamma (float) – Softmin temperature hyperparameter. Default: 1.0
step_sizes (list) – Alignment step sizes in [n,m] direction, given as list of tuples. Default: [[1,0], [0,1], [1,1]]
global_step_weights (list) – Step weights associated to the step sizes. Default: [1.0, 1.0, 1.0]
normalization (str) – Normalization of SDTW cost. Choose among (“M”, “N”, “NM”, “ctc”, “none”). “N”: divide by N. “M”: divide by M. “NM”: divide by (N*M). “ctc”: divide by (M-1)/2. “none”: no normalization. Default: “N”
backend (str) – Backend to use. Choose among (“auto”, “torch”, “cpu_numba”, “cuda_cpp”). Default: “auto”
dtype_float (torch.dtype) – Number format for internal computations. Default: torch.float32
cuda_device (str or torch.device, optional) – CUDA device to use, for example
"cuda:0". Default:None.store_debug (bool) – Whether to retain intermediate backend matrices for inspection. Default: False.
- forward(X=None, Y=None, C=None, list_N=None, list_M=None)[source]
Compute the SDTW loss.
- Parameters:
X (torch.tensor [shape=(B, N, D)]) – Input sequence, usually the DNN predictions.
Y (torch.tensor [shape=(B, M, D)]) – Input sequence, usually the weak targets.
C (torch.tensor [shape=(B, N, M)]) – Pre-computed local cost matrix C.
list_N (torch.tensor [shape=(B)]) – Sequence lengths of X (<= N) of the individual batch elements. If None, defaults to [N, N, …, N]
list_M (torch.tensor [shape=(B)]) – Sequence lengths of Y (<= M) of the individual batch elements. If None, defaults to [M, M, …, M]
- Returns:
Scalar batch-mean SDTW loss.
- Return type:
torch.Tensor
- class ddtw.ddtw_variants.DTW(*args: Any, **kwargs: Any)[source]
Bases:
dDTWInitialize DTW loss function, see [3].
[3]. Meinard Müller. Fundamentals of Music Processing - Using Python and Jupyter Notebooks. Springer Verlag, 2nd edition, 2021.
- Parameters:
cost_function (str) – Local cost function for pair-wise comparison of sequence elements. Choose among (“MSE”, “BCE”, “CTC”). Default: “MSE”.
step_sizes (list) – Alignment step sizes in [n,m] direction, given as list of tuples. Default: [[1,0], [0,1], [1,1]]
global_step_weights (list) – Step weights associated to the step sizes. Default: [1.0, 1.0, 1.0]
normalization (str) – Normalization of SDTW cost. Choose among (“M”, “N”, “NM”, “ctc”, “none”). “N”: divide by N. “M”: divide by M. “NM”: divide by (N*M). “ctc”: divide by (M-1)/2. “none”: no normalization. Default: “N”
backend (str) – Backend to use. Choose among (“auto”, “torch”, “cpu_numba”, “cuda_cpp”). Default: “auto”
dtype_float (torch.dtype) – Number format for internal computations. Default: torch.float32
cuda_device (str or torch.device, optional) – CUDA device to use, for example
"cuda:0". Default:None.store_debug (bool) – Whether to retain intermediate backend matrices for inspection. Default: False.
- forward(X=None, Y=None, C=None, list_N=None, list_M=None)[source]
Compute the DTW loss.
- Parameters:
X (torch.tensor [shape=(B, N, D)]) – Input sequence, usually the DNN predictions.
Y (torch.tensor [shape=(B, M, D)]) – Input sequence, usually the weak targets.
C (torch.tensor [shape=(B, N, M)]) – Pre-computed local cost matrix C.
list_N (torch.tensor [shape=(B)]) – Sequence lengths of X (<= N) of the individual batch elements. If None, defaults to [N, N, …, N]
list_M (torch.tensor [shape=(B)]) – Sequence lengths of Y (<= M) of the individual batch elements. If None, defaults to [M, M, …, M]
- Returns:
Scalar batch-mean DTW loss. With a precomputed
C, gradients with respect toCmark the selected hard warping path.- Return type:
torch.Tensor
- class ddtw.ddtw_variants.smoothDTW(*args: Any, **kwargs: Any)[source]
Bases:
dDTWInitialize smoothDTW loss function, see [4].
[4]. Isma Hadji, K. Derpanis, and A. Jepson. Representation learning via global temporal alignment and cycle-consistency. In IEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 11068-11077, 2021.
- Parameters:
cost_function (str) – Local cost function for pair-wise comparison of sequence elements. Choose among (“MSE”, “BCE”, “CTC”). Default: “MSE”.
gamma (float) – Softmin temperature hyperparameter. Default: 1.0
step_sizes (list) – Alignment step sizes in [n,m] direction, given as list of tuples. Default: [[1,0], [0,1], [1,1]]
global_step_weights (list) – Step weights associated to the step sizes. Default: [1.0, 1.0, 1.0]
normalization (str) – Normalization of SDTW cost. Choose among (“M”, “N”, “NM”, “ctc”, “none”). “N”: divide by N. “M”: divide by M. “NM”: divide by (N*M). “ctc”: divide by (M-1)/2. “none”: no normalization. Default: “N”
backend (str) – Backend to use. Choose among (“auto”, “torch”, “cpu_numba”, “cuda_cpp”). Default: “auto”
dtype_float (torch.dtype) – Number format for internal computations. Default: torch.float32
cuda_device (str or torch.device, optional) – CUDA device to use, for example
"cuda:0". Default:None.store_debug (bool) – Whether to retain intermediate backend matrices for inspection. Default: False.
- forward(X=None, Y=None, C=None, list_N=None, list_M=None)[source]
Compute the smoothDTW loss.
- Parameters:
X (torch.tensor [shape=(B, N, D)]) – Input sequence, usually the DNN predictions.
Y (torch.tensor [shape=(B, M, D)]) – Input sequence, usually the weak targets.
C (torch.tensor [shape=(B, N, M)]) – Pre-computed local cost matrix C.
list_N (torch.tensor [shape=(B)]) – Sequence lengths of X (<= N) of the individual batch elements. If None, defaults to [N, N, …, N]
list_M (torch.tensor [shape=(B)]) – Sequence lengths of Y (<= M) of the individual batch elements. If None, defaults to [M, M, …, M]
- Returns:
Scalar batch-mean smoothDTW loss.
- Return type:
torch.Tensor
- class ddtw.ddtw_variants.sparseDTW(*args: Any, **kwargs: Any)[source]
Bases:
dDTWInitialize sparseDTW loss function, see [5].
[5] Arthur Mensch and Mathieu Blondel. Differentiable Dynamic Programming for Structured Prediction and Attention. In Proceedings of the International Converence on Machine Learning (ICML), pages 3459-3468, Stockholm, Sweden, 2018.
- Parameters:
cost_function (str) – Local cost function for pair-wise comparison of sequence elements. Choose among (“MSE”, “BCE”, “CTC”). Default: “MSE”.
gamma (float) – Sparsemin temperature hyperparameter. Default: 1.0
step_sizes (list) – Alignment step sizes in [n,m] direction, given as list of tuples. Default: [[1,0], [0,1], [1,1]]
global_step_weights (list) – Step weights associated to the step sizes. Default: [1.0, 1.0, 1.0]
normalization (str) – Normalization of SDTW cost. Choose among (“M”, “N”, “NM”, “ctc”, “none”). “N”: divide by N. “M”: divide by M. “NM”: divide by (N*M). “ctc”: divide by (M-1)/2. “none”: no normalization. Default: “N”
backend (str) – Backend to use. Choose among (“auto”, “torch”, “cpu_numba”, “cuda_cpp”). Default: “auto”
dtype_float (torch.dtype) – Number format for internal computations. Default: torch.float32
cuda_device (str or torch.device, optional) – CUDA device to use, for example
"cuda:0". Default:None.store_debug (bool) – Whether to retain intermediate backend matrices for inspection. Default: False.
- forward(X=None, Y=None, C=None, list_N=None, list_M=None)[source]
Compute the sparseDTW loss.
- Parameters:
X (torch.tensor [shape=(B, N, D)]) – Input sequence, usually the DNN predictions.
Y (torch.tensor [shape=(B, M, D)]) – Input sequence, usually the weak targets.
C (torch.tensor [shape=(B, N, M)]) – Pre-computed local cost matrix C.
list_N (torch.tensor [shape=(B)]) – Sequence lengths of X (<= N) of the individual batch elements. If None, defaults to [N, N, …, N]
list_M (torch.tensor [shape=(B)]) – Sequence lengths of Y (<= M) of the individual batch elements. If None, defaults to [M, M, …, M]
- Returns:
Scalar batch-mean sparseDTW loss.
- Return type:
torch.Tensor
- class ddtw.ddtw_variants.subSDTW(*args: Any, **kwargs: Any)[source]
Bases:
dDTWInitialize subsequence SDTW loss function, see [7].
[7] Johannes Zeitler and Meinard Müller. Subsequence Soft Dynamic Time Warping. In Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing (CASSP), Barcelona, Spain, 2026.
- Parameters:
cost_function (str) – Local cost function for pair-wise comparison of sequence elements. Choose among (“MSE”, “BCE”, “CTC”). Default: “MSE”.
min_function (str) – Minimum function or approximation thereof. Choose among (“softmin”, “sparsemin”, “smoothmin”, “hardmin”). Default: “softmin”.
gamma (float) – Min. function temperature hyperparameter. Default: 1.0
step_sizes (list) – Alignment step sizes in [n,m] direction, given as list of tuples. Default: [[1,0], [0,1], [1,1]]
global_step_weights (list) – Step weights associated to the step sizes. Default: [1.0, 1.0, 1.0]
normalization (str) – Normalization of SDTW cost. Choose among (“M”, “N”, “NM”, “ctc”, “none”). “N”: divide by N. “M”: divide by M. “NM”: divide by (N*M). “ctc”: divide by (M-1)/2. “none”: no normalization. Default: “N”
backend (str) – Backend to use. Choose among (“auto”, “torch”, “cpu_numba”, “cuda_cpp”). Default: “auto”
dtype_float (torch.dtype) – Number format for internal computations. Default: torch.float32
cuda_device (str or torch.device, optional) – CUDA device to use, for example
"cuda:0". Default:None.store_debug (bool) – Whether to retain intermediate backend matrices for inspection. Default: False.
sub_X (bool) – Whether to allow subsequence starts and ends along the X-direction (row axis). Default: True.
sub_Y (bool) – Whether to allow subsequence starts and ends along the Y-direction (column axis). Default: True.
compensate_subseq (bool) – Whether to add start/end penalties for skipped prefixes or suffixes. Default: True.
- forward(X=None, Y=None, C=None, list_N=None, list_M=None)[source]
Compute the subsequence SDTW loss.
- Parameters:
X (torch.tensor [shape=(B, N, D)]) – Input sequence, usually the DNN predictions.
Y (torch.tensor [shape=(B, M, D)]) – Input sequence, usually the weak targets.
C (torch.tensor [shape=(B, N, M)]) – Pre-computed local cost matrix C.
list_N (torch.tensor [shape=(B)]) – Sequence lengths of X (<= N) of the individual batch elements. If None, defaults to [N, N, …, N]
list_M (torch.tensor [shape=(B)]) – Sequence lengths of Y (<= M) of the individual batch elements. If None, defaults to [M, M, …, M]
- Returns:
Scalar batch-mean subsequence SDTW loss.
- Return type:
torch.Tensor
- class ddtw.ddtw_variants.CTC(*args: Any, **kwargs: Any)[source]
Bases:
dDTWInitialize CTC loss function, parameterized within the dDTW framework, see [8, 2].
[2] Johannes Zeitler and Meinard Müller. A Unified Perspective on CTC and Soft-DTW Using Differentiable DTW. IEEE Transactions on Audio, Speech and Language Processing, vol. 34, pages 936-951, 2026.
[8] Alex Graves et al. Connectionist Temporal Classification: Labelling Unsegmented Sequence Data with Recurrent Neural Networks. In Proceedings of the International Conference on Machine Learning (ICML), pages 369-376, Pittsburgh, Pennsylvania, USA, 2006.
- Parameters:
blank_penalty_weight (float) – Penalty for alignment of the blank symbol. Default: 1.0 (no penalty).
blank_index (int) – Class index of the blank symbol. Default: 0.
gamma (float) – Softmin temperature hyperparameter. Default: 1.0
global_step_weights (list) – Step weights associated to the CTC step sizes
[[1, 0], [1, 1], [1, 2]]. Default: [1.0, 1.0, 1.0].backend (str) – Backend to use. Choose among (“auto”, “torch”, “cpu_numba”, “cuda_cpp”). Default: “auto”
dtype_float (torch.dtype) – Number format for internal computations. Default: torch.float32
cuda_device (str or torch.device, optional) – CUDA device to use, for example
"cuda:0". Default:None.store_debug (bool) – Whether to retain intermediate backend matrices for inspection. Default: False.
Notes
The CTC variant fixes
cost_function="CTC",min_function="softmin",step_sizes=[[1, 0], [1, 1], [1, 2]], andnormalization="ctc". The normalization divides each batch item by its unexpanded target length before averaging.- forward(X=None, Y=None, list_N=None, list_M=None)[source]
Compute the CTC loss in the dDTW graph formulation.
- Parameters:
X (torch.tensor [shape=(B, N, D)]) – Input log-probabilities. The CTC local cost is the negative log-probability of the active target or blank state.
Y (torch.tensor [shape=(B, M)]) – Integer target-label indices before blank expansion. Labels should use
blank_indexonly for padding beyondlist_M.list_N (torch.tensor [shape=(B)]) – Sequence lengths of X (<= N) of the individual batch elements. If None, defaults to [N, N, …, N]
list_M (torch.tensor [shape=(B)]) – Target-label lengths before CTC blank expansion. If None, defaults to [M, M, …, M].
- Returns:
Scalar batch-mean CTC loss.
- Return type:
torch.Tensor
- class ddtw.ddtw_variants.partial_matching(*args: Any, **kwargs: Any)[source]
Bases:
dDTWInitialize partial matching loss function, parameterized within the dDTW framework, see [9, 2].
[2] Johannes Zeitler and Meinard Müller. A Unified Perspective on CTC and Soft-DTW Using Differentiable DTW. IEEE Transactions on Audio, Speech and Language Processing, vol. 34, pages 936-951, 2026.
[9] Pavel A. Pevzner. Computational Molecular Biology: An Algorithmic Approach. MIT Press, 2000.
- Parameters:
cost_function (str) – Local cost function used when
XandYare supplied. Default: “CTC”.min_function (str) – Minimum function or differentiable approximation thereof. Default: “hardmin”.
gamma (float) – Temperature parameter for differentiable minimum functions. Default: 1.0.
normalization (str) – Normalization of PM cost. Choose among (“M”, “N”, “NM”, “ctc”, “none”). “N”: divide by N. “M”: divide by M. “NM”: divide by (N*M). “ctc”: divide by (M-1)/2. “none”: no normalization. Default: “none”.
backend (str) – Backend to use. Choose among (“auto”, “torch”, “cpu_numba”, “cuda_cpp”). Default: “auto”
dtype_float (torch.dtype) – Number format for internal computations. Default: torch.float32
cuda_device (str or torch.device, optional) – CUDA device to use, for example
"cuda:0". Default:None.store_debug (bool) – Whether to retain intermediate backend matrices for inspection. Default: False.
Notes
This variant fixes
step_sizes=[[1, 0], [0, 1], [1, 1]]andglobal_step_weights=[0.0, 0.0, 1.0]. Horizontal and vertical moves therefore do not accumulate local cost; only diagonal matches do.- forward(X=None, Y=None, C=None, list_N=None, list_M=None)[source]
Compute the partial matching loss.
- Parameters:
X (torch.tensor [shape=(B, N, D)]) – Input sequence, usually the DNN predictions.
Y (torch.tensor [shape=(B, M, D)]) – Input sequence, usually the weak targets.
C (torch.tensor [shape=(B, N, M)]) – Pre-computed local cost matrix C. To compare against a score-maximizing partial matching reference, pass
C=-Sfor score matrixS.list_N (torch.tensor [shape=(B)]) – Sequence lengths of X (<= N) of the individual batch elements. If None, defaults to [N, N, …, N]
list_M (torch.tensor [shape=(B)]) – Sequence lengths of Y (<= M) of the individual batch elements. If None, defaults to [M, M, …, M]
- Returns:
Scalar batch-mean partial matching loss.
- Return type:
torch.Tensor