Installation

Install the package from PyPI:

python -m pip install ddtw

For local development, clone the repository and install it in editable mode:

python -m pip install -e ".[test,benchmark]"

Runtime Requirements

The implemented losses require:

  • Python

  • PyTorch

  • NumPy

  • Numba for the cpu_numba backend

  • A CUDA-capable PyTorch setup and a working compiler toolchain for the cuda_cpp backend

CUDA Backend

The cuda_cpp backend uses PyTorch’s JIT extension loader, which invokes ninja, c++, and nvcc to compile the native extension lazily on first use. [1] For this build to work, the CUDA-related components must be compatible:

  • the NVIDIA driver must support the CUDA runtime used by PyTorch;

  • the installed PyTorch wheel must match the intended CUDA version, visible as torch.version.cuda;

  • the active nvcc must come from a compatible CUDA toolkit and must see the CUDA development headers;

  • the host C++ compiler must be recent enough for PyTorch’s extension build.

nvidia-smi reports the maximum CUDA version supported by the driver. It does not show which CUDA toolkit or nvcc is active inside the Python environment. Check the active setup with:

which nvcc
nvcc --version

python - <<'PY'
import torch
print(torch.__version__)
print(torch.version.cuda)
print(torch.cuda.is_available())
PY

Conda CUDA Environments

The repository contains tested conda environment files for CUDA 11.8, 12.8, and 13.2 in environments/. They install PyTorch with pip and the CUDA toolkit, nvcc, and host compilers with conda. For example:

conda env create -f environments/ddtw_cu128.yml
conda activate ddtw_cu128
bash environments/install_activation_hooks.sh
conda deactivate
conda activate ddtw_cu128

The activation hooks clear inherited CUDA/compiler flags such as NVCC_PREPEND_FLAGS, CFLAGS, and CXXFLAGS, then select the conda compiler wrappers through CC, CXX, and CUDAHOSTCXX. This avoids two common failure modes: duplicate nvcc host-compiler flags and accidental use of an old system c++.

We tested and verified the ddtw-cuda environments for the following architectures:

GPU

CUDA 11.8

CUDA 12.8

CUDA 13.2

RTX 1080 TI

✓

✗

✗

RTX 2080 TI

✗

✓

✓

RTX 4090

✗

✓

✓

RTX A5500

✗

✓

✓

RTX Pro 6000

✗

✓

✓

If PyTorch auto-detects the wrong GPU architecture, set TORCH_CUDA_ARCH_LIST before rebuilding the extension. Typical values are 6.1 for GTX 1080 Ti, 7.5 for RTX 2080, 8.9 for RTX 4090, and 12.0 for RTX PRO 6000 Blackwell.

After changing CUDA, compiler, or architecture settings, remove the cached extension build and run the tests again:

rm -rf ddtw/backend/_cpp_build
python -m pytest test

The repository README.md and environments/README.md contain more concrete setup examples and troubleshooting notes.

Documentation Requirements

The documentation dependencies are listed in docs/requirements.txt:

python -m pip install -r docs/requirements.txt

Build the HTML documentation with:

sphinx-build -M html docs docs/_build

or:

make -C docs html

The Sphinx configuration mocks heavy runtime imports such as torch and numba. This allows API documentation to build on machines that are not configured for GPU training.

References