Abstract
Quantum-safe security is rapidly moving from theoretical compliance to real-world deployment. However, migrating to NIST’s FIPS 204 ML-DSA standard introduces an unforeseen engineering hurdle: algorithmic timing variability.
Unlike classical primitives, ML-DSA relies on rejection sampling – a mathematical process that introduces variability, where a single signing operation can potentially take up to 20 times longer than the statistical average. In high-throughput environments, or real-time, safety-critical systems such as automotive control units or aerospace electronics, these rare timing spikes can trigger catastrophic system timeouts and failures.
To prevent system outages and ensure sound architectural decisions, cryptographic evaluators can no longer rely on generic benchmarking metrics like “minimum, average, or maximum” execution times. This white paper exposes the critical pitfalls in ML-DSA evaluation – from ambiguous measurement boundaries, to misleading deterministic assumptions.
We introduce a standardized, dataset-driven benchmarking methodology. Featuring publicly verifiable worst-case execution time results on real-world embedded hardware implementations, this paper provides engineers, system architects, and compliance officers with the framework needed to make reliable migration decisions.

