Transactions on Cryptographic Hardware and Embedded Systems 2026
Improving ML-KEM and ML-DSA on OpenTitan:
Efficient Multiplication Vector Instructions for OTBN
Ruben Niederhagen
Academia Sinica, Taipei, Taiwan; University of Southern Denmark, Odense, Denmark
Hoang Nguyen Hien Pham
Max Planck Institute for Security and Privacy, Bochum, Germany; Université Grenoble Alpes, Saint-Martin-d’Hères, France
Keywords: OpenTitan, PQC, ML-KEM, ML-DSA, ISE, HW/SW co-design
Abstract
This work improves upon the instruction set extension proposed in the paper “Towards ML-KEM and ML-DSA on OpenTitan”, in short OTBNTW, for OpenTitan’s big number coprocessor OTBN. OTBNTW introduces a dedicated vector instruction for prime-field Montgomery multiplication, with a high multi-cycle latency and a relatively low utilization of the underlying integer multiplication unit. The design targets post-quantum cryptographic schemes ML-KEM and ML-DSA, which rely on 12-bit and 23-bit prime field arithmetic, respectively. We improve the efficiency of the Montgomery multiplication by fully exploiting existing integer multiplication resources and move modular multiplication from hardware back to software by providing more powerful and versatile integer-multiplication vector instructions. This enables us not only to reduce the overall computational overhead through lazy reduction in software but also to improve performance in other functions beyond finite-field arithmetic. We provide two variants of our instruction set extension, each offering different trade-offs between resource usage and performance. For ML-KEM and ML-DSA, we achieve a speedup of up to 17% in cycle count, with an ASIC area increase of up to 6% and an FPGA resource usage increase of up to 4% more LUT, 20% more CARRY4, 1% more FF, and the same number of DSP compared to OTBNTW. Overall, we significantly reduce the ASIC time-area product, if the designs are clocked at their individual maximum frequency, and at least match that of OTBNTW, if the designs are clocked at the same frequency.
Publication
IACR Transactions on Cryptographic Hardware and Embedded Systems, Volume 2026, Issue 2
PaperArtifact
Artifact number
tches/2026/a17
Artifact published
September 21, 2026
Badge
✅ IACR CHES Artifacts Functional
License
This work is licensed under the Apache License, Version 2.0.
Note that license information is supplied by the authors and has not been confirmed by the IACR.
BibTeX How to cite
Ruben Niederhagen, Hoang Nguyen Hien Pham. (2026). Improving ML-KEM and ML-DSA on OpenTitan: Efficient Multiplication Vector Instructions for OTBN. IACR Transactions on Cryptographic Hardware and Embedded Systems, 2026(2), 495–519. https://doi.org/10.46586/tches.v2026.i2.495-519. Artifact at https://artifacts.iacr.org/tches/2026/a17.