International Association for Cryptologic Research

International Association
for Cryptologic Research

Transactions on Cryptographic Hardware and Embedded Systems 2026

Improving ML-KEM and ML-DSA on OpenTitan:

Efficient Multiplication Vector Instructions for OTBN


Ruben Niederhagen
Academia Sinica, Taipei, Taiwan; University of Southern Denmark, Odense, Denmark

Hoang Nguyen Hien Pham
Max Planck Institute for Security and Privacy, Bochum, Germany; Université Grenoble Alpes, Saint-Martin-d’Hères, France


Keywords: OpenTitan, PQC, ML-KEM, ML-DSA, ISE, HW/SW co-design


Abstract

This work improves upon the instruction set extension proposed in the paper “Towards ML-KEM and ML-DSA on OpenTitan”, in short OTBNTW, for OpenTitan’s big number coprocessor OTBN. OTBNTW introduces a dedicated vector instruction for prime-field Montgomery multiplication, with a high multi-cycle latency and a relatively low utilization of the underlying integer multiplication unit. The design targets post-quantum cryptographic schemes ML-KEM and ML-DSA, which rely on 12-bit and 23-bit prime field arithmetic, respectively. We improve the efficiency of the Montgomery multiplication by fully exploiting existing integer multiplication resources and move modular multiplication from hardware back to software by providing more powerful and versatile integer-multiplication vector instructions. This enables us not only to reduce the overall computational overhead through lazy reduction in software but also to improve performance in other functions beyond finite-field arithmetic. We provide two variants of our instruction set extension, each offering different trade-offs between resource usage and performance. For ML-KEM and ML-DSA, we achieve a speedup of up to 17% in cycle count, with an ASIC area increase of up to 6% and an FPGA resource usage increase of up to 4% more LUT, 20% more CARRY4, 1% more FF, and the same number of DSP compared to OTBNTW. Overall, we significantly reduce the ASIC time-area product, if the designs are clocked at their individual maximum frequency, and at least match that of OTBNTW, if the designs are clocked at the same frequency.

Publication

IACR Transactions on Cryptographic Hardware and Embedded Systems, Volume 2026, Issue 2

Paper

Artifact

Artifact number
tches/2026/a17

Artifact published
September 21, 2026

Badge
✅ IACR CHES Artifacts Functional

README

ZIP (146758995 Bytes)  

View on Github

License
This work is licensed under the Apache License, Version 2.0.

Note that license information is supplied by the authors and has not been confirmed by the IACR.


BibTeX How to cite

Ruben Niederhagen, Hoang Nguyen Hien Pham. (2026). Improving ML-KEM and ML-DSA on OpenTitan: Efficient Multiplication Vector Instructions for OTBN. IACR Transactions on Cryptographic Hardware and Embedded Systems, 2026(2), 495–519. https://doi.org/10.46586/tches.v2026.i2.495-519. Artifact at https://artifacts.iacr.org/tches/2026/a17.