International Association for Cryptologic Research

International Association
for Cryptologic Research

Transactions on Cryptographic Hardware and Embedded Systems 2026

GPU-Accelerated DPF-Based Private Information Retrieval for Large-Scale Database


README

GPU-Accelerated DPF-Based Private Information Retrieval

Project Overview

This repository contains a CUDA implementation of DPF-based Private Information Retrieval (PIR), together with small drivers for functional checks and benchmarking. The current artifact provides:

Repository Organization

The source tree is organized as follows.

.
├── CMakeLists.txt
├── Dockerfile
├── emp-ot/
├── emp-tool/
├── mpc_cuda/
│   ├── aes_cuda.h
│   ├── aes_prg_device.h
│   ├── fss_cuda_api.cu
│   ├── fss_cuda_kernels.cu
│   ├── fss_cuda_launch.h
│   ├── mpc_core.h
│   └── pir_context.h
├── mpc_keys/
│   ├── aes_prg_host.h
│   ├── fss_keygen.h
│   └── uint128_type.h
├── test/
│   ├── CMakeLists.txt
│   ├── bench_lut_only.cpp
│   ├── bench_pir.cpp
│   ├── pir_test_utils.h
│   ├── test_pir.cpp
│   └── test_pir_utils.cpp
├── run_bench_pir.sh
├── run_bench_lut.sh
└── README.md

The main directories are:

Tested Environment and Dependencies

Tested environment

The current repository has been exercised in the following environment:

Required dependencies

Build-essential packages (typically pre-installed on development machines, but listed for completeness):

Libraries that may need manual installation:

CUDA and GPU requirements:

Bundled library dependencies:

Configure CUDA environment

The CUDA toolkit must be discoverable at configure time. Verify that nvcc is on PATH:

nvcc --version

If nvcc is not found, add the CUDA toolkit to your environment. The typical install location is /usr/local/cuda:

export PATH=/usr/local/cuda/bin:$PATH
export LD_LIBRARY_PATH=/usr/local/cuda/lib64:$LD_LIBRARY_PATH

Add these lines to your ~/.bashrc (or equivalent shell profile) to make the setting persistent.

If you prefer not to modify PATH, you can pass the compiler directly to CMake:

cmake -S . -B build -DCMAKE_CUDA_COMPILER=/usr/local/cuda/bin/nvcc

GPU architecture

By default, the build automatically detects the GPU architecture via nvidia-smi at configure time and sets CMAKE_CUDA_ARCHITECTURES accordingly. You will see a log line like:

-- Auto-detected GPU architecture: sm_90

If nvidia-smi is not available or returns no GPU info, the build falls back to CMake's built-in native detection.

To override the auto-detected value (e.g., when cross-compiling for a different GPU), pass it explicitly:

cmake -S . -B build -DCMAKE_CUDA_ARCHITECTURES=90

Common architecture numbers:

Value GPU generation Examples
80 Ampere A100, A10
86 Ampere RTX 3090, A40
89 Ada Lovelace RTX 4090, L4, L40
90 Hopper H100, H20

Using a mismatched architecture value (e.g., compiling for sm_89 but running on a Hopper GPU) will cause a runtime error: the provided PTX was compiled with an unsupported toolchain.

Build Instructions

Option 1: Build with Docker

Requirements:

Build the Docker image from the repository root:

docker build -t gpu-pir-artifact .

Run the image:

docker run --rm -it --gpus all gpu-pir-artifact

If you want to mount the local checkout:

docker run --rm -it --gpus all -v "$(pwd)":/workspace gpu-pir-artifact

Inside the container, the project is already built under /workspace/build.

Note: When using Docker, the GPU is not visible during docker build, so the nvidia-smi auto-detection will fail. The Dockerfile uses ARG CUDA_ARCH=89 (Ada Lovelace) as a default. Override it to match your target GPU:

docker build --build-arg CUDA_ARCH=90 -t gpu-pir-artifact .

Common values: 80 (A100), 86 (RTX 3090, A40), 89 (RTX 4090, L4, L40), 90 (H100, H20).

Option 2: Manual build

  1. Install required packages (skip any that are already present on your system):

    sudo apt update
    sudo apt install -y build-essential cmake git pkg-config libssl-dev libeigen3-dev libgmp-dev libmpfr-dev
    
  2. Ensure the CUDA toolkit is on PATH (see Configure CUDA environment).

  3. Build and install the bundled EMP dependencies:

    cmake -S emp-tool -B emp-tool/build
    cmake --build emp-tool/build -j"$(nproc)"
    sudo cmake --install emp-tool/build
    
    cmake -S emp-ot -B emp-ot/build
    cmake --build emp-ot/build -j"$(nproc)"
    sudo cmake --install emp-ot/build
    

    If the EMP packages are not discovered automatically, export:

    export CMAKE_PREFIX_PATH="/usr/local/lib/cmake/emp-tool:/usr/local/lib/cmake/emp-ot:${CMAKE_PREFIX_PATH}"
    
  4. Build this repository:

    cmake -S . -B build
    cmake --build build -j"$(nproc)"
    

    If nvcc is not on PATH, use:

    cmake -S . -B build -DCMAKE_CUDA_COMPILER=/usr/local/cuda/bin/nvcc
    

    If you need to target a specific GPU architecture:

    cmake -S . -B build -DCMAKE_CUDA_ARCHITECTURES=90
    

How to Run the Artifact

The build produces four binaries under build/test (test_pir, test_pir_utils, bench_pir, and bench_lut_only). The three drivers below cover functional checks, the PIR functional path, and the all-in-one benchmark; the LUT-only benchmark (bench_lut_only) is described together with the reproduction scripts in Reproducing paper results.

1. Utility regression checks

./build/test/test_pir_utils

This binary checks helper logic and some source-level repository invariants. It does not require a visible CUDA device.

2. Functional PIR driver

./build/test/test_pir [n] [batch_size]

Examples:

./build/test/test_pir
./build/test/test_pir 24 512
./build/test/test_pir 20 128

Default values:

3. Benchmark driver

./build/test/bench_pir [n] [batch_size]

Examples:

./build/test/bench_pir
./build/test/bench_pir 22 512
./build/test/bench_pir 20 128

Default values:

Reproducing paper results

Two helper scripts in the repository root drive the benchmarks for reproducing and verifying the paper's experimental results. Each writes structured data plus figures into its own output directory. Both require a completed build (see Build Instructions) and a visible CUDA device, and they need Python 3 with matplotlib, pandas, and numpy for the figure-generation step.

Script Reproduces Configuration Output directory
run_bench_pir.sh Table 2 — throughput and CUDA memory of the pipeline vs. non-pipeline PIR implementations n = 19..24, batch_size = 512 bench_results_pir/
run_bench_lut.sh Figure 11 — DPF-PIR LUT throughput as a function of batch_size n ∈ {16, 18, 20, 24}, batch_size ∈ {1, 10, 100, 1000, 10000} bench_results_lut/

Usage (defaults reproduce the paper directly):

# Table 2 — pipeline vs non-pipeline throughput & memory (n=19..24, batch=512)
./run_bench_pir.sh

# Figure 11 — LUT throughput vs batch_size (n=16,18,20,24)
./run_bench_lut.sh

Both scripts also accept optional overrides:

./run_bench_pir.sh "19 20 21" 256     # n list + batch_size
./run_bench_lut.sh "16 18" "10 100"   # n list + batch list

run_bench_pir.sh (Table 2)

Runs bench_pir for n = 19, 20, ..., 24 at batch_size = 512 and parses the DPF-PIR (non-pipeline) and DPF-PIR pipeline segments. The LUT segment emitted by bench_pir is ignored here. For each (n, mode) point it records throughput (PIRs/s) and allocated CUDA memory (MB), then plots the two modes against n. Produces:

run_bench_lut.sh (Figure 11)

Runs the LUT-only benchmark bench_lut_only for n ∈ {16, 18, 20, 24} across batch_size ∈ {1, 10, 100, 1000, 10000} and plots LUT throughput (PIRs/s, log-x axis) versus batch_size. LUT and regular PIR are evaluated separately, so these figures contain only the LUT series. Produces:

The bench_lut_only binary can also be run directly:

./build/test/bench_lut_only [n] [batch_size]   # default n=20, batch_size=512

Configuration Options

The artifact can be configured at two levels.

Runtime configuration

The functional and benchmark drivers accept:

Current runtime constraints enforced by the drivers:

Build-time configuration

You can change these values and rebuild if you want to evaluate other GPU targets or compile-time constants.

How to Interpret the Output

test_pir_utils

Expected successful behavior:

If a check fails, the binary prints a short diagnostic message and exits with a non-zero status.

test_pir

Expected output on success with a visible CUDA device:

If a correctness issue is detected, the program prints a line starting with [FAIL] and includes a mismatch description.

If no CUDA device is available, the program prints:

and exits cleanly.

bench_pir

Expected output on success with a visible CUDA device:

Typical labels are:

If the initial smoke check fails, the program prints a line starting with [FAIL].

If no CUDA device is available, the program prints:

and exits cleanly.