International Association for Cryptologic Research

International Association
for Cryptologic Research

Transactions on Cryptographic Hardware and Embedded Systems 2026

SAMSEM – A Generic and Scalable Approach for IC Metal Line Segmentation


README

SAMSEM - A Generic and Scalable Approach for IC Metal Line Segmentation

SAMSEM fine-tunes SAM2 (Segment Anything 2) for binary segmentation of metal interconnect lines in SEM images of integrated circuits.

The SAMSEM segmentation pipeline comprises two different fine-tuned SAM2 models: a Model 1 operating on original-size images and Model 2 operating on 512x512-pixel patches of these images. The segmentation outputs of both models are then merged to form the final segmentation, see our publication for details. An optional Betti Matching topology loss encourages topologically correct connectivity of thin structures. Hyperparameters are tuned via Optuna across two separate search phases (pixel-loss and topology-loss).

In this README file, we list the software and hardware requirements to interact with SAMSEM, give an overview of the file structure of this artifact, provide a quickstart guide based on Jupyter notebooks to start interacting with SAMSEM, and give instructions on how to use your own data with SAMSEM.

We recommend making sure that your hardware adheres to our requirements and then starting with the quickstart guide to get a feeling of the individual steps involved with fine-tuning and evaluating SAMSEM. Those willing to get more involved can then proceed to apply SAMSEM to other datasets or execute actual, real-world runs of the hyperparameter search, fine-tuning, and the segmentation pipeline.

⚠ Disclaimer

For legal reasons, we cannot publish the full image dataset that our SAMSEM model was fine-tuned and evaluated on. In this artifact, we demonstrate the fine-tuning and evaluation process of SAMSEM using only a single dataset that was published alongside related work. Hence, you will not be able to reproduce the results reported in our paper, not even with the fine-tuned models that we provide. Consequently, this artifact cannot be used to evaluate SAMSEM's generalization capabilities without additional datasets. When fine-tuned only on the one available dataset, other existing methods will actually perform better than SAMSEM. This is for two reasons:

  • The available dataset used as an example in this artifact is of high quality, has very few imaging errors, and comes with reliable ground truth.
  • Existing, less complex approaches are better at adapting to the images of a single layer of a single IC, while SAMSEM's power comes from its generalization across different layers and ICs.

Requirements

To successfully execute all scripts in this artifact, your system needs to fulfill the following requirements:

The SAMSEM pipeline was developed and tested on a server with:

Repository Structure

The structure of the artifact repository is shown below.

SAMSEM/
├── CODE/                                    # Contains all Python code
│   ├── helpers/                             # Helper functions
│   │   ├── __init__.py
│   │   ├── combine_patches.py               # Reconstruct original-size masks from patches
│   │   ├── convert_masks.py                 # Converts masks into binary PNG images
│   │   ├── cut_patches.py                   # Extract overlapping 512 × 512 patches from original-size images
│   │   ├── ESD_error.py                     # Implementation of ESD metric
│   │   ├── evaluate_segmentation.py         # Evaluate using pixel-based and ESD metrics
│   │   ├── load_SEM_dataset.py              # Download the example dataset
│   │   ├── merge_subtiles.py                # Decide between patches from Model 1 and Model 2
│   │   ├── parse_eval_to_excel.py           # Combines multiple evaluation results into one table
│   │   ├── read_optuna_results.py           # Read parameters from an Optuna database
│   │   └── split_dataset.py                 # Deterministic dataset split into train, val, and test
│   │
│   ├── modules/                             # Shared building blocks used by train/ and test/
│   │   ├── __init__.py
│   │   ├── checkpoint.py                    # Save / resume full training-state checkpoints
│   │   ├── datasets.py                      # Loads image/mask pairs into RAM
│   │   ├── logging_utils.py                 # TensorBoard helpers and logging setup
│   │   ├── losses.py                        # PixelComboLoss (Dice + BCE)
│   │   ├── sam2_wrapper.py                  # HuggingFace-compatible SAM2 wrapper (gradient-enabled encoder)
│   │   ├── supervised_unet.py               # U-Net baseline architecture
│   │   └── transforms.py                    # Train / val augmentation pipelines
│   │
│   ├── test/                                # Testing and evaluation scripts
│   │   ├── 05_comparison_table.py           # Generate table like notebook 5
│   │   ├── inference_evaluation_pipeline.py # Full segmentation & evaluation pipeline
│   │   ├── inference_supervised_CNNs.py     # CNN baseline segmentation (DeepLabV3, FCN, U-Net)
│   │   └── inference.py                     # SAM2 batch segmentation with configurable point prompts
│   │
│   └── train/                               # Scripts for fine-tuning and hyperparameter optimization
│       ├── finetuning.py                    # Single finetuning run
│       └── hyper_parameter_search.py        # Optuna HPO — pixel-loss or topology-loss mode
│
├── CONFIG/                                  # YAML configuration files
│   ├── finetune_Model1.yaml                 # Configuration used for all Model 1 fine-tuning runs
│   ├── finetune_Model2.yaml                 # Configuration used for all Model 2 fine-tuning runs
│   ├── hparam_search_pixel.yaml             # Configuration used for the pixel-based hyperparameter optimization
│   └── hparam_search_topology.yaml          # Configuration used for the topology-based hyperparameter optimization
│
├── MODELS/                                  # Contains provided models
│   └── *.torch / *.pth                      # 9 pre-trained models (see 'Supplied Models' below for details)
│
├── 01_prepare_data.ipynb                    # Notebook 1: download & prepare dataset and SAM2 checkpoint
├── 02_hparam_optimization.ipynb             # Notebook 2: Run hyperparameter optimization
├── 03_finetuning.ipynb                      # Notebook 3: fine-tune Model 1 & Model 2
├── 04_segmentation.ipynb                    # Notebook 4: full inference & evaluation pipeline
├── 05_gen_table.ipynb                       # Notebook 5: Reproduce the paper comparison table
├── LICENSE
├── README.md
├── REPLICATION.md                           # Scripts that can be used to replicate our results on your data. 
└── requirements.txt

Supplied Models

Nine pre-trained models are supplied in MODELS/, six of which are based on SAM2 and three more on standard convolutional neural networks for comparison.

SAM2-based Models

File Pipeline role Training data Loss
SAMSEM_Model_1.torch Stage 1 — original-size segmentation ICs 1–7 Pixel (DiceBCE)
SAMSEM_Model_2.torch Stage 2 — 512×512-pixel patch segmentation ICs 1–7 Pixel + topology (Betti Matching)
SAMSEM_ICs1to14_Model_1.torch Stage 1 — original-size segmentation All 14 ICs Pixel (DiceBCE)
SAMSEM_ICs1to14_Model_2.torch Stage 2 — 512×512-pixel patch segmentation All 14 ICs Pixel + topology (Betti Matching)
OnlyPixelLoss_Model_2.torch Ablation — Stage 2 without topology loss ICs 1–7 Pixel only
SingleModel.torch Ablation — one model combining Stage 1 and Stage 2 ICs 1–7 Pixel + topology (Betti Matching)

Please note: SAMSEM_Model_1 and SAMSEM_Model_2 (trained on ICs 1–7) are the models discussed in Sections 3 and 4 of our the paper. The two ICs1to14 models are introduced in Section 4.5 of the paper and were trained on the full dataset of 14 integrated circuits for even better generalization across unseen IC designs.

CNN baselines

For those machine-learning approaches that made their training / fine-tuning code publicly accessible and that can be trained on our whole dataset at once, we also publish the models we used to generate our comparison results in Table 3 and 4 of our paper.

File Architecture Training data
CNN_deeplabv3.pth DeepLabV3 512 × 512 patches, ICs 1–7
CNN_fcn.pth FCN 512 × 512 patches, ICs 1–7
CNN_unet.pth U-Net 512 × 512 patches, ICs 1–7

Setup

Please follow these setup steps for both the quickstart guide and the real-world application below.

First, make sure that your system runs Python 3.10 or newer and has Git installed. There are also a couple dependencies that will not automatically be installed through the later steps.

For Ubuntu / Debian, run:

apt-get install libcairo2 libgdk-pixbuf2.0-0 libpango-1.0-0 libpangocairo-1.0-0

On macOS (via Homebrew), run:

brew install cairo pango gdk-pixbuf

Next, set up a virtual Python environment:

python -m venv venv-samsem
source venv-samsem/bin/activate

Within that environment, install all remaining dependencies, including SAM2:

pip install -r requirements.txt

This command will install the following dependencies (copied from the requirements.txt):

# Core ML framework
sam-2 @ git+https://github.com/facebookresearch/sam2.git
torch>=2.5.1
torchvision>=0.20.1
numpy>=1.26

# HuggingFace ecosystem
transformers>=4.40
accelerate>=1.0

# Training utilities
optuna>=3.0
tensorboard>=2.14
topolosses>=0.2.0

# Image processing
matplotlib>=3.7
opencv-python-headless>=4.9
scipy>=1.11
Pillow>=10.0
cairosvg>=2.9.0

# Data / evaluation
pandas>=2.0
openpyxl>=3.1  
shapely>=2.0         
tqdm>=4.65
requests>=2.31

# Config files
pyyaml>=6.0

# Jupyter notebook execution
ipykernel>=6.0
jupyter>=1.0

Quickstart Guide

This quickstart guide demonstrates the workflow that we used to create and evaluate SAMSEM. This workflow comprises five sequential steps:

  1. Download the SAM2 checkpoint and prepare the IC image dataset.
  2. Run a hyperparameter optimization to determine the best-fitting parameters for the fine-tuning process.
  3. Actually perform fine-tuning using the previously determined hyperparameters.
  4. Use the fine-tuned model for IC image segmentation and compare against the ground truth for evaluation.
  5. Compare results with existing models from previous work.

We have implemented each of these steps in a separate Jupyter notebook that can be executed on a server equipped with an Nvidia H100 right away. If you have another GPU with less GPU memory, you can try to reduce BATCH_SIZE in the notebooks and scripts. Some parameters have deliberately been chosen small to reduce the runtime and allow for a quick demonstration of our workflow.

Notebooks

To execute the individual steps of our SAMSEM pipeline, please run the notebooks in the presented order with SAMSEM/ as your working directory.

Notebook Description
01_prepare_data.ipynb Download SAM2 checkpoint + SEM dataset; convert SVG masks to binary PNG; split into train/val/test; cut into 512 × 512 patches.
02_hparam_optimization.ipynb Two-phase Optuna HPO: first pixel-loss parameters, then topology-loss parameters.
03_finetuning.ipynb Fine-tune Model 1 (original-size) and Model 2 (512x512-pixel patches).
04_segmentation.ipynb Run segmentation and evaluation pipeline on the test set using our pre-trained models.
05_gen_table.ipynb Reproduce the comparison table across all 8 methods / variants.

Comparison of Notebook Results

Below we present the expected results for notebook 5. You may compare against your notebook outputs to verify correct reproduction of the steps on your setup.

Method ESD total ↓ ESD % ↓ Opens Shorts FPs FNs PA ↑ Dice ↑ IoU ↑
SAMSEM [Notebook 04] 60 0.24 5 15 36 4 0.979 0.953 0.910
— only pixel loss 56 0.22 4 16 35 1 0.980 0.956 0.915
— single model 61 0.24 5 17 38 1 0.979 0.953 0.911
— only Model 1 1196 4.79 411 71 488 226 0.961 0.912 0.838
— only Model 2 374 1.50 5 15 351 3 0.977 0.947 0.900
CNN DeepLabV3 54 0.22 15 10 26 3 0.976 0.944 0.894
CNN FCN 77 0.31 20 23 30 4 0.977 0.947 0.900
CNN U-Net 76 0.30 10 44 13 9 0.976 0.944 0.894

Real-World Application

The notebooks above are only intended as a tutorial to understand how to interact with SAMSEM. If you want to use your own data to evaluate or fine-tune SAMSEM yourself, you need to directly interface with the Python scripts we provide. In the following, we give an overview of the respective workflow.

In addition to the generic description below, we also provide the scripts used to generate the results in our paper in REPLICATION.md. Note, however, that these scripts cannot be used to reproduce our exact results as not all datasets from our paper could be published. Nonetheless, the scripts can help you re-run our experiments on your own data.

General Notes

Step 1 – Prepare Your Dataset

  1. Make sure that mask images are in correct format: single-channel grayscale .png images that use white (255) for metal lines and black (0) for the background. The helper script CODE/helpers/convert_masks.py takes care of this conversion for various source image types. IC images and masks are later matched based on name, hence the names of images and masks should be identical and they should be stored in separate folders (one for images and one for masks). Please make sure these folders are named image/ and mask/.
python CODE/helpers/convert_masks.py --input /path/to/mask_folder --output /path/to/output
  1. Split your data into training, validation, and test sets using CODE/helpers/split_dataset.py. At the output location, three folders named test, train, and val will be created containing the respective images and masks. If you want to change the distribution of data between training, validation, and test datasets (defaults to 70% training, 10% validation, and 20% testing), please use the --ratios option (e.g. --ratios 0.70,0.10,0.20).
python CODE/helpers/split_dataset.py --images /path/to/image_folder --masks /path/to/mask_folder --output /path/to/output
  1. Cut your original-size images and masks into smaller patches (SAMSEM uses 512x512 pixels) using CODE/helpers/cut_patches.py. The script will recursively go through all folders starting from the input path. It creates the same folder structure at the output path.
python CODE/helpers/cut_patches.py --input /path/to/dataset_folder --output /path/to/output

Step 2 – Run Hyperparameter Optimization

If you want to fine-tune SAMSEM for a new dataset, you should run your own hyperparameter optimization first. However, the hyperparameter optimization will start dozens of fine-tuning runs and may take several days or weeks depending on the number of Cuda devices that you have available. Hence, you may also choose to skip this step and proceed using our hyperparameters as detailed in step 3.

  1. You can change the configuration of the hyperparameter optimization by editing the respective YAML files: CONFIG/hparam_search_pixel.yaml and CONFIG/hparam_search_topology.yaml. Using these YAML files, you can change paths to datasets and outputs or adjust the number of epochs and the batch size depending on your GPU memory. If it does not yet exist, the SQLite Optuna database will be created automatically at the location that you provide. There is currently no limit in the number of trials that are automatically iniated by the hyperparameter optimization algorithm, it will start new trials until you terminate the process. If desired, you may set a limit in the corresponding config yaml files. The specific pruning algorithm and its settings are also hardcoded in that script.

  2. Run hyperparameter optimization for the pixel-based loss using accelerate on a server or Slurm on a cluster as shown below. Please keep the --num_processes and --mixed_precision flags as shown below to reduce GPU memory usage.

#server
accelerate launch --num_processes 1 --mixed_precision bf16 CODE/train/hyper_parameter_search.py \
    --config CONFIG/hparam_search_pixel.yaml

# Slurm cluster (example: 2 GPUs)
sbatch --gres=gpu:2 --ntasks-per-node=2 --wrap \
  "accelerate launch --num_processes 1 --mixed_precision bf16 CODE/train/hyper_parameter_search.py \
    --config CONFIG/hparam_search_pixel.yaml"
  1. Before we proceed with the hyperparameter optimization for the topology-based loss function, we need to fill in the best parameters found for the pixel-based loss function in the previous run. Run the provided helper script CODE/helpers/read_optuna_results.py to read out these parameters from the respective Optuna database and manually copy them to the CONFIG/hparam_search_topology.yaml. You can find the path to the pixel.db database file in the CONFIG/hparam_search_pixel.yaml.
python CODE/helpers/read_optuna_results.py --database /path/to/pixel.db --output /path/to/output 
  1. Run hyperparameter optimization for the topology-based loss using accelerate on a server or Slurm on a cluster as shown below. Please keep the --num_processes and --mixed_precision flags as shown below to reduce GPU memory usage. For efficiency reasons, you may optionally add the --resume_from flag to start the topology-based hyperparameter optimization from the last checkpoint of the best run of the pixel-based hyperparameter optimization. The checkpoint_save_interval parameter in the respective YAML files determines how often (in epochs) a new checkpoint is saved. If you load a previous checkpoint, training will proceed from the respective epoch. For example, if you fine-tuned with a pixel-based loss for 35 epochs and now want to add 15 epochs with the topology-based loss, you should load the checkoint from epoch 35 and set the epoch limit for the topology-based loss hyperparameter optimization to 35+15=50.
# server
accelerate launch --num_processes 1 --mixed_precision bf16 CODE/train/hyper_parameter_search.py \
    --config CONFIG/hparam_search_topology.yaml --resume_from /path/to/checkpoint

# Slurm cluster (example: 2 GPUs)
sbatch --gres=gpu:2 --ntasks-per-node=2 --wrap \
  "accelerate launch --num_processes 1 --mixed_precision bf16 CODE/train/hyper_parameter_search.py \
    --config CONFIG/hparam_search_topology.yaml --resume_from /path/to/checkpoint"

Step 3 – Fine-Tuning

  1. If you have completed the hyperparameter optimization before, you now need to copy the parameters determined through that optimization into the model configs CONFIG/finetune_Model1.yaml and CONFIG/finetune_Model2.yaml using the provided helper script CODE/helpers/read_optuna_results.py. If you wish to work with the hyperparameters used in our paper, you can simply skip this step and proceed with the default parameters.
python CODE/helpers/read_optuna_results.py --database /path/to/optuna_database.db --output /path/to/output
  1. Run fine-tuning using accelerate on a server as shown below. The two-stage SAMSEM pipeline requires fine-tuning two models: Model 1 on full-resolution images (pixel loss only) and Model 2 on 512×512 patches (pixel + topology loss), each with its own config. Please keep the --num_processes and --mixed_precision flags as shown below to reduce GPU memory usage. If you want to activate the topology-based loss function on top of the pixel-based loss, you have to specify the epoch in which the topology-based loss should be introduced using START_EPOCH_TOPOLOSSFUNC in the YAML file. You can disable the topology-based loss entirely by not specifying the START_EPOCH_TOPOLOSSFUNC parameter altogether.
# Model 1 — full-resolution images, pixel loss only
accelerate launch --num_processes 1 --mixed_precision bf16 CODE/train/finetuning.py --config CONFIG/finetune_Model1.yaml

# Model 2 — 512×512 patches, pixel + topology loss
accelerate launch --num_processes 1 --mixed_precision bf16 CODE/train/finetuning.py --config CONFIG/finetune_Model2.yaml

Step 4 – Segmentation and Evaluation

To run our full segmentation and evaluation pipeline on SAMSEM, either using our models or the ones that you fine-tuned yourself, please use the CODE/test/inference_evaluation_pipeline.py script. Below, we differentiate between these two cases.

Using Our Models

When using our models, the path to the models is already correctly specified in the script. Hence, you just need to overwrite the paths to the image dataset to segment. To this end, you should already have your dataset prepared by following step 1 of this guide. Now, you just need to provide the path to the original-size images and the path to the 512x512-pixel patches of these images. Please also specify an output path at which the segmentation masks and evaluation results are stored. To run the full segmentation and evaluation pipeline, execute the script:

python CODE/test/inference_evaluation_pipeline.py --PATH_IMAGES_MODEL_1 /path/to/images_original_size --PATH_IMAGES_MODEL_2 /path/to/image_patches --PATH_OUTPUT /path/to/output

Using Your Own Models

When using your fine-tuned models, you must additionally specify the paths to these models. Hence, we have to extend the above command with the respective --PATH_MODEL_1 and --PATH_MODEL_2 flags as shown below.

python CODE/test/inference_evaluation_pipeline.py --PATH_MODEL_1 /path/to/model_1 --PATH_MODEL_2 /path/to/model_2 --PATH_IMAGES_MODEL_1 /path/to/images_original_size --PATH_IMAGES_MODEL_2 /path/to/image_patches --PATH_OUTPUT /path/to/output

Executing Only Individual Steps

If you want to execute only individual sub-steps of the segmentation and evaluation workflow, you can call the respective scripts directly:

python CODE/test/inference.py --path_images /path/to/images --path_finetuned_model /path/to/model --path_output path/to/output [--save_mode MODE]
python CODE/helpers/cut_patches.py --input /path/to/dataset --output /path/to/patches_dataset
python CODE/helpers/merge_subtiles.py --path_patches_Model_1 /path/to/segmentation_patches_Model_1 --path_patches_Model_2 /path/to/segmentation_patches_Model_2 --output /path/to/output
python CODE/helpers/combine_patches.py --input-folder /path/to/patches --output-folder /path/to/combined_patches [--combined true/false] [--combine-original true/false] [--original-folder PATH] 
python CODE/helpers/evaluate_segmentation.py --ground_truth /path/to/ground_truth_masks --segmentation /path/to/segmentation_masks --output /path/to/output [--combined_imgs true/false] [--name NAME]

Dataset Source

The IC image dataset used in this artifact was published by:

Rothaug, N.; Klix, S.; Auth, N.; Böcker, S.; Puschner, E.; Becker, S.; Paar, C. (2023). A Real-World Metal-Layer SEM Image Dataset with Partial Labels. Edmond, V1. DOI: 10.17617/3.HY5SYN License: CC-BY 4.0