Transactions on Cryptographic Hardware and Embedded Systems 2026
SAMSEM – A Generic and Scalable Approach for IC Metal Line Segmentation
README
SAMSEM - A Generic and Scalable Approach for IC Metal Line Segmentation
SAMSEM fine-tunes SAM2 (Segment Anything 2) for binary segmentation of metal interconnect lines in SEM images of integrated circuits.
The SAMSEM segmentation pipeline comprises two different fine-tuned SAM2 models: a Model 1 operating on original-size images and Model 2 operating on 512x512-pixel patches of these images. The segmentation outputs of both models are then merged to form the final segmentation, see our publication for details. An optional Betti Matching topology loss encourages topologically correct connectivity of thin structures. Hyperparameters are tuned via Optuna across two separate search phases (pixel-loss and topology-loss).
In this README file, we list the software and hardware requirements to interact with SAMSEM, give an overview of the file structure of this artifact, provide a quickstart guide based on Jupyter notebooks to start interacting with SAMSEM, and give instructions on how to use your own data with SAMSEM.
We recommend making sure that your hardware adheres to our requirements and then starting with the quickstart guide to get a feeling of the individual steps involved with fine-tuning and evaluating SAMSEM. Those willing to get more involved can then proceed to apply SAMSEM to other datasets or execute actual, real-world runs of the hyperparameter search, fine-tuning, and the segmentation pipeline.
⚠ Disclaimer
For legal reasons, we cannot publish the full image dataset that our SAMSEM model was fine-tuned and evaluated on. In this artifact, we demonstrate the fine-tuning and evaluation process of SAMSEM using only a single dataset that was published alongside related work. Hence, you will not be able to reproduce the results reported in our paper, not even with the fine-tuned models that we provide. Consequently, this artifact cannot be used to evaluate SAMSEM's generalization capabilities without additional datasets. When fine-tuned only on the one available dataset, other existing methods will actually perform better than SAMSEM. This is for two reasons:
- The available dataset used as an example in this artifact is of high quality, has very few imaging errors, and comes with reliable ground truth.
- Existing, less complex approaches are better at adapting to the images of a single layer of a single IC, while SAMSEM's power comes from its generalization across different layers and ICs.
Requirements
To successfully execute all scripts in this artifact, your system needs to fulfill the following requirements:
- Native Ubuntu 24.04 LTS installation
- Server with a NVIDIA H100 (or a similar CUDA-capable GPU)
- Internet connection to download around 6 GB of additional external resources (e.g., a large image dataset)
- See
requirements.txtor the setup instructions below for a full list of dependencies
The SAMSEM pipeline was developed and tested on a server with:
- CPU: AMD EPYC 9534 (64 cores)
- GPU: 2× NVIDIA H100 NVL (94 GB VRAM each)
- RAM: 1.13 TB DDR5 ECC
Repository Structure
The structure of the artifact repository is shown below.
SAMSEM/
├── CODE/ # Contains all Python code
│ ├── helpers/ # Helper functions
│ │ ├── __init__.py
│ │ ├── combine_patches.py # Reconstruct original-size masks from patches
│ │ ├── convert_masks.py # Converts masks into binary PNG images
│ │ ├── cut_patches.py # Extract overlapping 512 × 512 patches from original-size images
│ │ ├── ESD_error.py # Implementation of ESD metric
│ │ ├── evaluate_segmentation.py # Evaluate using pixel-based and ESD metrics
│ │ ├── load_SEM_dataset.py # Download the example dataset
│ │ ├── merge_subtiles.py # Decide between patches from Model 1 and Model 2
│ │ ├── parse_eval_to_excel.py # Combines multiple evaluation results into one table
│ │ ├── read_optuna_results.py # Read parameters from an Optuna database
│ │ └── split_dataset.py # Deterministic dataset split into train, val, and test
│ │
│ ├── modules/ # Shared building blocks used by train/ and test/
│ │ ├── __init__.py
│ │ ├── checkpoint.py # Save / resume full training-state checkpoints
│ │ ├── datasets.py # Loads image/mask pairs into RAM
│ │ ├── logging_utils.py # TensorBoard helpers and logging setup
│ │ ├── losses.py # PixelComboLoss (Dice + BCE)
│ │ ├── sam2_wrapper.py # HuggingFace-compatible SAM2 wrapper (gradient-enabled encoder)
│ │ ├── supervised_unet.py # U-Net baseline architecture
│ │ └── transforms.py # Train / val augmentation pipelines
│ │
│ ├── test/ # Testing and evaluation scripts
│ │ ├── 05_comparison_table.py # Generate table like notebook 5
│ │ ├── inference_evaluation_pipeline.py # Full segmentation & evaluation pipeline
│ │ ├── inference_supervised_CNNs.py # CNN baseline segmentation (DeepLabV3, FCN, U-Net)
│ │ └── inference.py # SAM2 batch segmentation with configurable point prompts
│ │
│ └── train/ # Scripts for fine-tuning and hyperparameter optimization
│ ├── finetuning.py # Single finetuning run
│ └── hyper_parameter_search.py # Optuna HPO — pixel-loss or topology-loss mode
│
├── CONFIG/ # YAML configuration files
│ ├── finetune_Model1.yaml # Configuration used for all Model 1 fine-tuning runs
│ ├── finetune_Model2.yaml # Configuration used for all Model 2 fine-tuning runs
│ ├── hparam_search_pixel.yaml # Configuration used for the pixel-based hyperparameter optimization
│ └── hparam_search_topology.yaml # Configuration used for the topology-based hyperparameter optimization
│
├── MODELS/ # Contains provided models
│ └── *.torch / *.pth # 9 pre-trained models (see 'Supplied Models' below for details)
│
├── 01_prepare_data.ipynb # Notebook 1: download & prepare dataset and SAM2 checkpoint
├── 02_hparam_optimization.ipynb # Notebook 2: Run hyperparameter optimization
├── 03_finetuning.ipynb # Notebook 3: fine-tune Model 1 & Model 2
├── 04_segmentation.ipynb # Notebook 4: full inference & evaluation pipeline
├── 05_gen_table.ipynb # Notebook 5: Reproduce the paper comparison table
├── LICENSE
├── README.md
├── REPLICATION.md # Scripts that can be used to replicate our results on your data.
└── requirements.txt
Supplied Models
Nine pre-trained models are supplied in MODELS/, six of which are based on SAM2 and three more on standard convolutional neural networks for comparison.
SAM2-based Models
| File | Pipeline role | Training data | Loss |
|---|---|---|---|
SAMSEM_Model_1.torch |
Stage 1 — original-size segmentation | ICs 1–7 | Pixel (DiceBCE) |
SAMSEM_Model_2.torch |
Stage 2 — 512×512-pixel patch segmentation | ICs 1–7 | Pixel + topology (Betti Matching) |
SAMSEM_ICs1to14_Model_1.torch |
Stage 1 — original-size segmentation | All 14 ICs | Pixel (DiceBCE) |
SAMSEM_ICs1to14_Model_2.torch |
Stage 2 — 512×512-pixel patch segmentation | All 14 ICs | Pixel + topology (Betti Matching) |
OnlyPixelLoss_Model_2.torch |
Ablation — Stage 2 without topology loss | ICs 1–7 | Pixel only |
SingleModel.torch |
Ablation — one model combining Stage 1 and Stage 2 | ICs 1–7 | Pixel + topology (Betti Matching) |
Please note: SAMSEM_Model_1 and SAMSEM_Model_2 (trained on ICs 1–7) are the models discussed in Sections 3 and 4 of our the paper. The two ICs1to14 models are introduced in Section 4.5 of the paper and were trained on the full dataset of 14 integrated circuits for even better generalization across unseen IC designs.
CNN baselines
For those machine-learning approaches that made their training / fine-tuning code publicly accessible and that can be trained on our whole dataset at once, we also publish the models we used to generate our comparison results in Table 3 and 4 of our paper.
| File | Architecture | Training data |
|---|---|---|
CNN_deeplabv3.pth |
DeepLabV3 | 512 × 512 patches, ICs 1–7 |
CNN_fcn.pth |
FCN | 512 × 512 patches, ICs 1–7 |
CNN_unet.pth |
U-Net | 512 × 512 patches, ICs 1–7 |
Setup
Please follow these setup steps for both the quickstart guide and the real-world application below.
First, make sure that your system runs Python 3.10 or newer and has Git installed. There are also a couple dependencies that will not automatically be installed through the later steps.
For Ubuntu / Debian, run:
apt-get install libcairo2 libgdk-pixbuf2.0-0 libpango-1.0-0 libpangocairo-1.0-0
On macOS (via Homebrew), run:
brew install cairo pango gdk-pixbuf
Next, set up a virtual Python environment:
python -m venv venv-samsem
source venv-samsem/bin/activate
Within that environment, install all remaining dependencies, including SAM2:
pip install -r requirements.txt
This command will install the following dependencies (copied from the requirements.txt):
# Core ML framework
sam-2 @ git+https://github.com/facebookresearch/sam2.git
torch>=2.5.1
torchvision>=0.20.1
numpy>=1.26
# HuggingFace ecosystem
transformers>=4.40
accelerate>=1.0
# Training utilities
optuna>=3.0
tensorboard>=2.14
topolosses>=0.2.0
# Image processing
matplotlib>=3.7
opencv-python-headless>=4.9
scipy>=1.11
Pillow>=10.0
cairosvg>=2.9.0
# Data / evaluation
pandas>=2.0
openpyxl>=3.1
shapely>=2.0
tqdm>=4.65
requests>=2.31
# Config files
pyyaml>=6.0
# Jupyter notebook execution
ipykernel>=6.0
jupyter>=1.0
Quickstart Guide
This quickstart guide demonstrates the workflow that we used to create and evaluate SAMSEM. This workflow comprises five sequential steps:
- Download the SAM2 checkpoint and prepare the IC image dataset.
- Run a hyperparameter optimization to determine the best-fitting parameters for the fine-tuning process.
- Actually perform fine-tuning using the previously determined hyperparameters.
- Use the fine-tuned model for IC image segmentation and compare against the ground truth for evaluation.
- Compare results with existing models from previous work.
We have implemented each of these steps in a separate Jupyter notebook that can be executed on a server equipped with an Nvidia H100 right away. If you have another GPU with less GPU memory, you can try to reduce BATCH_SIZE in the notebooks and scripts. Some parameters have deliberately been chosen small to reduce the runtime and allow for a quick demonstration of our workflow.
Notebooks
To execute the individual steps of our SAMSEM pipeline, please run the notebooks in the presented order with SAMSEM/ as your working directory.
| Notebook | Description |
|---|---|
01_prepare_data.ipynb |
Download SAM2 checkpoint + SEM dataset; convert SVG masks to binary PNG; split into train/val/test; cut into 512 × 512 patches. |
02_hparam_optimization.ipynb |
Two-phase Optuna HPO: first pixel-loss parameters, then topology-loss parameters. |
03_finetuning.ipynb |
Fine-tune Model 1 (original-size) and Model 2 (512x512-pixel patches). |
04_segmentation.ipynb |
Run segmentation and evaluation pipeline on the test set using our pre-trained models. |
05_gen_table.ipynb |
Reproduce the comparison table across all 8 methods / variants. |
Comparison of Notebook Results
Below we present the expected results for notebook 5. You may compare against your notebook outputs to verify correct reproduction of the steps on your setup.
| Method | ESD total ↓ | ESD % ↓ | Opens | Shorts | FPs | FNs | PA ↑ | Dice ↑ | IoU ↑ |
|---|---|---|---|---|---|---|---|---|---|
| SAMSEM [Notebook 04] | 60 | 0.24 | 5 | 15 | 36 | 4 | 0.979 | 0.953 | 0.910 |
| — only pixel loss | 56 | 0.22 | 4 | 16 | 35 | 1 | 0.980 | 0.956 | 0.915 |
| — single model | 61 | 0.24 | 5 | 17 | 38 | 1 | 0.979 | 0.953 | 0.911 |
| — only Model 1 | 1196 | 4.79 | 411 | 71 | 488 | 226 | 0.961 | 0.912 | 0.838 |
| — only Model 2 | 374 | 1.50 | 5 | 15 | 351 | 3 | 0.977 | 0.947 | 0.900 |
| CNN DeepLabV3 | 54 | 0.22 | 15 | 10 | 26 | 3 | 0.976 | 0.944 | 0.894 |
| CNN FCN | 77 | 0.31 | 20 | 23 | 30 | 4 | 0.977 | 0.947 | 0.900 |
| CNN U-Net | 76 | 0.30 | 10 | 44 | 13 | 9 | 0.976 | 0.944 | 0.894 |
Real-World Application
The notebooks above are only intended as a tutorial to understand how to interact with SAMSEM. If you want to use your own data to evaluate or fine-tune SAMSEM yourself, you need to directly interface with the Python scripts we provide. In the following, we give an overview of the respective workflow.
In addition to the generic description below, we also provide the scripts used to generate the results in our paper in REPLICATION.md. Note, however, that these scripts cannot be used to reproduce our exact results as not all datasets from our paper could be published. Nonetheless, the scripts can help you re-run our experiments on your own data.
General Notes
- We strongly recommend completing our quickstart guide before proceeding to work through the SAMSEM workflow for real-world application.
- If you just want to use our pre-trained models to segment your own IC images, you may skip steps 2 and 3 of this workflow.
- If you want to fine-tune SAMSEM using our hyperparameters and your training data, you may skip step 2 of this workflow.
- The hyperparameter optimization is always executed on 512x512-pixel patches, not the original-size images. Fine-tuning, however, is conducted on both original-size images and patches seperately, resulting in Model 1 and Model 2, respectively.
Step 1 – Prepare Your Dataset
- Make sure that mask images are in correct format: single-channel grayscale
.pngimages that use white (255) for metal lines and black (0) for the background. The helper scriptCODE/helpers/convert_masks.pytakes care of this conversion for various source image types. IC images and masks are later matched based on name, hence the names of images and masks should be identical and they should be stored in separate folders (one for images and one for masks). Please make sure these folders are namedimage/andmask/.
python CODE/helpers/convert_masks.py --input /path/to/mask_folder --output /path/to/output
- Split your data into training, validation, and test sets using
CODE/helpers/split_dataset.py. At the output location, three folders namedtest,train, andvalwill be created containing the respective images and masks. If you want to change the distribution of data between training, validation, and test datasets (defaults to 70% training, 10% validation, and 20% testing), please use the--ratiosoption (e.g.--ratios 0.70,0.10,0.20).
python CODE/helpers/split_dataset.py --images /path/to/image_folder --masks /path/to/mask_folder --output /path/to/output
- Cut your original-size images and masks into smaller patches (SAMSEM uses 512x512 pixels) using
CODE/helpers/cut_patches.py. The script will recursively go through all folders starting from the input path. It creates the same folder structure at the output path.
python CODE/helpers/cut_patches.py --input /path/to/dataset_folder --output /path/to/output
Step 2 – Run Hyperparameter Optimization
If you want to fine-tune SAMSEM for a new dataset, you should run your own hyperparameter optimization first. However, the hyperparameter optimization will start dozens of fine-tuning runs and may take several days or weeks depending on the number of Cuda devices that you have available. Hence, you may also choose to skip this step and proceed using our hyperparameters as detailed in step 3.
-
You can change the configuration of the hyperparameter optimization by editing the respective YAML files:
CONFIG/hparam_search_pixel.yamlandCONFIG/hparam_search_topology.yaml. Using these YAML files, you can change paths to datasets and outputs or adjust the number of epochs and the batch size depending on your GPU memory. If it does not yet exist, the SQLite Optuna database will be created automatically at the location that you provide. There is currently no limit in the number of trials that are automatically iniated by the hyperparameter optimization algorithm, it will start new trials until you terminate the process. If desired, you may set a limit in the corresponding config yaml files. The specific pruning algorithm and its settings are also hardcoded in that script. -
Run hyperparameter optimization for the pixel-based loss using
accelerateon a server or Slurm on a cluster as shown below. Please keep the--num_processesand--mixed_precisionflags as shown below to reduce GPU memory usage.
#server
accelerate launch --num_processes 1 --mixed_precision bf16 CODE/train/hyper_parameter_search.py \
--config CONFIG/hparam_search_pixel.yaml
# Slurm cluster (example: 2 GPUs)
sbatch --gres=gpu:2 --ntasks-per-node=2 --wrap \
"accelerate launch --num_processes 1 --mixed_precision bf16 CODE/train/hyper_parameter_search.py \
--config CONFIG/hparam_search_pixel.yaml"
- Before we proceed with the hyperparameter optimization for the topology-based loss function, we need to fill in the best parameters found for the pixel-based loss function in the previous run. Run the provided helper script
CODE/helpers/read_optuna_results.pyto read out these parameters from the respective Optuna database and manually copy them to theCONFIG/hparam_search_topology.yaml. You can find the path to thepixel.dbdatabase file in theCONFIG/hparam_search_pixel.yaml.
python CODE/helpers/read_optuna_results.py --database /path/to/pixel.db --output /path/to/output
- Run hyperparameter optimization for the topology-based loss using
accelerateon a server or Slurm on a cluster as shown below. Please keep the--num_processesand--mixed_precisionflags as shown below to reduce GPU memory usage. For efficiency reasons, you may optionally add the--resume_fromflag to start the topology-based hyperparameter optimization from the last checkpoint of the best run of the pixel-based hyperparameter optimization. Thecheckpoint_save_intervalparameter in the respective YAML files determines how often (in epochs) a new checkpoint is saved. If you load a previous checkpoint, training will proceed from the respective epoch. For example, if you fine-tuned with a pixel-based loss for 35 epochs and now want to add 15 epochs with the topology-based loss, you should load the checkoint from epoch 35 and set the epoch limit for the topology-based loss hyperparameter optimization to 35+15=50.
# server
accelerate launch --num_processes 1 --mixed_precision bf16 CODE/train/hyper_parameter_search.py \
--config CONFIG/hparam_search_topology.yaml --resume_from /path/to/checkpoint
# Slurm cluster (example: 2 GPUs)
sbatch --gres=gpu:2 --ntasks-per-node=2 --wrap \
"accelerate launch --num_processes 1 --mixed_precision bf16 CODE/train/hyper_parameter_search.py \
--config CONFIG/hparam_search_topology.yaml --resume_from /path/to/checkpoint"
Step 3 – Fine-Tuning
- If you have completed the hyperparameter optimization before, you now need to copy the parameters determined through that optimization into the model configs
CONFIG/finetune_Model1.yamlandCONFIG/finetune_Model2.yamlusing the provided helper scriptCODE/helpers/read_optuna_results.py. If you wish to work with the hyperparameters used in our paper, you can simply skip this step and proceed with the default parameters.
python CODE/helpers/read_optuna_results.py --database /path/to/optuna_database.db --output /path/to/output
- Run fine-tuning using
accelerateon a server as shown below. The two-stage SAMSEM pipeline requires fine-tuning two models: Model 1 on full-resolution images (pixel loss only) and Model 2 on 512×512 patches (pixel + topology loss), each with its own config. Please keep the--num_processesand--mixed_precisionflags as shown below to reduce GPU memory usage. If you want to activate the topology-based loss function on top of the pixel-based loss, you have to specify the epoch in which the topology-based loss should be introduced usingSTART_EPOCH_TOPOLOSSFUNCin the YAML file. You can disable the topology-based loss entirely by not specifying theSTART_EPOCH_TOPOLOSSFUNCparameter altogether.
# Model 1 — full-resolution images, pixel loss only
accelerate launch --num_processes 1 --mixed_precision bf16 CODE/train/finetuning.py --config CONFIG/finetune_Model1.yaml
# Model 2 — 512×512 patches, pixel + topology loss
accelerate launch --num_processes 1 --mixed_precision bf16 CODE/train/finetuning.py --config CONFIG/finetune_Model2.yaml
Step 4 – Segmentation and Evaluation
To run our full segmentation and evaluation pipeline on SAMSEM, either using our models or the ones that you fine-tuned yourself, please use the CODE/test/inference_evaluation_pipeline.py script. Below, we differentiate between these two cases.
Using Our Models
When using our models, the path to the models is already correctly specified in the script. Hence, you just need to overwrite the paths to the image dataset to segment. To this end, you should already have your dataset prepared by following step 1 of this guide. Now, you just need to provide the path to the original-size images and the path to the 512x512-pixel patches of these images. Please also specify an output path at which the segmentation masks and evaluation results are stored. To run the full segmentation and evaluation pipeline, execute the script:
python CODE/test/inference_evaluation_pipeline.py --PATH_IMAGES_MODEL_1 /path/to/images_original_size --PATH_IMAGES_MODEL_2 /path/to/image_patches --PATH_OUTPUT /path/to/output
Using Your Own Models
When using your fine-tuned models, you must additionally specify the paths to these models. Hence, we have to extend the above command with the respective --PATH_MODEL_1 and --PATH_MODEL_2 flags as shown below.
python CODE/test/inference_evaluation_pipeline.py --PATH_MODEL_1 /path/to/model_1 --PATH_MODEL_2 /path/to/model_2 --PATH_IMAGES_MODEL_1 /path/to/images_original_size --PATH_IMAGES_MODEL_2 /path/to/image_patches --PATH_OUTPUT /path/to/output
Executing Only Individual Steps
If you want to execute only individual sub-steps of the segmentation and evaluation workflow, you can call the respective scripts directly:
- To segment images using a single model, run
CODE/test/inference.py. If you work with image patches, you will have to stitch the resulting segmentation masks afterwards usingCODE/helpers/combine_patches.py. This also works if you do not have ground truth masks for the images. You may optionally define a--save_mode, which can be eithersingle_mask(default option; each output file just contains the segmentation mask) orconcat_img_mask(each output file contains the input image next to the segmentation mask; may be used for debugging). Note that the two-model merge workflow below (merge_subtiles.py) matches Model 1 and Model 2 patches by filename and therefore requiressingle_maskoutputs — useconcat_img_maskonly for visual debugging of individual predictions.
python CODE/test/inference.py --path_images /path/to/images --path_finetuned_model /path/to/model --path_output path/to/output [--save_mode MODE]
- To cut segmentation masks from Model 1 (operating on original-size images) into patches, please use
CODE/helpers/cut_patches.py.
python CODE/helpers/cut_patches.py --input /path/to/dataset --output /path/to/patches_dataset
- To decide between segmentation patches from Model 1 and 2 using our proposed decision algorithm, run
CODE/helpers/merge_subtiles.py.
python CODE/helpers/merge_subtiles.py --path_patches_Model_1 /path/to/segmentation_patches_Model_1 --path_patches_Model_2 /path/to/segmentation_patches_Model_2 --output /path/to/output
- To combine segmentation mask patches into original-size masks again, please run
CODE/helpers/combine_patches.py. If--save_mode concat_img_maskwas used withCODE/test/inference.py,--combined truemust be set to make sure that image and mask are correctly seperated again before combining the patches. You can also set--combine-original trueto again combine the original-size image and the (merged) segmentation mask into the same file for debugging. If you do so, you must also specify the path to the original-size images using--original-folder.
python CODE/helpers/combine_patches.py --input-folder /path/to/patches --output-folder /path/to/combined_patches [--combined true/false] [--combine-original true/false] [--original-folder PATH]
- To evaluate a segmentation against the ground truth, you can use
CODE/helpers/evaluate_segmentation.py. Set--combined_imgs trueif you are working with input files that combine the input image and the segmentation mask. You may optionally set a project name via--nameto change the name of the evaluation result files.
python CODE/helpers/evaluate_segmentation.py --ground_truth /path/to/ground_truth_masks --segmentation /path/to/segmentation_masks --output /path/to/output [--combined_imgs true/false] [--name NAME]
Dataset Source
The IC image dataset used in this artifact was published by:
Rothaug, N.; Klix, S.; Auth, N.; Böcker, S.; Puschner, E.; Becker, S.; Paar, C. (2023). A Real-World Metal-Layer SEM Image Dataset with Partial Labels. Edmond, V1. DOI: 10.17617/3.HY5SYN License: CC-BY 4.0