X-BQSR: Holistic acceleration of base quality recalibration for scalable genomic analysis

Marinelli, Eugenio; Donchev Kabadzhov, Ivan; Appuswamy, Raja
EURO-PAR 2026, 32nd International European Conference on Parallel and Distributed Computing, 24-28 August 2026, Pisa, Italy / Also published in "Lecture Notes in Computer Science"

With the continued decline in sequencing costs, genomics-based precision medicine is gaining increasing importance. The Genome Analysis Toolkit (GATK) is the standard framework for converting raw genomic data into actionable variants, but one of its most expensive steps is base quality score recalibration (BQSR), which corrects systematic sequencing errors. Despite significant work on accelerating other parts of the variant calling pipeline, BQSR has not been systematically explored on GPUs.

In this work, we analyze BQSR to identify computational and I/O bottlenecks and show that simply offloading key stages to an accelerator is insufficient to achieve overall gains. We then present X-BQSR, a holistic redesign of the entire BQSR pipeline, including I/O, compression, and kernel execution, to better leverage GPU parallelism. We implement X-BQSR using portable SYCL kernels for NVIDIA, AMD, and Intel GPUs, along with native CUDA and HIP backends for comparison, and we evaluate its performance against previous FPGA work. Our evaluation shows up to 141.7×" role="presentation" style="box-sizing: inherit; display: inline-block; line-height: normal; font-size-adjust: none; word-spacing: normal; overflow-wrap: normal; white-space: nowrap; float: none; direction: ltr; max-width: none; max-height: none; min-width: 0px; min-height: 0px; border: 0px; padding: 0px; margin: 0px; outline: 0px; scroll-margin-top: 74px; position: relative;">

acceleration compared to single-threaded GATK and 8.7×" role="presentation" style="box-sizing: inherit; display: inline-block; line-height: normal; font-size-adjust: none; word-spacing: normal; overflow-wrap: normal; white-space: nowrap; float: none; direction: ltr; max-width: none; max-height: none; min-width: 0px; min-height: 0px; border: 0px; padding: 0px; margin: 0px; outline: 0px; scroll-margin-top: 74px; position: relative;">

acceleration compared to a proprietary GPU-accelerated solution.


DOI
HAL
Type:
Conference
City:
Pisa
Date:
2026-08-24
Department:
Data Science
Eurecom Ref:
8907
Copyright:
© Springer. Personal use of this material is permitted. The definitive version of this paper was published in EURO-PAR 2026, 32nd International European Conference on Parallel and Distributed Computing, 24-28 August 2026, Pisa, Italy / Also published in "Lecture Notes in Computer Science" and is available at : https://doi.org/10.1007/978-3-032-35248-4_36

PERMALINK : https://www.eurecom.fr/publication/8907