Download PDF
Abstract
Short-read sequencing (SRS)-based disease-targeted NGS gene panels have revolutionized rare disease diagnostics but often leave autosomal recessive cases unsolved when only one pathogenic allele is detected. Missing variants may reside in deep intronic regions or involve structural variants (SVs) undetectable by SRS. To improve diagnostic yield, we implemented a cost-effective target capture-based long-read sequencing (LRS) assay covering 56 genes and retrospectively analyzed 78 patients suspected of autosomal recessive disorders who remained undiagnosed after SRS. Functional validation using reverse transcription PCR (RT-PCR) and minigene assays was performed to further determine pathogenicity. Target capture-based LRS solved 25.6% (20/78) of cases by identifying 10 SVs, 3 deep intronic variants experimentally confirmed to cause aberrant splicing, and 7 cases in which haplotype phasing confirmed that variants were in trans with the known pathogenic variant, leading to reclassification of the VUS as likely pathogenic. This study demonstrates that target capture-based LRS effectively detects diverse types of variants missed by SRS. Integrating this assay into stepwise diagnostic workflows offers a practical and cost-effective strategy to enhance diagnostic yield in autosomal recessive diseases. However, because this cohort was retrospectively defined based on a prior single-allele detection by SRS, this 25.6% (20/78) diagnostic yield reflects performance within a highly enriched population and should not be directly extrapolated to unselected rare disease cohorts.
Introduction
The use of short-read-based next-generation sequencing (NGS) has improved the diagnosis of patients with rare genetic diseases, yet more than half of individuals with suspected rare genetic disorders remain undiagnosed [1, 2]. Disease-targeted NGS gene panels, which are widely implemented in clinical settings, yield molecular diagnoses in only 20–30% of cases [3]. Short-read NGS is often restricted to detecting single-nucleotide variants and small insertions or deletions within coding regions, and is technically limited in its ability to detect structural variants (SVs) and other complex variants, with reported SV detection sensitivity as low as 10–70% and false-positive rates up to 89% [4]. In patients with autosomal recessive conditions, disease-targeted NGS gene panels often identify a single heterozygous pathogenic variant, representing a critical diagnostic gap since the second allele may reside in deep intronic regions or harbor structural variants (SVs) that are inaccessible to short-read sequencing (SRS). Therefore, for autosomal recessive diseases in which only a single heterozygous pathogenic variant has been identified, a more comprehensive and effective strategy is required to detect the missing variant on the opposite allele and confirm biallelic pathogenicity.
Long-read sequencing (LRS) technologies can generate sequence reads spanning tens to hundreds of kilobase pairs, with PacBio HiFi reads averaging 15–25 kb and Oxford Nanopore Technologies ultra-long reads exceeding 100 kb, enabling the direct detection of SVs, deep intronic alterations, and repeat expansions while providing haplotype phasing across extended genomic intervals [5, 6]. Several recent studies have demonstrated that targeted LRS focusing on candidate disease-causing genes can uncover previously undetected pathogenic variants and substantially increase diagnostic yield in unresolved Mendelian disorders [7,8,9]. These advances are particularly relevant for autosomal recessive diseases, in which SRS often detects only a single heterozygous pathogenic variant, leaving the second allele unresolved. Here, we applied a targeted LRS strategy using a custom hybridization capture-based panel [10] encompassing 56 autosomal recessive disease genes to a systematically defined clinical cohort of patients with a single heterozygous pathogenic or likely pathogenic variant identified by SRS, aiming to identify the missing pathogenic allele and thereby improve diagnostic yield in unresolved recessive disorders.
Materials and methods
Study design and patients
We reviewed the results of disease-targeted NGS gene panel testing at Seoul National University Hospital from October 2017 to October 2021. Detailed clinical phenotypes and family histories were assessed, and 78 patients whose features were consistent with an autosomal recessive condition but remained undiagnosed with only a single pathogenic variant identified in a gene associated with an autosomal recessive disease were included in this study. The 78 patients spanned multiple disease categories, with inherited retinal and other eye disorders comprising the largest subgroup (n = 37, 47.4%), followed by neuromuscular diseases (n = 10, 12.8%), renal disorders (n = 8, 10.3%), neurological conditions (n = 7, 9.0%), metabolic diseases (n = 6, 7.7%), hereditary hearing loss (n = 5, 6.4%), endocrine disorders (n = 3, 3.8%), and immunological and skeletal conditions (n = 1 each, 1.3%). The study was approved by the Institutional Review Board of Seoul National University Hospital (IRB No. H-2301-101-1396).
Custom target capture-based long-read sequencing panel design
The custom target capture-based LRS panel was specifically designed to improve the detection of SVs, haplotype-phased variants, and deep intronic variants in genes associated with autosomal recessive diseases. The panel targeted the entire coding regions, all intronic regions, and the 5′ and 3′ untranslated regions (UTRs) of 56 genes (Table S1), which were selected based on a retrospective review of disease-targeted NGS panel data from patients carrying a single heterozygous pathogenic variant, with or without an additional variant of uncertain significance, in an autosomal recessive disease gene where a second undetected variant was suspected to contribute to the phenotype. These 56 genes were selected based on the retrospective review described above, encompassing all genes in which a single heterozygous pathogenic variant had been identified without a confirmed second allele in our institutional patient population. These clinical NGS panels included a wide range of disorders, such as metabolic diseases (glycogen storage disease, metabolic myopathy, and calcium/phosphate/magnesium metabolism disorders), neuromuscular diseases (Charcot–Marie–Tooth disease, congenital myopathy, limb-girdle muscular dystrophy), inherited retinal diseases, hereditary hearing loss, renal disorders, and neurogenetic disorders (ataxia, Parkinson’s disease, and stroke).
Long-read sequencing library Preparation
Genomic DNA (500–700 ng) was fragmented using a g-TUBE (Covaris, Massachusetts, USA) to generate target fragment sizes of 8–10 kb. For the fragmented DNA, end-repair and A-tailing were performed, followed by adapter ligation according to the Twist Bio Long Read Library Preparation and Standard Hyb v2 Enrichment workflow to construct a pre-capture library. Target enrichment was performed by hybridizing the ligated DNA fragments to a custom-designed panel of 120 bp biotinylated DNA probes (Twist Bioscience), specifically designed to capture the entire regions of genes selected in this study. A post-capture PCR was then performed, and the amplified products were used for SMRTbell library preparation following the SMRTbell Prep Kit 3.0 (Pacific Biosciences, Menlo Park, CA). The final libraries were pooled and sequenced on the PacBio Sequel II platform (Pacific Biosciences).
Sequence analysis
Raw sequencing data were processed into circular consensus sequence reads using CCS v6.4.0 software. Reads were demultiplexed with Lima (v2.1.0) and aligned to the human reference genome (hg38) using pbmm2 (v1.13.1). SVs were identified using pbsv (v2.9.0) and annotated with AnnotSV (v3.4.1). For single-nucleotide variant and small insertion/deletion calling and annotation, DeepVariant (v1.5.0) and snpEff (v5.1) were used. Haplotype phasing was performed with WhatsHap v2.1. Downstream analysis included identifying SVs and phasing candidate variants with known pathogenic or likely pathogenic variants. Each SV and its breakpoints were manually reviewed using Integrative Genomics Viewer (IGV, version 2.16.0) [11]. For novel intronic variants, splicing effects were predicted using SpliceAI [12], and variants with maximum delta scores ≥ 0.2 were further evaluated by minigene assays or RT-PCR to experimentally assess their splicing effects. Sequence variants, including single-nucleotide variants, small insertions/deletions, and deep intronic variants, were classified according to the 2015 ACMG/AMP guidelines [13]. For intragenic SVs in which both breakpoints reside within a single gene the 2015 ACMG/AMP guidelines were applied, while for SVs extending beyond a single gene boundary or affecting upstream regulatory regions, the ACMG/ClinGen technical standards were applied [14].
Splicing analysis
For deep intronic variants, splicing effects were evaluated using either a minigene assay or RT-PCR, depending on gene expression levels and the availability of additional blood specimens. RT-PCR was performed when the candidate gene was sufficiently expressed in peripheral blood and additional blood samples were available. When gene expression in blood was low (TPM < 0.2) or no additional blood samples could be obtained, a minigene assay was performed
Minigene assay
To assess splice variants affecting splicing, we performed minigene assays for EYS, USH2A, and NEB. Genomic regions containing four deep intronic variants (two in EYS and one each in USH2A and NEB) were amplified by PCR. Each fragment included approximately 150 bp of flanking intronic sequence on both the 5′ and 3′ ends. Primers were designed with 5′ tails carrying restriction sites for XhoI and BamHI, facilitating cloning into the splicing reporter vector pSPL3. Wild-type and variant constructs were confirmed by Sanger sequencing before being separately transfected into HEK293T cells (ATCC CRL-3216). After 48 hours, total RNA was extracted, and RT-PCR was performed using universal primers targeting exonic regions of the minigene vector. RT-PCR products were analyzed by gel electrophoresis and Sanger sequencing to compare splicing patterns between wild-type and variant alleles.
RT-PCR
For DYSF, we performed RT-PCR to determine whether the deep intronic variant affected splicing. Total RNA was extracted from peripheral blood using the QIAamp RNA Blood Mini Kit (Qiagen). cDNA was synthesized using SuperScript IV (Invitrogen) according to the manufacturer’s instructions. PCR amplification was performed from DYSF exon 38 to exon 45. RT-PCR products were analyzed by gel electrophoresis and Sanger sequencing
Results
Target capture-based LRS was performed on 78 patient samples, achieving a mean coverage depth of ≥30× and a read N50 of 4.6 kb. On average, 99.9% of reads were successfully mapped to the reference genome, and 76.1% of mapped reads were aligned to the target regions. Among the 78 patients clinically diagnosed with autosomal recessive diseases, 65 had a single pathogenic or likely pathogenic (P/LP) variant identified by SRS, and 15 carried one P/LP variant together with an additional VUS requiring phasing. Target capture-based LRS additionally revealed a pathogenic or likely pathogenic variant in 25.6% (20/78) of affected individuals, and independent review of these 20 solved cases confirmed phenotype concordance with the established disease spectrum of each gene, thus representing a significant improvement in diagnostic yield for these unresolved patients (Fig. 1). Specifically, SVs were identified in 10 patients (10/78, 12.8%), and novel deep intronic variants were detected in five patients, three of which (3/78, 3.85%) were experimentally confirmed to result in aberrant splicing through minigene assays or RT-PCR. Among the 15 patients harboring both a pathogenic variant and a VUS, haplotype phasing resolved 8 cases: 7 VUSs (7/78, 8.97%) were found to be in trans with the known pathogenic variants, while 1 was in cis. Collectively, target capture-based LRS improved diagnostic yield by enabling detection of hidden structural and deep intronic variants and by resolving phasing in compound heterozygous cases undetermined by conventional SRS.
Cases in which the causative variants were identified by this assay are highlighted in blue boxes, whereas cases that remained unresolved despite this approach are shown in gray boxes. P pathogenic, LP likely pathogenic, VUS variant of uncertain significance
Characterization of structural variants
Ten patients were identified to have pathogenic SVs, including deletions ranging from 600 bp to 1 Mb, a tandem duplication, and a transposable element insertion previously undetected by SRS (Table S2). Target capture-based LRS enabled precise breakpoint mapping and clarification of their molecular consequences
Four patients (P03, P24, P26, and P33) carried deletions predicted to cause frameshifts leading to premature termination and nonsense-mediated decay (Fig. 2A). P03 had a 2753 bp deletion encompassing exon 16 (94 bp) of PHKB (NC_000016.10:g.47641304_47644056del), confirming a diagnosis of glycogen storage disease type IXb. P24 had a 1,409 bp deletion spanning exon 20 (139 bp) of RPGRIP1 (NC_000014.9:g.21329012_21330420del), consistent with RPGRIP1-related retinitis pigmentosa. P26 had a 1,252 bp deletion (NC_000015.10:g.44563677_44564928del) in SPG11 spanning exon 39 (152 bp), consistent with hereditary spastic paraplegia. P33 had a 680 bp deletion (NC_000008.11:g.10607808_10608487del) in exon 4 of RP1L1, resulting in a frameshift and premature stop codon (p.Ser1871Ilefs*25).
Integrative Genomics Viewer screenshots for each SV are shown. A Large out-of-frame deletions in PHKB, RPGRIP1, SPG11, and RP1L1. B Large in-frame deletions in SLC12A3, IFT140, and PDE6B. C Large deletion in noncoding region of EYS. D Tandem duplication in CEP290 and 3 kb retrotransposal insertion within intron 16 of SLC25A13
Three patients (P20, P30, and P69) exhibited in-frame deletions that removed critical functional domains (Fig. 2B). P20 had a 2128 bp deletion spanning exons 7 and 8 of SLC12A3 (NC_000016.10:g.56871873_56874001del) resulted in an in-frame loss of 81 amino acids within transmembrane domains, consistent with Gitelman syndrome. P30 carried a 9,587 bp deletion (NC_000016.10:g.1504371_1513957del) encompassing exon 31 in IFT140, which is the terminal exon of the gene. Loss of part of the C-terminal TPR domain likely resulted in a nonsyndromic RP phenotype P69 had a 3,885 bp deletion (NC_000004.12:g.633537-637421del) in PDE6B, the gene responsible for autosomal recessive RP 40, was identified. This deletion led to the loss of exon 2 and 3, which is expected to lack one of GAF domains of the noncatalytic cyclic guanosine monophosphate -binding, and have partial loss of PDE6B function.
Patient P60 had a large 1 Mb deletion (NC_000006.12:g.65501436_66562891del) encompassing the 5′ UTR, including exons 1 and 2 of EYS (Fig. 2C). This deleted region contains the putative EYS promoter that interacts with transcription factors such as OTX2, CRX, and NRL, implicating EYS dysfunction as the cause of RP
We also identified other types of SVs using target capture-based -LRS (Fig. 2D). P52, clinically diagnosed with RP, carried a large in-frame tandem duplication (39,126 bp) of exons 31–53 in CEP290 (NC_000012.12:g.88050037_88089162dup). This in-frame tandem duplication could result in changes in structural integrity of protein, or excessive posttranslational modification, impairing ciliary function. P62 harbored a 3 kb insertion within intron 16 of SLC25A13, which was considered as retrotransposal insertion of originating from C6orf68 on chromosome 6. This insertion, recurrently reported in patients with citrin deficiency and absent from population databases, was confirmed to be pathogenic.
Phasing of clinically identified variants
Among the 15 patients carrying a previously identified pathogenic or likely pathogenic variant and an additional VUS, haplotype phasing successfully resolved allele phase in eight cases (Table 1). Seven were confirmed to be in trans, and one was in cis. The distances between the two variants ranged from 41 bp to 375,559 bp. Phasing enabled reclassification of the VUSs as likely pathogenic according to American College of Medical Genetics and Genomics / Association for Molecular Pathology 2015 criteria (PM3) [13].
This approach allowed phasing directly from the proband’s DNA without requiring parental samples. These findings highlight the essential role of LRS in determining allelic configuration in cases where SRS cannot resolve phase. Accurate determination of compound heterozygosity facilitates variant classification and shortens the time to a definitive molecular diagnosis
Identifying novel deep-intronic variants and splicing analysis
Target capture-based LRS detected five novel deep intronic variants that were not identified by clinical SRS, which primarily targets coding regions and immediate exon–intron boundaries (Table 2). These five variants were functionally evaluated by either a minigene approach or RT-PCR to determine their effect on pre-mRNA splicing. Three of the five variants induced partial intron retention or pseudo-exon inclusion, leading to premature termination codon-containing transcripts that are not expected to produce functional proteins, and were subsequently reclassified as likely pathogenic.
Patient P35 was pathologically confirmed with Nemaline rod myopathy; however, only one heterozygous pathogenic variant in NEB (c.1674+1 G > T) had been identified by congenital myopathy gene panel. Target capture-based LRS revealed an additional novel deep-intronic variant, c.23742+164 A > G, predicted by Splice AI to alter splicing (Δ score = 0.98). A minigene assay demonstrated that c.23742+164 A > G allele induced retention of a 163 bp of intron 165 (Fig. 3A–C). Patient P67 presented with bilateral sensorineural hearing loss, and a hereditary hearing loss gene panel detected a heterozygous USH2A c.10712 C > T (p.Thr3571Met) variant. Using LRS, we identified a deep intronic variant, USH2A c.7120+1475 A > G, predicted to affect splicing (Δ score = 0.28); this compound heterozygous genotype has been previously reported [15], and a minigene assay confirmed that the c.7120+1475 A > G allele caused inclusion of a 266 bp pseudo-exon, leading to an aberrant transcript (Fig. 3D, E). Importantly, because this patient is identical to the one described in the previous literature, the successful identification of this variant demonstrates that our targeted LRS approach can reliably and accurately reproduce known pathogenic variants. Patient P65 was clinically diagnosed with dysferlinopathy; however, a limb-girdle muscular dystrophy panel had detected only a heterozygous DYSF c.1464del (p.Gly489Glufs*4) variant. An additional novel deep intronic variant in DYSF, c.4528-1779 T > G, predicted to affect splicing was identified (Δ score = 0.68). RT-PCR confirmed a 61 bp pseudo-exon insertion (r.4527_4528ins[4528-1840_4528-1780]) (Fig. 4A, B).
A HEK293T cells were transfected with wild-type or mutant NEB minigene or an empty pSPL3 vector. After reverse transcription, splicing products were amplified by polymerase chain reaction using vector exon–specific primers and B visualized by agarose gel electrophoresis. C Sanger sequencing confirmed the reverse transcription–polymerase chain reaction products, and the results showed that partial intron 165 (163 bp) was retained. D HEK293T cells were transfected with wild-type or mutant USH2A minigene or an empty pSPL3 vector. After reverse transcription, splicing products were amplified by polymerase chain reaction using vector exon–specific primers and E visualized by agarose gel electrophoresis. F Sanger sequencing confirmed the reverse transcription–polymerase chain reaction products, and the results showed that pseudoexon (266 bp) was included. EV empty vector, NC normal control, P proband, S A splice acceptor exon, SD splice donor exon, TE Tris-EDTA buffer.
A Normal and normal alternative splicing identified in normal control. B Schematic representation of aberrant splicing in P65. The DYSF c.4528-1779 T > G variant (red star) created cryptic donor and acceptor splice sites, resulting in the inclusion of a 61 bp pseudo-exon (red dashed line and orange box) between exons 41 and 42. Exon 41 (63 bp) skipping (green dashed line) was observed in both the normal control and patient samples. C RT-PCR analysis showing an additional 783 bp and 720 bp bands (red arrows) corresponding to the 61 bp pseudo-exon inclusion, compared with the 722 bp and 659 bp (green arrows) normal transcripts in the control.
Discussion
We applied a custom hybridization capture-based LRS panel to identify missing pathogenic alleles in patients with autosomal recessive diseases who remained undiagnosed after SRS. This approach increased diagnostic yield by 25.6% by detecting additional pathogenic or likely pathogenic variants, including structural variants, deep intronic variants, and phased compound heterozygous variants. The clinical contribution of this work lies in the systematic application to a well-defined single-allele autosomal recessive cohort spanning heterogeneous phenotypes, the quantified incremental diagnostic yield in a real-world patient population, and the integration of functional splicing validation through minigene assays and RT-PCR, which confirmed aberrant splicing in several deep intronic variants and supported their pathogenicity.
Six of the ten patients with additional SVs had inherited retinal diseases (IRDs), including those involving RPGRIP1, EYS, IFT140, RP1L1, CEP290, and PDE6B. This finding aligns with previous studies that structural variation is a major contributor to the genetic basis of IRDs [16,17,18]. In a recent large-scale study by Zampaglione et al., pathogenic copy number variants were detected in 8.8% of patients with IRDs, including RPGRIP1 and EYS [19]. These results highlight the importance of strategic approaches for detecting SVs in IRD disease groups.
The recurrent SLC25A13 3 kb insertion observed in P62 has an allele frequency of approximately 1 in 228 among East Asians and Koreans [20,21,22]. This high-frequency copy number variant underscores the importance of including recurrent population-specific SVs when designing capture probes
In our study, the large deletion identified in the 5′ UTR of EYS highlights the potential pathogenic role of noncoding structural variants in RP. The deleted region spans exons 1 and 2, which harbor the putative EYS promoter bound by photoreceptor-specific transcription factors [23]. Loss of this promoter region likely disrupts transcriptional regulation of EYS and may also influence the dynamics of upstream initiation codons, leading to altered gene expression. Such defects are consistent with the relatively mild form of RP observed in similar cases. These findings support the growing evidence that a proportion of unsolved RP cases may result from noncoding variants affecting cis-regulatory elements rather than coding mutations. Therefore, screening of the EYS 5′-UTR, particularly exons 1 and 2, should be considered in patients carrying a single heterozygous EYS pathogenic variant to enable a more complete molecular diagnosis.
Several recent studies have demonstrated additional diagnostic value of LRS in previously unsolved cases after SRS. Miller et al. performed targeted LRS using adaptive sampling on the Oxford Nanopore Technologies (ONT) platform in ten individuals, and achieved 60% yield of targeted ONT LRS to identify either a single pathogenic variant in an autosomal recessive disease gene, or no pathogenic variant for a specific suspected autosomal dominant or X-linked disorder had been identified by previous clinical testing [8]. AlAbdi et al. performed long-read whole genome sequencing (lrWGS) on the PacBio Sequel II platform in 34 families with exome-negative autosomal recessive diseases, achieving a 38% diagnostic yield [24]. Similarly, de la Morena-Barrio et al. used lrWGS with the ONT platform in 11 families with antithrombin deficiency who had negative SERPINC1 results, identifying SVs in 18.1% of cases [25]. However, the authors noted that lrWGS remains limited by high cost and data complexity.
Our results are consistent with these prior findings and extend them by showing that targeted LRS can achieve comparable diagnostic yields through a focused and cost-efficient approach. Several distinctive features differentiate this study from prior targeted LRS and lrWGS studies (Table S3). First, unlike single disease focused panels [25, 26] or computational enrichment through adaptive sampling [8, 26], we applied a 56-gene hybrid capture panel to a large, phenotypically heterogeneous cohort of 78 single-allele AR cases. Of note, because our cohort encompasses all consecutive AR cases that remained undiagnosed after initial routine disease-targeted NGS panel testing in a certified clinical laboratory, this study reflects an unbiased, consecutive subject enrollment that underscores the practical applicability of our assay in routine clinical practice. Second, while genome-wide lrWGS requires a high amount of resources to store, process, and analyze data [24, 25], our targeted assay serves as a scalable, cost-efficient alternative that simultaneously detects SVs, deep intronic variants, and phase-informative haplotypes. The current protocol multiplexes 16 samples per SMRT Cell on the PacBio Sequel II platform, with an estimated per-sample cost of approximately $150 for library preparation including Twist hybrid capture and SMRTbell preparation, and $110 for sequencing, totaling approximately $260 per sample. This compares favorably with lrWGS at approximately $2,700 per sample. Furthermore, this targeted approach is highly cost-effective compared to other selective long-read strategies, such as Oxford Nanopore Technologies adaptive sampling, which is estimated at approximately $600 per sample [27], while providing the distinct advantage of PacBio’s higher sequencing accuracy. Although this cost is higher than that of MLPA at approximately $80 per sample, it offers significantly broader target regions, providing the additional capabilities of deep intronic variant detection and haplotype phasing within a single assay. In addition, it presents a more cost-effective alternative than short-read WGS-based CNV analysis at approximately $730 per sample while providing the distinct advantage of haplotype phasing. Third, we systematically integrated functional splicing validation using minigene and RT-PCR assays to support variant reclassification through functional characterization, a crucial step often unaddressed in prior LRS approaches.
Among the unresolved cases in our cohort, the specificity between the observed clinical phenotypes and the candidate genes was generally well-supported. For metabolic, endocrine, renal, and neurological disorders, biochemically or clinically defined phenotypes provided strong gene-specific attribution, such as confirmed hypothyroidism involving DUOX2, DUOXA2, or TG, renal hypouricemia associated to SLC22A12, and dopa-responsive dystonia attributed to TH. However, for inherited retinal diseases, the inherent genetic heterogeneity of the disease group means that clinical presentation alone cannot always definitively differentiate among multiple candidate genes.
This study has several limitations. First, although targeted LRS approach improved diagnostic yield, the cohort size was relatively small. In addition, despite the clear advantages of detection of structural and deep intronic variants and providing haplotype phasing, phasing was successfully resolved in 8 of 15 cases (53.3%). In the 7 unresolved cases, inter-variant distances ranged from approximately 6 to 126 kb (median 30 kb), exceeding the 8–10 kb median fragment size of our current library protocol. Strategies to improve phasing efficiency include generating larger average fragment sizes (e.g., 15–20 kb) and incorporating ultra-long read protocols on Oxford Nanopore Technologies platforms. Alternatively, high-throughput platforms such as PacBio Revio can be utilized to yield longer continuous HiFi reads, or mixed-fragment libraries can be employed, which may further improve phasing across longer inter-variant distances. Third, our 56-gene panel covers only a limited subset of autosomal recessive genes based on institutional cohort characteristics. Expanding to a comprehensive gene set will be necessary to enhance diagnostic yield in broader populations.
Contemporary exome- and genome-first diagnostic paradigms render targeted LRS a practical, cost-efficient strategy tailored for autosomal recessive conditions with a single identified pathogenic allele. Within exome-first workflows, this assay can be utilized as a reflex second-tier test when WES fails to identify a second pathogenic variant, uncovering structural variants and deep intronic variants that fall beyond the scope of standard exome analysis. In srWGS-first settings, targeted LRS remains complementary rather than redundant. Because short-read technology inherently lacks the sensitivity to detect SVs in repetitive or low-complexity regions and cannot directly phase compound heterozygous variants across long genomic intervals, targeted LRS effectively addresses these diagnostic gaps at a substantially lower cost and with greater analytical simplicity than lrWGS. Consequently, we propose integrating targeted LRS as a standard second-tier diagnostic test following inconclusive SRS-based evaluation, whether by gene panel, WES, or srWGS, particularly in autosomal recessive disorders characterized by structural or deep intronic pathogenicity.
This study demonstrates that target capture-based LRS is a powerful and cost-effective complement to SRS for resolving unsolved autosomal recessive diseases. By enabling detection of structural and deep intronic variants and providing haplotype phasing, this approach effectively identifies missing pathogenic alleles undetectable by conventional methods. The integration of high-resolution variant detection with functional validation, such as RT-PCR and minigene assays, strengthens clinical interpretation and supports variant reclassification. Although technical limitations remain, particularly in phasing distant variants, continued optimization of library preparation and sequencing strategies will further improve efficiency. Incorporating target capture-based LRS into stepwise diagnostic workflows represents a practical and cost-efficient strategy to improve molecular diagnosis and expand the genetic understanding of rare autosomal recessive diseases.
Data availability
The data that support the findings of this study are available from the corresponding author upon request
References
Schuler BA, Nelson ET, Koziura M, Cogan JD, Hamid R, Phillips JA, 3rd Lessons learned: next-generation sequencing applied to undiagnosed genetic diseases. J Clin Invest. 2022; 132; https://doi.org/10.1172/JCI154942
Yang Y, Muzny DM, Reid JG, Bainbridge MN, Willis A, Ward PA, et al. Clinical whole-exome sequencing for the diagnosis of mendelian disorders. N Engl J Med. 2013;369:1502–11. https://doi.org/10.1056/NEJMoa1306555
ArticleÂ
CASÂ
PubMedÂ
PubMed CentralÂ
Google ScholarÂXue Y, Ankala A, Wilcox WR, Hegde MR. Solving the molecular diagnostic testing conundrum for Mendelian disorders in the era of next-generation sequencing: single-gene, gene panel, or exome/genome sequencing. Genet Med. 2015;17:444–51. https://doi.org/10.1038/gim.2014.122
ArticleÂ
CASÂ
PubMedÂ
Google ScholarÂLiu Z, Zhu L, Roberts R, Tong W. Toward clinical implementation of next-generation sequencing-based genetic testing in rare diseases: where are we?. Trends Genet. 2019;35:852–67. https://doi.org/10.1016/j.tig.2019.08.006
ArticleÂ
CASÂ
PubMedÂ
Google ScholarÂDel Gobbo GF, Boycott KM. The additional diagnostic yield of long-read sequencing in undiagnosed rare diseases. Genome Res. 2025;35:559–71. https://doi.org/10.1101/gr.279970.124
ArticleÂ
CASÂ
PubMedÂ
PubMed CentralÂ
Google ScholarÂConlin LK, Aref-Eshghi E, McEldrew DA, Luo M, Rajagopalan R. Long-read sequencing for molecular diagnostics in constitutional genetic disorders. Hum Mutat. 2022;43:1531–44. https://doi.org/10.1002/humu.24465
Miller DE, Lee L, Galey M, Kandhaya-Pillai R, Tischkowitz M, Amalnath D, et al. Targeted long-read sequencing identifies missing pathogenic variants in unsolved Werner syndrome cases. J Med Genet. 2022;59:1087–94. https://doi.org/10.1136/jmedgenet-2022-108485
ArticleÂ
CASÂ
PubMedÂ
PubMed CentralÂ
Google ScholarÂMiller DE, Sulovari A, Wang T, Loucks H, Hoekzema K, Munson KM, et al. Targeted long-read sequencing identifies missing disease-causing variation. Am J Hum Genet. 2021;108:1436–49. https://doi.org/10.1016/j.ajhg.2021.06.006
ArticleÂ
CASÂ
PubMedÂ
PubMed CentralÂ
Google ScholarÂSchuy J, Grochowski CM, Carvalho CMB, Lindstrand A. Complex genomic rearrangements: an underestimated cause of rare diseases. Trends Genet. 2022;38:1134–46. https://doi.org/10.1016/j.tig.2022.06.003
ArticleÂ
CASÂ
PubMedÂ
PubMed CentralÂ
Google ScholarÂSteiert TA, Fuss J, Juzenas S, Wittig M, Hoeppner MP, Vollstedt M, et al. High-throughput method for the hybridisation-based targeted enrichment of long genomic fragments for PacBio third-generation sequencing. NAR Genom Bioinform. 2022;4:lqac051. https://doi.org/10.1093/nargab/lqac051
ArticleÂ
CASÂ
PubMedÂ
PubMed CentralÂ
Google ScholarÂRobinson JT, Thorvaldsdottir H, Winckler W, Guttman M, Lander ES, Getz G, et al. Integrative genomics viewer. Nat Biotechnol. 2011;29:24–26. https://doi.org/10.1038/nbt.1754
ArticleÂ
CASÂ
PubMedÂ
PubMed CentralÂ
Google ScholarÂJaganathan K, Kyriazopoulou Panagiotopoulou S, McRae JF, Darbandi SF, Knowles D, Li YI, et al. Predicting splicing from primary sequence with deep learning. Cell. 2019;176:535–548 e524. https://doi.org/10.1016/j.cell.2018.12.015
ArticleÂ
CASÂ
PubMedÂ
Google ScholarÂRichards S, Aziz N, Bale S, Bick D, Das S, Gastier-Foster J, et al. Standards and guidelines for the interpretation of sequence variants: a joint consensus recommendation of the American College of Medical Genetics and Genomics and the Association for Molecular Pathology. Genet Med. 2015;17:405–24. https://doi.org/10.1038/gim.2015.30
Riggs ER, Andersen EF, Cherry AM, Kantarci S, Kearney H, Patel A, et al. Technical standards for the interpretation and reporting of constitutional copy-number variants: a joint consensus recommendation of the American College of Medical Genetics and Genomics (ACMG) and the Clinical Genome Re. Genet Med. 2020;22:245–57. https://doi.org/10.1038/s41436-019-0686-8
Nam DW, Song YK, Kim JH, Lee EK, Park KH, Cha J, et al. Allelic hierarchy for USH2A influences auditory and visual phenotypes in South Korean patients. Sci Rep. 2023;13:20239. https://doi.org/10.1038/s41598-023-47166-w
Van Schil K, Naessens S, Van de Sompele S, Carron M, Aslanidis A, Van Cauwenbergh C, et al. Mapping the genomic landscape of inherited retinal disease genes prioritizes genes prone to coding and noncoding copy-number variations. Genet Med. 2018;20:202–13. https://doi.org/10.1038/gim.2017.97
Huang XF, Mao JY, Huang ZQ, Rao FQ, Cheng FF, Li FF, et al. Genome-wide detection of copy number variations in unsolved inherited retinal disease. Invest Ophthalmol Vis Sci. 2017;58:424–9. https://doi.org/10.1167/iovs.16-20705
Bujakowska KM, Fernandez-Godino R, Place E, Consugar M, Navarro-Gomez D, White J, et al. Copy-number variation is an important contributor to the genetic causality of inherited retinal degenerations. Genet Med. 2017;19:643–51. https://doi.org/10.1038/gim.2016.158
Zampaglione E, Kinde B, Place EM, Navarro-Gomez D, Maher M, Jamshidi F, et al. Copy-number variation contributes 9% of pathogenicity in the inherited retinal degenerations. Genet Med. 2020;22:1079–87. https://doi.org/10.1038/s41436-020-0759-8
ArticleÂ
CASÂ
PubMedÂ
PubMed CentralÂ
Google ScholarÂTabata A, Sheng JS, Ushikai M, Song YZ, Gao HZ, Lu YB, et al. Identification of 13 novel mutations including a retrotransposal insertion in SLC25A13 gene and frequency of 30 mutations found in patients with citrin deficiency. J Hum Genet. 2008;53:534–45. https://doi.org/10.1007/s10038-008-0282-2
ArticleÂ
CASÂ
PubMedÂ
Google ScholarÂSong YZ, Zhang ZH, Lin WX, Zhao XJ, Deng M, Ma YL, et al. SLC25A13 gene analysis in citrin deficiency: sixteen novel mutations in East Asian patients, and the mutation distribution in a large pediatric cohort in China. PLoS One. 2013;8:e74544 https://doi.org/10.1371/journal.pone.0074544
ArticleÂ
CASÂ
PubMedÂ
PubMed CentralÂ
Google ScholarÂLin J, Lin W, Lin Y, Peng W, Zheng Z. Clinical and genetic analysis of 26 Chinese patients with neonatal intrahepatic cholestasis due to citrin deficiency. Clin Chim Acta. 2024;552:117617. https://doi.org/10.1016/j.cca.2023.117617
ArticleÂ
CASÂ
PubMedÂ
Google ScholarÂHayman T, Ovadia S, Krishnan J, Bouckaert M, Panneman DM, English M, et al. Non-coding single-nucleotide and structural variants affecting the EYS putative promoter cause autosomal recessive retinitis pigmentosa. Genet Med. 2025;27:101427. https://doi.org/10.1016/j.gim.2025.101427
ArticleÂ
CASÂ
PubMedÂ
PubMed CentralÂ
Google ScholarÂAlAbdi L, Shamseldin HE, Khouj E, Helaby R, Aljamal B, Alqahtani M, et al. Beyond the exome: utility of long-read whole genome sequencing in exome-negative autosomal recessive diseases. Genome Med. 2023;15:114. https://doi.org/10.1186/s13073-023-01270-8
ArticleÂ
CASÂ
PubMedÂ
PubMed CentralÂ
Google ScholarÂde la Morena-Barrio B, Stephens J, de la Morena-Barrio ME, Stefanucci L, Padilla J, Minano A, et al. Long-Read Sequencing Identifies the First Retrotransposon Insertion and Resolves Structural Variants Causing Antithrombin Deficiency. Thromb Haemost. 2022;122:1369–78. https://doi.org/10.1055/s-0042-1749345
Nakamichi K, Huey J, Sangermano R, Place EM, Bujakowska KM, Marra M et al. Targeted long-read sequencing enriches disease-relevant genomic regions of interest to provide complete Mendelian disease diagnostics. JCI Insight. 2024; 9; https://doi.org/10.1172/jci.insight.183902
Gueuning M, Thun GA, Koller S, Sigurdardottir S, Trost N, Wagner L, et al. Proof-of-principle: nanopore adaptive sampling enables full blood group genome analysis and resolution of hybrid alleles. Blood Adv. 2026;10:864–75. https://doi.org/10.1182/bloodadvances.2025017463
ArticleÂ
CASÂ
PubMedÂ
PubMed CentralÂ
Google ScholarÂ
Acknowledgements
We thank all the patients who have contributed to this study
Funding
This research was funded by the Seoul National University Hospital Kun-hee Lee Child Cancer & Rare Disease Project, Republic of Korea (grant number : 25B-001-0200). Open Access funding enabled and organized by Seoul National University
Author information
Authors and Affiliations
Department of Laboratory Medicine, Seoul National University Hospital, Seoul National University College of Medicine, Seoul, Republic of Korea
Jee-Soo Lee, Kyeong Seon Ryu, Hyesu Lee, Hara Lim, Seonhoo Youn, Hansol Lim, Seung Won Chae, Hobin Sung, Sung Im Cho, Yeseul Kim, Joo Won Jang, Hoyeon Lee & Moon-Woo Seong
Department of Pediatrics, Seoul National University Hospital Child Cancer and Rare Disease Administration, Seoul National University Children’s Hospital, Seoul, Republic of Korea
Jin Sook Lee
Department of Pediatrics, Seoul National University Children’s Hospital, Seoul, Republic of Korea
Jung Min Ko & Jong-Hee Chae
Department of Genomic Medicine, Rare Disease Center, Seoul National University Children’s Hospital and Seoul National University College of Medicine, Seoul, Republic of Korea
Jong-Hee Chae
Authors
- Jee-Soo LeeView author publications
Search author on:PubMed Google Scholar
- Kyeong Seon RyuView author publications
Search author on:PubMed Google Scholar
- Hyesu LeeView author publications
Search author on:PubMed Google Scholar
- Hara LimView author publications
Search author on:PubMed Google Scholar
- Seonhoo YounView author publications
Search author on:PubMed Google Scholar
- Hansol LimView author publications
Search author on:PubMed Google Scholar
- Seung Won ChaeView author publications
Search author on:PubMed Google Scholar
- Hobin SungView author publications
Search author on:PubMed Google Scholar
- Sung Im ChoView author publications
Search author on:PubMed Google Scholar
- Yeseul KimView author publications
Search author on:PubMed Google Scholar
- Joo Won JangView author publications
Search author on:PubMed Google Scholar
- Hoyeon LeeView author publications
Search author on:PubMed Google Scholar
- Jin Sook LeeView author publications
Search author on:PubMed Google Scholar
- Jung Min KoView author publications
Search author on:PubMed Google Scholar
- Jong-Hee ChaeView author publications
Search author on:PubMed Google Scholar
- Moon-Woo SeongView author publications
Search author on:PubMed Google Scholar
Contributions
Study design: MWS and JSL: Sequencing team: HSL, HL, SY, HL, SWC, SH, SIC: Data analysis: YK, JWJ, HL, KSR: Clinical data: JHC, JMK, JSL: Writing the manuscript: JSL: Supervision: MWS
Ethics declarations
Competing interests
The authors declare no competing interests
Ethical approval
The study was approved by the institutional review board of the Seoul National University Hospital (IRB No. H-2301-101-1396). Informed consent was obtained from all patients included in this study
Additional information
Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations
Supplementary information
Table S1. Gene included in the custom target capture-based long-read sequencing (download DOCX )
Table S2. Structural variants identified by custom target capture-based long-read sequencing. (download DOCX )
41431_2026_2197_MOESM3_ESM.docx (download DOCX )
Table S3. Comparison of targeted long-read sequencing and long-read whole-genome sequencing approaches for unsolved Mendelian disease diagnostics
Rights and permissions
Open Access This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by/4.0/.
About this article
Cite this article
Lee, JS., Ryu, K.S., Lee, H. et al. Unraveling missing variants through target capture-based long-read sequencing in autosomal recessive disorders.
Eur J Hum Genet (2026). https://doi.org/10.1038/s41431-026-02197-5
Received:10 March 2026
Revised:22 June 2026
Accepted:10 July 2026
Published:20 July 2026
Version of record:20 July 2026
DOI
:https://doi.org/10.1038/s41431-026-02197-5


