Download PDF
Abstract
Whole-genome sequencing (WGS) projects for rare disease diagnosis typically yield a diagnostic rate of 25–41%, depending on the methods for patient selection and the extent of prior genetic testing. The Scottish Genomes Partnership (SGP) is a collaborative programme using genome sequencing to diagnose rare disease patients with presumed monogenic aetiology in the Scottish NHS. Within SGP, short-read sequencing (SRS) had previously achieved a diagnostic rate of 23% in affected families. To increase diagnostic yield, we applied Oxford Nanopore Technologies (ONT) long-read sequencing (LRS) to a cohort of 24 SGP families (74 individuals) that remained undiagnosed after SRS-based SNV/indel analysis. We also retrospectively reviewed SRS-derived structural variant (SV) calls to assess whether LRS results could have been detected in SRS data. After quality, rarity, panel-based and inheritance-based filtering, 392 candidate SVs were retained across de novo, homozygous recessive, compound heterozygous and X-linked inheritance models, of which 8 overlapped PanelApp diagnostic-grade “green” genes. Pathogenic or likely pathogenic de novo SVs were identified in 3 of 24 families: an AUTS2 inversion, a DLX5/6 locus inversion, and an FN1 deletion. All three SVs were independently confirmed by retrospective SRS re-analysis using SVRare. These findings demonstrate that LRS-based SV analysis, supported by orthogonal SRS re-analysis, can resolve clinically significant SVs in families who remain unsolved after standard rare disease testing.
Subjects
- Data processing
- Medical genomics
Introduction
Rare diseases affect fewer than 1 in 2000 people in a population [12]. Up to 10,000 rare diseases [1] collectively affect 3–6% of the population worldwide [2]. Many individuals remain undiagnosed due to non-specific clinical presentation, resulting in prolonged diagnostic odysseys [3]. Early diagnosis can improve clinical management, guide access to targeted therapies, and reduce the prolonged uncertainty [4]
Genome sequencing in large, rare disease cohorts yields diagnoses in 25–41% of cases [5,6,7,8,9]. Traditional investigations have largely been superseded by WES/WGS [10,11,12,13,14], though most diagnoses arise from SNVs or small indels, with SVs less frequently implicated in diagnoses [681516]
SVs are ≥50 bp genomic alterations including insertions, deletions, duplications, inversions and translocations that can cause monogenic Mendelian and complex diseases [1718]. They have historically been detected by cytogenetic methods and, more recently, inferred from short-read sequencing (SRS) [1920]. However, SRS suffers from mapping ambiguity in low-complexity regions, reducing sensitivity for SV detection [2122]
Long-read sequencing (LRS) technologies, such as Oxford Nanopore Technologies (ONT) and PacBio, generate reads spanning kilobases to megabases [2223]. These technologies can improve mapping across repetitive and difficult-to-map, low complexity regions, enabling detection of large SVs that are often missed by SRS [2124,25,26,27]. Pathogenic SVs may therefore remain undetected in cases analysed only by SRS [28,29,30]. LRS also enables accurate short tandem repeat (STR) expansion detection, which contributes to a range of rare diseases, by spanning the full repeat and enabling direct resolution of repeat size and sequence composition in a single experiment [31]. Despite ongoing challenges, LRS increasingly identifies pathogenic SVs and improves diagnostic yield [283233].
The Scottish Genomes Partnership (SGP) [8] is aligned with the 100,000 Genomes Project [6], with the aim of diagnosing Scottish patients with rare Mendelian disordersP SRS analysis of 999 genomes achieved a 23% diagnostic yield [8]. We applied LRS to 24 previously unsolved families to identify pathogenic SVs and improve diagnostic yield
Subjects and methods
Patient recruitment and sample collection
Twenty-four families (74 individuals) recruited for the SGP LRS study were selected from those that did not receive a genetic diagnosis in the SGP SRS study [8]. Only families for whom samples were available from one or more affected probands and both unaffected parents (primarily trios or quads) were selected, thereby supporting the identification of de novo (autosomal dominant), recessive and X-linked variants. Selected families were prioritised based on a high clinical suspicion of a monogenic aetiology, indicated either by a shared phenotype among multiple family members, or a syndromic presentation suggestive of an underlying monogenic disorder.
All families had already undergone Illumina whole-genome SRS within the SGP SRS study [8]. This initial analysis was done in two stages and focused on the identification of only SNVs/indels, without systematic assessment of SVs as part of the clinical pipeline at that time. This initial panel-based review examined Tier 1 and Tier 2 variants (variants in PanelApp diagnostic-grade “green” genes), followed by an assessment of Tier 3 variants (predicted coding variants outside phenotype-matched panels) for trios. Only families who remained undiagnosed after both analytical stages were included in this LRS SV study and subsequent retrospective SRS SV re-analysis.
The 24 families in the LRS cohort were selected from the SGP SRS cohort. These families had already undergone both panel-based and broader protein-coding exome SNV/indel analysis but remained undiagnosed and were considered suitable for SV-focused analysis. In silico panel-based prioritisation of SVs was used to achieve a clinically practical workflow
DNA sample processing and sequencing
DNA extraction and sample processing were performed as previously described [8]. For inclusion, a minimum of 15 µg of genomic DNA, preferably with a DNA integrity number (DIN) > 8, was targeted from all family members. DNA shearing and library preparation are described in Supplementary Methods. Before sequencing the full cohort of 22 trios and 2 quads (n = 74), a pilot study using 7 samples from 4 patients was used to evaluate DNA quality and size selection and the suitability of clinical samples not originally collected for LRS. Sequencing was carried out at Edinburgh Genomics using ONT ligation kits (SQK-LSK109/114) on PromethION platforms with R9.4.1 and R10.4.1 flow cells (Supplementary Table 1).
SV analysis
SV analysis and filtering
For the pilot study (n = 7), reads from PromethION Beta were base-called and demultiplexed with Guppy v4.0.11+f1071ce. “pass” tagged reads were filtered (<500 bp, quality score <9) and trimmed (40 bp start, 20 bp end) using NanoFilt v2.8.0, then aligned to GRCh38 using minimap2 v2.23 [34]. SVs were called with cuteSV v1.0.13 [35]. In the family-based study (n = 74), 73 samples were initially sequenced on an R9.4.1 flow cell (PromethION Beta). Among these, 28 were additionally sequenced on an R10.4.1 flow cell (PromethION 24) to increase coverage. One additional sample was sequenced only on R10.4.1, giving a total of 29 samples with R10.4.1 data. Base-calling and demultiplexing used Guppy v5.1.13 + b292f4d (PromethION Beta) and Guppy v6.5.7+ca6d6af (PromethION 24). “pass” tagged reads were filtered (<1000 bp, quality score <9) and trimmed (40 bp start, 20 bp end) using NanoFilt v2.8.0, then aligned to GRCh38 using minimap2 v2.23. Multi-sample SV calling was performed across all SGP LRS samples using Sniffles v2.0.6 [36]. Tools were selected based on benchmarking against GIAB reference data (see Supplementary Methods). Filtering and annotation of SVs was carried out as described in the Supplementary Methods, using standard and cohort-specific quality filters to remove artefacts, excluding common variants based on population databases and incorporating clinical and functional annotations to support downstream interpretation. A schematic diagram of the LRS SV calling and filtering workflow is shown in Fig. 1.
Blue-grey boxes represent data generation and annotation steps. Green boxes represent quality and depth filtering steps applied at the cohort level. Purple boxes represent family-based inheritance filtering, clinical prioritisation via clinician-assigned PanelApp panels containing green (diagnostic-grade) genes, and manual IGV review steps. The bold green box indicates the final diagnostic outcome. SV counts (n=) at each stage are shown in blue. Four modes of inheritance were evaluated – de novo, homozygous autosomal recessive, X-linked and compound heterozygous (SV + SNV), reflecting the assumption that all parents were unaffected. Software tools are shown in italics. Three pathogenic or likely pathogenic SVs were retained as candidate diagnoses: in AUTS2, at the DLX5/6 locus, and in FN1.
SV prioritisation
Analysis was initially restricted to clinician-assigned gene panels based on HPO terms. The panels were downloaded from PanelApp [37] with analysis limited to diagnostic-grade “green” genes. SVs were then evaluated under four modes of inheritance (MOI) consistent with unaffected parents using slivar v0.3.1 [38]: de novo autosomal dominant, homozygous autosomal recessive, compound heterozygous (SV + SNV/small indel from SRS data) and X-linked (in male probands). Variants retained after technical review and phenotype/inheritance prioritisation were discussed at MDT meetings. More detailed variant prioritisation rules are provided in Supplementary Methods.
A secondary analysis was carried out to explore variants beyond pre-defined gene panels. Genome-wide putative de novo SVs annotated by VEP [39] with HIGH or MODERATE impact were reviewed without restriction to gene panels and a genome-wide compound heterozygous analysis was done by intersecting SVs with SNVs in OMIM-annotated biallelic disease genes, with filters applied for rarity, quality, VEP impact (HIGH or MODERATE) and phase (in trans)
SRS SV analysis using SVRare
Methods for genome sequencing and data processing of the SGP SRS samples in the 100,000 Genomes Project have been described previously [8]. Briefly, 150 bp paired-end Illumina HiSeqX reads were aligned to GRCh37/38 via iSAAC aligner [40], with SVs called via Manta [41] and Canvas [42]. However, the SV calls were not analysed as part of the original GEL pipeline. The SRS SV calls were only subsequently processed, once LRS had been undertaken, using SVRare, a genome-wide prioritisation framework that aggregates SV and CNV calls from multiple callers to identify rare, clinically significant variants [43]. SVRare filters candidates based on cohort-wide allele frequency (AF) and prioritises those affecting genes linked to participant HPO terms or predefined gene panels. SVRare was not available at the time of the initial SV analysis but was used here to retrospectively collate SVs from SRS data and provide orthogonal confirmation of LRS-derived SVs.
Results
Pilot study: effects of DNA quality and size selection on ONT sequencing and SV detection
To assess the effects of DNA quality, size selection and fragmentation on SV discovery, seven samples from four patients were sequenced on an R9.4.1 flow cell. Most samples achieved 6–11× median depth of coverage with yield correlating with median/mean coverage but not with read count (Table 1). Size selection had little effect on sequence coverage but substantially improved read length profiles, with most reads in size-selected samples exceeding 10 kb and some reaching 80–100 kb. A partially degraded sample, Sample 4 (size-selected), performed well and achieved high coverage, suggesting that size selection can enable successful LRS from stored DNA samples not specifically prepared for LRS (Supplementary Fig. 1A).
Between 10,000 and 25,000 putative SVs were detected across the 7 samples with greater sequencing depth yielding higher SV counts (Supplementary Fig. 1B). SV type distribution was consistent with previous studies [44]
After retaining SVs ≥50 bp, the union set contained ~32,000 SVs, with lengths >80 kb (Supplementary Fig. 1C). Peaks of SVs were observed at around 300 bp and 6000 bp corresponding to detection of Alu and LINE-1 elements, consistent with previous reports [45] (Supplementary Fig. 1C). Overall, this pilot demonstrated that clinical DNA samples, including partially degraded DNA, can be used for reliable SV detection using LRS
Sequencing results of family samples (n = 74)
Per sample read counts and QC statistics for R9.4.1 and R10.4.1 sequencing runs are provided in Supplementary Tables 2–5. Key QC metrics before and after trimming and filtering are shown in Supplementary Fig. 2. Filtering improved median read length, quality and N50, and R10.4.1 flow cells showed higher and more consistent read quality than R9.4.1 (Supplementary Fig. 2A–D)
Read alignment to GRCh38 resulted in a mean depth of coverage per sample ranging from 10× to 57×, with >99% mapping rate and 91.5–97.2% identity (Supplementary Table 6, Supplementary Fig. 3A). Coverage from R10.4.1 flow cells was noticeably higher than from R9.4.1 flow cells (Supplementary Fig. 3B)
SV-calling statistics in SGP LRS families
Multi-sample SV-calling identified 74,388 putative SVs across 74 individuals, representing all variants relative to the GRCh38 reference genome prior to filtering. To minimise artefactual calls from low or unusually high depth, the distribution of site mean read depth (mean depth per SV site averaged across individuals) was examined (Supplementary Fig. 4). A depth filter of 10–56× was applied based on this empirical distribution, removing SVs in low-coverage regions or those with inflated depth likely reflecting alignment artefacts or repetitive sequence. The upper threshold (56×) corresponds to approximately twice the mean site-level depth. After applying additional quality filters (see Methods and Supplementary Methods), 60,022 high-confidence SVs were retained across chromosomes 1–22, X and Y. Most SVs were insertions (52%) and deletions (~44%), followed by translocations (~3%), duplications (0.3%) and inversions (0.3%) (Fig. 2A). Most SVs (61.9%) were unique to the SGP LRS cohort (Max_AF = 0), with no matches in SVAFotate databases (Supplementary Fig. 5). Most SVs were either unique to the cohort or rare, as reflected in the absolute counts across frequency categories (Fig. 2B). The percentage distribution of common, less frequent, rare and unique SVs was similar across SV types, except for duplications, where rare SVs were more abundant (Supplementary Fig. 6). SV counts per family did not show any major outliers (Fig. 2C, Supplementary Table 7).
A Barplot showing the number of filtered SVs stratified by SV type. B Histogram showing the number of filtered SVs that are common (Max_AF or maximum allele frequency ≥ 0.05), less frequent (Max_AF < 0.05 and ≥ 0.01), rare (Max_AF < 0.01 and Max_AF ≠ 0) and unique to the SGP LRS cohort (Max_AF = 0). C Box and jitter plot showing the distribution of the number of filtered INSs, DELs, DUPs, INVs and TRAs per family (n = 24)
SV prioritisation and analytical strategy
An SV prioritisation pipeline was devised to reduce the calls to a manageable number of candidate SVs for manual review (Fig. 1). The cohort-level SV call set was converted to family-specific datasets, with each family retaining 23,024–25,009 genome-wide SVs (Supplementary Table 7). After quality filtering, and functional annotation with VEP, SVs were prioritised within clinician-assigned, family-specific PanelApp gene panels. Rarity and inheritance-based filtering yielded 392 candidate SVs across all families (Table 2, columns D-K, summed across families). This primary analysis reflects a clinically driven, panel-based prioritisation strategy.
Within this set, the number of genome-wide de novo SVs identified under the de novo inheritance model was higher than expected (up to 46 in a single family; Table 2, “De Novo: Genome-wide” column). Each generation is expected to have at most one true de novo SV [18]. This inflation is likely due to missed detection of variants in parents at lower coverage, leading to inherited variants being misclassified as de novo. Families sequenced using both R9.4.1 and R10.4.1 flow cells showed fewer such calls than those sequenced only on R9.4.1 (Supplementary Fig. 7). These calls were reduced through downstream filtering and manual curation.
Of the 392 candidate SVs, 8 SVs overlapped diagnostic-grade “green” genes in PanelApp gene panels before manual curation, comprising 4 de novo and 4 X-linked candidates. These candidates were reviewed in IGV [46] to exclude artefactual SVs. Following this review, one de novo SV was excluded, while the remaining three de novo SVs, corresponding to the three exemplar families described below, were prioritised for further evaluation. These were then discussed with referring clinicians at MDT meetings to confirm genotype-phenotype concordance. All X-linked candidates were excluded because of inconsistent segregation patterns or insufficient evidence for functional or regulatory relevance.
Genome-wide examination of X-linked SVs did not identify additional candidate diagnoses, as most variants were intergenic or intronic, or showed phenotype mismatch, although long-range regulatory effects cannot be excluded
No SVs overlapping biallelic panel genes with autosomal homozygous recessive MOI were identified in this cohort. Additionally, no compound heterozygous scenarios (SV+SNVs) were found in biallelic panel genes for any family
A secondary genome-wide analysis, independent of PanelApp panels, was also undertaken using a distinct filtering strategy focused on de novo SVs with predicted functional impact. In the first phase, genome-wide putative de novo SVs annotated by VEP as HIGH or MODERATE impact were investigated. This yielded 66 unique variants across 20 families (Supplementary Table 8), which were manually reviewed in IGV. Many SVs occurred in low-complexity or repetitive regions and showed alignment patterns suggestive of artefacts. Several putative de novo SVs were reclassified as inherited or present in all three family members, indicating false-positive calls. A subset of events affected lncRNAs, non-coding transcripts, or pseudogenes, limiting clinical relevance. Manual review substantially reduced the candidate set to 10 variants, labelled as “Retained” in Supplementary Table 8. However, none of the retained variants overlapped OMIM genes with established disease associations.
In the secondary genome-wide analysis, SV + SNV combinations in OMIM-annotated biallelic disease genes were evaluated, but all were excluded after curation because in each case at least one variant was benign (Supplementary Table 9)
A compilation of genome-wide de novo SVs and compound heterozygous instances (SV+SNVs) is shown in Supplementary Tables 10 and 11, respectively. These are provided for future gene discovery and investigation of SV regulatory effects
Families with SVs in PanelApp genes
Pathogenic or likely pathogenic SVs were identified in three exemplar families. IGV plots generated from the LRS data are shown in Supplementary Fig. 8. SVs in Families 1 and 2 were accepted as pathogenic at MDT review meetings while Family 3 SV was designated as likely pathogenic
Family 1 (recruited disease: intellectual disability)
The proband is a male and the firstborn child of nonconsanguineous and unaffected parents. He was referred to Clinical Genetics at the age of 6.9 years for an undiagnosed neurodevelopmental disorder and recent onset of seizures. His medical issues included unclassified epilepsy, developmental delay (nonverbal), microcephaly (−3.5 s.d.), sleep disturbance, skeletal abnormalities (finger contractures, bilateral calcaneal valgus with foot pain) and convergent squint. Examination showed microcephaly with flattened occiput, convergent squint, mid-face hypoplasia, micrognathia, prominent nose/flattened nasal bridge, short philtrum, prominent ears, thickened lips and joint contractures.
Brain MRI (age 8), electroencephalogram (EEG) (age 9) and metabolic screening were normal. Chromosome microarray and exome sequencing through Deciphering Developmental Disorders [15] (DDD) study were non-diagnostic. He was subsequently recruited to SGP, where SRS analysis did not identify any Tier 1, 2, or 3 variants
LRS detected a 1.4 Mb de novo inversion disrupting AUTS2 (MIM: 615834), a candidate gene for neurodevelopmental and neurological disorders encoding the Activator of Transcription and Developmental Regulator protein [47], which mediates neurogenesis, neuronal maturation, synapse formation and dendritic spine regulation [48]. One inversion breakpoint was located within intron 2 and was predicted to have a high functional impact on the protein-coding sequence, removing exons 1 and 2. This disruption is anticipated to impair gene expression and is consistent with the proband’s phenotype. This inversion was also detectable on retrospective analysis of the SRS data and was prioritised using SVRare, providing orthogonal support for the event [49]. A samplot [50] image of the LRS inversion is shown in Fig. 3A and an IGV plot is shown in Supplementary Fig. 8A. The SV was classified as pathogenic according to ACMG guidelines.
A A 1.4 Mb de novo inversion found from LRS data in the proband from exemplar family 1 that overlaps with AUTS2 and is absent in the parents. The inversion is denoted by a purple dotted line in the proband. Other smaller SVs can be ignored. The x-axis refers to the chromosome base position. The left y-axis refers to the SV length and the right y-axis represents the depth of coverage. B A 59.8 kb de novo deletion found from LRS data in the proband from exemplar family 3. The multi-exon deletion is found within FN1, with exons 6 to 41 of the MANE transcript deleted.
Family 2 (recruited disease: ultra-rare undescribed monogenic disorders)
The proband is a female with three-limb ectrodactyly (both hands and right foot), retro-micrognathia with dental overcrowding and overfolded ear helices. Audiometry revealed narrow ear canals and mild mixed hearing loss. Resting sinus bradycardia was noted with normal cardiac structure. Ectrodactyly was observed antenatally in the proband’s daughter. Additionally, patent ductus arteriosus, atrial septal defect and pulmonary valve stenosis were noted at birth. The proband’s parents are unaffected.
For this family, samples from the unaffected parents and proband underwent LRS. SRS was available for the trio and, additionally, the proband’s affected daughter. Trio LRS identified a 5.2 Mb de novo inversion in the split hand/foot malformation type 1 (SHFM1, MIM: 183600) region, previously reported in association with limb malformation (Fig. 4B) [51,52,53,54]. This event was prioritised because the panel gene COL1A2 was present within the inversion, although neither breakpoint directly disrupted the gene. Closer inspection of the region showed that DLX5 and DLX6, important genes in the WNT pathway involved in limb development [53], were located close to the distal breakpoint. DYNC1I1, located at one end of the inversion, contains enhancer elements that regulate DLX5/6 and the inversion is predicted to reposition these enhancers away from their target genes. Disruption of these enhancer-gene interactions is likely to reduce DLX5/6 expression, consistent with the split hand/foot phenotype [51,52,53,54]. SRS data from the proband’s affected daughter, who also has ectrodactyly, was found to carry the inversion when analysed retrospectively using SVRare (Fig. 4A, C), providing orthogonal confirmation of the inversion. IGV visualisation of the LRS data is shown in Supplementary Fig. 8B. The SV was classified as pathogenic according to ACMG guidelines.
A Pedigree diagram showing the three generations of exemplar family 2. Sequence was unavailable for family members II.2 and III.2. SRS and LRS labels along with the symbols “+” or “-” indicate the availability and unavailability of the type of sequencing data, respectively. The proband (II.1) carries a 5.2 Mb heterozygous de novo inversion (denoted as “dn”) near the DLX5/6 locus, which was inherited by her daughter (III.1). B The UCSC session displays the SHFM1 locus (7q21) with previously reported SVs in various publications, alongside the inversion identified in the SGP LRS patient reads. C IGV plot showing soft-clipped bases near the inversion breakpoint in the proband and her daughter, which are absent in the parents of the proband. The IGV plot was generated using SRS data available in the Genomics England (GEL) research environment and includes an additional family member, the daughter of the proband, sequenced by SRS.
Family 3 (recruited disease: unexplained skeletal dysplasia)
The proband is a male and the only child of nonconsanguineous healthy Scottish parents. He presented at age 8 with abnormal gait and bilateral hip pain (onset at 6–7.5 years) and was regarded as bilateral Perthes-like disease. Progressive contractures developed in the hips from age 12 and elbows from 14, along with pain and swelling of interphalangeal joints from age 14, requiring intra-articular injections
On examination, there was reduced hip movement and significant girdle weakness with normal muscle power and reflexes. He walked with thoracic kyphosis, lumbar lordosis, flexed knees and stiff hips. His height, proportion, vision, hearing and teeth were normal and he was non-dysmorphic. A skeletal survey (age 11.6 years) showed slightly ovoid vertebral bodies. Bone age was appropriate. Pelvic MRI revealed pelvic muscle atrophy without evidence of myopathy. Hip MRI revealed fragmented, collapsed, irregular proximal femoral epiphyses. Further review revealed bilateral symmetrical irregular sclerosis of the femoral epiphyses, mild metaphyseal broadening with minor cystic changes and subchondral fracturing and collapse, in addition to bilateral hip joint effusions and lumbar spine end-plate concavities. Muscle biopsy with ultrastructural analysis was unhelpful and serum creatine kinase, nerve conduction, electromyograms and lumbar punctures were normal. A diagnosis of multiple epiphyseal dysplasia was considered.
Sanger DNA sequencing of 5 MED-related genes (COMP, COL9A1, COL9A2, COL9A3 and MATN3), in addition to LMNA, had been performed, but no pathogenic variant was identified. Chromosomal microarray revealed a de novo heterozygous deletion (~55 kb) at 2q35 in the FN1 gene (MIM: 184255), involving multiple exons which was classified as a VUS. Additionally, the SGP SRS study utilised gene panels for various skeletal and neuromuscular disorders, but initial analysis found no causative variants. The deletion was likely present in the SRS data but was not prioritised due to the absence of systematic SV review in the clinical pipeline at that time.
LRS identified a 59.8 kb de novo multi-exon in-frame deletion in FN1, spanning exons 6–41 (out of a total of 46 exons in the MANE Select transcript) (Fig. 3B, Supplementary Fig. 8C), confirming the chromosomal microarray result and precisely defining the deletion breakpoints that the array could not provide. IGV visualisation showed a clear drop in the depth of coverage in the proband compared with both parents. The deletion was also identified in retrospective SVRare analysis of the SRS data. In the musculoskeletal system, FN1 encodes an extracellular-matrix multimeric glycoprotein involved in cell adhesion and migration [55]. The FN1 deletion was initially mis-genotyped as homozygous by the LRS SV caller, leading to its exclusion under a strict cohort AF threshold of 0.01. Visual inspection of read alignments clarified the heterozygous state, highlighting the sensitivity of variant prioritisation to genotyping errors and AF thresholds. Using a more lenient threshold of 0.015 retained the variant. The FN1 SV was reclassified as likely pathogenic according to ACMG guidelines and additional work is ongoing to establish a definitive connection between FN1 and bilateral Perthes-like disease.
A comparison of SV detection across the original SGP SRS analysis, LRS analysis and retrospective SRS re-analysis is summarised in Table 3
Discussion
LRS, supported by orthogonal SRS-based SV analysis, enabled a diagnosis or likely diagnosis in 3 out of 24 families that had remained undiagnosed after prior extensive genetic testing. This was achieved using a panel-based strategy directly comparable to a standard clinical workflow, integrating automated SV calling with manual IGV review and MDT discussion across 74 individuals from 24 families
Our pilot study examined how DNA quality and fragment size affect sequencing performance and SV discovery. Size-selection of partially degraded DNA produced comparable coverage and SV call profiles to those of high-quality samples and, although it increased the proportion of reads >10 kb, it did not significantly alter the SV calling rate. ONT protocols using R10.4.1 flow cells did not require DNA shearing at the time of this study, which is now recommended. A major difference was observed in sequence read depth and yield between the R9.4.1 and R10.4.1 flow cells, providing increased depth of coverage and more accurate SV calling and genotyping.
Distinguishing true pathogenic SVs from artefact and common variation was central to our analytical strategy. We first applied a combination of standard and cohort-specific quality filters (see Methods and Supplementary Methods) on genome-wide SV calls to remove artefacts and then utilised SV population databases to remove common SVs, while VEP, ClinVar and OMIM annotations were added to support downstream interpretation. SVs were prioritised using clinician-assigned PanelApp gene panels, and evaluated under de novo, recessive, compound heterozygous and X-linked models. Candidates were cross-validated with SVRare, reviewed in IGV and discussed at MDT meetings to confirm genotype-phenotype concordance.
These analytical decisions led to two pathogenic variant findings (AUTS2 and DLX5/6 locus) and one likely pathogenic finding (FN1). AUTS2 is well established in a range of neurodevelopmental disorders, dysmorphic features and skeletal abnormalities, supporting its role in the phenotype of the Family 1 proband [48]. The novel DLX5/6 inversion linked to ectrodactyly in Family 2 builds on literature defining critical regulatory regions for limb development, supporting a positional effect through disruption of enhancer-gene interactions at the DLX5/6 locus [51,52,53,54]. For Family 3, the FN1 deletion was initially rejected as radiological review did not show corner fractures typically associated with FN1 mutations but was classified as likely pathogenic following MDT review given the close phenotypic match. Although most reported FN1 variants are missense or splice-site changes, this de novo multi-exon deletion is predicted to cause loss-of-function. FN1 is highly constrained against such variation (gnomAD v4.1.1 pLI = 1), supporting haploinsufficiency. No additional tiered variants in FN1 were identified and further investigations are ongoing to clarify its involvement in the proband’s bilateral Perthes features.
While our primary analysis focused on coding SVs within panel genes to facilitate clinical interpretation, we also implemented a secondary genome-wide analysis of autosomal de novo SVs with predicted functional impact. However, this did not yield additional clinically relevant findings. Many SVs were filtered out following manual review, largely due to artefactual signals or localisation to repetitive or non-coding regions, highlighting the importance of stringent filtering and phenotype-driven prioritisation. In three examples, LRS provided improved alignment in repeat regions that were poorly resolved by SRS (Supplementary Table 8), highlighting the inherent technical limitations of SRS.
We acknowledge that this study has some limitations. Coverage was uneven across the cohort. Samples sequenced only on the R9.4.1 flow cell had lower depth and higher rates of genotyping error for de novo SVs and updating these to R10.4.1 could uncover additional diagnoses. Further analysis of genes outside the panels may yield additional diagnoses. We have catalogued the genome-wide de novo SVs and compound heterozygous (SV + SNV) events to support future gene discovery and studies of SV regulatory effects (Supplementary Tables 10 and 11, respectively). Our primary analysis was restricted to clinician-assigned PanelApp panels available at the time. Newer updates may include recently identified disease genes [37] and could yield additional diagnoses. We did not perform a genome-wide analysis of runs of homozygosity (ROH) or absence of heterozygosity (AOH), which identify regions of shared parental ancestry where recessive disease-causing variants are more likely to reside.
Finally, advances in SV calling algorithms, including new deep learning-based methods such as Cue [56], which leverage image representations of read alignments to improve SV detection and genotyping, may enable identification of additional pathogenic variants. However, many such methods do not yet support multi-sample or family-based analyses, which were essential for this study
The exemplar cases described in this cohort demonstrate the potential of LRS identify SVs and uncover clinically relevant variants in cases that remain unsolved after standard analysis. Further studies are ongoing to establish the clinical relevance of genome-wide de novo SVs for selected families in this cohort
Several studies have shown that LRS has resolved previously unsolved cases after SRS failed [2833]. The SGP LRS cohort comprised 24 families deliberately selected from the original SGP study [8] that remained unsolved after extensive prior genetic testing, and LRS enabled a diagnosis or likely diagnosis in 3 of these families, with pathogenic or likely pathogenic SVs identified in cases where prior microarray, exome and genome-based SNV analysis had been non-diagnostic: an AUTS2 inversion, an FN1 deletion and a DLX5/6 locus inversion. The ability to comprehensively identify SVs through LRS, in tandem with improving SV calling tools, creates new opportunities for molecular diagnoses in rare disease. Further analyses remain possible, but the exemplars demonstrate the added diagnostic value of LRS in rare disease diagnosis.
Data availability
Data from the National Genomic Research Library (NGRL) used in this research are available within the secure Genomics England Research Environment. Access to NGRL data is restricted to adhere to consent requirements and protect participant privacy. Data used in this research include: ONT data: Aligned BAM files for 73 samples, family-specific SV VCFs for 24 families and a multi-sample SV VCF containing all 73 samples are available within the Genomics England Research Environment (GEL RE). One participant from the original cohort of 74 was removed for ethical reasons. File paths for single-sample and family-level BAM and VCF files are provided in the LabKey table “rare_disease_ont_sgp” within the project main-programme_v19_2024-10-31. The multi-sample SV VCF file can be found in the folder: /gel_data_resources/LRS_cohort_genomes/Rare_Disease/SGP2_LRS/Aggregate_VCF/. Further details about the dataset are available at: https://re-docs.genomicsengland.co.uk/rd_sgp_ont/. Illumina data: SRS data for participants was accessed through the LabKey table “rare_disease_analysis” within the project main-programme_v19_2024-10-31. Access to NGRL data is provided to approved researchers who are members of the Genomics England Research Network, subject to institutional access agreements and research project approval under participant-led governance. For more information on data access, visit: https://www.genomicsengland.co.uk/research. GRCh37 human reference genome and index: https://ftp-trace.ncbi.nlm.nih.gov/ReferenceSamples/giab/release/references/GRCh37/hs37d5.fa.gz. https://ftp-trace.ncbi.nlm.nih.gov/ReferenceSamples/giab/release/references/GRCh37/hs37d5.fa.gz.fai. GRCh38 human reference genome and index: ftp://ftp.ncbi.nlm.nih.gov/genomes/all/GCA/000/001/405/GCA_000001405.15_GRCh38/seqs_for_alignment_pipelines.ucsc_ids/GCA_000001405.15_GRCh38_no_alt_plus_hs38d1_analysis_set.fna.gz. ftp://ftp.ncbi.nlm.nih.gov/genomes/all/GCA/000/001/405/GCA_000001405.15_GRCh38/seqs_for_alignment_pipelines.ucsc_ids/GCA_000001405.15_GRCh38_no_alt_plus_hs38d1_analysis_set.fna.fai. PanelApp website: https://panelapp.genomicsengland.co.uk/.
Code availability
The codes and custom scripts used in this manuscript will be made available upon request
References
Haendel M, Vasilevsky N, Unni D, Bologa C, Harris N, Rehm H, et al. How many rare diseases are there? Nat Rev Drug Discov. 2020;19:77–8
Nguengang Wakap S, Lambert DM, Olry A, Rodwell C, Gueydan C, Lanneau V, et al. Estimating cumulative point prevalence of rare diseases: analysis of the Orphanet database. Eur J Hum Genet. 2020;28:165–73
Marwaha S, Knowles JW, Ashley EA. A guide for the diagnosis of rare and undiagnosed disease: beyond the exome. Genome Med. 2022;14:23
Bauskis A, Strange C, Molster C, Fisher C. The diagnostic odyssey: insights from parents of children living with an undiagnosed condition. Orphanet J Rare Dis. 2022;17:233
Wojcik MH, Lemire G, Berger E, Zaki MS, Wissmann M, Win W, et al. Genome sequencing for diagnosing rare diseases. N Engl J Med. 2024;390:1985–97
Smedley D, Investigators GPP, Smith KR, Martin A, Thomas EA, McDonagh EM, et al. 100,000 genomes pilot on rare-disease diagnosis in health care – preliminary report. N Engl J Med. 2021;385:1868–80
Wright CF, Campbell P, Eberhardt RY, Aitken S, Perrett D, Brent S, et al. Genomic diagnosis of rare pediatric disease in the United Kingdom and Ireland. N Engl J Med. 2023;388:1559–71
Hocking LJ, Andrews C, Armstrong C, Ansari M, Baty D, Berg J, et al. Genome sequencing with gene panel-based analysis for rare inherited conditions in a publicly funded healthcare system: implications for future testing. Eur J Hum Genet. 2022;31:231–8
Pagnamenta AT, Camps C, Giacopuzzi E, Taylor JM, Hashim M, Calpena E, et al. Structural and non-coding variants increase the diagnostic yield of clinical whole genome sequencing for rare diseases. Genome Med. 2023;15:94
Franceschini N, Frick A, Kopp JB. Genetic testing in clinical settings. Am J Kidney Dis. 2018;72:569–81
Cheung SW, Shaw CA, Yu W, Li J, Ou Z, Patel A, et al. Development and validation of a CGH microarray for clinical cytogenetic diagnosis. Genet Med. 2005;7:422–32
Wang TL, Maierhofer C, Speicher MR, Lengauer C, Vogelstein B, Kinzler KW, et al. Digital karyotyping. Proc Natl Acad Sci USA. 2002;99:16156–61
Yang Y, Muzny DM, Reid JG, Bainbridge MN, Willis A, Ward PA, et al. Clinical whole-exome sequencing for the diagnosis of mendelian disorders. N Engl J Med. 2013;369:1502–11
Lionel AC, Costain G, Monfared N, Walker S, Reuter MS, Hosseini SM, et al. Improved diagnostic yield compared with targeted gene sequencing panels suggests a role for whole-genome sequencing as a first-tier genetic test. Genet Med. 2018;20:435–43
Wright CF, Fitzgerald TW, Jones WD, Clayton S, McRae JF, van Kogelenberg M, et al. Genetic diagnosis of developmental disorders in the DDD study: a scalable analysis of genome-wide research data. Lancet. 2015;385:1305–14
Lai G, Gu Q, Lai Z, Chen H, Chen J, Huang J. The application of whole-exome sequencing in the early diagnosis of rare genetic diseases in children: a study from Southeastern China. Front Pediatr. 2024;12:1448895
Stankiewicz P, Lupski JR. Structural variation in the human genome and its role in disease. Annu Rev Med. 2010;61:437–55
Collins RL, Brand H, Karczewski KJ, Zhao X, Alfoldi J, Francioli LC, et al. A structural variation reference for medical and population genetics. Nature. 2020;581:444–51
Kosugi S, Momozawa Y, Liu X, Terao C, Kubo M, Kamatani Y. Comprehensive evaluation of structural variation detection algorithms for whole genome sequencing. Genome Biol. 2019;20:117
Ho SS, Urban AE, Mills RE. Structural variation in the sequencing era. Nat Rev Genet. 2020;21:171–89
Sedlazeck FJ, Lee H, Darby CA, Schatz MC. Piercing the dark matter: bioinformatics of long-range sequencing and mapping. Nat Rev Genet. 2018;19:329–46
Warburton PE, Sebra RP. Long-read DNA sequencing: recent advances and remaining challenges. Annu Rev Genom Hum Genet. 2023;24:109–32
From kilobases to “whales”: a short history of ultra-long reads and high-throughput genome sequencing: Oxford Nanopore Technologies. 2021. https://nanoporetech.com/about-us/news/blog-kilobases-whales-short-history-ultra-long-reads-and-high-throughput-genome
De Coster W, Van Broeckhoven C. Newest methods for detecting structural variations. Trends Biotechnol. 2019;37:973–82
Shafin K, Pesout T, Chang PC, Nattestad M, Kolesnikov A, Goel S, et al. Haplotype-aware variant calling with PEPPER-Margin-DeepVariant enables high accuracy in nanopore long-reads. Nat Methods. 2021;18:1322–32
Wagner J, Olson ND, Harris L, McDaniel J, Cheng H, Fungtammasan A, et al. Curated variation benchmarks for challenging medically relevant autosomal genes. Nat Biotechnol. 2022;40:672–80
Delahaye C, Nicolas J. Sequencing DNA with nanopores: troubles and biases. PLoS One. 2021;16:e0257521
Miller DE, Sulovari A, Wang T, Loucks H, Hoekzema K, Munson KM, et al. Targeted long-read sequencing identifies missing disease-causing variation. Am J Hum Genet. 2021;108:1436–49
Eichler EE. Genetic variation, comparative genomics, and the diagnosis of disease. N Engl J Med. 2019;381:64–74
Abel HJ, Larson DE, Regier AA, Chiang C, Das I, Kanchi KL, et al. Mapping and characterization of structural variation in 17,795 human genomes. Nature. 2020;583:83–9
Del Gobbo GF, Boycott KM. The additional diagnostic yield of long-read sequencing in undiagnosed rare diseases. Genome Res. 2025;35:559–71
Steyaert W, Sagath L, Demidov G, Yepez VA, Esteve-Codina A, Gagneur J, et al. Unraveling undiagnosed rare disease cases by HiFi long-read genome sequencing. Genome Res. 2025;35:755–68
Sinha S, Rabea F, Ramaswamy S, Chekroun I, El Naofal M, Jain R, et al. Long read sequencing enhances pathogenic and novel variation discovery in patients with rare diseases. Nat Commun. 2025;16:2500
Li H. Minimap2: pairwise alignment for nucleotide sequences. Bioinformatics. 2018;34:3094–100
Jiang T, Liu Y, Jiang Y, Li J, Gao Y, Cui Z, et al. Long-read-based human genomic structural variation detection with cuteSV. Genome Biol. 2020;21:189
Smolka M, Paulin LF, Grochowski CM, Horner DW, Mahmoud M, Behera S, et al. Detection of mosaic and population-level structural variants with Sniffles2. Nat Biotechnol. 2024;42:1571–80
Martin AR, Williams E, Foulger RE, Leigh S, Daugherty LC, Niblock O, et al. PanelApp crowdnels. Nat Genet. 2019;51:1560–5
Pedersen BS, Brown JM, Dashnow H, Wallace AD, Velinder M, Tristani-Firouzi M, et al. Effective variant filtering and expected candidate variant yield in studies of rare human disease. NPJ Genom Med. 2021;6:60
McLaren W, Gil L, Hunt SE, Riat HS, Ritchie GR, Thormann A, et al. The Ensembl Variant Effect Predictor. Genome Biol. 2016;17:122
Raczy C, Petrovski R, Saunders CT, Chorny I, Kruglyak S, Margulies EH, et al. Isaac: ultra-fast whole-genome secondary analysis on Illumina sequencing platforms. Bioinformatics. 2013;29:2041–3
Chen X, Schulz-Trieglaff O, Shaw R, Barnes B, Schlesinger F, Kallberg M, et al. Manta: rapid detection of structural variants and indels for germline and cancer sequencing applications. Bioinformatics. 2016;32:1220–2
Roller E, Ivakhno S, Lee S, Royce T, Tanner S. Canvas: versatile and scalable detection of copy number variants. Bioinformatics. 2016;32:2375–7
Yu J, Szabo A, Pagnamenta AT, Shalaby A, Giacopuzzi E, Taylor J, et al. SVRare: discovering disease-causing structural variants in the 100K Genomes Project. medRxiv. 2022. https://www.medrxiv.org/content/10.1101/2021.10.15.21265069v1
Beyter D, Ingimundardottir H, Oddsson A, Eggertsson HP, Bjornsson E, Jonsson H, et al. Long-read sequencing of 3,622 Icelanders provides insight into the role of structural variants in human diseases and other traits. Nat Genet. 2021;53:779–86
Zook JM, Hansen NF, Olson ND, Chapman L, Mullikin JC, Xiao C, et al. A robust benchmark for detection of germline large deletions and insertions. Nat Biotechnol. 2020;38:1347–55
Robinson JT, Thorvaldsdottir H, Winckler W, Guttman M, Lander ES, Getz G, et al. Integrative genomics viewer. Nat Biotechnol. 2011;29:24–6
Hori K, Shimaoka K, Hoshino M. AUTS2 Gene: Keys to Understanding the Pathogenesis of Neurodevelopmental Disorders. Cells. 2011;11:11
Loberti L, Adamo L, Antolini E, Casamassima G, Destree A, Brunetti-Pierri N, et al. AUTS2-related syndrome: insights from a large European cohort. Genet Med. 2025;27:101375
Pagnamenta AT, Yu J, Walker S, Noble AJ, Lord J, Dutta P, et al. The impact of inversions across 33,924 families with rare disease from a national genome sequencing project. Am J Hum Genet. 2024;111:1140–64
Belyeu JR, Chowdhury M, Brown J, Pedersen BS, Cormier MJ, Quinlan AR, et al. Samplot: a platform for structural variant visual validation and automated filtering. Genome Biol. 2021;22:161
Lango Allen H, Caswell R, Xie W, Xu X, Wragg C, Turnpenny PD, et al. Next generation sequencing of chromosomal rearrangements in patients with split-hand/split-foot malformation provides evidence for DYNC1I1 exonic enhancers of DLX5/6 expression in humans. J Med Genet. 2014;51:264–7
Ramos-Zaldivar HM, Martinez-Irias DG, Espinoza-Moreno NA, Napky-Rajo JS, Bueso-Aguilar TA, Reyes-Perdomo KG, et al. A novel description of a syndrome consisting of 7q21.3 deletion including DYNC1I1 with preserved DLX5/6 without ectrodactyly: a case report. J Med Case Rep. 2016;10:156
Velinov M, Ahmad A, Brown-Kipphut B, Shafiq M, Blau J, Cooma R, et al. A 0.7 Mb de novo duplication at 7q21.3 including the genes DLX5 and DLX6 in a patient with split-hand/split-foot malformation. Am J Med Genet Part A. 2012;158A:3201–6
Rasmussen MB, Kreiborg S, Jensen P, Bak M, Mang Y, Lodahl M, et al. Phenotypic subregions within the split-hand/foot malformation 1 locus. Hum Genet. 2016;135:345–57
Yang C, Wang C, Zhou J, Liang Q, He F, Li F, et al. Fibronectin 1 activates WNT/beta-catenin signaling to induce osteogenic differentiation
Popic V, Rohlicek C, Cunial F, Hajirasouliha I, Meleshko D, Garimella K, et al. Cue: a deep-learning framework for structural variant discovery and genotyping. Nat Methods. 2023;20:559–68
Acknowledgements
This study would not have been possible without families, patients, clinicians, nurses, research scientists, laboratory staff, informaticians and the wider Scottish Genomes Partnership team, to whom we extend gratitude. We thank Asier Gonzalez, Olga Leonova and Elliot Gould from GEL for their inputs and support with data uploaded to the NGRL; Ailin Buzzi from GEL for inputs on data governance and ethics; and Adam Giess (GEL) for early technical discussions on data analysis. We thank Oxford Nanopore Technologies, UK (ONT) for support with consumables and Vania Costa (ONT) for support with DNA extraction and library preparation for R10.4.1 samples. We also thank Megan Baxter for her input for future studies regarding FN1 deletion. This work was supported by the Edinburgh International Data Facility (EIDF) and the Data-Driven Innovation Programme at the University of Edinburgh. We gratefully acknowledge the participants of the National Genomic Research Library (NGRL), whose contributions made this research possible. Secure access to the NGRL under project ID: RR566 was provided by Genomics England, which delivers the NGRL in partnership with NHS England and is wholly owned by the UK Department of Health and Social Care. The NGRL contains participants’ health data collected by the NHS as part of their care, along with samples and data from their participation in research, for which fully informed consent has been obtained. This includes genomic and clinical data provided through the NHS Genomic Medicine Service, as well as data obtained through research studies, including the 100,000 Genomes Project and the Generation Study, both of which are delivered in partnership with the NHS and from other research cohorts involving external collaborators. The views expressed are those of the authors and not necessarily those of the NHS, the NIHR or the Department of Health or Wellcome Trust.
Funding
The Scottish Genomes Partnership was funded by the Chief Scientist Office of the Scottish Government Health Directorates (SGP/1 and SGP/2) and the Medical Research Council Whole Genome Sequencing for Health and Wealth Initiative (MC/PC/15080). PD was additionally supported by funding awarded to JCT. (Grant Ref: MR/W01761X/1) and by the National Institute of Health Research (NIHR) Oxford Biomedical Research Centre (BRC)
Author information
Author notes
Alistair T. Pagnamenta
Present address: Institute of Biomedical and Clinical Science, University of Exeter Medical School, Royal Devon University Healthcare NHS Foundation Trust, Exeter, Devon, UK
Jing Yu
Present address: Sequoia Genetics, I-HUB, 84 Wood Ln, London, UK
These authors contributed equally: Jenny C. Taylor, Timothy J. Aitman
Authors and Affiliations
Centre for Genomic and Experimental Medicine, MRC Institute of Genetics and Cancer, University of Edinburgh, Edinburgh, UK
Prasun Dutta, Christelle Robert & Timothy J. Aitman
Centre for Human Genetics and NIHR Oxford Biomedical Research Centre, University of Oxford, Oxford, Oxfordshire, UK
Prasun Dutta, Alistair T. Pagnamenta, Anthony E. F. McGuigan, Jing Yu & Jenny C. Taylor
North of Scotland Medical Genetic Service, NHS Grampian, Polwarth Building, Foresterhill, Aberdeen, UK
Alison Ross & Zosia Miedzybrodzka
West of Scotland Centre for Genomic Medicine, NHS Greater Glasgow & Clyde and University of Glasgow, Queen Elizabeth University Hospital, Glasgow, Scotland, UK
Edward S. Tobias
West of Scotland Centre for Genomic Medicine, Queen Elizabeth University Hospital, Glasgow, Scotland, UK
Ruth McGowan
South East of Scotland Genetics Service, Western General Hospital, Edinburgh, UK
Morad Ansari, Austin Diamond & Anne Lampe
East of Scotland Regional Genetics Service, NHS Tayside, Ninewells Hospital, Dundee, UK
David Baty & Jonathan Berg
Laboratory Genetics, West of Scotland Centre for Genomic Medicine, Queen Elizabeth University Hospital, Glasgow, Scotland, UK
Therese Bradley, Vera Cerqueira & Nicola Williams
Bioinformatics Analysis Core, MRC Institute of Genetics and Cancer, University of Edinburgh, Edinburgh, UK
Mihail Halachev & Alison Meynert
Edinburgh Genomics, University of Edinburgh, Edinburgh, UK
Caitlin Newman, Marian Thomson, Urmi Trivedi & Javier Santoyo-Lopez
School of Medicine, Medical Sciences, Nutrition and Dentistry, University of Aberdeen, Aberdeen, UK
Zosia Miedzybrodzka
Authors
- Prasun DuttaView author publications
Search author on:PubMed Google Scholar
- Alistair T. PagnamentaView author publications
Search author on:PubMed Google Scholar
- Christelle RobertView author publications
Search author on:PubMed Google Scholar
- Anthony E. F. McGuiganView author publications
Search author on:PubMed Google Scholar
- Alison RossView author publications
Search author on:PubMed Google Scholar
- Edward S. TobiasView author publications
Search author on:PubMed Google Scholar
- Ruth McGowanView author publications
Search author on:PubMed Google Scholar
- David BatyView author publications
Search author on:PubMed Google Scholar
- Jonathan BergView author publications
Search author on:PubMed Google Scholar
- Therese BradleyView author publications
Search author on:PubMed Google Scholar
- Vera CerqueiraView author publications
Search author on:PubMed Google Scholar
- Austin DiamondView author publications
Search author on:PubMed Google Scholar
- Mihail HalachevView author publications
Search author on:PubMed Google Scholar
- Anne LampeView author publications
Search author on:PubMed Google Scholar
- Alison MeynertView author publications
Search author on:PubMed Google Scholar
- Caitlin NewmanView author publications
Search author on:PubMed Google Scholar
- Marian ThomsonView author publications
Search author on:PubMed Google Scholar
- Urmi TrivediView author publications
Search author on:PubMed Google Scholar
- Nicola WilliamsView author publications
Search author on:PubMed Google Scholar
- Jing YuView author publications
Search author on:PubMed Google Scholar
- Javier Santoyo-LopezView author publications
Search author on:PubMed Google Scholar
- Zosia MiedzybrodzkaView author publications
Search author on:PubMed Google Scholar
- Jenny C. TaylorView author publications
Search author on:PubMed Google Scholar
- Timothy J. AitmanView author publications
Search author on:PubMed Google Scholar
Contributions
TJA conceived the study, obtained funding and co-wrote the manuscript; PD designed the bioinformatics analysis methodology, performed data analysis on all LRS samples and wrote the manuscript with inputs from TJA, JCT, ATP, CR, AM, MH and JSL; JY developed SVRare software and JY and ATP performed the analysis of SRS SV data via SVRare; ZM oversaw the governance and permissions aspects of the study; AEFM performed data analysis on SRS and LRS SV data for the compound heterozygous analysis; CN, MT, UT and JSL undertook the library preparation and sequencing of all clinical samples; AR, EST and RM provided clinical and phenotypic details for the cases described in the manuscript; ZM, MA, DB, JB, TB, VC, AD, AL and NW recruited patients and obtained phenotypic details from study participants.
Ethics declarations
Competing interests
TJA is a Council Member and Trustee of the UK Academy of Medical Sciences and is co-founder and equity holder of BioCaptiva plc
Ethics approval and consent to participate
Samples were collected from 74 individuals for the long-read sequence study within the Scottish Genomes Partnership programme: “NHS Scotland in 100,000 genomes study”. All gave informed consent for use of samples and phenotype data for research purposes, including genome sequencing, data storage and de-identified data sharing. This research study was approved by North of Scotland Research Ethics Committee (16/NS/0137) and Scotland A Research Ethics Committee (17/SS/0113); the Public Benefit and Privacy Panel (1516-0377); and NHS Scotland health board Research and Development departments.
Consent for publication
Pedigree details and details of the clinical phenotype were included for three families, all of whom gave additional consent for inclusion of their data and clinical features in this publication
Additional information
Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations
Supplementary information
Supplementary Tables, Figures and Methods (download PDF )
Supplementary Table 1 (download XLSX )
Supplementary Table 2 (download XLSX )
Supplementary Table 3 (download XLSX )
Supplementary Table 6 (download XLSX )
Supplementary Table 8 (download XLSX )
Supplementary Table 9 (download XLSX )
Supplementary Table 10 (download XLSX )
Supplementary Table 11 (download XLSX )
Rights and permissions
Open Access This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by/4.0/.
About this article
Cite this article
Dutta, P., Pagnamenta, A.T., Robert, C. et al. Detecting pathogenic structural variation in families with undiagnosed rare disease in a national genome project.
Eur J Hum Genet (2026). https://doi.org/10.1038/s41431-026-02210-x
Received:22 December 2025
Revised:07 April 2026
Accepted:22 July 2026
Published:05 August 2026
Version of record:05 August 2026
DOI
:https://doi.org/10.1038/s41431-026-02210-x


