Close Menu
healthylife7.comhealthylife7.com

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    What's Hot

    An ‘Exercise Pill’ Just Cleared Its First Human Test

    August 19, 2026

    La Newyorkina, the Spanish granola brand inspired by a bad breakfast in New York

    August 19, 2026

    Family genetic designs in MoBa provide insights into health and functioning

    August 19, 2026
    Facebook X (Twitter) Instagram
    Trending
    • An ‘Exercise Pill’ Just Cleared Its First Human Test
    • La Newyorkina, the Spanish granola brand inspired by a bad breakfast in New York
    • Family genetic designs in MoBa provide insights into health and functioning
    • Meghan Trainor shows off enviable physique in black bikini during picturesque family glamping vacation
    • A Focus on More than Fitness: Axon Performance opens its doors
    • Food for modern warfare: The military’s new field rations are smaller, more nutritious, and tastier
    • Bridging the gap between healthcare and technology
    • This super popular food storage container from Rubbermaid is more than 25% off and less than $15
    Facebook X (Twitter) Instagram
    healthylife7.comhealthylife7.com
    • Home
    • Fitness
    • Health
    • Nutrition
    • Lifestyle
    • Conditions
    • Mental Health
    • Weight Loss
    • Wellness Tips
    Wednesday, August 19
    healthylife7.comhealthylife7.com
    Home»Conditions»Family genetic designs in MoBa provide insights into health and functioning
    Conditions

    Family genetic designs in MoBa provide insights into health and functioning

    healthylife7By healthylife7August 19, 2026No Comments66 Mins Read
    Facebook Twitter Pinterest LinkedIn Tumblr Reddit WhatsApp Email
    Family genetic designs in MoBa provide insights into health and functioning
    Share
    Facebook Twitter LinkedIn Pinterest WhatsApp Email

    Download PDF

    Abstract

    Genome-wide association studies using large, population-based samples of unrelated individuals have discovered thousands of genetic associations with health and disease1. These studies can help explain genetic and environmental risks. However, increasing evidence suggests that population-based estimates, while precise, can also reflect confounding that affects their use and interpretation. This confounding can be overcome using data from genotyped family members, such as nuclear mother–father–child trios2,3. However, samples of genotyped families are rare4,5,6,7,8,9,10,11. Here we illustrate some of the advantages of familial data using the Norwegian Mother, Father and Child Cohort Study (MoBa), a population-based cohort of parents and offspring with extensive genotype data (n≈ 230,000) (ref. 3), along with broad and longitudinal phenotyping of health and functioning. We provide an overview of MoBa and describe the quality control of genotype data tailored to this extensively related sample. We then use trio data to illustrate how family-based genomic designs can identify distinct direct and indirect sources of genetic influence and structural confounding. As examples, we analyse children’s height, educational achievement, depressive symptoms and sleep duration. These demonstrations highlight MoBa as a broadly valuable resource for advancing understanding of health and functioning across the lifecourse and generations.

    Subjects

    The field of genome-wide association studies (GWAS) has continued to advance through ever-increasing sample sizes and imputation with improved reference panels12,13,14,15,16,17,18,19,20,21,22,23. This has allowed the identification of thousands of single-nucleotide polymorphisms (SNPs) associated with a multitude of human traits1. Advancements in statistical methodologies and modelling software have allowed increasingly complex research questions to be addressed, for example, using multivariate GWAS24, polygenic score analysis25,26,27 and Mendelian randomization28,29.

    Most of the GWAS so far have used clinical or population-based samples of unrelated individuals3,30. However, increasing evidence suggests that genetic associations estimated from unrelated individuals reflect not only direct genetic effects—in which an individual’s genetic variants causally influence their own phenotype—but also indirect effects of other individuals’ genotypes, such as parental genotypes shaping the rearing environment and thereby influencing offspring outcomes31,32. Estimates of indirect genetic effects in parent–offspring models reflect the association between parental genotypes and offspring outcomes after conditioning on the offspring’s genotype. These estimates capture ‘genetic nurture’ effects, as well as components arising from broader population structure, such as assortative mating and population stratification. Disentangling direct from indirect genetic effect estimates is very challenging in samples of unrelated individuals3,33. By contrast, family-based analyses can directly control for parental genotypes and exploit the random inheritance of genetic variation from parents to offspring to preclude many forms of environmental confounding. Therefore, using large samples of genotyped, related individuals4,5,6,7,8,9,10,11 is one of the most compelling ways to estimate the contributions of direct and indirect effects of both genotypes and phenotypes.

    Genome-wide significant associations between SNPs and educational attainment in population-based GWAS of unrelated individuals are particularly attenuated in family-based studies34. A within-family GWAS of 179,086 siblings also suggested substantial attenuation of population-based estimates within families for height, age at first birth, number of children, cognitive ability, depressive symptoms and smoking2. Other studies of parent–offspring duos and trios have explicitly demonstrated that the parents’ non-transmitted genetic variants related to educational attainment associated with their offspring’s outcomes31,35. These findings provide strong evidence that within-family analyses can meaningfully disentangle direct and indirect genetic effects; a task that is beyond the reach of standard population-based approaches. Methods using familial relatedness in genotyped samples to estimate population parameters hold marked potential to advance our understanding of familial aggregation of health conditions and disabilities, co-occurrence between different health conditions and identification of causal risk and protective factors36,37,38,39,40,41,42.

    Despite the valuable insights possible from studying genotyped families, few cohorts or biobanks have implemented large-scale family-based sampling4,5,6,7,8,9,10,11. Here, we emphasize the important role of genotyped family cohorts in advancing genomic research on health and functioning in the years ahead, highlighting the Norwegian Mother, Father and Child Cohort Study (MoBa) as a prime resource. MoBa combines extensive family-based genetic data with rich prospective and longitudinal follow-up of more than 100,000 children from prenatal life into adulthood43.

    MoBa, a family-based genetic resource

    MoBa represents a uniquely powerful resource for family-based genetic analyses, offering a wide range of phenotypic measures on mothers, fathers and offspring, collected from early pregnancy and spanning more than 25 years. Within the main data collections covering the first 16 years of life alone (Fig. 1a), MoBa has amassed observations on neurodevelopment, mental health, physical health, health behaviours and medication use, social and environmental exposures (Fig. 1b), and broader aspects of functioning such as sleep, school achievement and social participation.

    Fig. 1: Overview of phenotypic and genetic data availability among MoBa participants across the first 16 years of data collection.
    Full size image

    a, Core MoBa questionnaires by respondent and child year during the first 16 years, points scaled in size relative to proportion of the full baseline sample of participating MoBa children with data available at each data collection wave. b, Summary of information obtained in different domains across core questionnaires administered in the first 14 years of core data collections; ‘n observations’ is calculated as the product of the number of responses by each participant type at each wave and the number of distinct measurement items in each domain at the same wave. c, Total n of mothers, fathers and children available in MoBa. Asterisk indicates the percentage of individuals with available genotype data passing quality control and with genetic similarity to the 1000 Genomes Project European superpopulation. The list is indicative of key linkable nationwide registries. d, Overlap among genotyped individuals within families in c—this is calculated at the level of individual MoBa children, meaning that genotyped mothers and fathers participating with more than one child are included more than once.

    Ongoing waves of data collection follow the children through adolescence and into early adulthood. The rich prospective and longitudinal observations are complemented with linkage to Norwegian registry data (Fig. 1c). Norway has an extensive, high-quality health data infrastructure, including national health and administrative registries with continuous information on contacts, diagnoses, functioning assessments, treatments, care and support from the primary health and care services, the specialized health care services, and educational and social services. MoBa also includes an extensive range of environmental data collected through repeated questionnaires and registry linkages, as well as diverse biological samples (including plasma, urine and primary teeth), enabling integrated studies of genetic, environmental and biological mechanisms across the lifecourse. Furthermore, numerous nested clinic studies in MoBa have collected additional diverse biological samples and deep phenotype data, further expanding the potential to derive insights into human health and functioning using family-based genetic analyses43.

    This potential3 is underpinned by the widespread availability of genotype data among MoBa parents and children. A considerable amount of relatedness between participating families in the cohort, combined with complexities in genotyping, necessitated the implementation of a quality control pipeline tailored to MoBa. Therefore, we developed and applied a family-based quality control and imputation pipeline that verifies and accounts for diverse types of relatedness in large population-based samples of individuals while appropriately handling differences arising from a varied genotyping strategy (Supplementary Note 1). The dataset passing post-imputation quality control included 77,634 mothers, 53,358 fathers and 76,577 children with the following relationships identified (Supplementary Note 2): 287 monozygotic twin pairs, 22,884 full-siblings, 117,006 parent–offspring pairs, 23,299 second-degree relative pairs, and 10,828 third-degree relative pairs. This included 44,017 full father–mother–child trios (Fig. 1d), which form a basis for refined genomic discovery and enhanced downstream genetic analyses.

    Here, we demonstrate the valuable insights that can be gained by incorporating maternal and paternal genotypes in the investigation of genetic effects on offspring health and functioning. We apply a range of trio-based methods to four example offspring phenotypes spanning different domains of health and functioning, all measured in middle childhood. Children’s height (age 7 years), sleep duration (age 7 years) and depressive symptoms (age 8 years) were reported by mothers in MoBa questionnaires, whereas educational achievement was assessed using a composite score of numeracy, reading and English skills derived from the mandatory national tests of the educational registry in grade 5 (age 10 years). Descriptive information for these measures is presented in Table 1. To provide additional context for interpreting the offspring genetic and phenotypic results, descriptive statistics and Pearson correlations for similar outcomes to the four offspring phenotypes are presented for mothers and fathers in Extended Data Table 1. Spousal Pearson correlations are presented in Extended Data Table 2.

    Table 1 Descriptive statistics and Pearson correlation matrix (with 95% CI) for the four middle childhood phenotypes
    Full size table

    Refined genomic discovery with trio-GWAS

    The availability of rich phenotypic and genotypic data from family cohorts such as MoBa opens new opportunities to identify specific common genetic variants causing variability in health and functioning. GWAS using parent–offspring data are particularly useful for disentangling the different sources of genetic association at the parent and offspring levels2. We performed population-based and trio-GWAS on the four offspring traits and compared SNP–phenotype associations between population-based and family-based analyses.

    In the population models, SNP-based heritability (({h}_{mathrm{SNP}}^{2})) for height and educational achievement was ({h}_{mathrm{SNP}}^{2}) = 35% (standard error (s.e.) = 3%) and ({h}_{mathrm{SNP}}^{2}) = 31% (s.e. = 2%), respectively. Estimates were lower for sleep duration (({h}_{mathrm{SNP}}^{2}) = 5%, s.e. = 2%) and depressive symptoms (({h}_{mathrm{SNP}}^{2}) = 6%, s.e. = 3%). In the trio models, which conditioned on parental genotypes, we saw the heritability estimates for height and education attenuate to ({h}_{mathrm{SNP}}^{2}) = 22% (5%) and ({h}_{mathrm{SNP}}^{2}) = 22% (3%), respectively (Fig. 2a). Heritability for sleep duration was reduced to near zero within-families (({h}_{mathrm{SNP}}^{2}) = 1%, s.e. = 4%), whereas the within-family estimate for depressive symptoms was higher (({h}_{mathrm{SNP}}^{2}) = 11%, s.e. = 5%).

    Fig. 2: GWAS conducted within-family trios refine genomic discovery and reduce bias.
    Full size image

    a, SNP heritability estimates calculated using LDSC based on summary statistics from population- and trio-based GWAS models of children’s height, educational achievement, sleep duration and depressive symptoms. Data are presented as LDSC model estimates with 95% CIs. b, Estimates of genetic correlations (rGs) between children’s height, educational achievement, sleep duration and depressive symptoms and five externally GWASed traits were calculated using LDSC based on either population- or trio-based GWAS models conducted in MoBa (note that sleep duration (trio model) – major depressive disorder (MDD)/BMI rGs could not be stably estimated). Data are presented as LDSC model estimates with 95% CIs. n trios by outcome: educational achievement, 41,761; height, 23,786; sleep duration, 24,376; and depressive symptoms, 18,701.

    Figure 2b presents the linkage disequilibrium score regression (LDSC) estimates of genetic correlations (rGs) between each MoBa GWAS (height, educational achievement, sleep duration and depressive symptoms) and large external population-based GWAS of height44, body mass index (BMI)45, educational attainment46, smoking initiation47 and depression48. For MoBa children’s educational achievement, the within-family estimated genetic correlations were attenuated for all phenotypes. The genetic correlation with educational attainment was lower in the family-based (rG = 0.62, s.e. = 0.04, P = 1.14 × 10−47) compared with the population-based analyses (rG = 0.76, s.e. = 0.02, P = 3.06 × 10−216). Similarly, its genetic correlations with BMI and smoking initiation were halved in the family-based (BMI rG = −0.07 (s.e. = 0.03, P = 0.03) and smoking rG = −0.05 (s.e. = 0.04, P = 0.23)) versus in the population-based analyses (rG = −0.16 (s.e. = 0.02, P = 7.82 × 10−12) and rG = −0.13 (s.e. = 0.03, P = 1.95 × 10−5), respectively). Similar attenuation was observed for the genetic correlation between MoBa educational achievement and height (rG = 0.06 (s.e. = 0.03, P = 0.06) in family-based and rG = 0.12 (s.e. = 0.02, P = 8.96 × 10−7) in population-based analyses).

    For MoBa children’s height, changes in rG estimates were less substantial, but notably, rG estimates with external GWAS of height and BMI (rG = 0.88 (s.e. = 0.11, P = 1.81 × 10−16) and rG = 0.24 (s.e. = 0.05, P = 9.06 × 10−6), respectively) increased when adjusting for parental genotypes

    For depressive symptoms in MoBa children, the genetic correlation with depression was attenuated in the family-based (rG = 0.15 (s.e. = 0.08, P = 0.05)) compared with the population-based analyses (rG = 0.40 (s.e. = 0.13, P = 0.002)). The genetic correlations with BMI and smoking initiation were also attenuated within-family (rG = 0.11 (s.e. = 0.07, P = 0.13) and rG = 0.08 (s.e. = 0.08, P = 0.34)) compared with the population-based analyses (rG = 0.22 (s.e. = 0.09, P = 0.01) and rG = 0.24 (s.e. = 0.10, P = 0.02), respectively), whereas the genetic correlation with educational attainment was marginally higher within-family rG = −0.26 (s.e. = 0.09, P = 0.004) compared with the population-based analyses rG = −0.20 (s.e. = 0.09, P = 0.03).

    For MoBa sleep duration, no genetic correlations with external GWAS traits were statistically significant in either the population-based or trio-based analyses (all P > 0.05). Estimates were modest and accompanied by large standard errors, particularly in the within-family model

    The estimates of associations of genome-wide significant variants previously reported to be associated with height and education were attenuated by 11.0% (95% confidence interval (CI) 8–15%) and 23.0% (95% CI 18–28%) after controlling for parental genotype, respectively. We also observed within-family attenuation for genetic variants associated with sleep duration and depressive symptoms, with association estimates decreasing by 29% (95% CI 9–49%) and 16% (95% CI 3–28%) from population to within-family analyses, respectively.

    Our population-based GWAS identified 98 and 26 genome-wide significant (P < 5 × 10−8) offspring variants associated with height and educational achievement, respectively. Of these, five variants for height and one for educational achievement remained associated at the genome-wide significance level after adjusting for parental genotype (see Miami plot Extended Data Fig. 1a for height and Extended Data Fig. 2a for educational achievement, Extended Data Fig. 3a for sleep duration and Extended Data Fig. 4a for depressive symptoms). The within-family standard errors were on average 38% larger than the population-based standard errors across traits (height 38%, educational achievement 37%, sleep duration 38% and depression 39%), indicating lower statistical power in the within-family analyses. Quantile–quantile plots indicate significant deviation from the expected distribution in the population-based models for each phenotype (see Extended Data Fig. 1b for height, Extended Data Fig. 2b for educational achievement, Extended Data Fig. 3b for sleep duration and Extended Data Fig. 4b for depressive symptoms). This is substantially reduced in the conditional models that control for parental genotype.

    Taken together, these results show that a substantial proportion of SNP-based heritability, and many reported SNP-trait and cross-trait associations in population-based GWAS, especially for educational achievement and depressive symptoms, reflect indirect effects and between-family structure rather than direct genetic effects on offspring traits. Trio-GWAS in MoBa reveals smaller but more likely causal within-family genetic effects, altered patterns of genetic overlap with other traits and fewer genome-wide significant loci, underscoring both the power and necessity of family-based designs for refining genomic discovery.

    Trio-polygenic score methods

    Apart from facilitating refined genomic discovery, genetic data in family cohorts allow downstream genetic analyses to be performed in new ways. We used two family-based approaches that use polygenic scores (PGS): trio-PGS and polygenic disequilibrium transmission testing. These complementary approaches were used in MoBa to analyse within-family transmission mechanisms for height, educational achievement, sleep duration and depressive symptoms

    The trio-PGS model allows the respective contributions of children’s own and each of their parents’ PGS for specific traits to explain variance in an outcome simultaneously38 (Fig. 3a). Using trio-PGS, we showed that the unadjusted association (β = 0.43, 95% CI 0.43–0.44) between a child’s PGS for height and their measured height is inflated by 9% because of the influence of parents’ PGS (Fig. 3b, left-most). Both paternal (β = 0.03, 95% CI 0.01–0.04) and maternal (β = 0.03, 95% CI 0.02–0.05) PGS independently contributed to this inflation (Fig. 3c, top). Similar evidence of parental PGS biasing estimates of children’s PGS associations with their own traits was evident across all outcomes. In most cases, unadjusted estimates were inflated, leading to seemingly stronger associations when parents’ PGS were not controlled for. However, there were notable examples of parents’ PGS suppressing associations between children’s PGS and their traits. For instance, the association between children’s PGS and educational attainment and their depressive symptoms more than doubled after adjustment for parents’ PGS (βunadjusted = −0.03, 95% CI −0.04 to −0.02; βadjusted = −0.06, 95% CI −0.08 to −0.04; Fig. 3b, right-most). Figure 3c (bottom) illustrates that this suppression effect, as well as the 130% inflation of the association between child PGS to depression and their own depressive symptoms in unadjusted models, were attributable to maternal, but not paternal, genetic effects. Sensitivity analyses confirmed that the inclusion of individuals related across trios was not responsible for driving the pattern of results (see Extended Data Fig. 5a for estimates of direct genetic effects and Extended Data Fig. 5b for effects of parental PGS).

    Fig. 3: Within-family PGS-based analyses provide estimates of direct genetic effects adjusted for systematic structural confounding.
    Full size image

    a, Schematic of links between mothers’, fathers’ and children’s genetic liabilities (GM/F/C) and phenotypes (PM/F/C). The solid lines indicate known relations and the dotted lines indicate unknown, but plausible relations. Trio-PGS analyses involve estimation of the direct path between GC and PC, with control for confounding by GM/F and PM/F. b, Estimates of direct genetic effects on MoBa children’s height, educational achievement, sleep duration and depressive symptoms from child-only and trio-PGS regression models. Data are presented as standardized regression model coefficients with 95% CIs. c, Effects of parental PGS that explain differences between direct effect estimates in the two models in b. Data are presented as standardized regression model coefficients with 95% CIs. d, Schematic of the expectation of polygenic transmission disequilibrium, in which genetic liabilities are systematically over- or under-transmitted when families are selected based on a child trait with a direct link to that genetic liability. e, Polygenic transmission disequilibrium (PTD), indexed by systematic differences in polygenic scores—as compared with a calculated mid-parental average polygenic score—among children selected for, respectively, height 2 standard deviations (s.d.) above the mean, educational achievement above 90th percentile, sleep duration below 10th percentile, and depressive symptoms 2 s.d. above the mean. Over- or under-transmission relative to mid-parental average polygenic score is usually interpreted as evidence of genetic influence on child trait free from structural confounding such as population stratification or assortative mating. Data are presented as means with 95% CIs. n trios by outcome in PGS analyses (b,c): educational achievement, 40,101; height, 21,008; sleep duration, 21,451; and depressive symptoms, 18,212. n trios by outcome in PTD analyses (e): educational achievement, 3,227; height, 527; sleep duration, 438; and depressive symptoms, 618.

    Polygenic transmission disequilibrium testing (PTDT; Fig. 3d) examines whether children selected for a given phenotype, for example height, receive a higher (or lower) polygenic load for this phenotype from their parents39. Under Mendelian inheritance, a child’s PGS should, on average, be equal to the mid-parent PGS. If the phenotype is influenced by genetic variants in the PGS, then children selected for extreme values on that phenotype should show over-transmission of the relevant PGS (on average, the children should have higher PGS than expected based on the average of their parents’ PGS). If, instead, the PGS–phenotype association is due to structural confounding (for example, from population stratification) that also influences parental PGS, this pattern of over-transmission would be absent or attenuated. Importantly, PGS of unselected siblings should not exceed the mid-parent average—regardless of the relationship between the PGS and the phenotype of selection.

    For these analyses (Fig. 3e), we used subsets of MoBa children based on extreme height (>2 s.d., above the sample mean; n = 527, plus n = 84 unselected sibling controls), educational achievement above the 90th percentile, (n = 3,227, nsiblings = 535), sleep duration below the 10th percentile (n = 438, nsiblings = 78), and extreme depressive symptoms (>2s.d. above the sample mean; n = 618, nsiblings = 95). Using PTDT, we observed significant over-transmission of PGS for height among extremely tall children (PFDR < 0.001) and educational attainment among children with high educational achievement (PFDR < 0.03). In both cases, siblings did not show evidence of similar over-transmission (Extended Data Fig. 5c), although estimates were considerably less precise. Neither sleep duration nor depressive symptoms PTDT analyses provided robust evidence for over- or under-transmission of any of the PGS. These findings suggest that, for height and educational achievement, the PGS signal reflects genuine within-family genetic transmission rather than solely between-family confounding, strengthening causal interpretation of these polygenic influences. By contrast, the lack of clear over-transmission for sleep duration and depressive symptoms suggests that current PGS for these traits either capture weaker within-family genetic effects or are more strongly influenced by confounding and measurement error, highlighting areas in which larger discovery samples and improved phenotyping are needed.

    Spousal PGS Pearson correlations indicated modest within-trait assortative mating for the PGS included in our analyses, most pronounced for educational attainment (r = 0.14) and weaker for height, BMI, smoking and depressive symptoms (r = 0.02–0.06) (Extended Data Table 3)

    Together, the results of the trio-PGS and PTDT analyses show that modelling parental genotypes, even in analyses using scores based on population-based GWAS, can reduce bias in the estimation of direct genetic effects. The results also show that this bias can influence PGS–phenotype associations by maternal and paternal genotypes independently, highlighting the importance of trio data for solving this problem

    Trio-genome-wide complex trait analysis

    Variance decomposition methods such as genome-wide complex trait analysis (GCTA)49 use observed genome-wide genetic relatedness to estimate heritability. Similar to genomic discovery and polygenic score-based methods, their estimates can be refined and improved in the context of genotyped family data. Trio-GCTA36 uses parent–offspring trios to estimate the genome-wide contributions of direct and indirect genetic effects. We applied trio-GCTA to the exemplar child phenotypes—height, educational achievement, sleep duration and depressive symptoms—to disentangle the relative roles of genome-wide direct and indirect genetic effects beyond trait-specific PGS.

    Trio-GCTA provided clear evidence for genetic effects on height and educational achievement, and statistically detectable but smaller genetic contributions to sleep duration and depressive symptoms (Fig. 4a). For educational achievement, the best-fitting model included covariances between maternal, paternal and offspring genetic effects, whereas for the other phenotypes a simpler model without covariances sufficed. Likelihood ratio tests indicated that parental genetic variance components improve model fit for height (P = 0.03), sleep duration (P = 3.4 × 10−4), and depressive symptoms (P = 0.02), but the estimated indirect effects are modest and imprecise and the contrast between model fit statistics (Extended Data Table 4) suggests only weak evidence for indirect effects.

    Fig. 4: Trio-GCTA variance decomposition allows estimation of the total contribution of direct and both maternal and paternal indirect genetic effects.
    Full size image

    a, Estimated variance components from trio-GCTA of MoBa children’s height, educational achievement, sleep duration and depressive symptoms. The proportion of variance explained by child, parental genetic effects and positive covariance between parental and child genetic effects, which inflates the total variability in the phenotype explained by genetic effects, are shown in stacked bars above zero. Negative covariance between parental and child genetic effects is shown in stacked bars below zero as these effects will deflate the total variability in the phenotype explained by genetic effects. The best-fitting model as selected on the basis of likelihood ratio model fit comparison tests (likelihood ratio test, LRT) and Akaike information criterion values is indicated with an asterisk. Models with no genetic effects were also tested but are not shown. P-values for the LRT were calculated from a single chi-square distribution when covariance parameters were constrained (two-sided) and a mixture of chi-square distributions when variance parameters were constrained (one-sided); for all outcomes, a model including indirect effects were preferred—for educational achievement only, the covariance between effects was also included in the best-fitting model. b, LRT P values from trio-GCTA of height, educational achievement, sleep duration and depressive symptoms. For both height and educational achievement, the decision to reject the model without genetic effects was clear. All other decisions were marginal, indicating that genetic effects overall were minimal for sleep duration and depression and indirect effects played a limited part in explaining variation in all outcomes. n trios by outcome in GCTA analyses: educational achievement, 23,221; height, 12,332; sleep duration, 12,799; and depressive symptoms, 10,794.

    Direct effects were the largest contributor to genetic variance in educational achievement and height (educational achievement: 17.2% (s.e. = 2.5%); height 32.0% (s.e. = 2.8%) in the best-fitting model). Reflecting conclusions from comparison of model fit indices, depressive symptoms and sleep duration had comparatively little variance explained by genetic effects in general (Fig. 4b), with weak evidence for primarily indirect effects. Indirect effects were also found for educational achievement and, to a lesser extent, height. For these traits, in contrast to depressive symptoms and sleep duration, the indirect effects were primarily estimated to originate with fathers. For educational achievement, significant positive covariances between parental and child effects indicate gene-environment correlation acting to increase the overall variance explained by genetic effects in the phenotype.

    The results indicate that when modelling genome-wide genetic effects in parent–offspring trios, children’s own genotypes (direct effects) can explain a substantial part of genetic variance for some traits (especially educational achievement and height), but not all of it. Ignoring parental genotypes can bias estimates of child genetic effects, particularly for educational achievement, for which family genetics may be especially intertwined

    Inclusion of parental genetic effects in downstream genetic analyses—whether based on specific genetic propensities indexed by PGS or based on decomposition of all genetic variance in the trait—gives new insights into genetic mechanisms. Neither approach conclusively proves why children’s traits associate with parental genotypes; so-called indirect ‘effects’ comprise many different sources of association3,33, but both can demonstrate that the associations between children’s own genotypes and phenotypes may be confounded. Moreover, both using specific genetic propensities (PGS) and genome-wide variation approaches can provide unbiased estimates of the direct effects of children’s own genotypes on their phenotypes.

    Conclusion

    Our work demonstrates the vast potential of MoBa, a genotyped family cohort with rich phenotype data, for refining genomic discoveries, clarifying genetic overlaps and investigating intergenerational transmission. With large-scale genotyping of parents and offspring, deep longitudinal phenotyping and linkage to comprehensive health and administrative registries, MoBa provides a unique resource for family-based genetic analyses across health and functioning. Here, we have brought together and harmonized genotype data from multiple projects in MoBa and applied quality control procedures tailored to the unique characteristics of the sample.

    Using these data, trio-GWAS embeds direct control for parental genotype into the genomic discovery process. This allows the most likely biologically mediated SNP–trait associations to be distinguished from those due to population stratification, genetic nurture and assortative mating. Modern population-based GWAS efforts, powered by ever-increasing sample sizes, now routinely detect hundreds of robust, independent SNP–outcome associations18,44,45,46,48,50. Our trio-GWAS results indicate that these associations cannot necessarily be attributed to the direct effects of genotype on phenotype. Moreover, we show not only inflation of heritability estimates in population-based compared with trio-GWAS but also substantial changes in the patterns of overlap with other traits. Polygenic-score-based and variance-decomposition-based approaches reveal these nuances, highlighting avenues for future investigation—such as parent-specific indirect effects, opposing effects of the same PGS in children and parents, and assortative mating. Family-based analyses and samples such as these represent a jumping-off point for the defining challenge of the post-GWAS era: explaining mechanistic links between common genetic variation and human traits. Estimates of direct genetic effects using trio data should be reasonably robust to all the sources of indirect effects, and therefore allow for prioritization of causal variants and effect sizes. However, estimates of indirect genetic effects may reflect varied components, including not only genetic nurture but also broader demographic factors, such as assortative mating, population stratification, dynastic effects and indirect genetic effects from relatives other than parents. Strategies to disentangle these factors involve triangulation across genetic designs—including those demonstrated here, and others such as mediation analyses51,52, extended pedigrees53 and assortative mating modelling40, as well as across multiple sources of family genetic data. Systematic investigation of assortative mating across cohorts, countries and historical periods is important for understanding how population-level mating patterns shape indirect genetic effects and their generalizability. Furthermore, as within-family GWAS incorporating samples such as MoBa become larger and more powerful, the potential to directly mitigate and control for confounding from assortative mating and population stratification in downstream within-family analyses will also increase.

    Beyond its family-based structure, MoBa has many other distinctive features that can facilitate the development and application of new methods and innovative uses of genotype data. For example, MoBa features a prospective, longitudinal follow-up of more than 100,000 children from prenatal life to adulthood, allowing for comprehensive tracking and genetic analysis of health and functioning. Moreover, the 10-year recruitment period (Fig. 1a) presents opportunities to study temporal interplay between genetics and shifting environmental and societal contexts that affect health and functioning across generations. Nevertheless, no single resource or study design is without limitations. Lengthy recruitment periods and follow-up periods can present challenges as well as opportunities, particularly where measurement cannot be kept consistent either across time or age. Furthermore, the effects of non-random ascertainment and attrition represent crucial considerations for within-family analyses using such data. Finally, family-based cohorts such as MoBa are expensive to recruit and maintain, and care must be taken to ensure that prioritization of family-based cohorts in the coming years does not undermine ongoing efforts to diversify the ancestral populations included in genetic samples and ensure insights can benefit the broadest possible population. At the same time, because within-family designs are robust to population stratification, they are particularly well suited to addressing concerns about stratification in ancestrally diverse and samples with recent admixture, underscoring the importance of investing in family-based studies across a broad range of populations. Despite these limitations, a combination of widespread implementation of existing within-family approaches to new data sources and the integration of these approaches with longitudinal designs and natural experiments means that research in family-based cohorts is likely to play a central part in promoting ground-breaking discoveries in genomics in the coming years.

    Methods

    The Norwegian MoBa cohort

    MoBa is a population-based pregnancy cohort study conducted by the Norwegian Institute of Public Health. Participants were recruited from all over Norway from 1999 to 2008. The women consented to participation in 41% of the pregnancies (n = 112,908 recruited pregnancies)43,54. The cohort includes approximately 114,500 children, 95,200 mothers and 75,200 fathers. The establishment of MoBa and initial data collection were based on a licence granted by the Norwegian Data Protection Agency and an approval from the Regional Committees for Medical and Health Research Ethics (REK). The MoBa cohort is currently regulated by the Norwegian Health Registry Act. The current study was approved by the Regional Committees for Medical and Health Research Ethics (14140 and 2016/1226), and all research was performed in accordance with relevant guidelines and regulations. All mothers and fathers provided written informed consent at recruitment. Mothers consented to participation on behalf of themselves and their children. At 18 years of age, the children become independent participants and receive an information letter about their rights, including how to withdraw. Participation is voluntary, and participants can withdraw their consent at any time. In accordance with REK regulations, individuals who withdraw consent are excluded. As shown in Fig. 1a, MoBa includes questionnaire data collections at many time points, including during the pregnancy, infancy, preschool-age, school-age, adolescence and beyond. Further details about the cohort representativeness and participation in specific waves of data collection are described elsewhere43,54,55,56. Detailed instrument documentation is available on the MoBa website (https://www.fhi.no/en/ch/studies/moba/). A wide range of national health and administrative registries can be linked to MoBa for further phenotyping without reliance on participant engagement in completing questionnaires (Fig. 1b).

    Phenotypic measures

    For our exemplar analyses, we included four phenotypes assessed in middle childhood (age 7–10 years), spanning physical health (height), mental health (depression symptoms) and aspects of functioning (sleep duration and educational achievement). Child height at 7 years of age was reported by mothers in the age-7 questionnaire for the question, ‘What is the child’s height and weight now at 7 years old?’ Mothers were asked to report their child’s current height in centimetres. Sleep duration at 7 years of age was maternally reported in the 7-year-questionnaire on the item ‘Approximately how many hours does the child usually sleep on a weeknight?’ with the following response categories: (1) 8 h or less, (2) 9 h, (3) 10 h, (4) 11 h and (5) 12 h or more. Depressive symptoms were measured using the 13-item Short Mood and Feelings Questionnaire57 reported by mothers in the 8-year-questionnaire. We prepared the questionnaire-assessed phenotypes using the phenotools R package (0.2.8) (ref. 58) in R 4.1.059. Educational achievement at age 10 was assessed as a composite score based on performance on three national standardized tests of skills in literacy, numeracy and English, derived from the Statistics Norway Educational Registry. The national tests are administered in the autumn of grade 5 (age 10), grade 8 (age 13) and grade 9 (age 14), and are mandatory with exemptions only on application on the grounds that the results will not be useful for assessing the child’s learning due to disability or lack of knowledge of the Norwegian language. Because the raw score ranges differed across subjects and test years, the raw scores were standardized within each test year and subject test. Before standardization, outliers defined as values more than 4 standard deviations (s.d.) from the mean were dropped. The four phenotype scores were standardized for use in the trio analyses to place them on a common scale and to facilitate interpretation and comparison of effect sizes.

    Parental phenotypes were not used in the analyses presented in this study. However, similar outcomes in parents to the four exemplar offspring phenotypes were defined to provide supplementary context for interpreting the offspring genetic and phenotypic results. Parental height was self-reported in centimetres in response to the question, ‘How tall are you?’ Educational achievement was derived from Statistics Norway data corresponding to the International Standard Classification of Education (ISCED)60 definitions. The ISCED levels from 0 to 8 corresponding to ‘early childhood education (“less than primary”)’, ‘Primary education’, ‘Lower secondary education’, ‘Upper secondary education’, ‘Post-secondary non-tertiary education’, ‘Short-cycle tertiary education’, ‘Bachelor’s or equivalent level’, ‘Master’s or equivalent level’ and ‘Doctoral or equivalent level’ were coded as 1, 7, 10, 11, 13, 14, 16, 18 and 21 years of education, respectively. Sleep problems were self-reported by mothers when the child was 14 years old and by fathers in 2015–2016, using the total score of three items modified from the Karolinska Sleep Questionnaire61: ‘How often do you find it difficult to get to sleep at night?’, ‘How often have you woken up repeatedly during the night?’ and ‘How often do you feel tired or sleepy during the day?’ The response options ‘Never’, ‘Less than once a week’, ‘Once per week’, ‘Twice per week’, ‘Three times per week’ and ‘Four times or more per week’, were coded from 0 to 5, respectively. Parental depression symptoms were self-reported by mothers at about week 30 of pregnancy and fathers around week 15 of pregnancy, using the depression subscale of the short eight-item Hopkins Symptoms Checklist (SCL-8)62.

    Information from the Medical Birth Registry of Norway (MBRN)63, a national health registry containing information about all births in Norway from 1967 onwards, and the MoBa questionnaire data were used to identify registered sex, year of birth, multiple births (in the offspring generation) and reported parent–offspring relationships. Before genotyping quality control, all pedigrees were constructed based on the reported parent–offspring relationships for each pregnancy. Each family contained all reported relationships (these included parent–offspring, full-sibling relationships from pregnancies with single and multiple births, and half-sibling relationships). Wherever possible, registered sex was assigned using information from the MBRN. In instances where sex was not registered in the MBRN, the sex registered at birth reported in the MoBa questionnaires was used.

    Biological data

    Blood samples64 were collected from participating mothers and fathers at approximately the 17th week of pregnancy during the ultrasound examination. A second blood sample was taken from the mother soon after birth. The blood sample for the child was taken from the umbilical cord after birth. Biological samples were sent to the Norwegian Institute of Public Health, where deoxyribonucleic acid (DNA) was extracted by standard methods and stored65

    Genotyping and quality control

    Genotyping of MoBa has been conducted through multiple research projects, spanning several years. Full details about MoBa genotyping are described in Supplementary Methods 1. The MoBaPsychGen pipeline for quality control and imputation was developed to ensure the complex relationship structure and varying selection criteria, genotyping batches and genotyping arrays were handled appropriately (pre-printed66). Quality control was performed based on current best-practice protocols67,68,69,70,71. Throughout the pipeline, both SNP and individual-level quality control were performed; with priority given to retaining individuals over SNPs, as SNPs can be imputed. The complete pipeline, with full details about quality control, phasing, imputation and post-imputation quality control, is described module by module in Supplementary Methods 2, and a pipeline overview is shown in Supplementary Fig. 1a.

    The primary software used throughout the pipeline was PLINK v.1.90b7.2 and v.1.90b6.18 (ref. 72) and the R Project for Statistical Computing v.4.0.5 (ref. 59) was used to produce figures. Moreover, LiftOver tool73,74, KING v.2.2.5 (ref. 75), FlashPCA2.0 (ref. 76), imputation preparation and checking script v.4.2.13 (ref. 77), SHAPEIT2 release 904 (ref. 78) with the duoHMM79 algorithm, IMPUTE4.1.2_r300.3 (ref. 80), QCTOOL v.2.0.8 and v.2.2.0 (ref. 81), ic tool v.1.0.8 (ref. 82), PLINK2 v2.00a2.3LM (ref. 72), IMPUTE2.3.2 (ref. 83), yHaplo 2.1.12 (ref. 84), the MitoImpute pipeline v.0.2 (refs. 85,86), BCFtools v.1.8 and v.1.9 (ref. 87), cat-bgen v.1.1.4, COSGAP v.1.0.0 containers (gwas.sif, python3.sif and r.sif from https://github.com/comorment/containers/releases/tag/v1.0.0; for details see the documentation at cosgap.readthedocs.io/en/latest/), Python 3.8.10 and perl v.5.32.1.

    Trio genetic analyses

    Exemplar trio analyses were conducted using the four offspring phenotypes. All analyses were carried out on complete case data and restricted to very well-imputed (imputation quality score ≥0.95) autosomal variants available after application of the quality control pipeline. All analyses adjust for offspring age and sex, and account for effects of technical covariates (genotyping batch and first 20 principal components; PC1–20). Where analyses estimate or account for ‘indirect genetic effects’, these refer exclusively to variation in child outcomes associated with parents’ genotypes (that is, not those originating with siblings or other family members). The term ‘indirect genetic effects’ is used for consistency with the standards in the field, despite many methods being agnostic as to the specific mechanism, which can include indirect effects of parents, population structure (the presence of subgroups within a population that can bias genetic associations) and assortative mating (the tendency for individuals with similar traits to mate non-randomly, influencing genetic associations in the next generation) among other sources.

    Trio-GWAS

    We conducted within-family GWAS (Supplementary Note 4) using genotype data from parent–offspring trios, focusing on four offspring traits: educational achievement (n = 41,761), height (n = 23,786), sleep duration (n = 24,376) and depressive symptoms (n = 18,701). GWAS analyses involved fitting both population and within-family models to the same sample in linkage-disequilibrium-adjusted kinship v.6 (ref. 88). To obtain population and within-family genetic association estimates, linear regression was first performed on offspring genotype without adjusting for parental genotype, and subsequently in a mutually adjusted analysis accounting for parental genotypes. The former is synonymous with a conventional principal-component-adjusted model. The latter extends this model by including the parental genotypes as additional covariates to estimate direct genetic effects, while accounting for (and estimating) indirect parental genetic effects. Standard errors were clustered at the family level to account for offspring relatedness. The models were specified as follows (omitting technical covariates for simplicity):

    Population model:

    $$text{Outcome},=,{beta }_{0}+{beta }_{1}{text{G}}_{text{child}}+{beta }_{2}text{Sex}+{beta }_{3}text{Age}+varepsilon $$

    Within-family model:

    $$text{Outcome},=,{beta }_{0}+{beta }_{1}{text{G}}_{text{child}}+{beta }_{2}{text{G}}_{text{mother}}+{beta }_{3}{text{G}}_{text{father}},+{beta }_{4}text{Sex}+{beta }_{5}text{Age}+varepsilon $$

    To quantify the attenuation of SNP associations in the within-family relative to population-based models, we compared estimates for genome-wide significant variants reported in previously published GWAS of height44, educational attainment46, sleep duration89 and depression48. For each phenotype, we extracted the set of independent genome-wide significant variants from the external GWAS and compared their association estimates in our MoBa population-based and within-family GWAS. Attenuation was calculated as the proportional decrease in SNP effect estimates between the population-based and within-family models, and 95% CIs were derived using a leave-one-out jackknife across the set of included variants. Furthermore, we illustrate the relationship between power and parental heterozygosity rate in Supplementary Fig. 2.

    Linkage disequilibrium score regression

    We used LDSC90 (v.1.0.1 with default parameters) to estimate SNP-based heritability (({h}_{mathrm{SNP}}^{2})) and genetic correlations (rG) between our population and within-family GWAS results and publicly available summary statistics from previous studies on height44, BMI45, educational attainment46, smoking initiation47 and depression48. We input summary-level GWAS data into LDSC, using precomputed linkage disequilibrium scores from the 1000 Genomes Project European superpopulation. For the LDSC n parameter, we used the effective sample size (based on standard errors) for each phenotype and model. Effective sample size was computed per SNP using the following formula:

    (text{Effective},n=left(frac{1}{{{rm{s}}.{rm{e}}.}^{2}}right)times left(frac{{{{rm{s}}.{rm{d}}.}_{{rm{Resid}}}}^{2}}{(2times {rm{MAF}}times (1-{rm{MAF}}))}right))

    where s.e. is the standard error of the SNP effect estimates, MAF is the minor allele frequency, and s.d.Resid is the residual s.d. of the phenotype after covariate adjustment. SNP-based heritability and genetic correlations were calculated separately for the population-based and within-family GWAS models to assess the impact of adjusting for parental genotypes. We examined the attenuation of genetic correlations in within-family analyses compared with population-based models

    Polygenic-score-based approaches

    Polygenic scores (PGSs) were calculated using LDpred291, implemented in the bigsnpr package v.1.12.21 (ref. 92), using the ‘LDPred2-auto’ option and, following established quality control procedures93, based on quality control passing, well-imputed variants in MoBa that were present in an extended set of HAPMAP3+ (ref. 94) variants. These restrictions resulted in a list of 1,310,867 SNPs. We used a precomputed linkage disequilibrium matrix from UK Biobank as the reference linkage disequilibrium panel94. PGS were calculated based on the same external summary statistics used in the LDSC (height44, BMI45, educational attainment46 and smoking initiation47). All PGS were adjusted for the first 20 PCs and genotyping batch.

    Trio-PGS

    We performed trio-PGS analyses in a multiple regression framework, incorporating mothers’, fathers’ and children’s PGS for a given trait in a single model to estimate direct and indirect genetic effects on children’s height (n = 21,008), educational achievement (n = 40,101), sleep duration (n = 21,451) and depressive symptoms (n = 18,212). The models were specified as follows:

    $$begin{array}{c}text{Outcome},=,{beta }_{0}+{beta }_{1}{text{PGS}}_{text{mother}}+{beta }_{2}{text{PGS}}_{text{father}}\ ,+{beta }_{3}{text{PGS}}_{text{child}}+{beta }_{4}text{Sex}+{beta }_{5}text{Age}+varepsilon end{array}$$

    Estimates were calculated with robust standard errors with clustering on maternal IDs to account for the presence of siblings in the data. In sensitivity analyses to assess the sufficiency of this clustering strategy, we restricted to unrelated trios, removing at random all but one trio from families in which children were siblings, half-siblings, or first cousins – leaving the following number of trios per outcome: 14,541 for height, 27,926 for educational achievement, 15,026 for sleep duration, and 12,632 for depressive symptoms. To estimate effects without adjustment for parental indirect genetic effects, we re-ran models without parental PGS.

    Polygenic transmission disequilibrium testing

    PTDT analyses were performed in sub-samples of MoBa families identified by selecting genotyped children with extreme values for height (>2 s.d., above the sample mean; n = 527) educational achievement (>90th percentile; n = 3,227), sleep duration (below 10th percentile; n = 438), and depressive symptoms (>2 s.d. above the sample mean; n = 618), where parental average values on equivalent measures were not similarly extreme. In PTDT, the genetic relationship between different traits or conditions is assessed based on the extent of over- or under-inheritance of genetic propensities by individuals selected based on a trait or condition39. This is calculated as the average deviation of selected children’s PGS from the mid-parental PGS average, scaled to the standard deviation of the mid-parental PGS. Siblings (n = 84 for height, 535 for educational achievement, 78 for sleep duration, and 95 for depressive symptoms) were used as negative controls in these analyses. In our PTDT analyses, we assessed the over- or under-inheritance of genetic variants associated with height44, BMI45, educational attainment46 and smoking initiation47.

    Trio-genome-wide complex trait analysis

    Trio-genome-wide complex trait analysis (trio-GCTA)36 was run to estimate the additive genetic variance attributable to mothers, fathers and children—across all measured SNPs—on measures of height, educational achievement, sleep duration and depressive symptoms. Genomic-related matrices were estimated for complete trios of genetically inferred European ancestry in which children had phenotypic data for at least one of the four phenotypes. In the case of more than one sibling having phenotypic data available, only one sibling was included at random. To remove individuals related across family trios a bottom-up algorithm from the OpenMendel package95 was used with a relatedness threshold of 0.1. The final number of trios, by outcome, was as follows: 12,332 for height, 23,221 for educational achievement, 12,799 for sleep duration and 10,794 for depressive symptoms. Four alternative models were run: (1) estimating the variance explained by the effects of all trio members and the covariance between them (full model); (2) dropping the covariance terms from the full model (no covariance model); (3) estimating the additive genetic effects of the child only (direct effects only model); and (4) estimating no genetic effects (null model). All models included child’s age (in years) at collection of the phenotype, sex, genotyping batch and parents’ first 20 PCs as fixed effect covariates. Akaike’s information criterion was primarily used to select the best model, whereas results of likelihood ratio tests (LRTs) was used to evaluate loss of fit when model parameters were removed. LRT and Bayesian information criterion were consulted for competing models with similar Akaike’s information criterion values. For the likelihood ratio test, we calculated P-values using the classical procedure with a single chi-square null distribution when covariance component parameters were removed. When variance component parameters were removed, we calculated P-values using the equation below, where (hat{lambda }) denotes the test statistic, and j is the number of parameters removed so that the null distribution is a mixture of chi-square distributions dependent on j (refs. 96,97):

    $$P={2}^{-j}mathop{sum }limits_{i=0}^{j}left(genfrac{}{}{0ex}{}{j}{i}right),text{Pr}({chi }_{i}^{2}ge hat{lambda })$$

    LDSC uses GWAS summary data, meaning its power is influenced by both the heritability of the trait and the precision of the underlying GWAS. By contrast, trio-GCTA does not use any summary data and estimates heritability as part of a variance decomposition, using individual-level genetic data. Therefore, although both methods can be used to estimate SNP heritability, their properties (and the underlying data they use) differ substantially

    Assortative mating

    Whereas trio estimates of direct genetic effects tend to be robust to assortative mating and population structure, estimates of indirect genetic effects can reflect not only genetic nurture effects but also assortative mating and population structure. To aid interpretation of the potential impact of assortative mating on the indirect genetic effects estimates, we calculated spousal Pearson correlations with 95% CIs for the four PGS used in the trio-PGS and PTDT analyses

    Reporting summary

    Further information on research design is available in the Nature Portfolio Reporting Summary linked to this article

    Data availability

    Data from MoBa are managed by the Norwegian Institute of Public Health. The MoBa cohort is currently regulated by the Norwegian Health Registry Act. The data that support the findings of this study are available from the Norwegian Institute of Public Health, but restrictions apply to the availability of these data. The MoBa website provides details on how to obtain access for research and the available variables (https://www.fhi.no/en/ch/studies/moba/for-forskere-artikler/research-and-data-access/). Access to individual-level data from MoBa and national registries requires approval from the Regional Committees for Medical and Health Research Ethics (REK), compliance with the General Data Protection Regulation (GDPR) and data holder approval. Researchers must apply through https://helsedata.no/en/. International researchers need to collaborate with a principal investigator employed at a Norwegian research institution—the MoBa administration team (mobaadm@fhi.no) can facilitate contact with relevant research environments. MoBa can be used for health research. The purpose of MoBa is to generate new knowledge about the causes, risk factors, courses and consequences of diseases and health outcomes, with the ultimate goal of improving prevention and treatment. The MoBa administration normally takes 3–5 weeks to reach a decision from the time a complete application has been received. The expected processing time from the decision is in place until data delivery, which is normally between 4 and 5 weeks. Delivery of linked data from health registries usually takes up to 8 weeks. Updated information about processing time is available on the MoBa website (https://www.fhi.no/en/ch/studies/moba/for-forskere-artikler/research-and-data-access/). Information on how to access the MoBaPsychGen post-imputation quality-controlled data is available by contacting the MoBa administration and on the following website: https://www.fhi.no/en/more/research-centres/psychgen/access-to-genetic-data-after-quality-control-by-the-mobapsychgen-pipeline-v/. All summary statistics from GWAS conducted in this study will be available on the MoBa website for GWAS summary statistics: https://www.fhi.no/en/ch/studies/moba/for-forskere-artikler/gwas-data-from-moba/. Participant consent does not allow individual-level data storage in repositories or journals. Moreover, the following external reference panels and linkage disequilibrium resources were used in the MoBaPsychGen pipeline and exemplar trio analyses: the publicly available (access controlled) European Genome–Phenome Archive (Study ID EGAS00001001710) Haplotype Reference Consortium (HRC) release 1.1 (https://ega-archive.org/studies/EGAS00001001710, https://ega-archive.org/access/request-data/how-to-request-data/); the Genome Aggregation Database (gnomAD) v.3.1.2 (https://gnomad.broadinstitute.org/downloads); the MitoImpute reference panel restricted to SNPs with minor allele frequency above 0.1% (https://github.com/sjfandrews/MitoImpute/tree/master); the 1000 Genomes Project Phases 1 and 3 (https://www.internationalgenome.org/data/, https://www.cog-genomics.org/plink/2.0/resources#phase3_1kg); pre-computed linkage disequilibrium scores from the 1000 Genomes Project European superpopulation (https://data.broadinstitute.org/alkesgroup/LDSCORE/eur_w_ld_chr.tar.bz2); and pre-computed linkage disequilibrium matrix from UK Biobank (https://figshare.com/articles/dataset/LD_reference_for_HapMap3_/21305061).

    Code availability

    Scripts used throughout the MoBaPsychGen pipeline are available on GitHub: https://github.com/psychgen/MoBaPsychGen-QC-pipeline. Similarly, the analytic code for the exemplar trio analyses is available in an additional GitHub repository: https://github.com/psychgen/moba-trio-analyses

    References

    1. Loos, R. J. F. 15 years of genome-wide association studies and no signs of slowing down. Nat. Commun.11, 5900 (2020)

      Article 
      ADS 
      CAS 
      PubMed 
      PubMed Central 
      Google Scholar 

    2. Howe, L. J. et al. Within-sibship genome-wide association analyses decrease bias in estimates of direct genetic effects. Nat. Genet.54, 581–592 https://doi.org/10.1038/s41588-022-01062-7 (2022)

      Article 
      CAS 
      PubMed 
      PubMed Central 
      Google Scholar 

    3. Davies, N. M. et al. The importance of family-based sampling for biobanks. Nature634, 795–803 (2024)

      Article 
      ADS 
      CAS 
      PubMed 
      PubMed Central 
      Google Scholar 

    4. McEachan, R. R. C. et al. Cohort profile update: born in Bradford. Int. J. Epidemiol.53, dyae037 (2024)

      Article 
      PubMed 
      PubMed Central 
      Google Scholar 

    5. Milbourn, H. et al. Generation Scotland: an update on Scotland’s longitudinal family health study. BMJ Open14, e084719 (2024)

      Article 
      PubMed 
      PubMed Central 
      Google Scholar 

    6. Fitzsimons, E. et al. Collection of genetic data at scale for a nationally representative population: the UK millennium cohort study. Longit. Life Course Stud.13, 169–187 (2022)

      Article 
      Google Scholar 

    7. Golding, J., Pembrey, M. & Jones, R. The ALSPAC Study Team. ALSPAC–The Avon Longitudinal Study of Parents and Children. Paediatr. Perinat. Epidemiol. 15, 74–87 (2001)

    8. Jaddoe, V. W. V. et al. The generation R study: design and cohort profile. Eur. J. Epidemiol.21, 475–484 (2006)

      Article 
      PubMed 
      Google Scholar 

    9. Brumpton, B. M. et al. The HUNT study: a population-based cohort for genetic research. Cell Genom.2, 100193 (2022)

      Article 
      CAS 
      PubMed 
      PubMed Central 
      Google Scholar 

    10. Gudbjartsson, D. F. et al. Large-scale whole-genome sequencing of the Icelandic population. Nat. Genet.47, 435–444 (2015)

      Article 
      CAS 
      PubMed 
      Google Scholar 

    11. Andersson, C., Johnson, A. D., Benjamin, E. J., Levy, D. & Vasan, R. S. 70-year legacy of the Framingham Heart Study. Nat. Rev. Cardiol.16, 687–698 (2019)

      Article 
      PubMed 
      Google Scholar 

    12. Ziyatdinov, A. et al. Genotyping, sequencing and analysis of 140,000 adults from Mexico City. Nature622, 784–793 (2023)

      Article 
      ADS 
      CAS 
      PubMed 
      PubMed Central 
      Google Scholar 

    13. Walters, R. G. et al. Genotyping and population characteristics of the China Kadoorie Biobank. Cell Genom.3, 100361 (2023)

      Article 
      CAS 
      PubMed 
      PubMed Central 
      Google Scholar 

    14. Kurki, M. I. et al. FinnGen provides genetic insights from a well-phenotyped isolated population. Nature613, 508–518 (2023)

      Article 
      ADS 
      CAS 
      PubMed 
      PubMed Central 
      Google Scholar 

    15. Bycroft, C. et al. The UK Biobank re2018)

      Article 
      ADS 
      CAS 
      PubMed 
      PubMed Central 
      Google Scholar 

    16. The All of Us Research Program Genomics Investigators. Genomic data in the All of Us Research Program. Nature627, 340–346 (2024)

      Article 
      CAS 
      Google Scholar 

    17. Verma, A. et al. Diversity and scale: genetic architecture of 2068 traits in the VA Million Veteran Program. Science385, eadj1182 (2024)

      Article 
      CAS 
      PubMed 
      PubMed Central 
      Google Scholar 

    18. Zhou, W. et al. Global Biobank Meta-analysis Initiative: powering genetic discovery across human disease. Cell Genom.2, 100192 (2022)

      Article 
      CAS 
      PubMed 
      PubMed Central 
      Google Scholar 

    19. Nagai, A. et al. Overview of the BioBank Japan project: study design and profile. J. Epidemiol.27, S2–S8 (2017)

      Article 
      PubMed 
      PubMed Central 
      Google Scholar 

    20. Dummer, T. J. B. et al. The Canadian Partnership for Tomorrow Project: a pan-Canadian platform for research on chronic disease prevention. Can. Med. Assoc. J.190, E710–E717 (2018)

      Article 
      Google Scholar 

    21. Kim, Y. & Han, B.-G. The KoGES Group. Cohort profile: the Korean Genome and Epidemiology Study (KoGES) consortium. Int. J. Epidemiol. 46, e20 (2017)

    22. Feng, Y.-C. A. et al. Taiwan biobank: a rich biomedical research database of the Taiwanese population. Cell Genom.2, 100197 (2022)

      Article 
      CAS 
      PubMed 
      PubMed Central 
      Google Scholar 

    23. Psaty, B. M. et al. Cohorts for Heart and Aging Research in Genomic Epidemiology (CHARGE) consortium: design of prospective meta-analyses of genome-wide association studies from 5 cohorts. Circ. Cardiovasc. Genet.2, 73–80 (2009)

      Article 
      PubMed 
      PubMed Central 
      Google Scholar 

    24. Kim, S. & Xing, E. P. Statistical Estimation of Correlated Genome Associations to a Quantitative Trait Network. PLoS Genet.5, e1000587 (2009)

      Article 
      PubMed 
      PubMed Central 
      Google Scholar 

    25. Dudbridge, F. Power and predictive accuracy of polygenic risk scores. PLoS Genet.9, e1003348 (2013)

      Article 
      CAS 
      PubMed 
      PubMed Central 
      Google Scholar 

    26. Wang, X. et al. Polygenic risk prediction: why and when out-of-sample prediction R2 can exceed SNP-based heritability. Am. J. Hum. Genet.110, 1207–1215 (2023)

      Article 
      CAS 
      PubMed 
      PubMed Central 
      Google Scholar 

    27. Ni, G. et al. A comparison of ten polygenic score methods for psychiatric disorders applied across multiple cohorts. Biol. Psychiatry90, 611–620 (2021)

      Article 
      PubMed 
      PubMed Central 
      Google Scholar 

    28. Davey Smith, G. & Hemani, G. Mendelian randomization: genetic anchors for causal inference in epidemiological studies. Hum. Mol. Genet.23, R89–R98 (2014)

      Article 
      CAS 
      PubMed 
      PubMed Central 
      Google Scholar 

    29. Davey Smith, G. & Ebrahim, S. ‘Mendelian randomization’: can genetic epidemiology contribute to understanding environmental determinants of disease? Int. J. Epidemiol.32, 1–22 (2003)

      Article 
      Google Scholar 

    30. Benyamin, B., Visscher, P. M. & McRae, A. F. Family-based genome-wide association studies. Pharmacogenomics10, 181–190 (2009)

      Article 
      CAS 
      PubMed 
      Google Scholar 

    31. Kong, A. et al. The nature of nurture: effects of parental genotypes. Science359, 424–428 (2018)

      Article 
      ADS 
      CAS 
      PubMed 
      Google Scholar 

    32. Zhang, G. et al. Assessing the causal relationship of maternal height on birth size and gestational age at birth: a Mendelian randomization analysis. PLoS Med.12, e1001865 (2015)

      Article 
      PubMed 
      PubMed Central 
      Google Scholar 

    33. Young, A. I., Benonisdottir, S., Przeworski, M. & Kong, A. Deconstructing the1396–1400 (2019)

      Article 
      ADS 
      CAS 
      PubMed 
      PubMed Central 
      Google Scholar 

    34. Lee, J. J. et al. Gene discovery and polygenic prediction from a genome-wide association study of educational attainment in 1.1 million individuals. Nat. Genet.50, 1112–1121 (2018)

      Article 
      CAS 
      PubMed 
      PubMed Central 
      Google Scholar 

    35. Tan, T. et al. Family-GWAS reveals effects of environment and mating on genetic associations. Preprint at medRxivhttps://doi.org/10.1101/2024.10.01.24314703 (2024)

    36. Eilertsen, E. M. et al. Direct and indirect effects of maternal, paternal, and offspring genotypes: Trio-GCTA. Behav. Genet.51, 154–161 (2021)

      Article 
      PubMed 
      Google Scholar 

    37. Eaves, L. J., St Pourcain, B., Smith, G. D., York, T. P. & Evans, D. M. Resolving the effects of maternal and offspring genotype on dyadic outcomes in genome wide complex trait analysis (“M-GCTA”). Behav. Genet.44, 445–455 (2014)

      Article 
      PubMed 
      PubMed Central 
      Google Scholar 

    38. Pingault, J.-B. et al. Genetic nurture versus genetic transmission of risk for ADHD traits in the Norwegian Mother Father and Child Cohort Study. Mol. Psychiatry21, 1731–1738 https://doi.org/10.1038/s41380-022-01863-6 (2022)

      Article 
      Google Scholar 

    39. Weiner, D. J. et al. Polygenic transmission disequilibrium confirms that common and rare variation act additively to create risk for autism spectrum disorders. Nat. Genet.49, 978–985 (2017)

      Article 
      CAS 
      PubMed 
      PubMed Central 
      Google Scholar 

    40. Torvik, F. A. et al. Modeling assortative mating and genetic similarities between partners, siblings, and in-laws. Nat. Commun.13, 1108 (2022)

      Article 
      ADS 
      CAS 
      PubMed 
      PubMed Central 
      Google Scholar 

    41. Davies, N. M. et al. Within family Mendelian randomization studies. Hum. Mol. Genet.28, R170–R179 (2019)

      Article 
      CAS 
      PubMed 
      Google Scholar 

    42. Hwang, L.-D. et al. DINGO: increasing the power of locus discovery in maternal and fetal genome-wide association studies of perinatal traits. Nat. Commun.15, 9255 (2024)

      Article 
      ADS 
      CAS 
      PubMed 
      PubMed Central 
      Google Scholar 

    43. Brandlistuen, R. E. et al. Cohort profile update: the Norwegian Mother Father and Child Cohort (MoBa). Int. J. Epidemiol.54, dyaf139 (2025)

      Article 
      PubMed 
      PubMed Central 
      Google Scholar 

    44. Wood, A. R. et al. Defining the role of common variation in the genomic and biological architecture of adult human height. Nat. Genet.46, 1173–1186 (2014)

      Article 
      CAS 
      PubMed 
      PubMed Central 
      Google Scholar 

    45. Yengo, L. et al. Meta-analysis of genome-wide association studies for height and body mass index in ∼700000 individuals of European ancestry. Hum. Mol. Genet.27, 3641–3649 (2018)

      Article 
      CAS 
      PubMed 
      PubMed Central 
      Google Scholar 

    46. Okbay, A. et al. Polygenic prediction of educational attainment within and between families from genome-wide association analyses in 3 million individuals. Nat. Genet.54, 437–449 (2022)

      Article 
      CAS 
      PubMed 
      PubMed Central 
      Google Scholar 

    47. Karlsson Linnér, R. et al. Genome-wide association analyses of risk tolerance and risky behaviors in over 1 million individuals identify hundreds of loci and shared genetic influences. Nat. Genet.51, 245–257 (2019)

      Article 
      PubMed 
      PubMed Central 
      Google Scholar 

    48. Adams, M. J. et al. Trans-ancestry genome-wide study of depression identifies 697 associations implicating cell types and pharmacotherapies. Cell188, 640–652 https://doi.org/10.1016/j.cell.2024.12.002 (2025)

      Article 
      CAS 
      Google Scholar 

    49. Yang, J., Lee, S. H., Goddard, M. E. & Visscher, P. M. GCTA: a tool for genome-wide complex trait analysis. Am. J. Hum. Genet.88, 76–82 (2011)

      Article 
      CAS 
      PubMed 
      Google Scholar 

    50. O’Connell, K. S., Genomics yields biological and phenotypic insights into bipolar disorder. Nature639, 968–975 https://doi.org/10.1038/s41586-024-08468-9 (2025)

    51. Cheesman, R. et al. How important are parents in the development of child anxiety and depression? A genomic analysis of parent-offspring trios in the Norwegian Mother Father and Child Cohort Study (MoBa). BMC Med.18, 284 (2020)

      Article 
      PubMed 
      PubMed Central 
      Google Scholar 

    52. Huang, Q. Q. et al. Examining the role of common variants in rare neurodevelopmental conditions. Nature636, 404–411 (2024)

      Article 
      ADS 
      CAS 
      PubMed 
      PubMed Central 
      Google Scholar 

    53. Nivard, M. G. et al. More than nature and nurture, indirect genetic effects on children’s academic achievement are consequences of dynastic social processes. Nat. Hum. Behav.8, 771–778 (2024)

      Article 
      PubMed 
      PubMed Central 
      Google Scholar 

    54. Magnus, P. et al. Cohort profile update: the Norwegian Mother and Child Cohort Study (MoBa). Int. J. Epidemiol.45, 1146–1150 (2016)

      Article 
      Google Scholar 

    55. Biele, G. et al. Bias from self selection and loss to follow-up in prospective cohort studies. Eur. J. Epidemiol.34, 927–938 (2019)

      Article 
      CAS 
      PubMed 
      Google Scholar 

    56. Rayner, C. et al. Quantifying and adjusting for selection biases in the Norwegian Mother, Father and Child Cohort Study using population-wide individual-level registry information. Int. J. Epidemiol.55, dyag122 https://doi.org/10.1093/ije/dyag122 (2026)

    57. Sharp, C., Goodyer, I. M. & Croudace, T. J. The Short Mood and Feelings Questionnaire (SMFQ): a unidimensional item response theory and categorical data factor analysis of self-report ratings from a community sample of 7-through 11-year-old children. J. Abnorm. Child Psychol.34, 365–377 (2006)

      Article 
      Google Scholar 

    58. Hannigan, L. J. et al. phenotools: an R package to facilitate efficient and reproducible use of phenotypic data from MoBa and linked registry/6G8BJ (2023)

    59. R Core Team. R: A Language and Environment for Statistical Computing (R Foundation for Statistical Computing, 2020)

    60. UNESCO Institute for Statistics. International Standard Classification of Education, ISCED 2011 (UNESCO, 2012)

    61. Nordin, M., Åkerstedt, T. & Nordin, S. Psychometric evaluation and normative data for the Karolinska Sleep Questionnaire. Sleep Biol. Rhythms11, 216–226 https://doi.org/10.1111/sbr.12024 (2013)

      Article 
      Google Scholar 

    62. Tambs, K. & Røysamb, E. Selection of questions to short-form versions of original psychometric instruments in MoBa. Nor. Epidemiol.24, 195–201 (2014)

      Google Scholar 

    63. Irgens, L. M. The Medical Birth Registry of Norway. Epidemiological research and surveillance throughout 30 years. Acta Obstet. Gynecol. Scand.79, 435–439 (2000)

      Article 
      CAS 
      PubMed 
      Google Scholar 

    64. Paltiel, L. et al. The biobank of the Norwegian mother and child cohort study – present status. Nor. Epidemiol.24, 29–35 (2014)

      Google Scholar 

    65. Norwegian Institute of Public Health. Protocols for the Norwegian Mother, Father and Child Cohort Study (MoBa). NIPHhttps://www.fhi.no/en/publ/2012/protocols-for-moba/ (2025)

    66. Corfield, E. C. et al. The Norwegian Mother, Father, and Child cohort study (MoBa) genotyping data reg/10.1101/2022.06.23.496289 (2022)

    67. Pedigree Imputation Consortium Pipeline (PICOPILI). Nealelab / picopili. GitHubhttps://github.com/Nealelab/picopili (2016)

    68. Anderson, C. A. et al. Data quality control in genetic case-control association studies. Nat. Protoc.5, 1564–1573 (2010)

      Article 
      CAS 
      PubMed 
      PubMed Central 
      Google Scholar 

    69. Lam, M. et al. RICOPILI: Rapid Imputation for COnsortias PIpeLIne. Bioinformatics36, 930–933 (2020)

      Article 
      CAS 
      PubMed 
      PubMed Central 
      Google Scholar 

    70. Turner, S. et al. Quality control procedures for genome-wide association studies. Curr. Protoc. Hum. Genet.68, 1.19.11–1.19.18 (2011)

      Google Scholar 

    71. König, I. R., Loley, C., Erdmann, J. & Ziegler, A. How to include chromosome X in your genome-wide association study. Genet. Epidemiol.38, 97–103 (2014)

      Article 
      PubMed 
      Google Scholar 

    72. Chang, C. C. et al. PLINK: rising to the challenge of larger and richer datasets. GigaScience4, s13742-015-0047-8 https://doi.org/10.1186/s13742-015-0047-8 (2015)

      Article 
      CAS 
      Google Scholar 

    73. Center for Statistical Genetics. LiftOver. Center for Statistical Geneticshttps://genome.sph.umich.edu/wiki/LiftOver#Lift_genome_positions (2015)

    74. Hinrichs, A. S. et al. The UCSC Genome Browser Database: update 2006. Nucleic Acids Res.34, D590–D598 (2006)

      Article 
      CAS 
      PubMed 
      PubMed Central 
      Google Scholar 

    75. Manichaikul, A. et al. Robust relationship inference in genome-wide association studies. Bioinformatics26, 2867–2873 (2010)

      Article 
      CAS 
      PubMed 
      PubMed Central 
      Google Scholar 

    76. Abraham, G., Qiu, Y. & Inouye, M. FlashPCA2: principal component analysis of Biobank-scale genotype datasets. Bioinformatics33, 2776–2778 (2017)

      Article 
      CAS 
      PubMed 
      Google Scholar 

    77. McCarthy Group Tools. HRC or 1000G imputation preparation and checking, v.4.2.13. McCarthy Group Toolshttps://www.well.ox.ac.uk/~wrayner/tools/ (2011)

    78. Delaneau, O., Marchini, J. & Zagury, J.-F. A linear complexity phasing method for thousands of genomes. Nat. Methods9, 179–181 (2011)

      Article 
      PubMed 
      Google Scholar 

    79. O’Connell, J. et al. A general approach for haplotype phasing across the full spectrum of relatedness. PLoS Genet.10, e1004234 (2014)

      Article 
      PubMed 
      PubMed Central 
      Google Scholar 

    80. IMPUTE 4, version impute4.1.2_r300.3. jmarchinihttps://jmarchini.org/software/#impute-4 (2016)

    81. QCTOOL v.2.08. qctool v2https://www.well.ox.ac.uk/~gav/qctool_v2/ (2020)

    82. ic, a post-imputation data checking program v.1.0.8. McCarthy Group Toolshttps://www.well.ox.ac.uk/~wrayner/tools/Post-Imputation.html (2020)

    83. Howie, B. N., Donnelly, P. & Marchini, J. A flexible and accurate genotype imputation method for the next generation of genome-wide association studies. PLoS Genet.5, e1000529 (2009)

      Article 
      PubMed 
      PubMed Central 
      Google Scholar 

    84. Poznik, G. D. Identifying Y-chromosome haplogroups in arbitrarily large samples of sequenced or genotyped men. Preprint at bioRxivhttps://doi.org/10.1101/088716 (2016)

    85. McInerney, T. W. et al. A globally diverse reference alignment and panel for imputation of mitochondrial DNA variants. BMC Bioinformatics22, 417 (2021)

      Article 
      CAS 
      PubMed 
      PubMed Central 
      Google Scholar 

    86. McInerney, T. W. et al. sjfandrews/MitoImpute: v0.2 v. v0.2. GitHubhttps://github.com/sjfandrews/MitoImpute (2020)

    87. Danecek, P. et al. Twelve years of SAMtools and BCFtools. GigaScience10, giab008 (2021)

      Article 
      PubMed 
      PubMed Central 
      Google Scholar 

    88. Speed, D., Hemani, G., Johnson, M. R. & Balding, D. J. Improved heritability estimation from genome-wide SNPs. Am. J. Hum. Genet.91, 1011–1021 (2012)

      Article 
      CAS 
      PubMed 
      PubMed Central 
      Google Scholar 

    89. Dashti, H. S. et al. Genome-wide association study identifies genetic loci for self-reported habitual sleep duration supported by accelerometer-derived estimates. Nat. Commun.10, 1100 (2019)

      Article 
      ADS 
      PubMed 
      PubMed Central 
      Google Scholar 

    90. Bulik-Sullivan, B. K. et al. LD score regression distinguishes confounding from polygenicity in genome-wide association studies. Nat. Genet.47, 291–295 (2015)

      Article 
      CAS 
      PubMed 
      PubMed Central 
      Google Scholar 

    91. Privé, F., Arbel, J. & Vilhjálmsson, B. J. LDpred2: better, faster, stronger. Bioinformatics36, 5424–5431 (2021)

      Article 
      PubMed 
      PubMed Central 
      Google Scholar 

    92. Privé, F., Aschard, H., Ziyatdinov, A. & Blum, M. G. B. Efficient analysis of large-scale genome-wide data with two R packages: bigstatsr and bigsnpr. Bioinformatics34, 2781–2787 (2018)

      Article 
      PubMed 
      PubMed Central 
      Google Scholar 

    93. Privé, F., Arbel, J., Aschard, H. & Vilhjálmsson, B. J. Identifying and correcting for misspecifications in GWAS summary statistics and polygenic scores. HGG Adv.3, 100136 (2022)

      PubMed 
      PubMed Central 
      Google Scholar 

    94. Privé, F., Albiñana, C., Arbel, J., Pasaniuc, B. & Vilhjálmsson, B. J. Inferring disease architecture and predictive ability with LDpred2-auto. Am. J.Hum. Genet.110, 2042–2055 (2023)

      Article 
      PubMed 
      PubMed Central 
      Google Scholar 

    95. Zhou, H. et al. OpenMendel: a cooperative programming project for statistical genetics. Hum. Genet.139, 61–71 (2020)

      Article 
      PubMed 
      Google Scholar 

    96. Verbeke, G. & Molenberghs, G. The use of score tests for inference on variance components. Biometrics59, 254–262 (2003)

      Article 
      ADS 
      MathSciNet 
      PubMed 
      Google Scholar 

    97. SAS Institute. The Glimmix Procedure. in SAS/STAT 9.1: User’s Guide 3233–3237 (SAS Institute, 2004)

    Download references

    Acknowledgements

    We are grateful to all the participating families in Norway who take part in this continuing cohort study. This work was performed on the Tjeneste for Sensitive Data (TSD) facilities, owned by the University of Oslo, operated and developed by the TSD service group at the University of Oslo, IT Department (USIT), using resources provided by Sigma2—the National Infrastructure for High Performance Computing and Data Storage in Norway (UNINETT). We thank V. Appadurai and S. Ripke for advising on quality control, phasing and imputation. We thank S. Barbo Valand for his valuable contributions to the data management. MoBa is supported by the Norwegian Ministry of Health and Care Services, and the Ministry of Education and Research. The genotype data were provided by the HARVEST Collaboration (supported by the Research Council of Norway (RCN) 229624), the NORMENT Centre (RCN 223273, South-Eastern Norway Regional Health Authority (HSØ) and Stiftelsen Kristian Gerhard Jebsen) in collaboration with deCODE Genetics, and the Center for Diabetes Research at the University of Bergen (funded by the ERC AdG project SELECTionPREDISPOSED, Stiftelsen Kristian Gerhard Jebsen, Trond Mohn Foundation, the RCN, the Novo Nordisk Foundation, the University of Bergen and the Western Norway Regional Health Authority).

    Funding

    The genotype data quality control and imputation described in this manuscript was supported by the RCN (nos. 223273, 274611, 273291, 296030, 324252, 324499, 335258 and 322672), the HSØ (nos. 2020022, 2021045, 2020024, 2018058, 2019097 and 2022073), the Horizon 2020 Research and Innovation programme of the European Union (no. 847776), and the NordForsk project (no. 164218). This work was supported by the Horizon 2020 Research and Innovation programme of the European Union (FAMILY, grant agreement no. 101057529). E.C.C. was supported by the RCN (no. 274611) and HSØ fellowship (no. 2021045). T.T.F. was supported by the RCN (no. 344121) and the Horizon 2020 Research and Innovation programme of the European Union (no. 964874). L.H. was supported by funding from RCN (no. 336085) and HSØ (no. 2020022). R.E.W. was supported by a postdoctoral fellowship from the HSØ (no. 2020024). C.A. was supported by a UKRI ESRC Studentship awarded by UCL, Bloomsbury and East London Doctoral Training Partnership (ES/P000592/1). L.T.W. was supported by the Horizon 2020 Research and Innovation programme of the European Union (ERC StG, no. 802998), the RCN (no. 300767), and the H SØ (no. 2019101). P.R.N. was funded by the ERC AdG project SELECTionPREDISPOSED (no. 293574), the Stiftelsen Kristian Gerhard Jebsen, Trond Mohn Foundation (TMS2022TMT01), the RCN (no. 240413), the Novo Nordisk Foundation (no. NNF18OC0054741), the Danish Diabetes and Endocrinology Agency (no. VP015-25), the University of Bergen, and the Western Norway Regional Health Authority. T.Z. was funded by R37MH107649-07S1 (NIH) and by the RCN (no. 288083). H.A. was supported by the RCN (no. 324620) and NordForsk (no. 156298). T.R.-K. was supported by RCN (no. 274611). G.H. was funded by the Wellcome Trust [208806/Z/17/Z]. N.M.D. was supported by the RCN (no. 295989), the MRC MR/V002147/1, and an NIMH grant (MH130448). L.J.H. was supported by funding from HSØ (nos. 2022083 and 2019097). O.A.A. was supported by the Stiftelsen Kristian Gerhard Jebsen (SKGJ-MED-021), RCN (no. 324252, 326813, 324499, 322672), Nordforsk (no. 164218), Horizon 2020 Research and Innovation programme of the European Union (RealMent, no. 964874). A.K.H. was supported by the RCN (nos. 274611, 288083, 336085 and 300668), HSØ (nos. 2020022, 2019097, 2018059 and 2026069) and the Horizon 2020 Research and Innovation programme of the European Union (FAMILY, grant agreement no. 101057529; Marie Skłodowska-Curie grant ESSGN no. 101073237). E.C.C., R.E.W., A.M.H., G.H., L.J.H. and A.K.H. are members of the Medical Research Council (MRC) Integrative Epidemiology Unit at the University of Bristol, which is supported by the MRC and the University of Bristol (MC_UU_00032/1, MC_UU_00032/7). A.K.H, N.M.D., G.H. and I.B. were supported by the MRC UKRI1510. Views and opinions expressed are those of the authors and do not necessarily reflect those of the funders.

    Author information

    Author notes

    1. These authors jointly supervised this work: Ole A. Andreassen, Alexandra Karoline Havdahl

    Authors and Affiliations

    1. PsychGen Centre for Genetic Epidemiology and Mental Health, Norwegian Institute of Public Health, Oslo, Norway

      Elizabeth C. Corfield, Laura Hegemann, Robyn E. Wootton, Ragnhild E. Brandlistuen, Helga Ask, Ted Reichborn-Kjennerud, Laurie J. Hannigan & Alexandra Karoline Havdahl

    2. Psychiatric Genetic Epidemiology Group, Research Department & Nic Waals Institute, Lovisenberg Diaconal Hospital, Oslo, Norway

      Elizabeth C. Corfield, Laura Hegemann, Robyn E. Wootton, Laurie J. Hannigan & Alexandra Karoline Havdahl

    3. MRC Integrative Epidemiology Unit, Bristol Medical School, University of Bristol, Bristol, UK

      Elizabeth C. Corfield, Robyn E. Wootton, Amanda M. Hughes, Gibran Hemani, Laurie J. Hannigan & Alexandra Karoline Havdahl

    4. Population Health Sciences, Bristol Medical School, University of Bristol, Bristol, UK

      Elizabeth C. Corfield, Amanda M. Hughes, Gibran Hemani, Laurie J. Hannigan & Alexandra Karoline Havdahl

    5. Centre for Precision Psychiatry, Division of Mental Health and Addiction, Oslo University Hospital and University of Oslo, Oslo, Norway

      Alexey A. Shadrin, Oleksandr Frei, Zillur Rahman, Tahir Tekin Filiz, Aihua Lin, Lavinia Athanasiu, Espen Hagen, Lars T. Westlye & Ole A. Andreassen

    6. KG Jebsen Centre for Neurodevelopmental disorders, University of Oslo, Oslo, Norway

      Alexey A. Shadrin, Lars T. Westlye & Ole A. Andreassen

    7. Department of Pharmacy, University of Oslo, Oslo, Norway

      Oleksandr Frei & Bayram Cevdet Akdeniz

    8. Department of Tumor biology, Institute for Cancer Research, Oslo University Hospital, Oslo, Norway

      Bayram Cevdet Akdeniz & Eivind Hovig

    9. Division of Psychiatry, University College London, London, UK

      Isabella Badini & Neil M. Davies

    10. School of Psychological Science, University of Bristol, Bristol, UK

      Robyn E. Wootton

    11. Research Department of Clinical, Educational and Health Psychology, University College London, London, UK

      Chloe Austerberry

    12. Centre for Child, Adolescent and Family Research, University of Cambridge, Cambridge, UK

      Chloe Austerberry

    13. Department of Mental Health and Suicide, Norwegian Institute of Public Health, Oslo, Norway

      Martin Tesli

    14. Division of Mental Health and Addiction, Oslo University Hospital, Oslo, Norway

      Martin Tesli

    15. Promenta Research Center, Department of Psychology, University of Oslo, Oslo, Norway

      Ragnhild E. Brandlistuen, Espen Moen Eilertsen, Tetyana Zayats, Helga Ask & Alexandra Karoline Havdahl

    16. Department of Psychology, University of Oslo, Oslo, Norway

      Lars T. Westlye

    17. Department of Clinical Science, University of Bergen, Bergen, Norway

      Pål R. Njølstad

    18. Department of Children and Adolescent Medicine, Haukeland University Hospital, Bergen, Norway

      Pål R. Njølstad

    19. Centre for Fertility and Health, Norwegian Institute of Public Health, Oslo, Norway

      Per Magnus

    20. Department of Informatics University of Oslo, Oslo, Norway

      Eivind Hovig

    21. Analytic and Translational Unit, Massachusetts General Hospital, Boston, MA, USA

      Tetyana Zayats

    22. Stanley Center for Psychiatric Disorders, Broad Institute of MIT and Harvard, Boston, MA, USA

      Tetyana Zayats

    23. Department of Child Health and Development, Norwegian Institute of Public Health, Oslo, Norway

      Helga Ask

    24. Institute of Clinical Medicine, University of Oslo, Oslo, Norway

      Ted Reichborn-Kjennerud

    25. Department of Statistical Science, University College London, London, UK

      Neil M. Davies

    26. Department of Public Health and Nursing, Norwegian University of Science and Technology, Trondheim, Norway

      Neil M. Davies

    Authors

    1. Elizabeth C. CorfieldView author publications

      Search author on:PubMed Google Scholar

    2. Alexey A. ShadrinView author publications

      Search author on:PubMed Google Scholar

    3. Oleksandr FreiView author publications

      Search author on:PubMed Google Scholar

    4. Zillur RahmanView author publications

      Search author on:PubMed Google Scholar

    5. Bayram Cevdet AkdenizView author publications

      Search author on:PubMed Google Scholar

    6. Tahir Tekin FilizView author publications

      Search author on:PubMed Google Scholar

    7. Aihua LinView author publications

      Search author on:PubMed Google Scholar

    8. Isabella BadiniView author publications

      Search author on:PubMed Google Scholar

    9. Laura HegemannView author publications

      Search author on:PubMed Google Scholar

    10. Lavinia AthanasiuView author publications

      Search author on:PubMed Google Scholar

    11. Robyn E. WoottonView author publications

      Search author on:PubMed Google Scholar

    12. Chloe AusterberryView author publications

      Search author on:PubMed Google Scholar

    13. Amanda M. HughesView author publications

      Search author on:PubMed Google Scholar

    14. Martin TesliView author publications

      Search author on:PubMed Google Scholar

    15. Espen HagenView author publications

      Search author on:PubMed Google Scholar

    16. Ragnhild E. BrandlistuenView author publications

      Search author on:PubMed Google Scholar

    17. Espen Moen EilertsenView author publications

      Search author on:PubMed Google Scholar

    18. Lars T. WestlyeView author publications

      Search author on:PubMed Google Scholar

    19. Per MagnusView author publications

      Search author on:PubMed Google Scholar

    20. Eivind HovigView author publications

      Search author on:PubMed Google Scholar

    21. Tetyana ZayatsView author publications

      Search author on:PubMed Google Scholar

    22. Helga AskView author publications

      Search author on:PubMed Google Scholar

    23. Ted Reichborn-KjennerudView author publications

      Search author on:PubMed Google Scholar

    24. Gibran HemaniView author publications

      Search author on:PubMed Google Scholar

    25. Neil M. DaviesView author publications

      Search author on:PubMed Google Scholar

    26. Laurie J. HanniganView author publications

      Search author on:PubMed Google Scholar

    27. Ole A. AndreassenView author publications

      Search author on:PubMed Google Scholar

    28. Alexandra Karoline HavdahlView author publications

      Search author on:PubMed Google Scholar

    Contributions

    Throughout the project, the International Committee of Medical Journal Editors authorship guidelines were used to identify authors for this paper. Authorship contributions (including duration of time working on the project) were used to determine authorship order and were defined using the Contributor Role Taxonomy. E.C.C., T.Z., G.H., N.M.D., L.J.H., O.A.A. and A.K.H. conceptualized the study. E.C.C., L.A., E. Hagen, H.A. and L.J.H. curated the data. E.C.C., A.A.S., O.F., Z.R., B.C.A., T.T.F., A.L., I.B., L.H., C.A., T.Z. and L.J.H. conducted the formal analysis. P.R.N., P.M., T.R.-K., O.A.A. and A.K.H. helped with funding acquisition. E.C.C., A.A.S., O.F., Z.R., T.T.F. and T.Z. conducted the investigation. E.C.C., A.A.S., T.T.F., A.M.H., T.Z., G.H., N.M.D., O.A.A. and A.K.H. devised the methodology. E.C.C., L.T.W., T.Z., O.A.A. and A.K.H. helped with project administration. P.R.N., P.M., E. Hovig, T.R.-K., O.A.A. and A.K.H. sourced the resources. E.C.C., A.A.S., O.F., Z.R., B.C.A., T.T.F., A.L., E.M.E. and T.Z. handled the software. E.C.C., A.A.S., O.F., Z.R., B.C.A., T.T.F., T.Z. and G.H. performed the validation. E.C.C., A.A.S., O.F., Z.R., A.L., I.B., L.H., C.A., M.T., R.E.B., T.Z., H.A. G.H., L.J.H. and N.M.D. conducted the visualization. E.C.C., A.A.S., O.F., T.T.F., I.B., L.H., R.E.W., H.A., N.M.D., L.J.H. and A.K.H. wrote the original draft. All authors reviewed and edited the manuscript. O.A.A. and A.K.H. jointly supervised the work.

    Ethics declarations

    Competing interests

    A.K.H. has served as Scientific Director of the MoBa cohort at the Norwegian Institute of Public Health (NIPH) since September 2025, and R.E.B. and P.M. have previously served as Scientific Director of the MoBa cohort at NIPH. O.A.A. has received speaker fees from Lundbeck, Janssen, BMS, Lilly, Otsuka and Medice and is a consultant to Precision Health and Ledidi. O.F. is a consultant to Precision Health. M.T. has received speaker fees from Lundbeck and Otsuka. All other authors report no competing interests.

    Peer review

    Peer review information

    Nature thanks Andrea Ganna, Scott Vrieze and Zhiyu Yang and the other, anonymous, reviewer(s) for their contribution to the peer review of this work. Peer reviewer reports are available

    Additional information

    Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations

    Extended data figures and tables

    Extended Data Fig. 1 Trio- and population genome-wide association studies (GWAS) of MoBa children’s height (n = 23,786 trios)

    Panel a, Miami plot showing -log10 p-values for associations between individual genetic variants and MoBa children’s height in trio- (blue) and population-based (orange) GWAS models. The blue line is the genome-wide significance threshold (p-value < 5 × 10−8) for common variants in samples with genetically inferred European ancestry, while the dashed grey line is the suggestive threshold (p-value < 1 × 10−5); b. Quantile-quantile (QQ) plots showing the distribution of p-values from the population-based (left) and conditional (right) GWAS models. The red diagonal line illustrates the expected distribution of p-values under the null hypothesis of no association.

    Extended Data Fig. 2 Trio- and population genome-wide association studies (GWAS) of MoBa children’s educational achievement (n = 41,761 trios)

    Panel a, Miami plot showing -log10 p-values for associations between individual genetic variants and MoBa children’s educational achievement in trio- (blue) and population-based (orange) GWAS models. The blue line is the genome-wide significance threshold (p-value < 5 × 10−8) for common variants in samples with genetically inferred European ancestry, while the dashed grey line is the suggestive threshold (p-value < 1 × 10−5); b. Quantile-quantile (QQ) plots showing the distribution of p-values from the population-based (left) and conditional (right) GWAS models for educational achievement. The red diagonal line illustrates the expected distribution of p-values under the null hypothesis of no association.

    Extended Data Fig. 3 Trio- and population genome-wide association studies (GWAS) of MoBa children’s sleep duration (n = 24,376 trios)

    Panel a, Miami plot showing -log10 p-values for associations between individual genetic variants and MoBa children’s sleep duration in trio- (blue) and population-based (orange) GWAS models. The blue line is the genome-wide significance threshold (p-value < 5 × 10−8) for common variants in samples with genetically inferred European ancestry, while the dashed grey line is the suggestive threshold (p-value < 1 × 10−5); b. Quantile-quantile (QQ) plots showing the distribution of p-values from the population-based (left) and conditional (right) GWAS models for sleep duration. The red diagonal line illustrates the expected distribution of p-values under the null hypothesis of no association.

    Extended Data Fig. 4 Trio- and population genome-wide association studies (GWAS) of MoBa children’s depression symptoms (n = 18,701 trios)

    Panel a, Miami plot showing -log10 p-values for associations between individual genetic variants and MoBa children’s depressive symptoms in trio- (blue) and population-based (orange) GWAS models. The blue line is the genome-wide significance threshold (p-value < 5 × 10−8) for common variants in samples with genetically inferred European ancestry, while the dashed grey line is the suggestive threshold (p-value < 1 × 10−5); b. Quantile-quantile (QQ) plots showing the distribution of p-values from the population-based (left) and conditional (right) GWAS models for depression symptoms. The red diagonal line illustrates the expected distribution of p-values under the null hypothesis of no association.

    Extended Data Fig. 5 Within-family polygenic score (PGS) based sensitivity analyses

    Panel a, Estimates of direct genetic effects on MoBa children’s height, educational achievement, sleep duration, and depressive symptoms from child only- and trio- polygenic score (PGS) models in unrelated trios. The error bars are 95% confidence intervals; b. Effects of parental polygenic scores (PGS) on MoBa children’s height, educational achievement, sleep duration, and depressive symptoms in trio-PGS models in unrelated trios. The error bars are 95% confidence intervals; c. Results of polygenic transmission disequilibrium (PTD) analysis, among the siblings of children ascertained on the basis of extreme trait scores. Data are presented as means with 95% confidence intervals. N trios by outcome in PGS analyses (panels a & b): educational achievement: 27,926; height: 14,541, sleep duration: 15,026; depressive symptoms: 12,632. N trios by outcome in PTD analyses (panel c): educational achievement: 535; height: 84, sleep duration: 78; depressive symptoms: 95.

    Extended Data Table 1 Descriptive statistics and Pearson correlation matrix (with 95% confidence interval) for similar outcomes to the four offspring phenotypes in mothers and fathers
    Full size table
    Extended Data Table 2 Spousal Pearson correlation matrix (with 95% confidence intervals) for similar outcomes to the four offspring phenotypes
    Full size table
    Extended Data Table 3 Spousal polygenic score (PGS) Pearson correlations (with 95% confidence intervals)
    Full size table
    Extended Data Table 4 Fit statistics for the Trio-GCTA models
    Full size table

    Supplementary information

    Supplementary Text (download PDF )

    This file contains Supplementary Notes 1–4, Supplementary Methods 1–15 and Supplementary References. The Supplementary Notes and Methods include the theoretical justification of the trio-GWAS approach, a description of the MoBa genotyping batches and the MoBaPsychGen pipeline methods

    Reporting Summary (download PDF )

    Supplementary Figures (download PDF )

    Supplementary Figs. 1–11 with figure legends. Supplementary Figs. 1–3 show an overview of the MoBaPsychGen pipeline, the relationship between power and parental heterozygosity rate, and a Directed Acyclic Graph relevant to the Theoretical justification of the trio-GWAS approach, respectively. Supplementary Figs. 4–11 are from the implementation of the MoBaPsychGen pipeline

    Supplementary Tables (download XLSX )

    The Supplementary Tables provide an overview of the MoBa genotyping batches and the number of SNPs (autosomal, chromosome X and pseudoautosomal) and individuals passing various stages of the MoBaPsychGen pipeline

    Peer Review file (download PDF )

    Rights and permissions

    Open Access This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by/4.0/.

    Reprints and permissions

    About this article

    Cite this article

    Corfield, E.C., Shadrin, A.A., Frei, O. et al. Family genetic designs in MoBa provide insights into health and functioning.
    Nature (2026). https://doi.org/10.1038/s41586-026-10926-5

    • Received:31 January 2025

    • Accepted:17 July 2026

    • Published:19 August 2026

    • Version of record:19 August 2026

    • DOI
      :https://doi.org/10.1038/s41586-026-10926-5

    designs family Genetic MoBa provide
    healthylife7
    • Website

    Related Posts

    Residents were told the air around Lineage was not dangerous. A new report found health concerns

    August 19, 2026

    Diet culture is back. Are GLP

    August 19, 2026

    LAMC To Offer Free Drive

    August 19, 2026
    Leave A Reply Cancel Reply

    Health

    An ‘Exercise Pill’ Just Cleared Its First Human Test

    By healthylife7August 19, 20260

    Scientists might be inching closer to a breakthrough that some of us lazy bums would very much appreciate: a pill that can safely provide the same healthy benefits of exercise

    La Newyorkina, the Spanish granola brand inspired by a bad breakfast in New York

    August 19, 2026

    Family genetic designs in MoBa provide insights into health and functioning

    August 19, 2026

    Meghan Trainor shows off enviable physique in black bikini during picturesque family glamping vacation

    August 19, 2026
    Stay In Touch
    • Facebook
    • Twitter
    • Pinterest
    • Instagram
    • YouTube
    • Vimeo
    Fitness

    Opinion: The FDA must put biotech at its center or continue to cede early research to China

    July 6, 2026

    Inside Elevance’s digital chronic disease management strategy

    July 6, 2026

    Best, Worst States For Well

    July 6, 2026

    What do the Middle Ages tell us about mental health then and now? VCU historian Leigh Ann Craig has answers

    July 6, 2026

    Subscribe to Updates

    Get the latest creative news from SmartMag about art & design.

    About Us

    Welcome to HealthyLife7.com, your trusted source for reliable health, wellness, fitness, and lifestyle information. Our mission is to help people make informed decisions about their health by providing clear, practical, and easy-to-understand content.

    At HealthyLife7.com, we believe that good health starts with the right knowledge. Whether you're looking for healthy eating tips, fitness advice, mental wellness strategies, weight management guidance, or information about common health conditions, our goal is to deliver valuable content that supports a healthier lifestyle.

    Fitness

    An ‘Exercise Pill’ Just Cleared Its First Human Test

    August 19, 2026

    La Newyorkina, the Spanish granola brand inspired by a bad breakfast in New York

    August 19, 2026

    Family genetic designs in MoBa provide insights into health and functioning

    August 19, 2026
    Health

    Opinion: The FDA must put biotech at its center or continue to cede early research to China

    July 6, 2026

    Inside Elevance’s digital chronic disease management strategy

    July 6, 2026

    Best, Worst States For Well

    July 6, 2026
    Facebook X (Twitter) Instagram Pinterest
    • About Us
    • Contact us
    • Disclaimer
    • Privacy Policy
    • Terms and Conditions
    © 2026 healthylife7.com. Designed by Pro.

    Type above and press Enter to search. Press Esc to cancel.