1 of 22

R9

Introduction

FinnGen research project is a public-private partnership combining genotype data from Finnish biobanks and digital health record data from Finnish health registries. FinnGen provides a unique opportunity to study genetic variation in relation to disease trajectories in an isolated population.

FinnGen is a growing project, aiming at 500,000 individuals in the end of 2023.

FinnGen results are subjected to one year embargo and, after that, available to the larger scientific community via the or through .

Data download

To download FinnGen summary statistics you will need to fill the online form at this link. You will then receive an email containing the detailed instructions for downloading the data.

Release 9 contains

GWAS summary association statistics
Fine-mapping results

When using these results in publications, please remember to:

1) Acknowledge the FinnGen study. You can use the following text:

“We want to acknowledge the participants and investigators of the FinnGen study”

2) Cite our latest publication:

Kurki M.I., et al. . Nature 2023 Jan;613(7944):508-518. doi: 10.1038/s41586-022-05473-8. Epub 2023 Jan 18.

Furthermore, if possible, include "FinnGen" as a keyword for your publication.

If you want to cite this website, use the following citation:

The manifest file with the link to all the downloadable summary stats is available at:

Data description

File naming pattern and file structure

GWAS summary statistics (tab-delimited, bgzipped, genome build 38, index files included) are named as {endpoint}.gz. For example, endpoint I9_CHD has I9_CHD.gz and I9_CHD.gz.tbi.

To learn more about the methods used, see section .

The {endpoint}.gz have the following structure:

Data releases

Timeline for releases:

Release

Date release to partners

Date release to public

Total sample size [1]

[1] samples used for PheWAS.

How to cite

Please use the following description when referring to our project:

The FinnGen study is a large-scale genomics initiative that has analyzed over 500,000 Finnish biobank samples and correlated genetic variation with health data to understand disease mechanisms and predispositions. The project is a collaboration between research organisations and biobanks within Finland and international industry partners.

When using these results in publications, please remember to:

Acknowledge the FinnGen study. You can use the following text:

Methods

Participating biobanks/cohorts

Genotypes

FinnGen individuals were genotyped with Illumina and Affymetrix chip arrays (Illumina Inc., San Diego, and Thermo Fisher Scientific, Santa Clara, CA, USA).

Chip genotype data were imputed using the population-specific SISu v4.0 imputation reference panel of 8,554 whole genomes.

Merged imputed genotype data is composed of 96 data sets that include samples from multiple cohorts.

Total number of individuals: 392,649
Total number of variants (merged set): 20,175,454
Reference assembly: GRCh38/hg38

Genotype data

Chip genotype data processing and QC Samples were genotyped with Illumina (Illumina Inc., San Diego, CA, USA) and Affymetrix arrays (Thermo Fisher Scientific, Santa Clara, CA, USA).

Genotype calls were made with GenCall and zCall algorithms for Illumina and AxiomGT1 algorithm for Affymetrix data.

Chip genotyping data produced with previous chip platforms and reference genome builds were lifted over to build version 38 (GRCh38/hg38) following the protocol described here: dx.doi.org/10.17504/protocols.io.xbhfij6.

Quality control

In sample-wise quality control steps, individuals with ambiguous gender, high genotype missingness (>5%), excess heterozygosity (+-4SD) and non-Finnish ancestry were excluded. In variant-wise quality control steps, variants with high missingness (>2%), low HWE P-value (<1e-6) and low minor allele count (MAC<3) were excluded.

Pre-phasing

Before imputation, chip-genotyped samples were pre-phased with Eagle 2.3.5 using the default parameters, except the number of conditioning haplotypes, which was set to 20,000.

Software used

Hail v0.2
Cromwell-42
Wdltool-0.14
Plink 1.9 and 2.0
BCFtools 1.7 and 1.9
Eagle 2.3.5
Beagle 4.1 (version 08Jun17.d8b)
R 3.4.1 (packages: data.table 1.10.4, sm 2.2-5.4)

LD estimation

The BCOR files were created using LDstore from the Finnish SISU panel v4.0.

The panel has been divided per chromosome. For example, to use the LD information in the first chromosome, FG_LD_chr1.bcor would be the file to use.

Settings used

number of samples: 3775
window size: 1500 kb
accuracy: low
number of threads: 96
LD threshold to include correlations: 0.05

can be downloaded via:

And an example to extract variant range 20 Mb - 50 Mb from chromosome 7 is as follows:

It is not preferred to use these LD estimate files for e.g. fine-mapping, since many of the fine-mapping methods (e.g. SuSiE) require in-sample LD information for good results!

Endpoints

Registries

The disease endpoints were defined using nationwide registries:

We harmonized over the International Classification of Diseases (ICD) revisions 8, 9 and 10, cancer-specific ICD-O-3, (NOMESCO) procedure codes, Finnish-specific Social Insurance Institute (KELA) drug reimbursement codes and ATC-codes.

These registries spanning decades were electronically linked to the cohort baseline data using the unique national personal identification numbers assigned to all Finnish citizens and residents.

A full list of FinnGen endpoints is for release 9.

The endpoints with fewer than 80 cases, and developmental “helper” endpoints were excluded from the final PheWas (“OMIT” tag in the endpoint definition file).

(Risteys = intersection in Finnish) allows browsing of the FinnGen data at the phenotype level, including endpoint definitions, statistics about number of individuals, gender distribution, and longitudinal relationships. Please also note the R9 specific page

GWAS

We used regenie for release 9. Regenie's main advantages are fast leave-one-chromosome-out relatedness calculation which avoids proximal contamination, and use of an approximate Firth test which gives more reliable effect size estimates for rare variants.

We used regenie version 2.2.4.

Links:

Association tests

Endpoint

We included 2,272 endpoints in the analysis, which consisted of 2,269 binary endpoints and 3 quantitative endpoints (HEIGHT_IRN, WEIGHT_IRN, BMI_IRN). Endpoints with less than 80 cases among the 377,277 samples were excluded, as well as endpoints labeled with an OMIT tag in the endpoint definition file.

The quantitative endpoints HEIGHT and WEIGHT were acquired from minimum phenotype data. After that, phenotype BMI was formed from them, and all of them were inverse normal transformed.

Null models

For regenie step 1 LOCO prediction computation for each endpoint, we used age, sex, 10 PCs, Finngen 1 or 2 chip or legacy genotyping batch as covariates. For sex-specific phenotypes, sample sex was left out from the covariates. We excluded covariates that had less than 10 cases.

For calculating genetic relatedness in regenie step 1, we included variants 1) imputed with an INFO score > 0.95 in all batches and 2) > 97 % non-missing genotypes and 3) MAF > 1 %. The remaining variants were LD pruned with a 1Mb window and r2 threshold of 0.1. This resulted in a set of 60,896 well-imputed not rare variants for relatedness calculation.

We used a genotype block size of 1,000 in regenie step 1.

Association tests

We ran association tests with regenie for each of the 2,272 endpoints for each variant with a minimum allele count of 5 among each phenotype’s cases and controls. We used the approximate Firth test for variants with an initial p-value of less than 0.01 and computed the standard error based on effect size and likelihood ratio test p-value (regenie options --firth --approx --pThresh 0.01 --firth-se).

Colocalization

Colocalizations in FinnGen

Our colocalization approach uses the probabilistic model for integrating GWAS and eQTL data presented in eCAVIAR (Hormozdiari et al. 2016). Compared to eCAVIAR, we are using SuSiE (Wang et al. 2019) to fine-map our inputs and provide an additional colocalization metric (CLPA).

Our goal is to extract a list of genomic regions that show colocalization between two phenotypes p1 and p2. Further, we assume that the summary statistics of p1 and p2 have been fine-mapped. The fine-mapping output for each phenotype contains three columns: the variant identifier (VAR), posterior inclusion probability (PIP), and the credible set (CS) identifier.

CLPP

The Causal Posterior Probability (CLPP) is computed between two credible sets cs1 and cs2, with cs1 coming from a given phenotype p1 and cs2 coming from phenotype p2. CLPP is defined as follows: For vectors x and y, containing the PIP for variants in cs1 and cs2, respectively, CLPP is calculated by

This CLPP calculation is similar to equation 8 in Hormozdiari et al. 2016.

CLPP is dependent on the credible set size. By definition, any credible set size > 1 will yield a CLPP < 1.

We derived another colocalization metric called causal posterior agreement (CLPA) that is independent of credible set size.

The picture below shows how colocalizations are defined.

This rough example shows why we mostly use CLPA since it is independent of sample size.

The colocalization is performed between FinnGen endpoints as well as between FinnGen endpoints and various QTL resources, as shown in the image below.

These resources are listed below:

The SuSiE finemapping results for the release were used as the FinnGen data.

GTEx v8: SuSiE fine-mapping, 49 tissues, donors of mixed ancestry, Aguet et al. (2019, BioRxiv) (49 tissues only involve tissues with a sample size of n >= 50). Fine-mapping performed by Hilary Finucane, Jacob Ulirsch, Masahiro Kanai from the . Effect size interpretation: change in normalised gene expression (sd units) per alternate allele. Normalization = inverse normal transformation.
EMBL-EBI (European Bioinformatics Institute) . eQTL data from 24 tissues/cell types, 16 RNAseq sources, 6 Microarray, SuSiE fine-mapping, donors of 88% European ancestry, Kerimov et al. (2020, BioRxiv). For RNAseq data, four quantification methods (gene expression, exon expression, transcript usage, txrevise event usage). Fine-mapping was performed by . Effect size interpretation: change in normalised gene expression (sd units) per alternate allele. Normalization = inverse normal transformation.

GeneRISK: 186 lipid species QTLs, SuSiE fine-mapping of Widen et al. (2020), 7632 Finnish samples. Effect size interpretation: change in standard deviation of the lipid species per alternate allele.

UK Biobank: 36 continuous endpoints, 57 biomarkers from UKBB prepared by , SuSiE fine-mapping. Effect size interpretation for quantitative traits: change in standard deviation of the normalized outcome per alternate allele. Effect size interpretation for binary traits increase in log(odds ratios) per alternate allele.

Only unique source1-source2-pheno1-pheno2-tissue2-quant2-locus_id1-locus_id2 combinations were included in the results. FinnGen endpoints with _COMORB-definition were left out of the results.

We thank the following people for helping us assembling the QTL resources:

Kaur Alasoo and Nurlan Kerimov provided us the fine-mapped EMBL-EBI eQTL catalogue datasets.
Hilary Finucane, Jacob Ulirsch, Masahiro Kanai gave us access to their fine-mapped GTEx data.

Colocalization

Colocalizations in FinnGen

CLPP

This CLPP calculation is similar to equation 8 in Hormozdiari et al. 2016.

CLPP is dependent on the credible set size. By definition, any credible set size > 1 will yield a CLPP < 1.

We derived another colocalization metric called causal posterior agreement (CLPA) that is independent of credible set size.

The picture below shows how colocalizations are defined.

This rough example shows why we mostly use CLPA since it is independent of sample size.

The colocalization is performed between FinnGen endpoints as well as between FinnGen endpoints and various QTL resources, as shown in the image below.

These resources are listed below:

The SuSiE finemapping results for the release were used as the FinnGen data.

GTEx v8: SuSiE fine-mapping, 49 tissues, donors of mixed ancestry, Aguet et al. (2019, BioRxiv) (49 tissues only involve tissues with a sample size of n >= 50). Fine-mapping performed by Hilary Finucane, Jacob Ulirsch, Masahiro Kanai from the . Effect size interpretation: change in normalised gene expression (sd units) per alternate allele. Normalization = inverse normal transformation.
EMBL-EBI (European Bioinformatics Institute) . eQTL data from 24 tissues/cell types, 16 RNAseq sources, 6 Microarray, SuSiE fine-mapping, donors of 88% European ancestry, Kerimov et al. (2020, BioRxiv). For RNAseq data, four quantification methods (gene expression, exon expression, transcript usage, txrevise event usage). Fine-mapping was performed by . Effect size interpretation: change in normalised gene expression (sd units) per alternate allele. Normalization = inverse normal transformation.

GeneRISK: 186 lipid species QTLs, SuSiE fine-mapping of Widen et al. (2020), 7632 Finnish samples. Effect size interpretation: change in standard deviation of the lipid species per alternate allele.

UK Biobank: 36 continuous endpoints, 57 biomarkers from UKBB prepared by , SuSiE fine-mapping. Effect size interpretation for quantitative traits: change in standard deviation of the normalized outcome per alternate allele. Effect size interpretation for binary traits increase in log(odds ratios) per alternate allele.

Only unique source1-source2-pheno1-pheno2-tissue2-quant2-locus_id1-locus_id2 combinations were included in the results. FinnGen endpoints with _COMORB-definition were left out of the results.

We thank the following people for helping us assembling the QTL resources:

Kaur Alasoo and Nurlan Kerimov provided us the fine-mapped EMBL-EBI eQTL catalogue datasets.
Hilary Finucane, Jacob Ulirsch, Masahiro Kanai gave us access to their fine-mapped GTEx data.

Sample QC and PCA

This is a description of the quality control procedures applied before running the GWAS.

PCA

The PCA for population structure has been run in the following way:

Variant filtering and LD pruning

The sisu version 4 imputation panel is pruned iteratively, until a target number of SNPS is reached:

9,385,753 starting variants: only variants with a minimum info score of 0.9 in all batches are kept.

The script starts with [500.0, 50.0, 0.9] params in plink (window,step,r2). It then decreases 0.05 in r2 iteratively pruning the imputation panel until the threshold of 200,000 snps is reached. Once the SNP count falls under 200,000 the closest pruning is returned.

If the higher r2 is closer, 200,000 snps are randomly selected, else the last pruned snps are returned.

Plink flags used: --snps-only --chr 1-22 --max-alleles 2 --maf 0.01 .

For this run the final ld params are --indep-pairwise 500.0 50.0 0.15 and 187,068 snps are returned.

Then, FinnGen data was merged with the 1k genome project (1kgp) data, using the variants mentioned above. This reduced the number of variants from 187,068 to 185,517. A round of PCA was performed and a bayesian algorithm was used to spot outliers. This process got rid of 12,639 FinnGen samples. The figure below shows the scatter plots for the first 3 PCs. Outliers, in green, are separated from the FinnGen red cluster.

While the method automatically detected as being outliers the 1kg samples with non European and southern European ancestries, it did not manage to exclude some samples with Western European origins. Since the signal from these samples would have been too small to allow a second round to be performed without detecting substructures of the Finnish population, another approach was used. The FinnGen samples that survived the first round were used to compute another PCA. The EUR and FIN 1kg samples were then projected onto the space generated by the first 3 PCs. Then, the centroid of each cluster was calculated and used to calculate the squared mahalanobis distance of each FinnGen sample to each of the centroids. Being the squared distance a sum of squared variables (with unitary variance, due to the mahalanobis distance), we could see it as a sum of 3 independent squared variables. This allowed to map the squared distance into a probability (chi squared with 3 degrees of freedom). Therefore, for each cluster, a probability of being part of it was computed. Then, a threshold of 0.95 was used to exclude FinnGen samples whose relative chance of being part of the Finnish cluster was below the level. This method produced another 64 outliers. The figure below shows the first three principal components.

FIN 1kgp samples are in purple, while EUR 1kgp samples are in Blue. Samples in green are FinnGen samples who are flagged as being non Finnish, while red ones are considered Finnish.

Then all pairs of FinnGen samples up to second degree were returned. The figure below shows the distribution of kinship values.

Then, the previously defined “non Finnish” samples were excluded and 2 algorithms were used to return a unique subset of unrelated samples:

one called greedy would continuously remove the highest degree node from the network of relations, until no more links are left in the network.
one called native, based on a native implementation of python’s networkx package, performed on each subgraph of the network.

The largest independent set of either algorithm would be used to keep those sample, while flagging the others as “outliers” for the final PCA.

Then, the subset of outliers who also belong to the set of duplicates/twins was identified.

To compute the final step the Finngen samples were ultimately separated in three groups:

233,371 inliers: unrelated samples with Finnish ancestry.
144,127 outliers: non duplicate samples with Finnish ancestries, but who are also related to the inliers.
15,151 rejected samples: either of non Finnish ancestry or are twins/duplicates with relations to other samples.

Finally, the PCA for the inliers was calculated, and then outliers were projected on the same PC space, allowing to calculate covariates for a total of 377,498 samples.

Of the 377,498 non-duplicate population inlier samples from PCA, we excluded 215 samples from analysis because of missing minimum phenotype data, and 6 samples because of failing sex check with F thresholds of 0.4 and 0.7. Sex matched between genotype data and phenotype data for all individuals! A total of 377,277 samples were used for core analysis. There are 210,870 females and 166,407 males among these samples.

Documentation from the original developers of the algorithm can be found here: .

R9

Introduction

Data download

Data description

Data releases

How to cite

Methods

Participating biobanks/cohorts

Genotypes

Genotype data

Quality control

Pre-phasing

Software used

LD estimation

Settings used

Endpoints

Registries

GWAS

Association tests

Endpoint

Null models

Association tests

Colocalization

Colocalizations in FinnGen

CLPP

Introduction

Genotypes

Data releases

Endpoints

Registries

Excluded endpoints

Risteys

LD estimation

Settings used

Example usage

Note

Association tests

Endpoint

Null models

Association tests

Genotype data

Quality control

Pre-phasing

Participating biobanks/cohorts

Data download

Using FinnGen data for publications

Manifest

Colocalization

Colocalizations in FinnGen

CLPP

Data description

How to cite

Software used

CLPA

Example Comparison

Data

FinnGen resources

Expression QTL datasets

Metabolon QTL datasets

Biomarkers

Post-colocalization QC

Acknowledgements

Summary association statistics

Fine-mapping results

LD estimation

Variant annotation

GWAS

SISu reference panel

Genotype imputation

Fine-mapping

PheWeb

1. Preprocessing

2. LD computation

3. Fine-mapping

Notes

Integration to PheWeb

Sample QC and PCA

PCA

Variant filtering and LD pruning

PCA outlier detection