Skip to content

Comparing National Genomic Databases

A national genomic database is public research infrastructure that connects genomic data, biospecimens, medical records, and lifestyle information from many people so researchers can analyze them. Its inputs are participants’ DNA and health data. Its output is not a single personal test report, but a large cohort1 and a secure analysis environment for finding associations with disease.

Korea, the United States, and Europe are all building this infrastructure, but they are not building the same product. The United States and the United Kingdom lead in making large research cohorts broadly available. Finland and Estonia have strong long-term links to national health records. France focuses on bringing whole-genome sequencing into hospital care. Korea has operated cohorts and biobanks for more than 20 years, and in 2024 began a national program that connects these assets to large-scale WGS and public-sector data.

Date of figures: The status of ongoing programs in this document was checked on August 16, 2026. Budgets remain in their original currencies, and announced plans are distinguished from actual spending.

Region and flagship programMain purposeVerifiable investmentConstruction and access outcomes
Korea, National Integrated Bio Big Data ProjectLink Korean clinical, genomic, and public data at the individual levelKRW 606.58 billion total for Phase 1, 2024 to 2028Target of 772,000 participants, 340,000 blood-derived whole genomes, and 41,000 cancer tissue genomes. By June 2025, recruitment had reached 17.3% of the cumulative target at that time
United States, All of UsResearch genomic, health, and lifestyle data from a diverse populationFY2025 presidential budget request of USD 541 million; actual FY2026 appropriation of USD 153 millionMore than 883,000 enrolled, with 535,000 whole genomes and about 482,000 electronic health records available for research
EU, 1+MG, GDI, and Genome of EuropeAnalyze data across borders without moving it out of each countryEUR 40 million for GDI and about EUR 45 million for Genome of EuropeInfrastructure under construction in 26 member states, with a target of more than 100,000 whole genomes representative of European populations
United Kingdom, UK BiobankConnect genetics, lifestyle, imaging, and long-term health outcomes in a research cohortGBP 200 million in joint public, charitable, and pharmaceutical investment to add WGS for 500,000 peopleWGS for 500,000 participants available, with more than 18,000 related peer-reviewed papers
Finland, FinnGenLink biobank genotypes to decades of national health registersAbout EUR 150 million cumulatively, including EUR 52 million in Phase 3 investmentGenotypes linked to medical histories for 500,000 people, with more than 1,000 medically relevant variants identified
Estonia, Estonian BiobankGenotype a large share of the adult population and return resultsEUR 5 million in government funding to recruit 100,000 more people in 2018; total cumulative cost is largerMore than 212,000 participants, over 20% of the adult population. More than 110,000 people have visited the personal results portal
France, PFMG2025Integrate genomic testing into rare-disease and cancer careInitial five-year investment plan of EUR 670 million, including about EUR 230 million in planned industry contributions58,211 prescriptions approved and 44,165 reports returned from 2019 to 2025, across 82 eligible indications

Do not read these amounts as a ranking. Korea’s figure includes recruitment, samples, data production, and a platform, while the US figure is a single year’s budget. The EU amounts cover two shared infrastructure programs but exclude national biobank investment. The UK’s GBP 200 million is not UK Biobank’s full operating cost, but the public-private investment that added 500,000 whole genomes. Estonia’s EUR 5 million likewise covers one expansion round in 2018, not the entire program.

Producing one whole-genome file for a person does not immediately make disease research possible. Researchers need four connected layers of data.

LayerActual dataQuestion it can answer
GenomeSNP arrays, whole exomes, whole genomes, transcriptomes, and related measurementsWhat variants are present?
PhenotypeDiagnoses, prescriptions, laboratory values, physical measurements, and surveysWhat health traits does this person have?
TimeHistorical records and prospective follow-upWas a variant present before disease, and does it predict an outcome?
Access environmentApproval process, secure analysis room, and common data modelCan multiple researchers perform reproducible analyses?

Participant count, whole-genome count, and the number actually available to researchers are therefore different metrics. The counts can decrease at each stage: registered participants, participants who supplied samples, participants whose genomes have been produced, and participants whose linked medical records are available for analysis.

Saying that data has been released does not mean anyone can download individual WGS files. National genomic resources are usually divided into three levels according to reidentification risk and whether a resource is consumed when used.

Access levelWhat you can seeTypical conditions
Public and aggregateData dictionaries, participant counts, condition frequencies, GWAS summary statistics, and teaching dataLogin or simple registration may be required. No individual-level rows
Controlled individual-level dataPseudonymized individual genotypes, clinical and epidemiological data, and medical recordsVerification of researcher and institution, a research plan, ethics or data-access review, a use agreement, and security training
BiospecimensDNA, serum, plasma, and tissue that are depleted when usedControlled-data conditions plus IRB approval, scientific justification for quantity, distribution review, and transport and disposal plans

Exploring variant frequencies in a browser, analyzing individual genotypes from 100,000 people, and receiving 500 serum samples for a new assay are therefore three entirely different applications.

Who can access individual-level data in each country?

Section titled “Who can access individual-level data in each country?”
Country and programEligible people and institutionsMain review and agreementDelivery method
Korea, CODA and KoGESResearchers who can submit a research plan and an IRB decision. The public guidance does not state a separate degree or nationality requirementIdentity verification, research plan, IRB decision, and approval by the Data Disclosure and Use CommitteeRemote analysis by default. Restricted variables or special software require analysis on site in Osong; only approved deidentified results may leave
Korea, National Biobank of Korea and KBNKorean principal investigators residing in Korea with an IRB-approved or exempt project. Biospecimens are available only to funded projectsUse plan, research plan, IRB decision, undertaking, distribution committee review, and fees when applicableData or physical specimens are distributed. Specimens must be disposed of and outcomes reported when use ends
Korea, National Integrated Bio Big Data ProjectAcademic, medical, and industry use is planned, but detailed eligibility has not yet been publishedResearch-plan and data-distribution review. IRB approval is also required when identifiable information is involvedWeb delivery or a VDI security environment depending on sensitivity; only results may leave for high-risk data
United States, All of UsResearchers at universities, health care institutions, nonprofits, and government institutions that have signed a DURA. Independent researchers and people employed only by general for-profit companies are not currently directly eligibleInstitutional DURA, institutional email, real-name identity verification, required training, and the Data User Code of Conduct. Controlled tier requires additional trainingAnalysis inside the cloud-based Researcher Workbench
United Kingdom, UK BiobankVerifiable researchers worldwide from universities, hospitals, government, charities, and companies. A recognized institution and research record are required, and sanctioned countries are excludedA health-related public-interest proposal, CV, institutional profile and publication record, access review, MTA, and feeAnalysis on the UKB Research Analysis Platform. Depletable samples receive separate, stricter review
Finland, FinnGenPublic summary statistics are open to all. Researchers and employees at FinnGen partner institutions have priority for individual-level dataPartner researchers submit an analysis proposal and pass a security test. External researchers apply separately to Findata and each biobankIndividual data stays in the Sandbox; only results aggregated over at least five people may leave
Estonia, Estonian BiobankDomestic and international academic, public-sector, and industry researchersPreliminary feasibility check, Scientific Advisory Committee, national research ethics committee, data or sample transfer agreement, and feeIndividual data is analyzed in the University of Tartu’s SAPU secure environment
France, PFMG2025 CADResearch professionals and institutions proposing public-interest health researchScientific and ethics committee review, regulatory permission, collaboration agreement, and named researchersAnalysis in project-specific secure spaces inside CAD
EU, GDIA single shared European researcher account is not yet completeResearchers apply to national nodes and receive permission from the relevant data access committees under national lawApproved analyses run in national secure environments rather than moving raw data to a central store

These differences do not rank countries from strict to permissive. All of Us emphasizes institutional agreements, identity checks, and training instead of requiring separate project-level IRB documents in every case. The UK accepts commercial researchers but evaluates track record and public benefit. Finland makes summary statistics broadly public, while direct use of FinnGen individual data is centered on partners. In Korea, project-level IRB and distribution review are prominent for both CODA data and biospecimens.

What do you actually need to prepare in Korea?

Section titled “What do you actually need to prepare in Korea?”

For individual clinical and genomic data: CODA

Section titled “For individual clinical and genomic data: CODA”

The CODA distribution process has seven steps.

  1. Verify your identity through Any-ID and create an account.
  2. Check populations, variables, and files such as WGS, WES, and SNP arrays in the catalog.
  3. Submit a research plan, IRB decision, and project information.
  4. Receive distribution approval from the Data Disclosure and Use Committee.
  5. Use the approved resources in a remote or on-site analysis environment.
  6. Submit a certificate of destruction when the use period ends.
  7. Register research outputs such as papers.

In the CODA analysis environment, only research results without personal identifiers may leave after a separate export review. Remote analysis is the default, but researchers must visit the CODA analysis room in Osong when KoGES restricted variables are included or unsupported software is required.

The public process does not require a particular degree or faculty rank. It does, however, require a concrete research project approved or exempted by an IRB and an accountable research identity. It is not a route for a member of the public to download individual WGS data out of curiosity. Students and data analysts usually participate as collaborators in an approved project led by an investigator at their institution.

The National Biobank of Korea distribution rules state eligibility more specifically.

  • The researcher must be a Korean national residing in Korea.
  • The applicant must be the principal investigator of a project that has IRB approval or exemption.
  • The project must receive support from a national R&D program, a government or government-funded institution, or a private research institution.
  • Biospecimens in particular are available only to funded projects.
  • The applicant submits a resource-use plan, research plan, IRB decision, undertaking, and personal-information consent form.
  • The distribution committee evaluates the purpose, requested quantity, and calculation behind that quantity, so not every requested sample is necessarily approved.

Population resources such as KoGES and KNHANES are requested through HuBIS_Desk, while disease-based KBN resources are requested through the KBN Portal. Physical samples generally must be collected in person. At the end of the use period, they must be destroyed and a completion and destruction report submitted within 30 days.

For the new National Integrated Bio Big Data resource: check the detailed call when it is published

Section titled “For the new National Integrated Bio Big Data resource: check the detailed call when it is published”

The program plans three data tiers. Anonymized teaching data will be available on the web. Data with reidentification risk will be available after research-plan and distribution review inside a restricted analysis environment. For data with high identification risk, only results may leave. According to the program’s privacy and security explanation, studies that may involve identifiable information require both IRB approval of the research plan and distribution approval from the data bank or biobank, and analysis remains inside a VDI environment after approval.

As of August 2026, however, the public material does not provide a single final user guide that answers which institutions, nationalities, or job levels are eligible, what agreement a company must sign, or how long review takes. Do not assume that current CODA or biobank rules will apply unchanged. When the first data release is announced, check eligibility, IRB scope, fees, and result-export rules again.

Applicant situationCurrent assessment
Principal investigator at a Korean university, hospital, or government-funded instituteCan apply to CODA with a research plan and IRB decision; can also apply for biobank samples when the project is funded
Graduate student, postdoctoral researcher, or data analystNo rule found prohibits the job level itself, but the principal investigator must submit a biobank application. Usually participates as a collaborator or analyst on an approved project
Researcher at a Korean companyA principal investigator with support from a private research institution and IRB approval is eligible for biobank resources. The national program also plans industry access, but its final user guide has not been released
Freelancer or independent developerIndividual-level data and samples are difficult to obtain without a research institution, an IRB-reviewed project, and an accountable investigator. Public aggregates and teaching data remain available
Researcher outside KoreaCannot directly request National Biobank of Korea specimens because of the Korean researcher residing in Korea rule. CODA’s public guidance gives no nationality rule, so contact the program before applying; collaboration with a Korean institution is the clearest route

IRB approval is not a permission that the data provider obtains for you. First submit the research plan to your own institution’s IRB and receive an approval or exemption decision. Then submit that decision to CODA or the biobank for a separate distribution review. IRB review and data-access approval are two different gates.

You can learn analysis methods with aggregate and teaching data even without an institutional project. The CODA PheWeb service lets users explore variant and phenotype associations from about 187,000 Korean participants. The KoGES teaching dataset provides modified survey data based on 10,000 baseline participants and 1,000 follow-up participants. These resources do not contain raw personal WGS files or reidentifiable medical records.

In Korea, eligibility is not created by simply claiming genomic analysis skills. It depends on an accountable institution and principal investigator, a reviewable research plan, an IRB decision, and a secure data environment. Analytical skill makes a plan more feasible, but it does not itself confer access.

Korea already built several layers of infrastructure separately

Section titled “Korea already built several layers of infrastructure separately”

Seeing Korea’s national genomic infrastructure as one new database started in 2024 misses the previous 20 years of assets. Cohorts, biobanks, population WGS, and rare-disease diagnosis developed first through separate programs.

Start or periodResourceWhat it builtVerifiable outcomes
2001 to presentKorean Genome and Epidemiology Study (KoGES)Long-term cohort repeatedly collecting surveys, examinations, biospecimens, and genetic data, primarily from the general populationAbout 235,000 participants followed for more than 20 years; more than 1,680 related papers by the end of 2023
2008 to presentNational Biobank of Korea and Korea Biobank Network (KBN)Storage and distribution of serum, plasma, DNA, tissue, and linked information from population and disease cohortsBy the end of 2024, 1.23 million cumulative donors, 20.87 million vials, 5,446 distributed research projects, and 2,025 papers
2016 to 2021Ulsan 10,000 Genomes ProjectHigh-depth whole genomes from healthy and affected participants, with selected clinical and multi-omics dataMore than KRW 18 billion invested and 10,044 genomes completed. Korea1K, the first public study, identified about 39 million variants in 1,094 participants
2020 to 2022National Bio Big Data pilotWGS for people with undiagnosed rare diseases plus clinical and genomic data from existing national projects14,905 rare-disease participants recruited and about 25,000 records opened including existing projects. Genomic diagnosis was reported as possible for more than 30% of a previously undiagnosed study group
2024 to presentNational Integrated Bio Big Data ProjectLink clinical data, genomes, public data, and samples around each participantPhase 1 is working toward 772,000 participants and staged research access

KoGES contributes a long timeline, not merely a large table

Section titled “KoGES contributes a long timeline, not merely a large table”

KoGES, the Korean Genome and Epidemiology Study, is a national longitudinal cohort for studying genetic and environmental factors in chronic conditions such as diabetes, hypertension, and cardiovascular disease. It collected disease history, prescriptions, diet and lifestyle, body measurements, blood and urine tests, and biospecimens from about 235,000 people. Some groups have been measured repeatedly since 2001.

OPEN KoGES allows researchers to search and analyze clinical and epidemiological data from about 210,000 participants in community, urban, and rural cohorts inside a secure environment. The Korea National Institute of Health reports that KoGES resources had supported more than 1,680 papers by the end of 2023. Korea’s asset that a new program cannot quickly buy is not merely participant count, but more than 20 years of follow-up.

The biobank preserves samples outside genomic files

Section titled “The biobank preserves samples outside genomic files”

The National Biobank of Korea manages population resources such as KoGES and KNHANES, while the hospital-centered KBN gathers disease-based resources. The official cumulative figures for 2024 list 1.23 million donors, 20.87 million vials of human-derived materials, and epidemiological information for 620,000 people across both systems. They distributed 1.55 million vials and 4,102 cohort-level information sets to 5,446 research projects, which produced 2,025 papers and 206 patent applications or registrations.

1.23 million donors does not mean 1.23 million distinct people with WGS. A single donor may supply several vials of serum, plasma, and DNA, and the available genetic and epidemiological data differ by resource. The value of this infrastructure is that it preserves specimens that can be measured again and phenotypes that have already accumulated.

Ulsan 10K showed the value of a Korean reference panel

Section titled “Ulsan 10K showed the value of a Korean reference panel”

The project led by Ulsan Metropolitan City and UNIST sequenced 10,044 whole genomes by 2021, including about 4,700 healthy participants and 5,300 affected participants. Investment exceeded KRW 18 billion.

The first research release, Korea1K, combined 1,094 WGS datasets at an average depth of 31x with 79 clinical traits. The original paper found about 39 million single-nucleotide variants and short insertions and deletions. It also showed that genotype imputation for Korean samples was more accurate than when only international reference panels were used. The 2025 paper describing the full Korea10K result was still a preprint when this page was checked, so read it separately from the peer-reviewed Korea1K result.

The pilot tested rare-disease diagnosis and release procedures

Section titled “The pilot tested rare-disease diagnosis and release procedures”

The pilot recruited 14,905 people with rare diseases and linked existing datasets from 5,000 general-population KoGES participants, 2,504 general-population Ulsan genomes, 892 autism-spectrum participants, 322 colorectal cancer participants, 84 lung cancer participants, and 995 dementia participants. The total is about 25,000. In the current CODA holdings, the 14,905 rare-disease genomes appear in three sets of 3,886, 10,550, and 469 participants. KoGES WGS for 2,500 people and Ulsan WGS for 2,504 people are also listed.

The pilot also produced a clinical result. Its announcement reported that genomic data made diagnosis possible for more than 30% of a study group that had not been diagnosed by existing methods. This is not the confirmed diagnostic rate across all 14,905 participants, but a result from a specific undiagnosed group.

Korea began a national program for 772,000 participants

Section titled “Korea began a national program for 772,000 participants”

Korea’s National Integrated Bio Big Data Project is a cross-ministerial program led by the Ministry of Health and Welfare, Ministry of Science and ICT, Ministry of Trade, Industry and Energy, and Korea Disease Control and Prevention Agency. Its goal is to connect clinical information, samples, genomic and other omics2 data, and public-sector data around each participant.

The Phase 1 plan summarized by the National Assembly Budget Office has the following scale.

  • Period: 2024 to 2028
  • Total project cost: KRW 606.58 billion
  • Participants: 585,000 general population, 140,000 with severe diseases, and 47,000 with rare diseases, for a total of 772,000
  • Data production: blood-derived WGS for 340,000 people and cancer tissue WGS for 41,000 people
  • Long-term plan: after an interim review, expand to 1 million participants in Phase 2 from 2029 to 2032

This does not mean all 772,000 participants receive whole-genome sequencing. The design first builds clinical information and biospecimen resources for nearly all participants, then produces whole genomes and other omics for specified groups. Without this distinction, 772,000 participants is easily misread as 772,000 whole genomes.

The original nine-year proposal submitted for preliminary feasibility review was KRW 998.8 billion. Because of uncertainty in a long program, only Phase 1, with KRW 606.58 billion from 2024 to 2028, was initially approved. The scale and unit costs of Phase 2 are to be recalculated after an interim review. 1 million participants and KRW 998.8 billion is therefore the long-term plan, while the currently approved and comparable scope is 772,000 participants and KRW 606.58 billion in Phase 1.

A public dashboard with the latest cumulative figures on one screen was not available when this page was checked. The latest independently verifiable recruitment figures are the June 2025 counts compiled by the National Assembly Budget Office from material submitted by the Ministry of Trade, Industry and Energy.

Participant groupCumulative target for 2024 to 2025Actual by June 2025Completion
General population161,20032,90020.4%
Severe disease36,0001,2113.3%
Rare disease8,8001,47517.1%
Total206,000About 35,60017.3%

The National Assembly Budget Office settlement analysis explains that selection of recruiting institutions was completed only at the end of 2024 after three calls, delaying the schedule. Of KRW 51.445 billion in R&D program funds excluding operating costs, KRW 2.41 billion also remained unspent. These numbers are not a verdict that the program failed. They form a baseline for measuring how much of the first two years’ target must be recovered later. As of August 2026, the program website lists 26 general-population recruiting institutions and 19 institutions recruiting severe-disease and cancer participants, but it does not publish a later cumulative participant count in the same format.

It is too early to compare this program with the United States or the UK by papers and clinical outcomes. The government’s 2024 self-evaluation also described the program launch as a major outcome while noting the need for performance indicators that capture actual data use, not just construction volume. The important Korean metrics now are not the announced target of 772,000 itself, but these three questions.

  1. Are recruitment, sample processing, and genomic production each progressing to plan?
  2. Are clinical and public-sector data linked at research-ready quality?
  3. How long does it take an approved researcher to discover and analyze the data?

Korea already has K-BDS, which collects R&D outputs; CODA, which provides controlled access to clinical and omics data; OPEN KoGES, which specializes in KoGES analysis; and the National Biobank of Korea, which distributes human-derived materials. The National Integrated Bio Big Data Project is not a renaming of these repositories. It aims to connect participant recruitment, clinical and genomic production, and research use in one long-term cohort.

The United States built a large cohort and cloud analysis environment together

Section titled “The United States built a large cohort and cloud analysis environment together”

The US National Institutes of Health’s All of Us Research Program aims to follow more than one million US residents over time. A central goal is to include racial, geographic, age, and social groups that have been underrepresented in previous genetic research.

Participants contribute surveys, physical measurements, biospecimens, electronic health records (EHR), and wearable data. Researchers generally analyze the data inside a secure cloud called the Researcher Workbench instead of downloading raw copies to their own servers. At petabyte scale, this approach reduces both security risk and the cost of storing duplicate datasets at every institution.

All of Us is a cohort that generates new data under common standards. NIH’s dbGaP, the database of Genotypes and Phenotypes, is instead an archive that stores and distributes genotype and phenotype data produced by many studies. In August 2026, dbGaP listed about 3,300 public studies and 8.1 million study participants. This does not mean 8.1 million people belong to one cohort with identical variables or all have WGS. Study designs, data types, and consent terms differ, and individual-level data requires separate approval.

NIH’s FY2025 presidential budget request for the program was USD 541 million. A request is not the same as actual spending. NIH states that the FY2026 appropriation was USD 153 million, about 72% below the 2023 level, which led to reduced new recruitment and some data collection. The case shows that a national database continues to require operating, follow-up, and security funding after construction.

The NIH release from June 2026 reported:

  • More than 883,000 enrolled participants
  • Data from more than 747,000 participants available for research
  • More than 535,000 whole genomes
  • Linked electronic health records for about 482,000 participants
  • About 23,000 researchers
  • More than 1,400 peer-reviewed papers
  • More than 733,000 health-related DNA results returned to 277,000 participants

Scale is not the only outcome. A 2024 analysis of about 240,000 genomes found more than one billion variants, including over 275 million not previously reported in major variant resources at the time. Median time from researcher registration to controlled data access was 29 hours. The results appear in the Nature paper.

Europe is building a federation, not one central database

Section titled “Europe is building a federation, not one central database”

Europe’s 1+ Million Genomes (1+MG) initiative is not copying national genomes to one server in Brussels. It aims for a federated approach3, keeping data in national and institutional secure environments while sending standardized queries and analyses to multiple nodes and combining permitted results.

This structure is necessary because countries already have different biobanks, consent systems, health-information laws, and clinical systems. Instead of creating one new cohort, the EU is pursuing two complementary programs.

  • Genomic Data Infrastructure (GDI): EUR 40 million total budget from 2022 to 2026. Twenty-six member states are building common access infrastructure, with a goal of having structures operational in 15 countries by the end of 2026. In 2024, the program demonstrated federated queries with synthetic data in secure environments in Finland, Portugal, Spain, and Sweden.
  • Genome of Europe: About EUR 45 million total, including EUR 20 million from Digital Europe. Fifty-one institutions in 27 countries plan to build a whole-genome reference cohort of more than 100,000 people representing Europe’s population and ancestry diversity by early 2028.

Budgets and progress are documented in the European Commission’s GDI interim results and Genome of Europe launch.

1+ Million is not a current count of one million standardized genomes available for download. It is a policy objective to make data held by multiple countries securely discoverable and analyzable. In 2026, the actual access network was still under construction, and the 100,000-person Genome of Europe resource remained a 2028 target.

The UK developed separate research and clinical foundations

Section titled “The UK developed separate research and clinical foundations”

The United Kingdom has two complementary pillars.

UK Biobank is a research cohort of 500,000 people recruited from 2006 to 2010, linking genetic data, lifestyle, blood tests, imaging, and later medical records. Government, Wellcome, and pharmaceutical companies jointly invested GBP 200 million to add whole-genome sequencing for all 500,000 participants. The full dataset has been available to approved researchers since 2023.

As of 2026, UK Biobank provides whole genomes for 500,000 participants and reports more than 18,000 peer-reviewed papers that directly use or extend its results. The case shows how the value of a cohort grows when new medical records and repeat measurements continue to accumulate.

The 100,000 Genomes Project was a clinical and research program for NHS patients with rare diseases and cancer. It produced more than 100,000 whole genomes from about 85,000 people and led into the NHS Genomic Medicine Service. In a study of about 4,000 early rare-disease participants, whole-genome sequencing provided a new diagnosis for 25%. That number is a result from the specific study group, not an average diagnostic rate across all participants.

Finland linked 500,000 genotypes to decades of registers

Section titled “Finland linked 500,000 genotypes to decades of registers”

FinnGen genotypes samples from nine biobanks with SNP arrays and links them to hospital and outpatient diagnoses, prescriptions, procedures, cancer, and death registers dating back to 1969. It did not produce 500,000 whole genomes. Its strength is linking 500,000 people, about 10% of Finland’s population, to long-term phenotypes.

The University of Helsinki reports about EUR 150 million in cumulative investment, including EUR 20 million from Business Finland. The total also includes EUR 52 million in Phase 3 investment to study disease progression and drug targets after resource construction through 2023.

An analysis of about 220,000 participants found 2,733 genome-wide significant associations across 1,932 disease phenotypes. Low-frequency variants enriched in Finland helped reveal disease biology. The consortium now reports more than 1,000 medically relevant risk or protective variants in the full 500,000-person resource.

Estonia stands out in population share and return of results

Section titled “Estonia stands out in population share and return of results”

The Estonian Biobank includes more than 212,000 people, over 20% of Estonia’s adult population. After recruiting about 52,000 people from 2002 to 2011, the government invested EUR 5 million in 2018 to recruit and genotype another 100,000 people. Later recruitment brought the resource to its current scale.

Its data does not flow only to researchers. Through the national electronic-ID-linked MyGenome portal, participants can view polygenic risk scores for type 2 diabetes and coronary artery disease, selected pharmacogenomic information, and trait results. By January 2026, more than 110,000 people had visited the portal.

Estonian data has supported more than 800 papers, and a clinical study using genetic risk to identify candidates for preventive statin therapy began in 2025. Its 212,000 participants are fewer than in the US or UK cohorts, but they represent more than one in five adults in the country. Population share and return of results are different forms of scale.

France connected data production to clinical prescriptions

Section titled “France connected data production to clinical prescriptions”

Plan France Médecine Génomique 2025 (PFMG2025) integrates whole-genome and transcriptome analysis into care pathways for rare-disease and cancer patients rather than building a general-population research cohort. The plan announced in 2016 proposed EUR 670 million of investment over the first five years, including about EUR 230 million expected from industry partnerships. This was an announced investment plan, not an actual cumulative spending figure.

France built two high-throughput sequencing laboratories, SeqOIA and AURAGEN; CRefIX, a reference center for standards, innovation, expertise, and technology transfer; and CAD, a data collection and analysis center. Disease-specific multidisciplinary meetings approve tests, the laboratories perform the analysis, and reports return to physicians through the care pathway.

The 2025 activity totals show 58,211 prescriptions approved since 2019 and 44,165 reports returned to physicians. Eligible indications totaled 82: 67 rare diseases, 11 cancers, and four hereditary cancer predispositions. The research environment opened later. In 2025, the program transferred data for five approved studies to a secure cloud and opened the first two secure analysis spaces.

France demonstrates that some programs are better evaluated by how many cases reached clinical prescription and return of results than by how many people were recruited.

A national program cannot be judged by participant count alone. Ask five questions together.

1. Is there as much analyzable data as the recruitment number suggests?

Section titled “1. Is there as much analyzable data as the recruitment number suggests?”

All of Us has 883,000 registrants and 535,000 whole genomes. Korea’s target of 772,000 participants differs from its target of 340,000 blood-derived whole genomes. Count recruitment, sample quality control, genomic production, medical-record linkage, and research availability separately.

2. Is the genome connected to a long timeline?

Section titled “2. Is the genome connected to a long timeline?”

To ask whether a genotype measured in a healthy person predicts disease ten years later, medical records must continue to be linked. The strength of UK Biobank and Nordic biobanks is not merely the initial genetic data, but the diagnoses, prescriptions, laboratory tests, and images added afterward.

Existing data produces little research when the search catalog, eligibility, review time, analysis environment, and costs are unclear. The All of Us Workbench, UK Biobank Research Analysis Platform, and EU’s federated GDI all move toward computing where the data resides instead of copying sensitive large-scale data to a researcher’s laptop.

Publication count shows how open a resource is to researchers but does not capture all participant value. The United States returned health-related DNA findings at large scale, and more than 110,000 Estonian participants used their personal portal. The UK and France created pathways that return rare-disease diagnoses and clinical reports.

Genomes can be produced once, but medical-record updates, renewed consent, security, standardization, and analysis environments cost money every year. The US budget reduction in 2026 shows that sustainable operations remain a separate problem after a database becomes large.

How does Korea differ from other countries?

Section titled “How does Korea differ from other countries?”

Korea’s participant target and planned WGS production place it among the world’s large programs. What it lacks is not experience in genomic research, but enough time to integrate separate assets into one research flow and demonstrate that flow’s performance publicly.

Comparison axisLeading exampleKorea’s current positionMeaning
Long-term follow-upUK Biobank and FinnGenKoGES has more than 20 years of follow-upLinking KoGES to the new program can turn existing time depth into a Korean strength
Population WGS535,000 in All of Us and 500,000 in UK BiobankTarget of 340,000 blood-derived genomes, plus Korea10K and pilot WGSThe target is large, but production and release are not yet mature enough for a direct comparison
National health-data linkageFinnish registers and Estonia’s electronic health systemThe program is designed to link national insurance and public-sector dataNational coverage has great potential, but linkage rate, update interval, and missingness should become public metrics
Research accessAll of Us Workbench and UK Biobank analysis platformSeparate services exist for CODA, OPEN KoGES, and biobank resourcesExisting services are useful, but discovery and combination across resources remain fragmented
Participant returnAll of Us DNA results, Estonia’s MyGenome, and French clinical reportsKorea has rare-disease diagnostic experience and plans participant reportsComparability requires reporting what was returned, to how many people, and how quickly
Cross-border useEU GDI’s federated accessDomestic integration comes firstIn the long term, international joint analysis needs standards and federated execution without moving raw data abroad

First, follow-up time already exists. KoGES has more than 20 years of longitudinal data, an asset that a cohort begun in 2024 cannot immediately buy. Second, the specimen base is large. The National Biobank of Korea and KBN preserve serum, plasma, tissue, and DNA that cannot be reconstructed from genomic files. Third, a single national health insurance system and public-data linkage provide unusual potential. Connecting genomes to diagnoses, prescriptions, tests, and long-term outcomes could reproduce the benefits of Finnish register research in a larger population.

These strengths do not automatically become outcomes of the new program. Historical consent scope, personal linkage keys, variable definitions, and quality must align. The gap between institutions hold the data and approved researchers can analyze it as one cohort is what the Korean program must solve.

Separate confirmed shortcomings from questions that cannot yet be answered

Section titled “Separate confirmed shortcomings from questions that cannot yet be answered”

The currently available evidence directly supports four shortcomings.

  1. Early recruitment delay: By June 2025, recruitment had reached 17.3% of the cumulative target at that time, and the severe-disease group was at 3.3%.
  2. Insufficient use metrics: The government’s self-evaluation identified the absence of indicators representing data use.
  3. Discontinuous progress reporting: Recruiting institutions and targets are public, but no public dashboard was found that reports recruitment, sample QC, WGS completion, linkage, and research release for the same reference date.
  4. Fragmented resource discovery: KoGES, CODA, the biobank, and K-BDS serve different purposes and use different application processes. Separate roles are necessary, but without an integrated catalog researchers must investigate potential linkage themselves.

Conversely, current public evidence is not sufficient to conclude that Korean clinical data is lower quality than overseas data or that public-data linkage has failed. The linkage rate, variable-level missingness, update frequency, access review time, and number of approved studies are not publicly reported. That makes them unknown, not necessarily poor. Convert what is not disclosed into measurable requests instead of treating it as evidence of low quality.

The future: make data flow better, not merely collect more

Section titled “The future: make data flow better, not merely collect more”

Korea’s next step is not to add another database name. It is to complete the flow from participant to research outcome and return of results. The following order is reasonable.

1. Publish a funnel, not one participant number

Section titled “1. Publish a funnel, not one participant number”

Each quarter, report general, severe-disease, and rare-disease participants by region, age, and sex, then publish counts for consent → sample received → quality passed → genome produced → clinical and public data linked → released for research on the same reference date. The goal is not the large number of 772,000, but managing attrition and representation at each stage.

2. Build an integrated catalog and an access-time objective

Section titled “2. Build an integrated catalog and an access-time objective”

The original KoGES, CODA, K-BDS, and biobank systems do not need to be merged immediately. Instead, one catalog should describe participant scope, variables, consent, file formats, QC versions, and linkable resources for every dataset. Publishing median and 90th-percentile time from application to use would turn access into a service level. All of Us provides a useful reference by reporting a 29-hour median for controlled access.

3. Version phenotype sources more rigorously than WGS

Section titled “3. Version phenotype sources more rigorously than WGS”

The label diabetes can come from a hospital diagnosis code, a drug prescription, a laboratory result, or a survey answer. Each variable needs a source institution and time, code-change history, and reason for missingness. Health insurance data should be refreshed regularly. Repeated KoGES measurements and new program examinations should preserve their original methods while mapping to a shared semantic layer.

4. Make analysis without moving data the default

Section titled “4. Make analysis without moving data the default”

Letting every institution download hundreds of thousands of genomes and medical records increases security risk and storage cost. Approved code should execute where the data resides, with the analysis environment, reference genome, and pipeline versions recorded together. In the long term, this can expand to EU GDI-style federated studies that run the same analysis across countries without exporting raw data.

5. Count outcomes returned to participants separately

Section titled “5. Count outcomes returned to participants separately”

For rare diseases, useful metrics include new diagnoses, diagnoses changed after reanalysis, and time to diagnosis. For general participants, programs can count what evidence level was returned, how many people received it, and whether they viewed an explanation or requested follow-up counseling. Publication counts and returned results should remain separate measures of research output and public value.

6. Plan operating costs and reanalysis beyond 2028 now

Section titled “6. Plan operating costs and reanalysis beyond 2028 now”

Sequencing finishes, but health-record refreshes, security patches, user support, and reanalysis against new reference genomes repeat. Phase 2 funding should not depend only on participant count. It should also track active researchers, access time, reproduced analyses, clinical findings, participant recontact, and data refresh rates.

Korea’s first scorecard needs how many people were recruited. Its second needs how many linked records became research-ready. Its third needs what was reproduced or discovered, and who received a result. If this sequence holds, Korea can use its distinctive combination of old cohorts, a large biobank, and national health data instead of simply following overseas programs.

  1. A cohort is a group recruited under common criteria and observed over time so that measurements can be linked to later health outcomes.

  2. Omics data measures one class of biological molecules at comprehensive scale, such as the genome, transcriptome, proteome, or metabolome.

  3. A federated approach runs the same analysis within each institution’s secure environment and combines only permitted results instead of collecting sensitive raw data in one place.