Genetics Data Engineer 1
Merck KGaA
Senior Principle Data Engineer - Genetics
Your Role
You will advance our human quantitative genetics strategy by providing the data engineering fundamentals that enable downstream analysis. You will collaborate with quantitative geneticists, data scientists, data engineers, platform experts, IT, and others to ensure that data availability and quality are never the bottlenecks for our analyses. In that collaboration, you will provide the vision and implementation for how our FAIR data environment works across internal and external platforms such as biobank trusted research environments (TREs).
You will work with other scientists to:
- Bring software tools, reference datasets, and genetic data into internal environments.
- Bring tools, containers, and reference data into TREs (e.g., UK Biobank Research Analysis Platform, All of Us Researcher Workbench) and manage their deployment and versioning.
- Ingest and maintain connections to genetic reference databases (OpenTargets, GWAS Catalog, ClinVar, dbSNP, OMIM, HGMD, ChEMBL, DrugBank) and integrate them with the internal knowledge graph (Synaptix) and analytics platforms.
- Automate data QC for diverse genomic data types.
- Develop, test, and execute reproducible analysis pipelines using workflow managers and containerized environments (Docker, Singularity).
- Return results from TREs to our internal platforms in accordance with each biobank's privacy and data protection policies.
- Optimize query performance and pipeline execution to support rapid-turnaround target assessments and in-licensing due diligence (20-25 targets per year requiring fast genetic evaluation).
- Contribute to the design and implementation of agentic AI workflows for automated genetic evidence generation, integrating genetics pipelines with the broader agentic AI platform.
- Build and maintain interactive dashboards and data services that expose genetic evidence to project teams, leadership, and due diligence committees.
- Link genetic data to our AI tools and platforms, ensuring seamless data flow between genetic analyses and downstream decision-support systems. Experience building production-grade data pipelines for genetic and genomics datasets,including familiarity with common formats ( VCF, PLINK/BED/BIM/FAM/BGEN,GWAS summary statistics).
Don't want to miss the next one?
Subscribe to daily email alerts for roles matching your interests.

