Senior Staff Data Engineer (London, UK, United Kingdom, Aberdeen City)
About this opportunity
AI ACCELERATOR
Most diseases are still poorly understood at a biological level. Despite decades of research, the causal mechanisms driving many conditions remain unclear, limiting our ability to identify the right targets, design the right interventions and bring the right medicines to patients.
The AI Accelerator exists to change that. Based in London and sitting within Computational Innovation (@computationalinnovation), a global organisation spanning computational biology, human genetics, data excellence and AI, the Accelerator's mission is to build production-quality AI capabilities that deepen our understanding of disease biology and increase probability of success.
We do this by applying neural-based methods across the biomedical data landscape to integrate heterogeneous, multimodal data sources, infer biological relationships and embed causal thinking into what we build. The goal is not just to predict but to explain and understand why disease occurs.
It could be electronic health records and medical imaging to support patient segmentation. It could be 'omics data to identify novel therapeutic targets. It could be predicting transcriptional change for a given disease-causing variant. It could be simulating the effect of modulating a target of interest.
None of this is possible without data. The quality, accessibility and engineering of biomedical data ultimately determine the quality of the models built upon them. The Data Engineering function provides the foundation on which the AI Accelerator operates by integrating diverse clinical, biological and real-world datasets into trustworthy, AI-ready assets that can support large-scale foundation model development and downstream scientific discovery.
THE POSITION
We are seeking a Senior Staff Data Engineer to join Computational Innovation's AI Accelerator. In this role, you will own the data engineering architecture, standards and technical direction for the AI Accelerator, ensuring that our foundation-model ambitions are built upon high-quality data. You will define how diverse, multimodal, biomedical datasets are integrated and provisioned, creating the data foundation that enables model training, evaluation and deployment, supporting current research priorities and future scientific opportunities.
This is both a strategic and hands-on technical leadership role. You will define architectural standards and engineering practices, solve the most challenging data engineering problems yourself, and serve as the senior technical authority for data engineering within the AI Accelerator. As a senior member of the AI Enablement leadership team, you will help shape the direction of the function alongside the Senior Staff AI Infrastructure Engineer and Senior Staff MLOps Engineer.
This is a unique opportunity to be part of a critical strategic initiative for a pharmaceutical company that invests heavily in research and development to discover and develop innovative therapies that can improve and extend lives in areas of high unmet medical need. Your work will directly influence how effectively we can understand disease, identify therapeutic opportunities and accelerate the development of innovative medicines for patients.
Key Responsibilities
Set the technical direction, strategy and roadmap for data engineering, aligned to AI Accelerator priorities.
Own the data engineering architecture for the AI Accelerator, defining and evolving a layered ("medallion"-style) architecture, feature and embedding provisioning patterns, and the harmonised multimodal data foundation that underpins model development.
Establish data engineering standards and engineering practices, including data quality controls, testing, CI/CD for data, reproducibility standards, data contracts, metadata management and dataset versioning.
Deliver hands-on leadership on the most difficult and highest-impact data engineering challenges, including integrating novel and complex modalities such as genomics, transcriptomics, imaging, clinical and real-world datasets.
Partner with the Data Excellence community and IT to align on governance, ontologies, metadata harmonisation, stewardship and shared enterprise data foundations.
Establish ways of working and coach other team members, onboarding, mentoring and technically leading data engineers as the team scales, while acting as the senior escalation point for complex data engineering challenges.
Requirements
PhD or MSc and equivalent experience in a STEM subject.
Extensive experience operating at the senior staff level within a data engineering function, including defining architecture, standards and technical direction, along with mentoring data engineers and establishing strong engineering principles withing a growing team
Deep expertise in large-scale data engineering, including distributed processing frameworks, pipeline orchestration, cloud data platforms and data integration at scale, with experience integrating internal and third-party datasets and working effectively with external data providers and technology partners.
Strong understanding of biomedical and healthcare data domains, such as genomics, transcriptomics, multi-omics, imaging, clinical data, electronic health records or real-world data, alongside an understanding of machine learning data requirements, including versioning, reproducibility and tensor-based workflows for large-scale AI systems.
Experience implementing data governance in practice, including metadata management, lineage, provenance, ontologies, cataloguing and FAIR principles and familiarity with Trusted Research Environments (TREs) and controlled-access research data environments.
Strong collaboration and influencing skills across technical and non-technical stakeholders, with the ability to communicate complex technical concepts clearly.
This is a hybrid role with approximately 4 days a week in the office.
WHY THIS IS A GREAT PLACE TO WORK
Boehringer Ingelheim has been recognised as a Top Employer in the UK, demonstrating our commitment to building an exceptional workplace through strong people practices and supportive HR policies.
To learn more about why BI is a great place to work, visit:
https://www.boehringer-ingelheim.co.uk/careers/uk-careers/why-great-place-work
]]>
Job details
How this role compares
Computed from every other active Data & Digital role in our database, not just this employer's listings.
We currently track 88 comparable Senior Data & Digital roles across 24 biopharma companies.
Salary context
22 of 88 peers report a salary range (USD, annualized)
Peers share this role's job function and a matching or adjacent seniority level -- not necessarily the same therapeutic area or country.
Where these roles are based
Top locations among the 88 comparable roles
+ 5 more countries
Seniority mix
88 of 88 peers have a known seniority level
Therapeutic area mix
1 of 88 peers have a known therapeutic area; the rest are genuinely unlabeled, not hidden
Similar opportunities
The closest matches from our peer group, ranked by how similar they are, not how well you'd qualify for them -- treat this as market context, not a guaranteed shortlist; a weak match is labeled as one below.
Notify me about similar jobs
Get an email when we spot other openings like this one – same job function, comparable seniority, roles you'd actually want to see.
How we calculate "similar"
No black box, no LLM guesswork: a deterministic score built from four normalized attributes. Here's this role's own peer group at different match levels, so you can see the mechanism, not just the result.
Every comparison starts from the same 100-point budget: 25 for working in the same function, 40 for the same therapeutic area, 20 for the same or adjacent seniority, 15 for the same country. A dimension we can't confirm on both sides contributes nothing, never a guess, never a free pass.
0 points, never a partial guess. A role we know almost nothing about beyond its function bottoms out at 25%; it never inflates to 100% just because there's little to compare against. Seniority uses a defined ladder (Associate → Manager → Associate Director → Senior → Principal → Director → Senior Director → Executive/VP) so "Director" and "Senior Director" count as adjacent, but "Director" and "Executive/VP" do not.