Lilly Posted August 28, 2026

Senior Scientific Data Curator

Indianapolis, Indiana, United States of America FULL_TIME
Notify me about similar jobs

Lilly is the source of truth for this posting and owns the application process. We surface normalized context and market comparison you won't find on the original listing.

About this opportunity

At Lilly, the work is demanding because patients are waiting. We unite caring with discovery to help make life better for people around the world, knowing that every decision, every detail, and every day matters. Headquartered in Indianapolis, Indiana, our over 50,000 employees around the globe take on complex challenges to discover and deliver life-changing medicines, strengthen how health is understood and managed, and support the communities we serve. This is hard, urgent, selfless work, but it’s work worth doing. If you’re driven by purpose and ready to bring your best to work that truly matters for patients, we invite you to join us.

Position Summary

The Senior Scientific Data Curator will lead the systematic discovery, assessment, harmonization, and quality assurance of Lilly’s scientific datasets across the full breadth of TuneLab’s modeling domains, small-molecule ADME/ADMET, safety and secondary pharmacology, in vivo pharmacokinetics and toxicology, antibody and biologics developability, and clinical PK/PD, in support of a strategic, cross-modality data unification initiative. This role sits at the intersection of biological and pharmacological domain expertise and data science, translating decades of fragmented, heterogeneous datasets spanning discovery through the clinic into a unified, AI-ready data infrastructure. The curator will partner closely with computational scientists, DMPK scientists, pharmacometricians, antibody engineers, and external consortium collaborators to ensure that the data substrate underpinning TuneLab’s federated AI/ML models is comprehensive, well-documented, and scientifically sound.

Core Responsibilities

Data Discovery & Assessment

Conduct comprehensive inventory of historical and ongoing datasets across TuneLab’s modeling domains, small-molecule ADME/ADMET, safety and secondary pharmacology, in vivo PK and toxicology, antibody and biologics developability, and clinical PK/PD, spanning therapeutic areas (oncology, immunology, metabolic diseases, neuroscience, etc.) and 20+ years of discovery, preclinical, and clinical data

Assess and score data quality, completeness, and integration feasibility for each dataset, accounting for the distinct data structures of each domain, including assay and dose–response measurements, concentration–time profiles and dosing regimens, in vivo study readouts, sequence- and structure-derived features for biologics, and biomarker and clinical covariate data

Map metadata gaps across legacy systems and source platforms, documenting study contexts, assay and protocol methods, protocol deviations, data quality flags, and provenance information

Develop automated pipelines (including LLM-assisted extraction where appropriate) to identify and extract domain-relevant data from internal documents, assay databases, and study reports into standardized, model-ready formats

Produce a prioritized data assessment report recommending which domains, therapeutic areas, and indications to integrate first, based on data volume, complexity, portfolio relevance, and model feasibility

Data Harmonization & Integration

Design and implement standardized, extensible schemas for the integrated multi-domain database, working with computational partners to ensure AI/ML readiness across small-molecule and biologics modalities

Build and maintain data harmonization pipelines: label normalization, unit and assay-condition standardization across studies and sources, time-point and dose alignment, sequence and structure normalization for biologics, and covariate encoding

Apply domain-driven quality control practices, sequence validation, hidden duplicate detection, cross-source discrepancy resolution, and cross-species dataset integration using allometric scaling where applicable

Develop and execute outlier detection protocols, flagging and adjudicating anomalous values in collaboration with clinical pharmacologists, DMPK scientists, toxicologists, and antibody engineers as appropriate to the domain

Create reproducible data quality assurance workflows with documented acceptance criteria and audit trails

Curate and enrich metadata to enable cross-study and cross-domain querying, linking compound and molecule identifiers, sequence and construct identifiers, assay methods, formulation details, and study design parameters

Cross-functional Partnership

Serve as the primary data domain expert for external consortium partners working within Lilly’s controlled cloud environment

Collaborate with pharmacometricians, DMPK scientists, toxicologists, and antibody engineers to validate harmonized datasets against legacy models and established analyses (e.g., NONMEM/Monolix outputs for the clinical PK/PD domain)

Work with the TuneLab ML team to ensure curated datasets meet the input specifications for the platform’s multi-task ML models, representation and foundation-model embeddings, and mechanistic/hybrid PK/PD frameworks (e.g., Neural ODE, SINDy)

Contribute to platform deployment by supporting the development of data dictionaries, user documentation, and training materials for internal and consortium end users

Required Qualifications

M.S. or Ph.D. in Computational Data Science, Pharmacometrics, Pharmaceutical Sciences, Computational Biology, Cheminformatics, Biomedical Informatics, or a related quantitative discipline

1+ year of hands-on experience curating, harmonizing, or building analysis-ready datasets from biological, chemical, or clinical data sources

Demonstrated skill in scientific dataset construction with domain-driven QC: sequence validation, duplicate detection, cross-source discrepancy resolution, or equivalent rigor applied to noisy real-world data

Proficiency in Python and/or R for data wrangling, transformation, and quality checks at scale

Working knowledge of pharmacological or chemical data structures across one or more TuneLab domains, for example, ADME/ADMET assay data, in vivo PK and toxicology readouts, antibody and biologics developability measurements, or clinical concentration–time and covariate data

Track record of producing clear data documentation, quality reports, and data dictionaries

Preferred Qualifications

Experience integrating cross-species datasets (e.g., allometric scaling) or multi-source public/internal data to expand training sets for ML models

Familiarity with cheminformatics and computational biology tooling, including biologics-specific tools (e.g., ANARCI, protein language model embeddings, molecular operating environment software)

Exposure to LLM-assisted workflows for information extraction, document parsing, or automated data-pipeline development

Experience with the data conventions of one or more TuneLab domains, population PK/PD modeling tools (NONMEM, Monolix, nlmixr) or CDISC standards (SDTM, ADaM) for clinical PK/PD; ADMET/DMPK assay conventions for small molecules; or developability assays for biologics

Familiarity with cloud-based data infrastructure (AWS, Azure, or GCP) and version-controlled, reproducible analysis environments (Git, Docker, Conda)

Prior experience providing curated data to federated learning or collaborative ML initiatives

Lilly is dedicated to helping individuals with disabilities to actively engage in the workforce, ensuring equal opportunities when vying for positions. If you require accommodation to submit a resume for a position at Lilly, please complete the accommodation request form ( https://careers.lilly.com/us/en/workplace-accommodation ) for further assistance. Please note this is for individuals to request an accommodation as part of the application process and any other correspondence will not receive a response.

Lilly is proud to be an EEO Employer and does not discriminate on the basis of age, race, color, religion, gender identity, sex, gender expression, sexual orientation, genetic information, ancestry, national origin, protected veteran status, disability, or any other legally protected status.

Our employee resource groups (ERGs) offer strong support networks for their members and are open to all employees. Our current groups include: Africa, Middle East, Central Asia (AMECA), Black Employees at Lilly (BE@Lilly), Chinese Culture Network (CCN), EnAble, Evolve, Lilly Indian Network (LIN), Organization of Latinx at Lilly (OLA), Pride (LGBTQ+ Allies), Veterans Leadership Network (VLN) and Women’s Initiative for Leading at Lilly (WILL).

Actual compensation will depend on a candidate’s education, experience, skills, and geographic location.  The anticipated wage for this position is

$132,000 - $244,200

Full-time equivalent employees also will be eligible for a company bonus (depending, in part, on company and individual performance). In addition, Lilly offers a comprehensive benefit program to eligible employees, including eligibility to participate in a company-sponsored 401(k); pension; vacation benefits; eligibility for medical, dental, vision and prescription drug benefits; flexible benefits (e.g., healthcare and/or dependent day care flexible spending accounts); life insurance and death benefits; certain time off and leave of absence benefits; and well-being benefits (e.g., employee assistance program, fitness benefits, and employee clubs and activities).Lilly reserves the right to amend, modify, or terminate its compensation and benefit programs in its sole discretion and Lilly’s compensation practices and guidelines will apply regarding the details of any promotion or transfer of Lilly employees.

#WeAreLilly

Job details

Seniority
Senior
Function
Other / Corporate Functions
Therapeutic area
Not listed
Location
Indianapolis, Indiana, United States of America
Employment type
FULL_TIME

How this role compares

Computed from every other active Other / Corporate Functions role in our database, not just this employer's listings.

We don't have enough classified peer data for this role yet, so there's no comparison to show. This happens when a posting's title/category doesn't match any taxonomy rule -- it's excluded rather than compared against the wrong peer group.

Notify me about similar jobs

Get an email when we spot other openings like this one – same job function, comparable seniority, roles you'd actually want to see.