Amgen Technology Pvt Ltd. Posted September 9, 2026

Senior Data Scientist - Protein Data Pipelines

Hyderabad, India Full time
Notify me about similar jobs

Amgen Technology Pvt Ltd. is the source of truth for this posting and owns the application process. We surface normalized context and market comparison you won't find on the original listing.

About this opportunity

Career Category

Research

Job Description

Senior Data Scientist - Protein Data Pipelines

Role Summary

The Senior Data Scientist - Protein Data Pipelines will play a critical role in enabling predictive modeling for protein sequence, structure, and function by building scalable, reliable, and reproducible data pipelines. This role will focus on transforming protein property data and related scientific outputs into ML-amenable assets that support model training, inference, deployment, and ongoing use across research programs.

Working at the intersection of data engineering, MLOps, computational biology, and applied machine learning, this individual will partner with ML developers, wet-lab scientists, domain experts, and distributed technical teams to translate scientific and engineering needs into robust data and inference solutions. The successful candidate will develop reusable frameworks for data engineering, model inference, deployment, validation, testing, and monitoring across in-house and external machine learning models.

This role is ideal for someone who enjoys building production-ready scientific data systems, collaborating across disciplines, and converting complex domain needs into maintainable technical solutions that scale across discovery pipelines.

Key Responsibilities

Scalable Data Pipelines for model training

Design and maintain scalable data pipelines that support predictive model training, with emphasis on protein sequence or structure-to-function applications.

Build ML-model amenable data assets for protein property data that are readable, quality-controlled, reproducible, and suitable for reuse across programs.

Translate scientific and engineering needs into reliable data solutions that support ongoing research and model-development workflows.

Model Deployment & Inference Pipelines

Develop deployment strategies and pipelines to embed trained models into ongoing projects.

Develop reusable inference, deployment, and testing frameworks for in-house and external machine learning models.

Convert model-development outputs into maintainable technical solutions that can be used reliably by research teams.

Data Quality, Validation & Reproducibility

Establish data quality, validation, monitoring, and reproducibility practices for protein property and related scientific datasets.

Implement validation and monitoring approaches that improve confidence in downstream model training, inference, and deployment.

Document data lineage, assumptions, validation outcomes, and reproducibility practices to support long-term reuse.

Cross-Functional Collaboration & Technical Coordination

Serve as a liaison between machine-learning developers and domain experts, including wet-lab collaborators where applicable.

Own and mediate collaborations between ML developers and wet-lab teams to ensure that data, modeling, and experimental needs are aligned.

Coordinate technical work across distributed teams and help align implementation plans, dependencies, and delivery timelines.

Documentation & Knowledge Sharing

Document systems, pipeline behavior, operational expectations, and technical decisions to support adoption and maintenance.

Support knowledge sharing across research, ML, and engineering teams through clear documentation, examples, and reusable implementation patterns.

Scale data and modeling infrastructure practices across research programs and pipelines.

Basic Qualifications

Bachelor’s degree in Computational Biology, Bioinformatics, Life Sciences, Computational Chemistry, Chemical Engineering, Materials Science, Data Science, or a related quantitative field and relevant professional experience.

Experience Requirements

Bachelor’s degree and 6+ years of relevant experience, OR

Master’s degree and 4+ years of relevant experience, OR

PhD

Preferred Qualifications

Scalable Data Engineering

Strong experience building scalable data pipelines in Python and/or SQL.

Experience designing readable, reusable, and maintainable data-processing workflows for scientific or machine-learning applications.

Experience with data pipeline automation, preferably using Databricks.

MLOps, Inference & Deployment

Hands-on experience owning reusable, end-to-end MLOps for at least one machine learning model.

Experience with MLflow is preferred; experience with other model-lifecycle, deployment, or tracking frameworks is also welcome.

Experience developing deployment, inference, validation, or testing workflows that support production-like use of machine learning models.

Data Quality, Monitoring & Reproducibility

Knowledge of data quality control, validation, and monitoring practices.

Experience applying reproducibility practices to scientific data, model-training datasets, or inference workflows.

Ability to identify data quality risks and develop practical controls for downstream model use.

Scientific Domain Experience

Familiarity with computational biology, computational chemistry, computational materials science, or related fields.

Experience working with protein sequence, protein structure, protein property, or related scientific datasets is beneficial.

Preferred experience collaborating with wet-lab teams and translating experimental needs into data or modeling workflows.

Communication & Collaboration

Ability to communicate effectively with machine-learning developers, software and data engineers, domain experts, and research scientists.

Experience coordinating technical work across distributed or cross-functional teams.

Strong documentation habits and commitment to knowledge sharing, maintainability, and long-term adoption.

Success Measures

Success in this role will be demonstrated through:

Delivery of readable, quality-controlled, and reproducible data pipelines for protein property data.

Successful embedding of trained models into ongoing projects through reliable deployment and inference pipelines.

Increased reuse of data-engineering, inference, deployment, validation, and testing frameworks across research programs.

Improved confidence in model-training and inference data through practical quality, monitoring, and reproducibility practices.

Effective collaboration between ML developers, wet-lab teams, domain experts, and distributed technical partners.

Expansion of scalable data and modeling infrastructure across research programs and pipelines.

Typical Candidate Profile

The ideal candidate combines strong data-engineering and MLOps expertise with enough scientific domain fluency to work effectively with ML developers, and experimental collaborators. They enjoy building reusable systems that make complex scientific data reliable, reproducible, and actionable for predictive modeling.

Candidates may come from data science, data engineering, machine learning infrastructure, computational biology, computational chemistry, computational materials science, bioinformatics, or research informatics backgrounds. They are motivated by bridging scientific and engineering needs, supporting production-ready model use, and scaling technical solutions across discovery programs.

Organizational Impact

This role will build ML-amenable data pipelines for protein property data, mediate collaborations between ML developers and wet-lab teams, and scale data and modeling infrastructure across research programs and pipelines. By converting complex domain needs into maintainable, production-ready technical solutions, the Senior Data Scientist - Protein Data Pipelines will help accelerate reliable model development, deployment, and adoption across Large Molecule Discovery.

.

Job details

Seniority
Senior
Function
Information Technology
Therapeutic area
Not listed
Location
Hyderabad, India
Employment type
Full time

How this role compares

Computed from every other active Information Technology role in our database, not just this employer's listings.

We currently track 393 comparable Senior Information Technology roles across 49 biopharma companies.

393Comparable roles tracked
367Currently active
49Companies hiring similar roles
28Countries represented

Salary context

61 of 393 peers report a salary range (USD, annualized)

Peers share this role's job function and a matching or adjacent seniority level -- not necessarily the same therapeutic area or country.

This roleSubject Not listed on this posting
Highest disclosed · Sr. AI Science Lead · Lilly $267,000/yr – $391,600/yr
Peer group range $98,050 – $329,300 (median $183,750)

Where these roles are based

Top locations among the 393 comparable roles

India182
United States113
Spain22
Poland10
Portugal8
Japan5

+ 22 more countries

Seniority mix

393 of 393 peers have a known seniority level

Senior271
Principal67
Associate Director55

Therapeutic area mix

1 of 393 peers have a known therapeutic area; the rest are genuinely unlabeled, not hidden

Vaccines & Infectious Disease1

Similar opportunities

The closest matches from our peer group, ranked by how similar they are, not how well you'd qualify for them -- treat this as market context, not a guaranteed shortlist; a weak match is labeled as one below.

60%similar
Bristol-Myers Squibb Business Services India Private Limited Hyderabad, India Senior
Same function Same seniority Same country
60%similar
Bristol-Myers Squibb Business Services India Private Limited Hyderabad, India Senior
Same function Same seniority Same country
60%similar
Amgen Technology Pvt Ltd. Hyderabad, India Senior
Same function Same seniority Same country
60%similar
Amgen Technology Pvt Ltd. Hyderabad, India Senior
Same function Same seniority Same country
60%similar
Amgen Technology Pvt Ltd. Hyderabad, India Senior
Same function Same seniority Same country
60%similar
Amgen Technology Pvt Ltd. Hyderabad, India Senior
Same function Same seniority Same country

Notify me about similar jobs

Get an email when we spot other openings like this one – same job function, comparable seniority, roles you'd actually want to see.

How we calculate "similar"

No black box, no LLM guesswork: a deterministic score built from four normalized attributes. Here's this role's own peer group at different match levels, so you can see the mechanism, not just the result.

Every comparison starts from the same 100-point budget: 25 for working in the same function, 40 for the same therapeutic area, 20 for the same or adjacent seniority, 15 for the same country. A dimension we can't confirm on both sides contributes nothing, never a guess, never a free pass.

60%
Senior Analyst, Identity Governance & Administration
Bristol-Myers Squibb Business Services India Private Limited · Hyderabad, India · Senior
Function Therapeutic area Seniority Country
60%
Sr Associate Scrum Master
Amgen Technology Pvt Ltd. · Hyderabad, India · Senior
Function Therapeutic area Seniority Country
60%
Sr Associate Software Engineer – Full stack & AI
Amgen Technology Pvt Ltd. · Hyderabad, India · Senior
Function Therapeutic area Seniority Country
60%
Senior Associate BC Engineer - AI & Automation
Amgen Technology Pvt Ltd. · Hyderabad, India · Senior
Function Therapeutic area Seniority Country
Unmatched or unknown dimensions score exactly the same: 0 points, never a partial guess. A role we know almost nothing about beyond its function bottoms out at 25%; it never inflates to 100% just because there's little to compare against. Seniority uses a defined ladder (Associate → Manager → Associate Director → Senior → Principal → Director → Senior Director → Executive/VP) so "Director" and "Senior Director" count as adjacent, but "Director" and "Executive/VP" do not.