Process Data Engineer I - Pharmaceutical Product Development
Notify me about similar jobsWe couldn't confirm this posting on Bristol-Myers Squibb Business Services India Private Limited's site in our latest checks -- see current similar openings below.
About this opportunity
At Bristol Myers Squibb, our employees often ask, “Who are you working for?”, a question that fuels collaboration, accountability, and urgency in our work. Our purpose-driven culture inspires us to discover, develop, and deliver innovative medicines to prevail over serious diseases. We offer uniquely interesting and meaningful work, opportunities for growth, and a supportive environment that values inclusion, wellbeing, flexibility, and comprehensive benefits. This is work that transforms the lives of patients, and the careers of those who do it.
Build the data foundation that helps us get to an AI native state
Advanced AI is only as powerful as the data that enables it.
We are seeking an early-career Data Engineer passionate about creating high-quality scientific data products and contributing to the development of data fabric to support advanced analytics, AI, machine learning, and scientific decision-making across Pharmaceutical Product Development.
The ideal candidate will be an engineer who sees data not as a collection of tables, but as a strategic asset that powers scientific discovery and AI innovation. This role will provide an opportunity to work on complex datasets, modern cloud technologies, and cutting-edge digital transformation initiatives, with applications across chemical process development, biologics development, drug product development, and analytical development.
What You Will Do
Build scientific data products for product development focused on datasets for molecular features, material properties, laboratory data, process and manufacturing parameters, stability and product performance data.
Contribute to initiatives for data structuring and data contextualization.
Develop and maintain scalable data pipelines supporting analytics, AI, and scientific modelling initiatives.
Transform raw scientific and operational data into trusted, model-ready data products.
Work with cross-functional teams to improve data accessibility, reliability, and quality.
Support enterprise initiatives involving Databricks, Data Fabric, and cloud-native architectures.
Automate data workflows and reduce manual effort through engineering best practices.
Contribute to development of reusable data assets supporting Product Development innovation across US, Europe and India.
What Makes You Successful
You are naturally curious and:
Ask why before asking how.
Investigate root causes rather than treating symptoms.
Enjoy solving messy, ambiguous data problems.
Balance technical rigor with practical execution.
Work effectively with scientists, analysts, data scientists, and engineers.
Qualifications
Bachelor's or Master’s degree in Computer Science, Chemical Engineering, Information Systems, Bioinformatics, Biotechnology or related field with 2+ years of industry experience.
Technical hands-on experience with modern data technologies.
Data Engineering: SQL, Python, ETL/ELT Development, Delta Lake, Lakehouse Architecture, Medallion Architecture, Data Modelling, Data Warehousing, Distributed Computing, Data Validation
Cloud & Modern Data Platforms: Databricks, dbt, AWS
Modern Data Practices: Data Product Design, Vector Databases, Data Quality Engineering, Data Observability, Metadata Management, Master Data Management, Data Lineage, Data Governance principles, Data Cataloguing
Emerging Technologies: AI-Ready Data Foundations, Vector Database fundamentals, Semantic Layer Design, Knowledge Graph Concepts, Data Foundations for GenAI & Agentic AI Applications
Software Engineering & Delivery Practices: Git & Version Control, API Integration, Workflow Automation, CI/CD Fundamentals, Agile Delivery
Strong analytical and problem-solving skills.
Excellent communication and collaboration abilities.
Preferred
Experience working with pharmaceutical product development datasets in the scientific, manufacturing, or laboratory data domains
Exposure to scientific, analytical, or laboratory-based methods through academic coursework, research with an aptitude for understanding the data generated by these techniques
Familiarity with Data products and AI/ML workflows
Experience with PySpark, Python, SQL, dbt, Databricks
We hire for skills and capabilities, not just credentials – if this role excites you, but doesn’t perfectly match your resume, we encourage you to apply anyway.
How We Work
Where you work matters – because collaboration, innovation and patient impact happen in many settings. Our roles are structured across four work models: site-essential, site-by-design, field-based and remote-by-design. The model assigned to this role is based on its core responsibilities. Learn more at https://careers.bms.com/ways-of-working.
Supporting People with Disabilities
BMS is dedicated to ensuring that people with disabilities can excel through a transparent recruitment process, reasonable workplace accommodations/adjustments and ongoing support in their roles. Applicants can request a reasonable workplace accommodation/adjustment prior to accepting a job offer. If you require reasonable accommodations/adjustments in completing this application, or in any part of the recruitment process, direct your inquiries to adastaffingsupport@bms.com . Visit careers.bms.com/eeo-accessibility to access our complete Equal Employment Opportunity statement.
Candidate Rights
BMS will consider qualified applicants with arrest and conviction records, pursuant to applicable laws in your area.
For roles based in Los Angeles County only: If you live in or expect to work from Los Angeles County if hired for this position, please visit this page for important additional information: https://careers.bms.com/california-residents/
Data Protection
We will never request payments, financial information, or social security numbers during our application or recruitment process. Learn more about protecting yourself at https://careers.bms.com/fraud-protection .
Any data processed in connection with role applications will be treated in accordance with applicable data privacy policies and regulations.
If this posting is missing required information required by local law or incorrect, contact BMS at TAEnablement@bms.com with the Job Title and Requisition number. Do not send application-related inquiries to this email. To check your application status, please login to your Candidate Home Account.
R1605604 : Process Data Engineer I - Pharmaceutical Product Development
Job details
How this role compares
Computed from every other active Data & Digital role in our database, not just this employer's listings.
We currently track 251 comparable Data & Digital roles across 45 biopharma companies.
Salary context
72 of 251 peers report a salary range (USD, annualized)
Peers share this role's job function. This posting doesn't list a seniority level, so peers aren't narrowed by seniority either -- the range below may span more levels than usual.
Where these roles are based
Top locations among the 251 comparable roles
+ 16 more countries
Seniority mix
142 of 251 peers have a known seniority level
Therapeutic area mix
4 of 251 peers have a known therapeutic area; the rest are genuinely unlabeled, not hidden
Currently open similar roles
This posting is closed, but here are live openings closest to it -- ranked by similarity, not by how well you'd qualify for them.
Notify me about similar jobs
Get an email when we spot other openings like this one – same job function, comparable seniority, roles you'd actually want to see.
How we calculate "similar"
No black box, no LLM guesswork: a deterministic score built from four normalized attributes. Here's this role's own peer group at different match levels, so you can see the mechanism, not just the result.
Every comparison starts from the same 100-point budget: 25 for working in the same function, 40 for the same therapeutic area, 20 for the same or adjacent seniority, 15 for the same country. A dimension we can't confirm on both sides contributes nothing, never a guess, never a free pass.
0 points, never a partial guess. A role we know almost nothing about beyond its function bottoms out at 25%; it never inflates to 100% just because there's little to compare against. Seniority uses a defined ladder (Associate → Manager → Associate Director → Senior → Principal → Director → Senior Director → Executive/VP) so "Director" and "Senior Director" count as adjacent, but "Director" and "Executive/VP" do not.