Johnson & Johnson Posted September 1, 2026

Senior Data Scientist, Biologics Discovery - Madrid, ES

Madrid, Spain Full time
Notify me about similar jobs

Johnson & Johnson is the source of truth for this posting and owns the application process. We surface normalized context and market comparison you won't find on the original listing.

About this opportunity

At Johnson & Johnson, we believe health is everything. Our strength in healthcare innovation empowers us to build a world where complex diseases are prevented, treated, and cured, where treatments are smarter and less invasive, and solutions are personal. Through our expertise in Innovative Medicine and MedTech, we are uniquely positioned to innovate across the full spectrum of healthcare solutions today to deliver the breakthroughs of tomorrow, and profoundly impact health for humanity. Learn more at jnj.com .

As guided by Our Credo, Johnson & Johnson is responsible to our employees who work with us throughout the world. We provide an inclusive work environment where each person is considered as an individual. At Johnson & Johnson, we respect the diversity and dignity of our employees and recognize their merit.

Job Function:

Data Analytics & Computational Sciences

Job Sub Function:

Data Science

Job Category:

Scientific/Technology

All Job Posting Locations:

Madrid, Spain

Job Description:

Our expertise in Innovative Medicine is informed and inspired by patients, whose insights fuel our science-based advancements. Visionaries like you work on teams that save lives by developing the medicines of tomorrow.

Join us in developing treatments, finding cures, and pioneering the path from lab to life while championing patients every step of the way.

Learn more at https://www.jnj.com/innovative-medicine

About the opportunity

Johnson & Johnson Innovative Medicine is seeking a Senior Data Scientist dedicated to our Biologics Discovery organization. This role sits within our Data Science team and partners closely with our In Silico Discovery (ISD) organization - the group that builds the molecular design and property-prediction models (for example, developability, affinity and binding, and other molecular-property and liability-risk models) that guide which biologic molecules to design, make, and advance. ISD owns core molecular model development; you will build the data-facing ML capabilities (featurization, model-ready datasets) that make ISD's models faster to build and better to trust.

This position will be based at one of our office locations in either Spring House, PA (strongly preferred), Titusville, NJ, or Raritan, NJ, USA; or Madrid, Spain. (No remote option.)

Please note that this role is available across multiple countries and may be posted under different requisition numbers to comply with local requirements. While you are welcome to apply to any or all of the postings, we recommend focusing on the specific country(s) that align with your preferred location(s):

USA - Requisition Number: R-095854

Spain - Requisition Number: R-096793

Why this role matters: Biologics Discovery is generating rich, fast-growing data across assays, sequences, and modalities, and the opportunity now is to make that data fully model-ready and seamlessly available for ML. This role ensures biologics data is structured for training, and that applied ML on discovery data helps scientists prioritize molecules, flag risks, and generate hypotheses earlier - strengthening the interface to ISD's models rather than duplicating them.

Position Summary

You will design robust featurization approaches and curate standardized, traceable, model-ready datasets from biologics assay, biophysical, sequence, and construct data. You operate at the interface between our data-generating and data-infrastructure partners and In Silico Discovery (ISD), ensuring the datasets and features you create strengthen ISD's molecular property models. This is an opportunity to shape how AI learns from every biologics experiment.

Key Responsibilities

Scientific Data Analysis & Enablement

Apply modern data science tools to explore, integrate, and characterize heterogeneous biologics discovery data (e.g., antibody/protein sequence, construct, assay, and biophysical data).

Work with discovery scientists to translate scientific questions and DMTL (design-make-test-learn) decision points into clear data and analytical requirements.

Identify data quality issues, biases, gaps, and risks such as leakage or distribution shift that could affect downstream modeling or scientific interpretation.

Support the effective use of data for molecule prioritization, risk identification, and hypothesis generation.

Featurization & Model-Ready Data

Develop featurization and model-ready datasets from biologics discovery data.

Work with data engineers to specify the features, labels, and levels of aggregation that models need, preserving raw representations where information matters.

Curate, document, and version datasets so modeling is reproducible and traceable.

Partnership, Rigor & Growth

Collaborate with ISD to hand off standardized, traceable training datasets and align on where Data Science enables versus where ISD owns modeling.

Partner with Discovery scientists to frame ML problems around real decision points in the design-make-test-learn (DMTL) cycle.

Work closely with ontology and MLOps colleagues so datasets carry consistent semantics and models move reliably from development into use.

Champion reproducibility, documentation, and responsible AI.

Why This Role Is Unique

This is an opportunity to apply ML where it truly moves the needle in biologics discovery - grounded in real assay and sequence data, tightly partnered with world-class molecular modeling, and with real room to grow your scope, technical leadership, and impact as you build a track record of delivery.

Qualifications

Required

Master's or Ph.D. in Computer Science, Machine Learning, Computational Biology, Bioinformatics, Statistics, or a related field.

At least 2 years of applied ML experience, including model development, evaluation, and dataset curation on complex scientific or biomedical data.

Strong proficiency with Python and the modern ML stack (e.g., PyTorch, scikit-learn) and SQL.

Experience turning complex, heterogeneous experimental data into robust features and training sets, with exposure to cloud training and data infrastructure.

Sound understanding of evaluation, validation, and the risks of leakage and distribution shift.

Ability to collaborate effectively with experimental scientists and modeling partners in a matrixed R&D environment.

Preferred

Experience with biologics, antibody/protein sequence models, or protein language models.

Experience with active learning, Bayesian optimization, or sequence-based generative models for molecular design.

Familiarity with biophysical/assay data and developability endpoints.

Experience with MLOps, experiment tracking, and model monitoring.

Familiarity with how ontologies or knowledge graphs support data reuse and AI-ready datasets.

This position will be based at one of our office locations in either Spring House, PA (strongly preferred), Titusville, NJ, or Raritan, NJ, USA; or Madrid, Spain. (No remote option.)

#LI-SL

#JNJDataScience

#JNJIMRND-DS

#JRDDS

#LI-Hyrbid

# 3

Required Skills:

Preferred Skills:

Advanced Analytics, Business Intelligence (BI), Coaching, Collaboration, Critical Thinking, Data Analysis, Database Management, Data Privacy Standards, Data Reporting, Data Savvy, Data Science, Data Visualization, Econometric Models, Process Improvements, Technical Credibility, Technologically Savvy, Workflow Analysis

The anticipated base pay range for this position is:

€55,400.00 - €87,860.00

Benefits:

In addition to base pay, we offer the following benefits*: an annual bonus with set target (% of pay) depending on pay grade / location, where the actual amount is based on the employees’ and companies’ performance of the previous calendar year, or sales commissions. Moreover, we offer vacation days, parental leave for a minimum of 12 weeks, bereavement leave, caregiver leave, volunteer leave, well-being reimbursement, programs for financial, physical and mental health. We also offer service anniversary and recognition awards, and subject to the terms of their respective plans, employees - and in some location’s eligible dependents - can participate in several insurance plans. For more information, visit Employee benefits | Supporting well-being & career growth | Johnson & Johnson Careers.

*This is for informative purposes only. Amounts and actual benefits may vary by location and are subject to change.

Job details

Seniority
Senior
Function
Biostatistics & Data Science
Therapeutic area
Not listed
Location
Madrid, Spain
Employment type
Full time

How this role compares

Computed from every other active Biostatistics & Data Science role in our database, not just this employer's listings.

We currently track 103 comparable Senior Biostatistics & Data Science roles across 28 biopharma companies.

103Comparable roles tracked
102Currently active
28Companies hiring similar roles
14Countries represented

Salary context

41 of 103 peers report a salary range (USD, annualized)

Peers share this role's job function and a matching or adjacent seniority level -- not necessarily the same therapeutic area or country.

This roleSubject Not listed on this posting
Lowest disclosed · Principal Computational Statistician · Lilly $66,000/yr – $193,600/yr
Highest disclosed · Biostatistics Associate Director (HYBRID) · Vertex Pharmaceuticals Inc (US) $174,800/yr – $262,200/yr
Peer group range $129,800 – $218,500 (median $184,000)

Where these roles are based

Top locations among the 103 comparable roles

United States62
United Kingdom8
India8
China6
Belgium4
Japan3

+ 8 more countries

Seniority mix

103 of 103 peers have a known seniority level

Senior43
Associate Director32
Principal28

Therapeutic area mix

8 of 103 peers have a known therapeutic area; the rest are genuinely unlabeled, not hidden

Oncology4
Rare Disease2
Immunology2

Similar opportunities

The closest matches from our peer group, ranked by how similar they are, not how well you'd qualify for them -- treat this as market context, not a guaranteed shortlist; a weak match is labeled as one below.

60%similar
Bayer Barcelona, Spain Senior
Same function Same seniority Same country
50%similar
AstraZeneca Barcelona, Spain Associate Director
Same function Adjacent seniority Same country
50%similar
Johnson & Johnson Madrid, Spain Principal
Same function Adjacent seniority Same country
45%similar
Lilly Indianapolis, Indiana, United States of America Senior
Same function Same seniority
45%similar
Lilly Indianapolis, Indiana, United States of America Senior
Same function Same seniority
45%similar
AbbVie Worcester, MA Senior
Same function Same seniority

Notify me about similar jobs

Get an email when we spot other openings like this one – same job function, comparable seniority, roles you'd actually want to see.

How we calculate "similar"

No black box, no LLM guesswork: a deterministic score built from four normalized attributes. Here's this role's own peer group at different match levels, so you can see the mechanism, not just the result.

Every comparison starts from the same 100-point budget: 25 for working in the same function, 40 for the same therapeutic area, 20 for the same or adjacent seniority, 15 for the same country. A dimension we can't confirm on both sides contributes nothing, never a guess, never a free pass.

60%
Senior Data Scientist (Barcelona, Cataluña, ES)
Bayer · Barcelona, Spain · Senior
Function Therapeutic area Seniority Country
45%
Sr. Advisor - Statistics (RWE)
Lilly · Indianapolis, Indiana, United States of America · Senior
Function Therapeutic area Seniority Country
45%
Sr Data Scientist (Tianjin, Tianjin, CN)
Novo Nordisk · Tianjin, China · Senior
Function Therapeutic area Seniority Country
45%
Senior Real World Evidence Expert (UK)
UCB · Slough, Berkshire, United Kingdom · Senior
Function Therapeutic area Seniority Country
Unmatched or unknown dimensions score exactly the same: 0 points, never a partial guess. A role we know almost nothing about beyond its function bottoms out at 25%; it never inflates to 100% just because there's little to compare against. Seniority uses a defined ladder (Associate → Manager → Associate Director → Senior → Principal → Director → Senior Director → Executive/VP) so "Director" and "Senior Director" count as adjacent, but "Director" and "Executive/VP" do not.