Senior Data Engineer
About this opportunity
Career Category
Engineering
Job Description
Role Description
We are seeking an experienced Senior Data Engineer to lead the design, development, and delivery of scalable enterprise data solutions. This role will build and optimize batch and real-time data pipelines, reusable integration frameworks, and governed data platforms that support analytics, AI, and self-service data access.
The ideal candidate has deep expertise in Databricks, Apache Spark, AWS, data modeling, governance, and production operations. The role also provides technical leadership, defines engineering standards, mentors other engineers, and partners with architecture, business, analytics, data science, and DevOps teams.
Experience in manufacturing, biotechnology, pharmaceutical, life sciences, or another regulated industry is preferred.
Roles and Responsibilities
Lead the design and development of scalable batch and real-time ETL/ELT pipelines.
Own complex data solutions from requirements and design through deployment and production support.
Build cloud-based lakehouse and data-platform solutions using Databricks and AWS.
Develop reusable, metadata-driven data integration frameworks.
Integrate structured, semi-structured, and unstructured data from enterprise, manufacturing, API, and third-party sources.
Optimize Spark workloads, Databricks compute, SQL queries, partitioning, storage, and caching for performance and cost.
Implement workflow orchestration, monitoring, alerting, data-quality controls, and recovery processes.
Develop CI/CD pipelines and automated testing for data solutions.
Implement metadata management, lineage, cataloging, governance, RBAC, and data-security controls.
Define data models, data contracts, integration patterns, and reusable engineering standards.
Lead architecture reviews, code reviews, troubleshooting, and root-cause analysis.
Mentor engineers and provide technical guidance across delivery teams.
Collaborate with architects, analysts, data scientists, product teams, and DevOps teams.
Support estimation, sprint planning, technical roadmaps, and delivery-risk management.
Evaluate emerging technologies and recommend solutions based on scalability, security, maintainability, and business value.
Participate in operational support, including occasional off-hours support.
Basic Qualifications
One of the following:
Master’s degree in Computer Science , Engineering, Information Technology, Data Science, or a related field and at least 7 years of relevant experience.
OR
Bachelor’s degree in a related field and at least 9 years of relevant experience.
Must-Have Skills
Advanced hands-on experience with Databricks, Apache Spark, PySpark , Spark SQL, Delta Lake, Python, and SQL.
Experience designing and operating production-grade batch and streaming pipelines.
Strong understanding of distributed computing, lakehouse architecture, data warehousing, and data integration.
Experience with Databricks Workflows or comparable orchestration tools.
Strong experience with AWS data, compute, storage, security, and monitoring services.
Experience with Spark performance tuning, cluster optimization, partitioning, and cost management.
Experience with Git, CI/CD, automated testing, monitoring, and production deployment.
Experience implementing data quality, metadata management, lineage, governance, and access controls.
Strong understanding of RBAC, least privilege, encryption, auditability, and regulated-data requirements.
Experience leading technical design, code reviews, and complex production implementations.
Ability to define reusable patterns, engineering standards, and development best practices.
Strong communication, collaboration, mentoring, and problem-solving skills.
Experience working in Agile or Scaled Agile delivery environments.
Preferred Qualifications
Experience in biotechnology, pharmaceutical, life sciences, manufacturing, or another regulated industry.
Experience with Unity Catalog, data products, Data Fabric, Data Mesh, or similar enterprise data architectures.
Experience building APIs and secure data services.
Experience with relational, NoSQL, analytical, operational, or vector databases.
Experience with OLAP and OLTP data modeling and performance tuning.
Experience with Kafka, Kinesis, or other streaming technologies.
Experience supporting AI and Generative AI solutions, including RAG, embeddings, vector search, and governed enterprise data access.
Familiarity with AI-assisted development tools such as GitHub Copilot, OpenAI Codex, or equivalent platforms.
Preferred Certifications
AWS Certified Data Engineer or another relevant AWS certification.
Databricks Certified Data Engineer Associate or Professional.
Scaled Agile Framework certification.
Soft Skills
Excellent analytical and troubleshooting skills.
Strong written, verbal, presentation, and stakeholder-communication skills.
High degree of ownership, initiative, and self-motivation.
Ability to manage multiple priorities in a fast-paced environment.
Ability to work effectively with global and virtual teams.
Strong attention to detail and commitment to engineering quality.
Ability to influence technical decisions and drive work to completion.
.
Job details
How this role compares
Computed from every other active Data & Digital role in our database, not just this employer's listings.
We currently track 92 comparable Senior Data & Digital roles across 23 biopharma companies.
Salary context
16 of 92 peers report a salary range (USD, annualized)
Peers share this role's job function and a matching or adjacent seniority level -- not necessarily the same therapeutic area or country.
Where these roles are based
Top locations among the 92 comparable roles
+ 7 more countries
Seniority mix
92 of 92 peers have a known seniority level
Therapeutic area mix
2 of 92 peers have a known therapeutic area; the rest are genuinely unlabeled, not hidden
Similar opportunities
The closest matches from our peer group, ranked by how similar they are, not how well you'd qualify for them -- treat this as market context, not a guaranteed shortlist; a weak match is labeled as one below.
How we calculate "similar"
No black box, no LLM guesswork: a deterministic score built from four normalized attributes. Here's this role's own peer group at different match levels, so you can see the mechanism, not just the result.
Every comparison starts from the same 100-point budget: 25 for working in the same function, 40 for the same therapeutic area, 20 for the same or adjacent seniority, 15 for the same country. A dimension we can't confirm on both sides contributes nothing, never a guess, never a free pass.
0 points, never a partial guess. A role we know almost nothing about beyond its function bottoms out at 25%; it never inflates to 100% just because there's little to compare against. Seniority uses a defined ladder (Associate → Manager → Associate Director → Senior → Principal → Director → Senior Director → Executive/VP) so "Director" and "Senior Director" count as adjacent, but "Director" and "Executive/VP" do not.