Data Engineer
About this opportunity
Career Category
Engineering
Job Description
Role Description:
The role is responsible for designing, building, maintaining, analyzing, and interpreting data to provide actionable insights that drive business decisions. This role involves working with large datasets, developing reports, supporting and executing data governance initiatives and visualizing data to ensure data is accessible, reliable, and efficiently managed. The ideal candidate has strong technical skills, experience with big data technologies, and a deep understanding of data architecture and ETL processes
Roles & Responsibilities:
Design, develop, and maintain data solutions for data generation, collection, and processing
Be a key team member that assists in design and development of the data pipeline
Create data pipelines and ensure data quality by implementing ETL processes to migrate and deploy data across systems
Contribute to the design, development, and implementation of data pipelines, ETL/ELT processes, and data integration solutions
Take ownership of data pipeline projects from inception to deployment, manage scope, timelines, and risks
Collaborate with cross-functional teams to understand data requirements and design solutions that meet business needs
Develop and maintain data models, data dictionaries, and other documentation to ensure data accuracy and consistency
Implement data security and privacy measures to protect sensitive data
Leverage cloud platforms (AWS preferred) to build scalable and efficient data solutions
Collaborate and communicate effectively with product teams
Collaborate with Data Architects, Business SMEs, and Data Scientists to design and develop end-to-end data pipelines to meet fast-paced business needs across geographic regions
Identify and resolve complex data-related challenges
Adhere to best practices for coding, testing, and designing reusable code/component
Explore new tools and technologies that will help to improve ETL platform performance
Participate in sprint planning meetings and provide estimations on technical implementation
Design and develop data pipelines leveraging Databricks, PySpark, and SQL to ingest, transform, and process large-scale datasets.
Engineer solutions for both structured and unstructured data to enable advanced analytics and insights.
Implement automated workflows for data ingestion, transformation, and deployment using Databricks Jobs and notebooks, with ongoing monitoring and scheduling.
Apply performance optimization techniques, including Spark job tuning, caching, partitioning, and indexing, to improve scalability and efficiency.
Build integrations with multiple data sources, such as SQL databases, APIs, and cloud storage platforms, ensuring seamless connectivity and reliability.
Collaborate effectively with global teams across time zones to maintain alignment, resolve issues, and deliver on shared objectives.
Basic Qualifications and Experience:
Bachelor’s / Master’s degree and 4 to 8 years of Computer Science, IT or related field experience
Functional Skills:
Must-Have Skills
Hands-on experience with big data technologies and platforms, such as Databricks, Apache Spark (PySpark, SparkSQL), workflow orchestration, performance tuning on big data processing
Proficiency in data analysis tools (e.g. SQL) and experience with data visualization tools
Excellent problem-solving skills and the ability to work with large, complex datasets
Strong understanding of data governance frameworks, tools, and best practices.
Good-to-Have Skills:
Knowledge of data protection regulations and compliance requirements (e.g., GDPR, CCPA) processing
Experience with ETL tools such as Apache Spark, and various Python packages related to data processing, machine learning model development
Strong understanding of data modeling, data warehousing, and data integration concepts
Knowledge of Python/R, Databricks, SageMaker, cloud data platforms
Experience implementing automated orchestration and monitoring of data pipelines using Databricks Jobs, Apache Airflow, or similar workflow tools.
Familiarity with performance optimization techniques for big data processing, such as Spark job tuning, caching, partitioning, and indexing.
Exposure to multi-source integration involving APIs, SQL databases, and cloud storage platforms.
Demonstrated ability to collaborate across global teams and time zones, ensuring alignment and delivery in distributed environments.
Professional Certifications (Preferred):
Certified Data Engineer / Data Analyst (preferred on Databricks or cloud environments)
Soft Skills:
Excellent critical-thinking and problem-solving skills
Strong communication and collaboration skills
Demonstrated awareness of how to function in a team setting
Demonstrated presentation skills
.
Job details
How this role compares
Computed from every other active Data & Digital role in our database, not just this employer's listings.
We currently track 292 comparable Data & Digital roles across 41 biopharma companies.
Salary context
60 of 292 peers report a salary range (USD, annualized)
Peers share this role's job function. This posting doesn't list a seniority level, so peers aren't narrowed by seniority either -- the range below may span more levels than usual.
Where these roles are based
Top locations among the 292 comparable roles
+ 14 more countries
Seniority mix
158 of 292 peers have a known seniority level
Therapeutic area mix
6 of 292 peers have a known therapeutic area; the rest are genuinely unlabeled, not hidden
Similar opportunities
The closest matches from our peer group, ranked by how similar they are, not how well you'd qualify for them -- treat this as market context, not a guaranteed shortlist; a weak match is labeled as one below.
How we calculate "similar"
No black box, no LLM guesswork: a deterministic score built from four normalized attributes. Here's this role's own peer group at different match levels, so you can see the mechanism, not just the result.
Every comparison starts from the same 100-point budget: 25 for working in the same function, 40 for the same therapeutic area, 20 for the same or adjacent seniority, 15 for the same country. A dimension we can't confirm on both sides contributes nothing, never a guess, never a free pass.
0 points, never a partial guess. A role we know almost nothing about beyond its function bottoms out at 25%; it never inflates to 100% just because there's little to compare against. Seniority uses a defined ladder (Associate → Manager → Associate Director → Senior → Principal → Director → Senior Director → Executive/VP) so "Director" and "Senior Director" count as adjacent, but "Director" and "Executive/VP" do not.