Roche Posted September 12, 2026

Fullstack Data Engineer - Applied & Agentic AI Systems

Hyderabad, Telangāna, India FULL_TIME
Notify me about similar jobs

Roche is the source of truth for this posting and owns the application process. We surface normalized context and market comparison you won't find on the original listing.

About this opportunity

At Roche you can show up as yourself, embraced for the unique qualities you bring. Our culture encourages personal expression, open dialogue, and genuine connections,  where you are valued, accepted and respected for who you are, allowing you to thrive both personally and professionally. This is how we aim to prevent, stop and cure diseases and ensure everyone has access to healthcare today and for generations to come. Join Roche, where every voice matters.

The Position

Job description

Our Applied AI Engineering Team is seeking a fullstack AI Data Engineer to join a newly formed, autonomous development squad dedicated to pioneering agentic, LLM-based solutions. Operating across the entire lifecycle, from initial ideation and rapid prototyping to production-grade deployment and ongoing operations, you will architect the resilient data infrastructure required to power next-generation AI. We are looking for an expert capable of orchestrating both structured and unstructured datasets, implementing high-performance vector databases, and managing real-time streams within cloud-native environments.

Description of the area

At Roche Digital Technology, we are advancing the boundaries of Applied AI. The Applied AI Engineering Team focuses on architecting, building, and operating high-value AI solutions and services to solve complex business challenges in healthcare.

In the 2026 tech landscape, we operate in small, highly autonomous agile teams (e.g., 9 members) powered by advanced coding agents (like Claude Code) to develop and ship solutions faster than ever before. In this highly regulated environment, quality cannot be an afterthought. You will be the foundational pillar ensuring our rapidly developed agentic workflows and AI, GenAI, and agentic applications are safe, compliant, and robust before they reach the clinical or enterprise user.

Job Responsibilities

Generative AI Application Co-creation : Collaborate with AI engineers, data scientists, product owners, and other developers in Agile teams to integrate LLMs into scalable, robust, fair, and ethical end-user applications, focusing on user experience, relevance, and real-time performance

Data Infrastructure Development and Data Integration : Design and implement scalable, high-performance data pipelines for AI/GenAI applications, ensuring efficient data ingestion, transformation, storage and retrieval; integrate different databases, requiring understanding of data architectures / Domain data ecosystem

Vector Databases : work with vector databases (e.g., AWS OpenSearch, Azure AI Search) to facilitate scalable, high-speed similarity search and RAG for generative AI applications with high-dimensional data.

Graph Databases : Work with graph databases (e.g., Neo4j, AWS Neptune) to enable GraphRAG, support multi-hop logical reasoning for agentic workflows, and provide auditable explainability for enterprise AI decision-making.

Cloud-Based Data Engineering: Build and maintain cloud-based data solutions using AWS (OpenSearch, S3) or Azure (Azure AI Search, Azure Blob Storage)

Snowflake Implementation: Design and optimize data storage and processing using Snowflake for scalable, cloud-native analytics solutions

Data Processing & Transformation : Develop ETL/ELT pipelines to enable real-time and batch data processing

Support AI Model Workflows : Collaborate with AI/ML Engineers and Data Scientists to ensure seamless integration of data pipelines with AI finetuning, inference and training workflows

Performance Optimization : Optimize data storage, retrieval, and processing strategies for efficiency, scalability, and cost-effectiveness

Software Development Lifecycle : understand and leverage an agentic software development lifecycle (SDLC) in day to day work

Security & Compliance : Implement data governance, security best practices, and compliance measures aligned with Roche’s standards

Monitoring & Maintenance : Set up monitoring, alerting, and logging for data pipelines, ensuring high availability and reliability

Skills

Must have:

Experience : 7+ years in data engineering, preferably supporting AI/ML applications

Advanced Programming & SQL: Writing production-grade code in Python alongside highly optimized, complex SQL queries.

Advanced System Architecture & Modeling: Designing scalable, fault-tolerant ETL/ELT data pipelines and Lakehouse architectures (e.g., Snowflake).

Orchestration: Hands-on expertise with orchestration tools (like Airflow).

Data Engineering in AI: Developing Retrieval-Augmented Generation (RAG), AI systems powered by Vector Databases and/or LLM fine-tuning, and data preparation

Document Processing Proficiency : Extracting, transforming, and loading data from diverse file formats (PDF, DOCX, CSV, JSON, etc.), including automated parsing and information retrieval from unstructured and semi-structured documents

Version Control & DevOps : Hands-on experience with Git, CI/CD, containerization (Docker, Kubernetes), and Infrastructure as Code (Terraform, CloudFormation)

Problem Solving: Excellent analytical skills and the ability to tackle complex challenges with innovative solutions

Should have:

AWS Cloud Platforms: Hands-on experience with AWS (OpenSearch, S3, Lambda, AWS fundamentals)

GraphDB: Experience in building solutions with GraphDB

APIs & Microservices : Ability to design and integrate RESTful APIs for data exchange

Data Security & Governance : Understanding of encryption and role-based access controls

Working in an SDLC environment meeting regulatory requirements

Proficiency in best practices of software engineering

Agentic SDLC & Engineering Excellence: Leverage AI coding assistants and autonomous agents (e.g., Claude Code, Ona) daily to accelerate full-stack development and testing cycles. Conduct rigorous code reviews for both human-written and AI-generated code.

Could have:

Regulatory Compliance: Proven experience in working within highly regulated industries.

Data Science & Classical Machine Learning : Practical background in Data Science, encompassing feature engineering, model training, and data preparation leveraging traditional ML techniques.

Distributed Data Processing: Hands-on expertise with big data frameworks (like Apache Spark or Flink).

Capabilities:

Problem-Solving Skills: Excellent analytical skills to tackle complex engineering and statistical challenges.

Ownership & Leadership: Deep sense of accountability, eager to define architectural patterns, and able to step into a Tech Lead role when necessary.

Consulting: Ability to work closely with stakeholders across the enterprise to consult on the technological approaches to their business problems.

Ethics: Strong understanding of biases, fairness, hallucination mitigation, and responsible AI deployment.

Qualifications

hold B.Sc., B.Eng., or higher, or equivalent in Computer Science, Data Engineering or related fields

have an interest in AI and stay up to date with the latest advancements in data engineering

be team-oriented, proactive, and collaborative

have strong analytical and problem-solving skills

have excellent verbal and written communication skills

be detail-oriented and highly organized

be willing to learn and expand their skill set

have the ability to work collaboratively in a fast-paced, dynamic environment

be able to communicate in English at the level of C1+

Located in Hyderabad, India, with working hours structured to capture the 'golden hours' of overlap with Central European Time (typically running through the IST evening).

#Hyd2026

 

 

Who we are

A healthier future drives us to innovate. Together, more than 100’000 employees across the globe are dedicated to advance science, ensuring everyone has access to healthcare today and for generations to come. Our efforts result in more than 26 million people treated with our medicines and over 30 billion tests conducted using our Diagnostics products. We empower each other to explore new possibilities, foster creativity, and keep our ambitions high, so we can deliver life-changing healthcare solutions that make a global impact.

Let’s build a healthier future, together.

Roche is an Equal Opportunity Employer.

Job details

Seniority
Not listed
Function
Data & Digital
Therapeutic area
Not listed
Location
Hyderabad, Telangāna, India
Employment type
FULL_TIME

How this role compares

Computed from every other active Data & Digital role in our database, not just this employer's listings.

We currently track 270 comparable Data & Digital roles across 46 biopharma companies.

270Comparable roles tracked
246Currently active
46Companies hiring similar roles
17Countries represented

Salary context

79 of 270 peers report a salary range (USD, annualized)

Peers share this role's job function. This posting doesn't list a seniority level, so peers aren't narrowed by seniority either -- the range below may span more levels than usual.

This roleSubject Not listed on this posting
Lowest disclosed · Data Engineer · Lilly $66,000/yr – $121,000/yr
Peer group range $93,500 – $339,950 (median $198,000)

Where these roles are based

Top locations among the 270 comparable roles

United States100
India83
United Kingdom19
France17
Spain13
Canada9

+ 11 more countries

Seniority mix

161 of 270 peers have a known seniority level

Senior44
Director27
Associate Director26
Manager25
Principal19
Executive/VP10
Associate5
Intern/Fellow/Postdoc3
Senior Director2

Therapeutic area mix

4 of 270 peers have a known therapeutic area; the rest are genuinely unlabeled, not hidden

Immunology3
Ophthalmology1

Similar opportunities

The closest matches from our peer group, ranked by how similar they are, not how well you'd qualify for them -- treat this as market context, not a guaranteed shortlist; a weak match is labeled as one below.

40%similar
Amgen Technology Pvt Ltd. Hyderabad, India
Same function Same country
40%similar
Regeneron India Private Limited Hyderabad, India Senior Director
Same function Same country
40%similar
Regeneron India Private Limited Hyderabad, India Manager
Same function Same country
40%similar
Regeneron India Private Limited Hyderabad, India
Same function Same country
40%similar
Regeneron India Private Limited Hyderabad, India
Same function Same country
40%similar
Regeneron India Private Limited Hyderabad, India Senior
Same function Same country

Notify me about similar jobs

Get an email when we spot other openings like this one – same job function, comparable seniority, roles you'd actually want to see.

How we calculate "similar"

No black box, no LLM guesswork: a deterministic score built from four normalized attributes. Here's this role's own peer group at different match levels, so you can see the mechanism, not just the result.

Every comparison starts from the same 100-point budget: 25 for working in the same function, 40 for the same therapeutic area, 20 for the same or adjacent seniority, 15 for the same country. A dimension we can't confirm on both sides contributes nothing, never a guess, never a free pass.

40%
Data Engineer, Translational Data Management, Automation & AI
Amgen Technology Pvt Ltd. · Hyderabad, India · Seniority not listed
Function Therapeutic area Seniority Country
40%
Specialist, Clinical Data Reporting
Regeneron India Private Limited · Hyderabad, India · Seniority not listed
Function Therapeutic area Seniority Country
40%
Associate Director Clinical Data Reporting
Regeneron India Private Limited · Hyderabad, India · Associate Director
Function Therapeutic area Seniority Country
40%
Principal Statistical Programmer
Sanofi Healthcare India Private Limited · Hyderabad, India · Principal
Function Therapeutic area Seniority Country
Unmatched or unknown dimensions score exactly the same: 0 points, never a partial guess. A role we know almost nothing about beyond its function bottoms out at 25%; it never inflates to 100% just because there's little to compare against. Seniority uses a defined ladder (Associate → Manager → Associate Director → Senior → Principal → Director → Senior Director → Executive/VP) so "Director" and "Senior Director" count as adjacent, but "Director" and "Executive/VP" do not.