Reference Data Specialist
About this opportunity
Career Category
Information Systems
Job Description
ABOUT AMGEN
Amgen harnesses the best of biology and technology to fight the world’s toughest diseases, and make people’s lives easier, fuller and longer. We discover, develop, manufacture and deliver innovative medicines to help millions of patients. Amgen helped establish the biotechnology industry more than 40 years ago and remains on the cutting-edge of innovation, using technology and human genetic data to push beyond what’s known today.
ABOUT THE ROLE
Role Description:
Amgen is seeking an experienced Reference Data specialist to support enterprise semantic data, ontology, taxonomy, and knowledge graph initiatives across business and scientific domains. The role will focus on designing, developing, governing, and optimizing enterprise reference data and semantic models supporting analytics, interoperability, and data governance initiatives.
As the Reference Data Product team member of the Data Foundations & Governance organization, you will be responsible for managing and promoting the use of reference data, partnering with business Subject Mater Experts on creation of vocabularies / taxonomies and ontologies, and developing analytic solutions using semantic technologies.
Roles & Responsibilities:
Design, build and maintain enterprise ontologies, taxonomies, and semantic reference data models
Collaborate with business SMEs, data architects, and governance teams to define semantic standards
Contribute towards defining Reference data management framework for Enterprise
Develop and optimize RDF/OWL/SKOS-based semantic frameworks and SPARQL/SQL queries.
Support semantic data integration, metadata harmonization, and semantic publishing workflows.
Manage ontology lifecycle activities including modeling, validation, versioning, and publishing.
Collaborate with multiple enterprise & functional teams to elicit, structure and formalize knowledge from domain experts and diverse sources to build CVs, taxonomies, and ontologies.
Troubleshoot semantic data load and mapping issues across Semantic Layer, CDL, and graph platforms.
Develop and optimize automated data ingestion / pipelines through Python/PySpark when APIs are available
Support enterprise Data Foundations, Knowledge Graph and Metadata Management initiatives.
Mentor junior engineers and provide technical leadership on semantic technologies.
Identify and resolve complex data-related challenges
Participate in sprint planning meetings and provide estimations on technical implementation.
Basic Qualifications and Experience:
Master’s degree with 7- 10 years of experience in Business, Engineering, IT or related field OR
Bachelor’s degree with 8 - 12 years of experience in Business, Engineering, IT or related field OR
Functional Skills:
Must-Have Skills:
Mandatory: 8–12 years of experience in Reference Data Engineering, Semantic Technologies, or Knowledge Graph implementations.
Advanced knowledge of ontologies and taxonomies with proficiency in Semantic Web technologies, standards and tools such as RDF/s, OWL, SKOS, SPARQL, SHACL, and Linked Data standards.
Hands-on experience with GraphDBs, TopBraid EDG, CenTree, MarkLogic, Stardog, or similar semantic platforms.
Strong expertise in Pharma Domain CVs/Taxonomies/Ontologies such as CDISC, MedDRA, WHODrug, NCIT, IDMP SPOR, SNOMED CT, ICD 10/11 etc. and integrating them into enterprise-wide applications.
Strong understanding of metadata management, reference data governance, and semantic interoperability.
Hands-on experience with modern data platforms such as Databricks and cloud engineering platforms such as AWS.
Experience with APIs, JSON-LD, XML, SQL, and semantic integration frameworks.
Strong analytical, troubleshooting, and stakeholder management skills.
Familiarity with FAIR Data Principles, Semantic Layer, metadata repositories, data catalogs and semantic publishing pipelines.
Exposure to enterprise knowledge graph and AI/ML initiatives.
Experience working in Agile delivery models and cloud-based ecosystems.
Technical Skills:
Semantic Technologies: RDF, OWL, SKOS, SPARQL, SHACL, Linked Data, JSON-LD
Platforms & Tools: GraphDB, TopBraid, CenTree, MarkLogic, Protégé, Semaphore, Databricks
Data & Integration: SQL, PySpark, ETL, Git, REST APIs, XML, JSON, Metadata & Data Governance
Professional Certifications :
Databricks Certificate preferred
SAFe® Practitioner Certificate preferred
Any Data Analysis certification (SQL, Python)
Any cloud certification (AWS or AZURE)
Soft Skills:
Strong analytical abilities to assess and improve master data processes and solutions.
Excellent verbal and written communication skills, with the ability to convey complex data concepts clearly to technical and non-technical stakeholders.
Effective problem-solving skills to address data-related issues and implement scalable solutions.
Ability to work effectively with global, virtual teams
.
Job details
How this role compares
Computed from every other active Information Technology role in our database, not just this employer's listings.
We currently track 1097 comparable Information Technology roles across 55 biopharma companies.
Salary context
143 of 1097 peers report a salary range (USD, annualized)
Peers share this role's job function. This posting doesn't list a seniority level, so peers aren't narrowed by seniority either -- the range below may span more levels than usual.
Where these roles are based
Top locations among the 1097 comparable roles
+ 23 more countries
Seniority mix
612 of 1097 peers have a known seniority level
Therapeutic area mix
1 of 1097 peers have a known therapeutic area; the rest are genuinely unlabeled, not hidden
Similar opportunities
The closest matches from our peer group, ranked by how similar they are, not how well you'd qualify for them -- treat this as market context, not a guaranteed shortlist; a weak match is labeled as one below.
How we calculate "similar"
No black box, no LLM guesswork: a deterministic score built from four normalized attributes. Here's this role's own peer group at different match levels, so you can see the mechanism, not just the result.
Every comparison starts from the same 100-point budget: 25 for working in the same function, 40 for the same therapeutic area, 20 for the same or adjacent seniority, 15 for the same country. A dimension we can't confirm on both sides contributes nothing, never a guess, never a free pass.
0 points, never a partial guess. A role we know almost nothing about beyond its function bottoms out at 25%; it never inflates to 100% just because there's little to compare against. Seniority uses a defined ladder (Associate → Manager → Associate Director → Senior → Principal → Director → Senior Director → Executive/VP) so "Director" and "Senior Director" count as adjacent, but "Director" and "Executive/VP" do not.