Data Engineer, PDS&T CMC
Responsibilities: Data Ingestion & Integration Design and implement scalable, robust data ingestion pipelines that connect CMC and manufacturing source systems — including MES (Manufacturing Execution Systems), process historians, LIMS, QMS, ERP platforms, and instrument data sources — to centralized and federated data environments.  Build connectors, adapters, and integration layers that handle the heterogeneous data formats, protocols, and latency profiles characteristic of pharmaceutical manufacturing environments.  Support both batch and real-time/streaming data patterns, selecting appropriate architectures based on use case requirements.    Data Harmonization & Semantic Modeling  Develop and maintain harmonized data models and ontologies that bring consistency to CMC and manufacturing data across sites, systems, and modalities.  Execute semantic mapping efforts that align source system fields, units, and identifiers to enterprise data standards and scientific meaning.  Collaborate with process scientists, analytical chemists, and manufacturing engineers to ensure data models accurately reflect domain reality.    Data Quality, Observability & Governance  Implement automated data quality controls, validation frameworks, and anomaly detection mechanisms across pipeline layers.  Build and maintain data lineage documentation and metadata infrastructure, enabling full traceability from source system to AI model input.  Establish pipeline observability practices — monitoring, alerting, SLA tracking — to ensure data product reliability in production.  Support data governance practices aligned with GxP requirements, 21 CFR Part 11, and AbbVie data standards.    AI/ML Enablement & Data Product Development  Architect and deliver governed, versioned, reusable data products purpose-built for AI/ML consumption, including feature stores, curated datasets, and vector-ready data layers for RAG and LLM applications.  Partner closely with data scientists, ML engineers, and process modelers to understand model data requirements and translate them into reliable, scalable data infrastructure.  Accelerate AI program delivery by eliminating data bottlenecks — not by workarounds, but by solving root causes structurally.    Platform & Operational Enablement  Contribute to the design and evolution of PDST's cloud-based data platform, including lakehouse architecture, data cataloging, access control, and compute infrastructure.  Write and maintain infrastructure-as-code, CI/CD pipelines, and automated testing frameworks for data systems.  Support platform onboarding of new CMC data domains and manufacturing sites, ensuring consistent application of standards and patterns.  Provide operational support for production data pipelines, maintaining uptime and data freshness commitments.    Stakeholder Engagement & Scientific Leadership  Influence technical decision-making without formal authority — earning trust through scientific rigor, transparent methodology, and demonstrated business impact.  Required: Bachelor's Degree in Computer Science, Data Engineering, Information Systems, Software Engineering, Bioinformatics, or a closely related technical field plus 2 years’ experience OR Master’s Degree with 0 years' experience. Respective years of hands-on experience designing and building enterprise-grade data pipelines, integration workflows, and data products in complex, multi-source environments.  Expert-level proficiency in Python for data engineering tasks — pipeline development, transformation logic, data validation, and automation.  Strong SQL skills across modern analytical and transactional databases; comfort with both ANSI SQL and platform-specific dialects.  Demonstrated experience with cloud data platforms (AWS, Azure, or GCP) and modern data stack components — including tools such as dbt, Spark, Airflow, Databricks, Snowflake, or equivalents.  Develop ETL/ELT pipelines using tools such as Informatica, Talend, Apache NiFi, and cloud-native services (e.g., AWS Glue, Azure Data Factory).  Implement master data management (MDM), metadata management, and data cataloging solutions to ensure proper data lineage, accessibility, and compliance.  Set and enforce standards for API development and data integration (REST, GraphQL, OData), enabling seamless integration using microservices architectures.  Design logical, physical, and conceptual data models using modeling tools (e.g., Erwin, PowerDesigner, dbt).  Ownership orientation: you define your own problem space, drive solutions to completion, and hold yourself accountable to outcomes — not just outputs.  Solution-architect instinct: you think before you build, consider the full landscape of available approaches, and choose tools based on fit-for-purpose reasoning rather than familiarity or trend.  Scientific integrity: you build models you can explain, defend, and improve — and you apply the same standard to the work of others.  Influence through credibility: you earn the confidence of scientists, engineers, and quality professionals by being right, being clear, and being useful — not by title or volume.  Bias for impact: you are drawn to problems where the stakes are high and the analytical opportunity is real, and you are energized rather than intimidated by ambiguity.  Preferred: Experience in pharmaceutical, biotech, or other regulated life sciences manufacturing environments.  Familiarity with GxP data principles, 21 CFR Part 11 compliance, or data integrity requirements in regulated industries.  Prior exposure to manufacturing source systems such as MES, process historians (e.g., OSIsoft PI/AV
Findigo hittar jobben och fyller i ansökan. Du klickar Skicka.
Visa jobbet och ansökUrsprunglig annons: jobs.smartrecruiters.com