Remotely
pythonsqlrestawssparkdatabricksdbtairflow
Job Description
📋 Description
- Design and operate production-grade data pipelines and data products powering AI/ML analytics
- Integrate data from MES, historians, LIMS, QMS, ERP, and instrument platforms into
- Develop scalable ingestion pipelines for batch and real-time data, with connectors for diverse data
- Harmonize data models and ontologies to ensure consistency across sites and modalities.
- Establish data quality, lineage, and governance with observability and compliance to GxP/21 CFR
- Create governed data products (feature stores, curated datasets, vector-ready layers) for AI/ML and
🎯 Requirements
- Bachelor’s Degree in CS/Data Eng/Software Eng/Bioinformatics or related field plus 6 years’
- Hands-on experience designing enterprise-grade data pipelines and data products in multi-source
- Expert-level Python for data engineering tasks.
- Strong SQL across modern databases; ANSI SQL and dialects.
- Experience with cloud data platforms (AWS/Azure/GCP) and data stack tools like dbt, Spark, Airflow
- ETL/ELT pipelines with Informatica, Talend, NiFi, or cloud-native services (e.g., AWS Glue, Azure
🎁 Benefits
- Comprehensive benefits package including paid time off, medical/dental/vision, and 401(k).
- Eligibility for short-term incentive programs.
- Equal opportunity employer; inclusive environment and opportunities for growth.
Back to all jobs