Remotely
awsterraformdatadatabrickspysparkdbtairflow
Job Description
📋 Description
- Design, build, and optimize ETL pipelines powering analytics and ML workflows.
- Develop labeling and retraining pipelines for ML models with quality and observability.
- Implement MLOps practices: model versioning, CI/CD for ML, model monitoring.
- Collaborate with data scientists to productionize training, inference, and evaluation.
- Contribute to data lakehouse design: schema, partitioning, performance.
- Document data architecture, lineage, and dependencies for transparency.
🎯 Requirements
- Bachelor’s degree in CS/Engineering or related field.
- 2–4 years building/operating large-scale data systems for analytics/ML.
- Proficiency in Python and SQL; experience with PySpark, pandas.
- Experience with DBT.
- Experience with modern data warehousing/lakehouse platforms, preferably Databricks.
- Hands-on with orchestration tools like Airflow, Dagster, or Prefect.
🎁 Benefits
- Competitive pay, health insurance, 401(k) matching, equity plan.
- Flexible time off, holidays, parental leave, year-end break.
- Work-from-home funds and wellness benefits to support growth and wellbeing.