San JoseFull TimeEngineering
Remotely
sparks3dbtairflowicebergflinkparquetpaimon
Job Description
📋 Description
- Design, build, and operate data infrastructure powering large-scale model training and inference.
- Own pipelines, storage, and data quality for ML workloads upstream of the platform.
- Ensure data is clean, high-throughput, and well-governed for training and evaluation.
- Collaborate with ML researchers, platform engineers, and software engineers on data access patterns
- Partner with ML engineers to define feature stores, dataset versioning, and data contracts
🎯 Requirements
- 5+ years of professional data engineering experience.
- BS/MS/PhD in CS, Data Eng, Software Eng, or related field.
- Hands-on with production pipelines using Spark, Flink, Airflow, dbt or similar batch/streaming
- Deep proficiency with Parquet and open table formats (Iceberg or Paimon); strong data lakehouse
- Experience with high-throughput messaging (Kafka, Pulsar) for real-time ingestion; strong SQL
- Experience with AWS data services (S3, Glue, EMR) and containerized workloads (Docker/Kubernetes)
🎁 Benefits
- Exposure to CDC replication (Debezium, Airbyte).
- Experience with data lineage, versioning, and data contracts for reproducible experiments.
- Opportunity to own early-stage data products in a startup-like AI org within an aerospace company.