Lead Data Engineer
JobgetherRemotely
pythonpostgresqlapache airflowdatanosqlpandasapache sparkvector databasespolars
Job Description
📋 Description
- Architect, build, and scale robust data pipelines and infrastructure supporting both AI products
- Design and maintain data ingestion, transformation, processing, and storage architectures for batch
- Develop scalable systems for vector search, retrieval, and machine-learning data workflows.
- Build and optimise data models, distributed processing systems, and large-scale query
- Establish frameworks for data reliability, quality, security, governance, monitoring, and
- Collaborate with AI and backend engineering teams to support model training, inference, and
🎯 Requirements
- 7+ years of professional experience, with significant experience in dedicated data engineering
- Strong experience designing and building scalable data pipelines and distributed data systems.
- Solid experience with relational databases, preferably PostgreSQL; experience with MySQL or
- Experience working with NoSQL databases and modern vector databases used in AI applications.
- Strong Python programming skills, including experience with data-processing libraries such as
- Demonstrated ability to make, communicate, and justify architectural decisions rather than simply
🎁 Benefits
- Remote-first working environment from locations across India.
- Opportunity to take significant ownership as the first senior data engineering hire.
- High-impact role working at the intersection of data engineering, AI, and intelligent automation.
- Direct collaboration with AI engineers, backend engineers, product teams, founders, and leadership.
- Opportunity to shape the data architecture, engineering standards, culture, and future team.
- Exposure to modern technologies including real-time data pipelines, vector databases