Lead Data Engineer
Smart WorkingPakistanContractSoftware Development
Remotely
pythonpostgresqldatanosqlapache sparkvector databasespgvectormilvusqdrant
Job Description
📋 Description
- Architect and build scalable data pipelines and infrastructure to support AI and product systems.
- Design and maintain data ingestion, transformation and storage architectures for operational and AI
- Develop and manage batch and real-time data pipelines.
- Build and optimise systems for vector search, retrieval and machine learning data pipelines.
- Ensure data reliability, security and governance across the platform.
- Collaborate with AI and backend engineering teams to support training, inference and product
🎯 Requirements
- 7+ years of professional experience, with the majority of that experience in dedicated data
- Strong experience designing and building data pipelines and distributed data systems.
- Experience working with relational databases, with PostgreSQL preferred, although MySQL or similar
- Experience working with NoSQL databases.
- Experience with vector databases used in modern AI systems.
- Strong programming experience in Python.
Nice to Have
- Experience working on AI or machine learning platforms.
- Familiarity with stream processing and event-driven architectures.
- Experience with cloud infrastructure such as GCP, AWS or Azure.
- Experience working in high-growth startups or early-stage companies.