Senior Data Scientist (UAE)
CloudPSO IncJob Description
This is a remote position.
•
Location: Remote – UAE
•
Requirement: A Valid UAE work permit/employment visa is mandatory.
•
Employment type: Independent Contractor
Key Responsibilities
- Generative AI & NLP for Engineering
•
Data Exploration and Analysis: Query and analyse large domain- or topic-specific data sets from both structured and unstructured sources, identify patterns and features. Ensure data meets quality standards and requirements before model development.
•
Regulation Text Interpretation: Design and fine-tune Large Language Models (LLMs) to parse complex regulatory texts (e.g., building codes, military standards) and extract structured rules for automated compliance checking.
•
Rule Formalization: Convert interpreted regulations into computer-processable formats (e.g., object-property-condition-value tuples) that can be executed by downstream compliance engines.
•
Querying via NLP: Architect methods for LLMs to map natural language requirements directly to specific metadata entities within various schemas (e.g., mapping "systems design" to specified attributes).
•
RAG Architecture: Implement Retrieval-Augmented Generation (RAG) pipelines that allow systems to query vast repositories of technical documentation and historical project data with high accuracy and low hallucination rates.
- Predictive Modeling & Optimization (Supply Chain)
•
Forecasting Engines: Develop time-series forecasting models to predict spend categories and material demand by correlating internal ERP data with external macroeconomic signals.
•
Classification & Risk Scoring: Build machine learning classifiers to categorize supplier risks and operational anomalies, integrating data from diverse sources to create dynamic risk scores.
•
Data Extraction Pipelines: Design robust pipelines to extract and transform raw data (from Data Lakehouse, external web sources, or SAP and other databases) into features required for predictive modeling and automated rule checking.
- System Integration & Performance
•
Model Orchestration: Work with Back End Engineers to integrate AI models into a cohesive "compliance engine" or "risk engine" that can be invoked programmatically via robust APIs.
•
Optimization: Streamline model performance to ensure complex checks (e.g., analyzing large datasets or processing thousands of supplier records) can be executed within reasonable timeframes, potentially using batching or asynchronous processing.
•
Quality Assurance: Validate model outputs against known test cases and historical data, debugging false positives/negatives to refine algorithms and ensure "defense-grade" reliability.
Requirements
•
Core AI/ML: Expert proficiency in Python and standard ML libraries (TensorFlow/PyTorch, Scikit-learn, Pandas, NumPy). Strong grasp of both supervised and unsupervised learning techniques.
•
NLP & LLMs: Deep experience with transformer-based models (GPT, BERT, Llama) and prompt engineering techniques (few-shot learning, fine-tuning) for domain-specific tasks.
•
Data Engineering: Proficiency in handling complex data structures (JSON, XML) and familiarity with database querying (SQL/NoSQL) or graph data structures. Experience with data extraction from specialized formats is a significant plus.
•
Backend Awareness: Understanding of how to expose models via RESTful APIs (Flask/FastAPI) and integrate them into larger software architectures.
•
Statistics:Solid understanding of statistics, probability distribution, A/B testing. Adept at identifying and mitigating biases in datasets
Originally posted on Himalayas