We are seeking a highly skilled Data Scientist with strong expertise in Machine Learning services, Recommender Systems (RS), and Generative AI (LLM & LVM). The role collaborates closely with Data Engineering, Data Science, and ML Engineering teams to design, develop, deploy, and scale intelligent data products and AI solutions.
Key Responsibilities
Data Science & Advanced Analytics
Develop and deploy end-to-end ML models from ideation to production.
Perform EDA, feature engineering, and model evaluation.
Build predictive and prescriptive models using statistical and ML techniques.
Machine Learning Services (Primary Focus)
Design and implement scalable ML pipelines for training, testing, and deployment.
Work with ML platforms: Azure ML, AWS SageMaker, GCP Vertex AI.
Implement model lifecycle management: versioning, monitoring, retraining.
Optimize models for performance, scalability, and reliability.
Recommender Systems (RS)
Design and build recommendation engines: collaborative, content-based, hybrid.
Work with large-scale datasets for ranking, personalization, user segmentation.
Evaluate models using precision@k, recall@k, NDCG.
Generative AI (GenAI – LLM & LVM)
Build and deploy LLM-powered solutions: chatbots, copilots, document intelligence.
Implement RAG (Retrieval-Augmented Generation) architectures.
Work with models such as OpenAI, Azure OpenAI, Hugging Face.
Develop use cases for text generation, summarization, classification, and image/video understanding (LVM).
Optimize prompts and manage prompt engineering workflows.
Collaboration with Data Engineering
Define data requirements; collaborate on data pipeline design.
Ensure data quality, governance, and availability.
Work with Spark, Databricks, Hadoop.
ML Engineering & Deployment Support
Deploy models via APIs/microservices; containerize with Docker, Kubernetes.
Integrate models into production systems and CI/CD pipelines.
Model Monitoring & Governance
Monitor model drift, performance degradation, and bias.
Implement logging, alerting, and explainability tools.
Ensure Responsible AI: fairness, transparency, interpretability.
Qualifications & Experience
5–8 years in Data Science, Machine Learning, or AI.
Proven end-to-end ML model development, deployment, and production support.
Hands-on predictive modeling using Python and modern ML frameworks.
Experience with cloud ML platforms (Azure ML, AWS SageMaker, or Google Vertex AI).
Experience developing GenAI solutions using LLMs, prompt engineering, and RAG frameworks.
Experience collaborating with Data Engineering and ML Engineering teams.
Agile experience delivering enterprise-scale AI applications.
YOUR PROFILE
Education
B.Tech/B.S./M.S. in Computer Science, Statistics, Mathematics, or related field.
Technical Skills
Programming: Python (mandatory), SQL
ML Libraries: Scikit-learn, TensorFlow, PyTorch
Data Processing: Pandas, NumPy, Spark
ML & AI: Supervised & unsupervised learning, model optimization
Recommender Systems: engine design, ranking algorithms, personalization
Generative AI: LLMs (GPT, Llama, etc.), prompt engineering, RAG (LangChain, LlamaIndex); multimodal AI (LVM) a plus
MLOps: CI/CD for ML, Docker, Kubernetes, model monitoring
Data Engineering: pipelines, ETL, data warehousing
Preferred Qualifications
Strong communication, stakeholder management, organizational skills
Self-motivated, customer-focused, detail-oriented
Azure ecosystem experience (Azure ML, Databricks)
Exposure to real-time data processing
ML/AI/Cloud certifications
SAP ERP knowledge strongly preferred
Six Sigma Yellow/Green Belt a plus
ITIL certification a plus
Data Science, Machine Learning, Python, SQL, Predictive Modeling, Prescriptive Modeling, Exploratory Data Analysis, Feature Engineering, Scikit-Learn, TensorFlow, PyTorch, Pandas, NumPy, Spark, Azure ML, AWS SageMaker, Google Vertex AI, MLOps, ML Pipelines, Model Lifecycle Management, Model Monitoring, Recommender Systems, Collaborative Filtering, Content-Based Recommendation, Hybrid Recommendation Systems, Ranking Algorithms, Personalization, User Segmentation, Generative AI, LLM, LVM, GPT, Llama, Prompt Engineering, RAG, LangChain, LlamaIndex, Multimodal AI, Azure OpenAI, Hugging Face, Data Engineering, ETL, Data Pipelines, Databricks, Hadoop, Docker, Kubernetes, CI/CD, Model Deployment, Model Drift Detection, Explainable AI, Responsible AI, Data Governance, Real-Time Data Processing, Agile, SAP ERP, Six Sigma, ITIL