We are seeking an experienced Lead Data Scientist with strong expertise in Machine Learning, Recommender Systems, and Generative AI (LLM & Large Vision Models). The role provides technical leadership in designing, developing, deploying, and scaling enterprise AI solutions while collaborating closely with Data Engineering, ML Engineering, Platform Engineering, and business stakeholders. The ideal candidate will drive AI solution architecture, mentor technical teams, establish best practices, and ensure scalable, secure, production-ready AI implementations.
Key Responsibilities
Data Science & Advanced Analytics — Lead end-to-end ML solutions from ideation to production. Drive EDA, feature engineering, model evaluation, and optimization. Develop predictive and prescriptive analytics models. Translate business problems into scalable AI/data science solutions.
Machine Learning Services (Primary Focus) — Design and implement scalable ML pipelines for training, testing, and deployment. Architect enterprise ML platforms using Azure Machine Learning, AWS SageMaker, or Google Vertex AI. Define model lifecycle management (versioning, monitoring, retraining, governance). Optimize models for performance, scalability, and reliability.
Recommender Systems — Lead design of recommendation engines (collaborative filtering, content-based, hybrid). Design ranking, personalization, and segmentation strategies. Define evaluation metrics (Precision@K, Recall@K, MAP, NDCG). Guide scalable recommendation architectures.
Generative AI (LLM & LVM) — Design enterprise GenAI solutions (chatbots, copilots, document intelligence). Architect RAG solutions using OpenAI, Azure OpenAI, or Hugging Face models. Design use cases for text generation, summarization, classification, and image/video understanding. Lead prompt engineering and LLM optimization.
Collaboration with Data Engineering — Define data requirements, ensure data quality, governance, and availability. Work with Spark, Databricks, Hadoop.
ML Engineering & Deployment Support — Deploy models via APIs/microservices, containerize (Docker, Kubernetes), and integrate into production systems and CI/CD pipelines.
Model Monitoring & Governance — Monitor model drift, performance degradation, and bias. Implement logging, alerting, explainability tools, and Responsible AI practices (fairness, transparency, interpretability).
Qualifications & Experience
8+ years in Data Science, ML, AI, or Advanced Analytics, with 5+ years hands-on deploying enterprise ML solutions. Proven experience leading end-to-end AI initiatives from concept to production, designing scalable AI architectures, mentoring technical teams, and delivering AI solutions in Agile environments on enterprise cloud platforms.
YOUR PROFILE
Education: B.Tech/B.S./M.S. in Computer Science, Statistics, Mathematics, or related field.
Technical Skills
Programming: Python (mandatory), SQL. ML Libraries: Scikit-learn, TensorFlow, PyTorch. Data Processing: Pandas, NumPy, Spark.
ML & AI Expertise: Supervised & Unsupervised Learning, model optimization, hands-on experience with ML platforms/services.
Recommender Systems: Designing and deploying recommendation engines; knowledge of ranking algorithms and personalization.
Generative AI: LLMs (GPT, Llama, etc.), prompt engineering, RAG frameworks (LangChain, LlamaIndex). Exposure to multimodal AI (LVM) is a strong plus.
MLOps & Deployment: CI/CD for ML pipelines, Docker, Kubernetes, model monitoring tools.
Data Engineering Understanding: Data pipelines, ETL processes, data warehousing concepts.
Additional Skills & Preferred Qualifications
Strong communication and stakeholder management skills. Self-motivated, customer-focused, detail-oriented. Azure ecosystem experience (Azure ML, Databricks) preferred. Exposure to real-time data processing. ML/AI/Cloud certifications a plus. SAP ERP knowledge strongly preferred. Six Sigma or ITIL certification a plus.
Data Science, Machine Learning, Artificial Intelligence, Python, SQL, Predictive Analytics, Prescriptive Analytics, Exploratory Data Analysis, Feature Engineering, Model Optimization, Scikit-Learn, TensorFlow, PyTorch, Pandas, NumPy, Spark, Azure Machine Learning, AWS SageMaker, Google Vertex AI, MLOps, ML Pipelines, Model Lifecycle Management, Model Versioning, Model Monitoring, Model Retraining, Recommender Systems, Collaborative Filtering, Content-Based Filtering, Hybrid Recommendation Systems, Ranking Algorithms, Personalization, User Segmentation, Generative AI, Large Language Models, Large Vision Models, GPT, Llama, Prompt Engineering, RAG, LangChain, LlamaIndex, Multimodal AI, OpenAI, Azure OpenAI, Hugging Face, Data Engineering, Data Pipelines, ETL, Data Warehousing, Databricks, Hadoop, Docker, Kubernetes, CI/CD, API Development, Microservices, Model Deployment, Model Drift Detection, Explainable AI, Responsible AI, Data Governance, Real-Time Data Processing, Agile, Enterprise AI Architecture, Technical Leadership, Team Mentoring, Stakeholder Management, SAP ERP, Six Sigma, ITIL