The Tech Lead / Voice AI Architect is responsible for designing and leading end-to-end production-grade voice AI solutions, covering Automatic Speech Recognition (ASR), Natural Language Understanding (NLU), Large Language Models (LLMs), dialogue management, Text-to-Speech (TTS), and telephony integrations. The role focuses on building scalable, low-latency voice pipelines, evaluating speech technology platforms, optimizing production performance, and mentoring engineering teams while driving technical excellence. The ideal candidate has 7–10 years of experience in ML/Speech Engineering, including 2+ years of architecting production voice pipelines, with expertise in conversational AI, speech technologies, and large-scale voicebot deployments.
|
Location |
Hyderabad (primary hub — deepest voice-AI talent pool in India outside Bangalore) |
|
Experience |
7-10 years in ML/speech engineering, 2+ years architecting production voice pipelines |
• Own end-to-end architecture: ASR → NLU/LLM → dialogue manager → TTS → telephony
• Make build-vs-buy calls on ASR/TTS vendors (Deepgram, Azure Speech, ElevenLabs, or open-source Whisper/Coqui)
• Own latency budget across the pipeline (target sub-1.5s round-trip for live calls)
• Mentor Conversational AI and ML/Voice engineers; own technical hiring bar
• Production experience with at least one ASR engine (Whisper, Deepgram, Google STT, Azure Speech) at scale
• Hands-on experience with LLM-based dialogue systems (function calling, RAG, or fine-tuning)
• Strong grasp of telephony integration: SIP, WebRTC, Twilio/Exotel/Ozonetel or similar
• Experience optimizing for latency and cost simultaneously in a live voice pipeline
• Comfortable working across Hindi/Hinglish and English accent handling for Indian/UAE markets
• Track record of hitting production accuracy benchmarks: 95%+ transcription (WER-based) accuracy and 90%+ intent classification accuracy at scale — not just in a demo
• Experience handling dialect variation and code-switching (e.g. Hinglish, Gulf Arabic variants) in production ASR/NLU
• Successfully deployed and managed Voicebots in a production environment.
• Experience handling 5000+ live customer calls per day.
• Prior experience at Amazon Alexa, Microsoft Cortana/Speech, Google Assistant, or a voice-AI startup
• Exposure to Arabic ASR/TTS for UAE market
• 2+ years hands-on with LLMs/Generative AI in a production (not POC) setting
Voice AI, Conversational AI, Speech Recognition, Automatic Speech Recognition (ASR), Natural Language Understanding (NLU), Large Language Models (LLMs), Generative AI, Dialogue Management, Text-to-Speech (TTS), Voicebot Development, Production Voice Pipelines, Whisper, Deepgram, Google Speech-to-Text, Azure Speech, ElevenLabs, Coqui, Retrieval-Augmented Generation (RAG), Function Calling, LLM Fine-Tuning, Prompt Engineering, Telephony Integration, SIP, WebRTC, Twilio, Exotel, Ozonetel, API Integration, Low-Latency System Design, Performance Optimization, Cost Optimization, Real-Time AI Systems, Machine Learning, Speech Engineering, Voice Pipeline Architecture, Intent Classification, Speech-to-Text (STT), Multilingual AI, Hindi, Hinglish, English, Dialect Handling, Code-Switching, Voice Analytics, Production Deployment, Model Monitoring, Scalability, Technical Leadership, System Architecture, Mentoring, Technical Hiring, Problem Solving, Communication, Cross-Functional Collaboration.