About the Role
Masan Group's AI team is extending our multi-agent AI platform — currently text-based and in production — into voice. You will own the speech layer of the agentic system: enabling agents to listen, understand, and speak naturally in real-world consumer and operational contexts across Masan's ecosystem.
The platform is established and in production; voice is its next frontier. You will shape how voice works in our agent stack — new problems, new technology choices, fast project cycles — building on a proven foundation rather than starting from zero.
What You Will Do
- Build voice capabilities into the multi-agent platform: speech-to-text, text-to-speech, and (as the field evolves) speech-native/real-time voice agent pipelines.
- Design the voice-agent interaction layer: turn-taking, latency budgets, interruption handling, streaming inference, and integration with agent orchestration/tool-calling.
- Evaluate and adapt Vietnamese-language speech models — a core requirement given our market — including fine-tuning and benchmarking for accents, noise, and domain vocabulary (retail, F&B, operations).
- Prototype fast, then productionize: take voice features from feasibility study to reliable deployed services.
- Contribute to the shared agentic platform (evaluation, observability, agent memory/tools) beyond the voice specialty.
Must-Have
- 3+ years in ML/AI engineering, with hands-on speech AI experience: ASR and/or TTS model integration, fine-tuning, or deployment (e.g., Whisper-family, commercial speech APIs, open-source TTS, or speech LLMs).
- Experience building on LLM/agentic systems (function calling, RAG, multi-agent frameworks) or strong evidence of ability to ramp quickly — the agent platform is the foundation this role builds on.
- Strong Python and production ML engineering: serving models, streaming pipelines, latency optimization.
- Comfort with ambiguity: able to take a new capability from concept to production without a playbook.
Strong Plus
- Real-time/streaming voice systems (WebRTC, streaming ASR/TTS, duplex voice agents).
- Vietnamese speech processing experience.
- Data engineering fundamentals (audio data pipelines, dataset curation at scale).
- Contributions to open-source speech or agent projects.