Search by job, company or skills

Senior AI Engineer (Voice / Speech)

  • Posted an hour ago
  • Be among the first 10 applicants

Job Description

About the Role

Masan Group's AI team is extending our multi-agent AI platform — currently text-based and in production — into voice. You will own the speech layer of the agentic system: enabling agents to listen, understand, and speak naturally in real-world consumer and operational contexts across Masan's ecosystem.

The platform is established and in production; voice is its next frontier. You will shape how voice works in our agent stack — new problems, new technology choices, fast project cycles — building on a proven foundation rather than starting from zero.

What You Will Do

  • Build voice capabilities into the multi-agent platform: speech-to-text, text-to-speech, and (as the field evolves) speech-native/real-time voice agent pipelines.
  • Design the voice-agent interaction layer: turn-taking, latency budgets, interruption handling, streaming inference, and integration with agent orchestration/tool-calling.
  • Evaluate and adapt Vietnamese-language speech models — a core requirement given our market — including fine-tuning and benchmarking for accents, noise, and domain vocabulary (retail, F&B, operations).
  • Prototype fast, then productionize: take voice features from feasibility study to reliable deployed services.
  • Contribute to the shared agentic platform (evaluation, observability, agent memory/tools) beyond the voice specialty.

Must-Have

  • 3+ years in ML/AI engineering, with hands-on speech AI experience: ASR and/or TTS model integration, fine-tuning, or deployment (e.g., Whisper-family, commercial speech APIs, open-source TTS, or speech LLMs).
  • Experience building on LLM/agentic systems (function calling, RAG, multi-agent frameworks) or strong evidence of ability to ramp quickly — the agent platform is the foundation this role builds on.
  • Strong Python and production ML engineering: serving models, streaming pipelines, latency optimization.
  • Comfort with ambiguity: able to take a new capability from concept to production without a playbook.

Strong Plus

  • Real-time/streaming voice systems (WebRTC, streaming ASR/TTS, duplex voice agents).
  • Vietnamese speech processing experience.
  • Data engineering fundamentals (audio data pipelines, dataset curation at scale).
  • Contributions to open-source speech or agent projects.

More Info

Job Type:
Industry:
Employment Type:

About Company

Job ID: 151756405