Search by job, company or skills

AI Expert Voice Bot Systems

  • Posted 3 hours ago
  • Be among the first 10 applicants

Job Description

About the Role

We are looking for an AI Expert to join our team as the technical anchor for next-generation real-time voice bot systems. While our Software Engineers focus on integration and infrastructure, this role owns the intelligence layer — designing, training, evaluating, and optimizing the AI models that make voice interactions actually work: from acoustic understanding to natural speech synthesis, speaker identification, and language comprehension in noisy, real-world conditions.

The ideal candidate has built voice AI systems in production — not just integrated APIs, but understood what's happening inside them and made deliberate choices about model architecture, training data, evaluation methodology, and latency trade-offs. You will work directly with Software Engineers to move models from research into production and with product stakeholders to define what good means for a voice bot.

Key Responsibilities

Voice AI Research & Development

  • Design and develop end-to-end voice pipelines: ASR → NLU → dialogue management → TTS, with full understanding of the trade-offs at each stage
  • Research, evaluate, and fine-tune models for Vietnamese-language voice tasks: ASR accuracy on domain-specific vocabulary (vehicle names, locations, plate numbers), TTS naturalness, speaker diarization in multi-party calls
  • Build and maintain evaluation frameworks for voice AI quality: Word Error Rate (WER), Real-Time Factor (RTF), MOS scoring, speaker turn accuracy, end-to-end latency benchmarks

Model Fine-tuning & Training

  • Fine-tune ASR models (Whisper, Gipformer, wav2vec 2.0) on domain-specific and accent-specific Vietnamese data
  • Fine-tune TTS models (F5-TTS, VoxCPM2, VITS) to produce natural, brand-consistent voice output for CSKH scenarios
  • Build and curate training datasets: audio collection pipelines, annotation tooling, data quality validation
  • Apply speaker diarization and voice activity detection (VAD) to segment and attribute multi-speaker call audio

LLM & NLU Integration

  • Design prompt strategies and context management for LLM-driven dialogue in voice-specific constraints (no markdown, short sentences, natural prosody cues)
  • Develop intent classification, entity extraction, and slot-filling components tuned for spoken Vietnamese, including noisy transcripts from imperfect ASR
  • Evaluate and mitigate hallucination in voice LLM pipelines, including ASR-grounded prompting and transcript correction techniques

Production & Optimization

  • Optimize inference latency for real-time constraints: quantization, batching, streaming inference, model distillation
  • Design and run A/B experiments on model variants in production; define statistical evaluation criteria for voice quality metrics
  • Collaborate with Software Engineers on model serving architecture (GPU inference servers, ONNX runtime, TorchServe, Triton)
  • Build monitoring pipelines for model performance drift, ASR degradation, and TTS quality regression in production

  • Required Qualifications

    Education

    • Bachelor's or Master's degree in Computer Science, Electrical Engineering, Applied Mathematics, Linguistics, or a directly related field. PhD is a plus for research-heavy responsibilities.

    Hands-on voice AI experience (required)

    • 3–5+ years of hands-on experience building and deploying voice AI systems — not just integrating cloud APIs, but working with the models themselves
    • Demonstrated experience with at least two of: ASR fine-tuning, TTS model training, speaker diarization, VAD, voice conversion, or acoustic modeling
    • Prior production experience with real-time voice pipelines where latency, reliability, and audio quality were hard constraints

    Deep learning & NLP foundations

    • Solid understanding of sequence modeling (Transformer, CTC, RNN), attention mechanisms, and how they apply to speech and language tasks
    • Familiarity with acoustic features (MFCCs, mel spectrograms, filterbanks) and audio signal processing fundamentals
    • Experience with speech-specific deep learning frameworks: ESPnet, SpeechBrain, NeMo, Hugging Face transformers + datasets for audio

    Engineering skills

    • Strong Python; comfortable with PyTorch or TensorFlow for model development and fine-tuning
    • Ability to write production-quality code, not just research scripts: modular, testable, observable
    • Experience with ML experiment tracking (MLflow, W&B), model versioning, and reproducibility practices
    • Familiarity with Docker and GPU infrastructure for model serving

    Evaluation mindset

    • Knows that a model is only as good as its evaluation — can design rigorous offline and online evaluation protocols for voice AI
    • Experience with human evaluation studies (MOS, MUSHRA, preference tests) and automated proxies

    Nice to Have
    • Experience with multilingual or low-resource language ASR/TTS, especially Vietnamese
    • Familiarity with speaker diarization systems (pyannote, NeMo MSDD) applied to call center audio
    • Knowledge of telephony audio constraints: G.711 µ-law/a-law codecs, 8kHz narrowband limitations, packet loss concealment
    • Experience with streaming ASR (chunk-based inference, endpointing) for real-time applications
    • Contributions to open-source speech/NLP projects or published research in voice AI

    What This Role Is Not
    • This is not an API integration role. We need someone who understands what happens inside the model, not just how to call it.
    • This is not a pure research role. Models need to ship into production voice systems with real latency and reliability requirements.
    • This is not a data annotation role. You will design the data strategy and evaluation criteria — others will execute annotation at scale.
    What We Offer
    • Ownership of the AI layer in a production voice system serving real customers
    • Direct collaboration with Software Engineers, DevOps, and product teams — short feedback loops between research and production
    • Access to real call audio data for training and evaluation (with appropriate privacy controls)
    • Competitive compensation with equity

    More Info

    Job Type:
    Industry:
    Employment Type:

    About Company

    Job ID: 151927627

    Beware of Scammers

    We don’t charge money for job offers