About Vindynamics
At
VinDynamics, we design safe, affordable, and intelligent humanoid robots to assist in everyday life —
robots for everyone. Backed by
Vingroup, Vietnam's leading technology conglomerate, we are on a mission to make advanced robotics accessible, reliable, and beneficial for billions of people worldwide. By combining cutting-edge AI, world-class engineering, and human-centered design, we aim to seamlessly integrate robots into daily life — enhancing safety, productivity, and happiness at home and beyond.
We are looking for a Senior QC Engineer specializing in testing Conversational AI systems (chatbots/voicebots). This role requires a different mindset from traditional software QC: beyond verifying functional correctness, it also involves evaluating response quality, semantic accuracy, safety, and the natural conversational experience of the AI system.
Job Description
- Conversational AI Test Strategy
- Develop test strategies for conversation flows, intents, and entities of the chatbot/voicebot.
- Design test case sets covering a wide range of conversational scenarios: happy path, edge cases, multi-turn conversations, flow interruptions, and context switching.
- Establish evaluation criteria for AI response quality (accuracy, relevance, tone, naturalness).
- Functional & Response Quality Testing
- Test the accuracy of NLU (Natural Language Understanding): intent recognition and entity extraction.
- Evaluate the quality of AI responses: correctness, coherence, contextual relevance, and avoidance of repetition or rambling.
- Test the system's ability to handle ambiguous cases, out-of-scope questions, and multiple languages (if applicable).
- Test the integration between the AI conversation engine and backend systems (API, CRM, database).
- AI Safety & Ethics Testing
- Assess the risk of the AI producing incorrect, biased, or inappropriate (harmful/toxic) content.
- Test the AI's ability to appropriately decline sensitive or unsafe requests.
- Work with the AI/ML team to report and improve cases of AI hallucination.
- Automated Testing & Large-Scale Evaluation
- Build automated test suites for conversation flows.
- Design and operate large-scale response quality evaluation processes (batch evaluation, sampling review).
- Collaborate on building a golden dataset/benchmark to measure quality over time.
- Defect Management & Reporting
- Log and classify conversation defects (incorrect intent, incorrect context, inappropriate responses, safety issues).
- Collaborate with AI/ML Engineers, Prompt Engineers, and Conversation Designers to improve quality.
- Provide periodic reports on AI system quality (accuracy rate, CSAT, escalation rate to human agents).
- Process Improvement
- Propose conversational AI quality evaluation processes tailored to the product's specific needs.
- Train and mentor junior/middle QC staff on AI conversation testing methods.
Requirements
- Minimum 5 years of software QC/QA experience, including experience testing chatbots/voicebots/conversational AI systems.
- Application-level understanding of how NLU, NLP, and LLMs (Large Language Models) work.
- Experience evaluating the output quality of AI models (prompt-response evaluation).
- Understanding of concepts such as intent, entity, context, multi-turn conversation, and RAG (Retrieval-Augmented Generation) is a strong plus.
- Able to write basic test scripts (Python) to automate API calls and evaluate responses.
- Proficient with test case & bug management tools: Jira, TestRail, or equivalent.
- Experience using AI evaluation tools (LLM-as-judge, evaluation frameworks) is a plus.
- Basic understanding of prompt engineering.
- Strong language analysis skills, sensitive to nuance and conversational context.
- Ability to ask adversarial questions to test the limits of the AI system (basic red-teaming).
- Good communication skills, able to collaborate with the AI/ML team, Conversation Designers, and Product.
- Careful and highly responsible regarding AI safety and ethics.
- Good reading comprehension of technical/AI research documents in English.
- Nice to have: Experience with common chatbot platforms (Dialogflow, Rasa, Botpress) or LLM API integration (OpenAI, Anthropic, etc.).
Benefits
- Competitive compensation package based on experience and qualifications
- Opportunity to build a strategic global data marketplace for robotics and AI training data from zero to one.
- Work in a high-speed technology environment backed by Vingroup and VinDynamics leadership.
- Competitive compensation package aligned with capability and business impact.
- Clear ownership, measurable KPIs, and exposure to global partners, US platform models, and frontier robotics businesses.