Search by job, company or skills

Senior ML Engineer (ML Platform & Production)

  • Posted 7 hours ago
  • Be among the first 10 applicants

Job Description

Company Description:

VinSmart Future (VSF) is the leading technology company within the Vingroup Corporation, formed by the merger of the group's entire technology ecosystem, including VinApp, VinIT, VinBigdata, and other tech units. As a core driver of Vingroup's future growth, VSF is at the forefront of technological development, with artificial intelligence (AI) as its foundation. With a talented team of nearly 4,000 local and international technology experts, VSF focuses on creating high-utility technologies that enhance lives and connect data, models, and infrastructure to unlock new possibilities.

Role Overview:

We are building a large-scale AI/ML platform for a next-generation messaging application, enabling intelligent experiences across personalization, recommendations, growth, engagement, trust & safety, and communication.

As a Senior ML Engineer – ML Platform & Production, you will design and build the infrastructure that takes machine learning from experimentation to reliable production systems. You will work across the full ML lifecycle, including model development, training pipelines, model management, deployment, online inference, monitoring, and continuous improvement.

You will partner closely with Data Scientists, Applied AI Engineers, Data Engineers, Backend Engineers, and Product teams to make ML systems scalable, reliable, observable, and easy to operate.

This is a hands-on engineering role for someone who wants to build ML infrastructure serving millions of users today and tens of millions globally in the future.

Key Responsibilities:

Build ML Platform & Infrastructure

  • Design and implement scalable infrastructure supporting the complete ML lifecycle from development and training to production inference and continuous improvement.
  • Build reusable pipelines for data preparation, training, evaluation, model registration, deployment, and monitoring.
  • Establish standardized workflows for experimentation, model versioning, release, rollback, and continuous delivery.
  • Build production-grade ML services and inference pipelines for real-time and batch workloads.
  • Design ML systems that support both cloud inference and on-device/edge inference.

Production ML & MLOps

  • Build and operate reliable ML systems with strong requirements for latency, scalability, availability, and cost efficiency.
  • Implement CI/CD and automation for ML training and deployment workflows.
  • Establish model lifecycle management, including model registry, versioning, validation, release, and rollback.
  • Develop systems for safe and controlled model promotion from development to production.
  • Build reusable tooling that enables Data Scientists and Applied AI teams to deploy models efficiently.

ML Serving & Model Optimization

  • Design and optimize online inference services for real-time applications.
  • Improve inference performance through model optimization, quantization, distillation, batching, caching, and hardware acceleration.
  • Support deployment and lifecycle management of lightweight models on mobile and edge devices.
  • Design mechanisms for distributing, updating, and rolling back edge model versions.
  • Work with Backend and Client teams to integrate ML inference into production user experiences.

Data, Features & ML Observability

  • Partner with Data Platform engineers to build reliable training and inference data pipelines.
  • Develop systems supporting offline and online features while maintaining training-serving consistency.
  • Build monitoring for model performance, data quality, feature quality, drift, latency, errors, and resource utilization.
  • Establish production telemetry and feedback loops to continuously improve ML models.
  • Diagnose and resolve complex production ML issues.

Technical Leadership

  • Drive engineering best practices for production machine learning.
  • Review architecture and code and mentor other engineers.
  • Help define technical standards for model deployment, monitoring, testing, and operational readiness.
  • Work with cross-functional teams to translate ML requirements into scalable technical solutions.

Key Requirements:

Required

  • 5+ years of experience in Machine Learning Engineering, Software Engineering, MLOps, or a related field.
  • Strong Python programming and software engineering fundamentals.
  • Strong understanding of production ML systems and the ML lifecycle.
  • Hands-on experience deploying and operating machine learning models in production.
  • Experience building scalable data or ML pipelines.
  • Experience with model versioning, experiment tracking, deployment, monitoring, and rollback.
  • Experience with cloud infrastructure and containerized applications.
  • Strong understanding of distributed systems, APIs, CI/CD, testing, and observability.
  • Ability to troubleshoot complex production systems.
  • Strong communication and collaboration skills.

Nice to Have

  • Experience with ML platforms such as MLflow, Kubeflow, Vertex AI, SageMaker, Databricks, or equivalent.
  • Experience with Kubernetes and distributed computing.
  • Experience with Spark, Kafka, Airflow, or similar technologies.
  • Experience with feature stores or online feature serving.
  • Experience with recommendation, ranking, personalization, or social-network ML systems.
  • Experience with model serving technologies such as Triton, TensorFlow Serving, TorchServe, or equivalent.
  • Experience with mobile/edge ML deployment.
  • Experience with model quantization, distillation, or inference optimization.
  • Experience operating ML systems at millions-of-users scale.
  • Experience in consumer internet, messaging, social networking, or other high-scale applications.

Benefits:

  • Flexible working hours and attendance policy (Work from Home on working Saturdays).
  • Attractive compensation and bonus packages, highly competitive in the market.
  • Exclusive employee benefits across the Group's ecosystem in accordance with company policies.
  • Opportunity to work on large-scale and strategic technology projects.
  • Professional technology environment with leading scientists, experts, and engineers from top technology companies in Vietnam and around the world.
  • Free access to learning platforms such as Udemy, Coursera, and O'Reilly; internal workshops; sponsorship for professional certifications; and exclusive mentoring programs from the Group and Company leadership team.
  • Full statutory insurance coverage in accordance with Vietnamese Labor Law (Social Insurance, Health Insurance, Unemployment Insurance), along with private healthcare insurance based on job grade and annual health check-ups at reputable hospitals and healthcare centers nationwide.
  • Participation in internal activities, team-building programs, and annual company events.

Location: Vincom Dong Khoi, District 1, Ho Chi Minh City

Contact: Ms. Huyen

Zalo/Call: 0963 957 235

Mail: [Confidential Information]

More Info

Job Type:
Industry:
Function:
Employment Type:

About Company

Job ID: 152545965

Beware of Scammers

We don’t charge money for job offers