Job Description
About the Role
We are looking for Senior MLOps & Infrastructure Engineers to build and operate the hybrid AI computing platform behind VinFast's ADAS and autonomous driving programme - on-premise GPU clusters running perception model training and large-scale inference, a multi-petabyte sensor data archive, and the platform services used daily by our engineering and annotation teams.
What the Team Covers
You will contribute across all of the following over time. We do not expect one person to master every part on day one, but we do expect you to be willing to work in any of them.
GPU and HPC clusters running model training, large-scale inference and annotation workloads
MLOps: model serving, pipeline orchestration, experiment and model registry
Storage at multi-petabyte scale: tiering, lifecycle, backup and recovery
Ingestion of large volumes of recorded sensor data, with integrity verification and metadata extraction
CI/CD, GitOps, infrastructure-as-code, monitoring and observability
Access control, audit logging and data protection for internal and external users
Key Responsibilities
AI Compute & Serving
Operate and scale GPU clusters across training, inference and annotation workloads; manage scheduling, utilisation and capacity
Deploy and scale high-throughput model serving; tune GPU memory and runtime performance
Manage multi-GPU distributed training and reproducible environment sandboxing
Scale pipeline orchestration on Kubernetes for large data processing jobs
Platform Automation & Delivery
Drive GitOps-based deployment and maintain infrastructure-as-code across the platform
Build and maintain CI pipelines, secure container builds and release automation
Build monitoring, logging and alerting so that failures are detected and actionable, never silent
Lead incident response and drive follow-up actions to closure
Data Infrastructure & Security
Operate storage at multi-petabyte scale: object storage, local storage, tiering, backup and recovery
Operate the ingestion path for large sensor data deliveries, with integrity verification and metadata extraction
Implement access control and single sign-on across platform services, with audit logging
Apply data protection measures to sensitive content before it reaches external users
- 4+ years in MLOps, DevOps, SRE or HPC platform engineering, with production ownership of AI/ML infrastructure
- Kubernetes administration at production scale: Helm, ingress (Traefik or Envoy), CNI, and GitOps (Flux or ArgoCD)
- GPU & HPC: SLURM, NVIDIA Container Toolkit, CUDA runtime tuning, multi-GPU memory debugging
- Strong Linux systems skills and infrastructure-as-code (SaltStack or Ansible)
- Python and Bash for automation, including Airflow DAGs and custom operators
- Object storage, and a monitoring and logging stack (Prometheus, Grafana or equivalent)
- Willingness to work across the full stack - compute, storage, networking, automation and security - rather than within a single specialty
- Good communication in English - technical documentation and working with international partners
- Competitive salary
- Premium healthcare package, including PVI insurance & annual health check-ups
- 13th-month salary & performance bonuses to reward your contributions
- Enjoy preferential pricing for services within the Vingroup ecosystem including Vinmec, Vinpearl, and Vinschool...
- Opportunity to collaborate with and learn from industry-leading professionals in the automotive domain
- Work Location: Technopark Tower, Gia Lam, Ha Noi
