Search by job, company or skills

Senior Data Engineer

  • Posted 4 hours ago
  • Be among the first 10 applicants

Job Description

About the team

Masan is building a centralized data platform that unifies data across Vietnam's largest consumer-retail ecosystem — WinCommerce's 3,000+ stores, Masan Consumer's FMCG brands, Phúc Long, and the WiN loyalty program serving millions of members. The Data Centralization team owns the Group's lakehouse on Azure Databricks — the ingestion, governance, and serving layers that turn fragmented BU-level data into a single, trusted source of truth powering analytics, AI/ML, and real-time business decisioning at Group scale.

You will be the technical lead for the platform's core pillars — data governance, platform billing & cost management, data ingestion, and data serving on Azure Databricks — owning architecture decisions, setting engineering standards, and leading a team of data engineers while remaining hands-on in the most critical builds. You will work directly with BU data teams, the AI platform team, and business stakeholders across one of the largest and most diverse data estates in Vietnam.

What You Will Do

  • Own the end-to-end architecture and operation of the Group lakehouse on Azure Databricks — Delta Lake, Unity Catalog, and medallion architecture — across ingestion, governance, and serving
  • Data governance: design and enforce the Unity Catalog governance model — catalog/schema standards, fine-grained access control (row/column-level security, dynamic PII masking), data classification, lineage, audit, and data quality monitoring (DLT expectations, Lakehouse Monitoring)
  • Billing & cost management: own platform spend end-to-end — monitor DBU and compute usage via Databricks system tables and Azure Cost Management, enforce tagging and cluster policies, run budgets and alerts, and deliver BU-level chargeback/showback reporting
  • Drive continuous cost optimization: job and SQL warehouse right-sizing, Photon, spot instances, serverless adoption, and storage optimization (OPTIMIZE, liquid clustering, retention policies)
  • Ingestion: design and operate batch, CDC, and streaming pipelines into the lakehouse — Azure Data Factory, Auto Loader, Delta Live Tables, Kafka/Event Hubs — from POS, ERP, loyalty, and e-commerce sources
  • Serving: own the serving layer for analytics and ML — Databricks SQL Warehouses for BI (Power BI), governed data products and Delta Sharing to BUs, feature-ready data for AI/ML — with SLAs for freshness, performance, and concurrency
  • Lead, mentor, and grow a team of 4–8 data engineers; set standards for code review, testing, documentation, and CI/CD (Databricks Asset Bundles, Terraform, Azure DevOps)
  • Partner with BU data teams, the AI platform team, and security/compliance to onboard new data domains and consumers under a governed operating model

Must-Have

  • 5+ years in data engineering, with 2+ years leading engineers or owning platform architecture as a tech lead — ideally in high-volume environments (retail, e-commerce, fintech, telco)
  • Deep, hands-on production experience with Databricks on Azure: Spark, Delta Lake, Unity Catalog, Delta Live Tables, Workflows, and Databricks SQL
  • Data governance at scale: Unity Catalog access model (RBAC, row/column-level security, dynamic masking), PII classification and protection, lineage, audit, and data quality frameworks
  • Platform billing & FinOps: DBU/cost analysis with Databricks system tables (system.billing), cluster policies and tagging standards, chargeback/showback models, Azure Cost Management
  • Solid Azure data stack: ADLS Gen2, Azure Data Factory, Event Hubs, Entra ID, Key Vault; familiarity with VNet injection and Private Link
  • Strong batch and streaming ingestion: Structured Streaming, Auto Loader, Kafka/Event Hubs, and CDC patterns
  • Strong SQL and Python; solid data modeling fundamentals (dimensional, medallion, domain-oriented)
  • CI/CD for data and infrastructure-as-code (Terraform, Databricks Asset Bundles); orchestration with Databricks Workflows or Airflow
  • Able to communicate architecture and trade-offs to non-technical stakeholders; English working proficiency

Nice to have

  • Databricks certifications (Data Engineer Professional, Platform Administrator)
  • Delta Sharing, Lakehouse Federation, or data mesh / data domain operating models at multi-BU scale
  • FinOps practices or certification; serverless and Photon cost tuning at scale
  • Experience supporting ML feature platforms (MLflow, Feature Engineering on Unity Catalog)

More Info

Job Type:
Industry:
Employment Type:

About Company

Job ID: 151924301

Similar Jobs

Ho Chi Minh, Vietnam

Skills:

snowflake Web ServicesCsvPysparkKafkaPostgresJsonSqlDevopsSpark StreamingXmlDatabricksPythonAWSParquetweb API frameworksSpark structured streamingTextCI CDAgentic AIDeltaMLFlowgit-based version control

Vietnam, Ho Chi Minh

Skills:

Data WarehousingData SecurityData IntegrationDatabase ManagementCollaborationETL ProcessesClickHouseData Quality and GovernanceDocumentation and ReportingData Pipeline DevelopmentAirFlowPerformance OptimizationAirByte

Remote

Skills:

data engineering PythonEtl ProcessPysparkSparkSqlAzure CloudGitTerraformMicrosoft FabricETL Developerbicep

India, Remote

Skills:

GithubSqlAzure SynapseApache SparkScala

Ho Chi Minh, Vietnam

Skills:

JavaTest Automation ToolsApi TestingTest automationTypescriptPythonData integrity verificationPrompt evaluationCI CD and DevOps practicesAI ML testingData pipeline testingSQL and dataset analysisBig data platformsCloud platforms AWSData quality validation

Beware of Scammers

We don’t charge money for job offers