Search by job, company or skills

Expert, Site Reliability Engineering Data Platform

  • Posted 14 hours ago
  • Be among the first 10 applicants

Job Description

About the Role

We're hiring a Site Reliability Engineer to make Techcombank's data platform reliable. You'll write real code, own production systems end to end, and replace manual toil with automation — increasingly with AI in the loop.

Our platform runs on AWS: Databricks and Unity Catalog, EKS, Glue, Lambda, MSK/Kafka, Flink, Airflow, provisioned with Terraform and delivered through Jenkins, GitLab and ArgoCD. Data engineers, scientists and analysts across the bank depend on it.

We're deliberately open on background. Strong candidates come from SRE, backend/platform software engineering, or infrastructure and DevOps. What matters is that you build software to solve operational problems rather than absorbing them manually.

What You'll Do

- Own reliability for production data platform services: define SLOs, run incident response, drive blameless postmortems, and fix root causes.

- Build automation and internal tooling that removes toil — self-service provisioning, automated remediation, guardrails in CI/CD.

- Improve observability across ingestion, transformation and delivery so failures surface before users notice.

- Harden the delivery path: safer deployments, faster rollbacks, change risk caught before production.

- Apply AI where it pays off — incident summarisation, log and metric triage, RCA assistance, runbooks that execute instead of instruct. We'll support you in learning this if it's new.

- Partner with data engineering squads and mentor engineers on reliability practice.

What We're Looking For

- 8+ years building and operating production systems as an SRE, software engineer, or infrastructure/DevOps engineer.

- Strong coding ability in Python (or Go/Scala/Java) — you ship tools and services, not just scripts.

- Solid experience on AWS and with Infrastructure as Code (Terraform preferred).

- Hands-on with containers and Kubernetes, and with CI/CD pipelines.

- Practical observability experience: metrics, logs, tracing, alerting that people trust.

- Good at English

Nice to Have

Any of these will make you stand out; none are required.

- Data platform experience: Databricks, Spark, Kafka, Flink, Airflow, Glue.

- Using or building with GenAI — LLM APIs, prompt engineering, RAG, agents in production, AI-assisted coding at scale.

- Data quality, lineage, or governance tooling (Great Expectations, Deequ, Unity Catalog).

- Incident management practice in a regulated environment.

- Clear writing. Postmortems and runbooks are engineering artefacts here.

More Info

Job Type:
Industry:
Employment Type:

About Company

Job ID: 152170469

Similar Jobs

Ho Chi Minh, Vietnam

Skills:

KubernetesPythonLangChainaiohttpVector databasesQdrantDocker ComposePineconeLangSmithLangGraphasyncioLangfusePydanticMilvusW B

Ho Chi Minh, Vietnam

Skills:

CloudformationScalaPrometheusKafkaGrafanaDatadogTerraformSparkDatabricksPythonAirflowGenerative AIFlinkOpenTelemetryFAISSGlueWeaviate

Beware of Scammers

We don’t charge money for job offers