Search by job, company or skills

Site Reliability Engineer (SRE)

  • Posted 5 hours ago
  • Be among the first 10 applicants

Job Description

We are looking for a Site Reliability Engineer (SRE) to join our engineering team, responsible for improving system reliability, scalability, and operational efficiency across our platform.

This role focuses on engineering solutions to prevent incidents and improve long-term system reliability, rather than only reacting to operational issues. You will work closely with L2 Support, DevSecOps, and development teams to analyze production issues, eliminate recurring problems, and enhance system resilience in a distributed, cloud-based environment.

You will play a key role in transforming operational insights into long-term improvements, automation, and reliability engineering practices.

Responsibilities

Your key responsibilities will include, but are not limited to:

Reliability Engineering & System Improvement

  • Design and implement solutions to improve system reliability, availability, and performance
  • Identify and eliminate recurring incidents and systemic bottlenecks
  • Define and track service level indicators (SLIs) and service level objectives (SLOs)

Root Cause Analysis & Problem Management

  • Lead root cause analysis (RCA) for major and recurring incidents
  • Identify systemic issues and define long-term corrective actions
  • Ensure RCA outcomes are translated into engineering improvements

Observability & Monitoring

  • Design and enhance monitoring, alerting, and observability frameworks
  • Build dashboards and metrics to provide visibility into system health
  • Analyze system performance (latency, throughput, resource usage)
  • Identify bottlenecks and optimization opportunities
  • Contribute to capacity planning and scaling strategies

Automation & Operational Efficiency

  • Develop and maintenance automation to reduce manual operational effort and improve recovery time (e.g., self-healing, auto-remediation)
  • Continuously optimize operational processes

Qualifications

We are looking for candidates who demonstrate:

  • Bachelor's degree in IT, Computer Science, or a related field (or equivalent experience)
  • 6 years experience or more in SRE or Production Engineering roles
  • Strong troubleshooting and analytical skills across distributed systems
  • Experience with:
  • Linux systems
  • Cloud platforms (AWS / Azure)
  • Microservices architecture
  • Monitoring and observability tools
  • CI/CD pipelines
  • Strong understanding of:
  • System design and failure scenarios
  • Distributed system behavior (timeouts, retries, cascading failures)
  • Good communication skills in English

 

Technical skills

  • Scripting/programming (Python, Bash, or similar)
  • Experience with container platforms (Docker, Kubernetes)
  • Familiarity with databases (MySQL, PostgreSQL)
  • Understanding of networking fundamentals

 

Nice-to-have

  • Experience with high-availability or high-transaction systems (e.g., fintech, payments)
  • Familiarity with:
  • Infrastructure as Code (Terraform, CloudFormation)
  • Chaos engineering or resilience testing
  • Knowledge of security best practices in cloud environments

 

Success in this role looks like

  • System reliability and availability improve over time
  • Recurring incidents are significantly reduced
  • MTTR decreases due to automation and improved processes
  • Monitoring becomes more accurate with fewer false alerts
  • Production releases become safer and more predictable

Benefits

STYL Solutions will give you the favorable conditions you need to tackle difficult problems and learn cutting-edge technologies:

  • An international working environment with friendly and passionate colleagues
  • Onsite opportunity to Japan, and Singapore for training and supporting customer
  • Meaningful work with experienced & strong technical veterans
  • Flat structure, simple processes & transparency

In addition to providing you with professional growth, STYL Solutions is also committed to taking care of our employees, personally. As a full-time employee, you are automatically enrolled in our benefits program, which includes:

  • Attractive compensation, regular assessments, and salary reviews
  • 19 paid days off per year
  • 100% salary & full social insurance during the probation period
  • Premium health care insurance
  • Free coffee, tea, and parking.
  • Special celebration on 8/3, 1/6, Xmas, Tet holiday...
  • Outing/team-building activities (trip, sport, dinner...)

Salary: Negotiation

Employment Type: Full-time

Important Note:

By submitting an application, or sending your CV to us:

a) you acknowledge that you have read, understood, and agreed to STYL's Candidate Privacy Notice, and consent to the collection, use, and/or disclosure of your personal data by us for the purposes set out in the Notice; and

(b) in the event that we have received your job application or personal data from any third party pursuant to the purposes set out in the Notice, you warrant that such third party has been duly

authorized by you to disclose your personal data to us for the purposes set out in the Notice.

 

More Info

Job Type:
Industry:
Employment Type:

Job ID: 152482145

Beware of Scammers

We don’t charge money for job offers