Search Jobs

Search by job, company or skills

Senior Site Reliability Engineer

Senior Site Reliability Engineer

talentiser
Early Applicant
  • Posted 2 months ago
  • Be among the first 20 applicants

Job Description

Role Overview

We are hiring backend engineers who specialize in reliability.

SRE at our company is a software engineering role — focused on designing, building, and improving highly available distributed systems through code.

This is not a DevOps / CI-CD / Terraform-heavy role.

What You'll Do

  • Design and build reliable, scalable distributed services
  • Engineer fault tolerance (rate limiting, retries, circuit breakers, backpressure)
  • Improve system resilience through architectural changes
  • Define and enforce SLOs and error budgets
  • Lead production incident deep dives and implement permanent fixes
  • Build automation to eliminate operational toil

Must-Have

  • 4+ years of backend software engineering experience
  • Strong coding skills in Go / Java / C++ / Rust / Python
  • Solid understanding of data structures & algorithms
  • Experience designing distributed systems at scale
  • Strong debugging and production troubleshooting skills

This Role Is NOT

  • Primarily Terraform / IaC
  • CI/CD pipeline management
  • Kubernetes administration
  • Monitoring dashboard setup

Infrastructure knowledge is useful — but software engineering depth is mandatory.

More Info

Job Type:
Industry:
Employment Type:

Key Skills

About Company

Similar Jobs

8-12 yrs
Bengaluru, India
Skills:
Jfrog Artifactory, PostgreSQL, Prometheus, Bash, Grafana, Redis, Rabbitmq, Jenkins, Docker, Terraform, Ansible, Azure, Python, Kubernetes, Loki, Go, OpenTelemetry
5-7 yrs
Bengaluru, India
Skills:
RDS, Elk, PostgreSQL, Prometheus, Grafana, Redis, Gcp, Terraform, Ansible, MySQL, ECS, Python, AWS, GitOps, HashiCorp Vault, Loki, Go, Kubernetes EKS, Google Secret Manager, Memorystore, OpenTelemetry, Cloud SQL, AWS Secrets Manager
3-5 yrs
Bengaluru, India
Skills:
Unix, C, Continuous Integration, Infrastructure Management, Software Architecture, Javascript, Linux, Distributed Systems, Python, disaster recovery planning and implementation, security standards and compliance, Java-based systems, Site Reliability Engineering, Root Cause Analysis, site and system administration, performance optimization techniques, scalability design patterns, Troubleshooting, continuous delivery automation
8-10 yrs
Bengaluru, India
Skills:
Change Management, Terraform, Ansible, Incident Management, Problem Management, Python, Azure, AWS, Alerting, Troubleshooting, Observability, Monitoring, SRE concepts, ITIL principles
4-6 yrs
Bengaluru, India
Skills:
Ibm Cloud, Prometheus, Bash, Grafana, Terraform, Docker, Ansible, Helm, Kubernetes, Python, AWS, Jenkins, Elk, Cloudwatch, Vmware Vsphere, Jira, Datadog, Red Hat OpenShift, OpenTelemetry, Argo CD, Opsgenie, EFK, PagerDuty