
Search by job, company or skills
Showing 6 jobs
Skills:
Change Management, Terraform, Ansible, Incident Management, Problem Management, Python, Azure, AWS, Alerting, Troubleshooting, Observability, Monitoring, SRE concepts, ITIL principles
Skills:
Incident Management, Infrastructure optimization, Forecasting, Cloud cost optimization, AWS billing analysis, FinOps, Budget Tracking, Production Support, Capacity Planning, Linux systems knowledge, Performance Analysis, Tagging, cost allocation, Reliability improvements

Skills:
Gcp, Terraform, Kubernetes, Python, Logging, AWS, Infrastructure as Code, Go, Observability, Monitoring
Skills:
Incident Management, Forecasting, Cloud cost optimization, AWS billing analysis, FinOps, Budget Tracking, Production Support, Linux systems knowledge, Tagging, cost allocation, Infrastructure optimization, Capacity Planning, Performance Analysis, Reliability improvements
Skills:
Ipmi, Prometheus, Grafana, Gcp, Terraform, Ansible, AWS, automated anomaly detection, production-grade automation, telemetry pipelines, Redfish, Temporal.io, Victoria Metrics, Go, Linux Unix systems, Cadence, distributed tracing, AI-driven predictive scaling, infrastructure-as-code, hardware lifecycle operations, Telegraf, AIOps
Skills:
Unix, PostgreSQL, Prometheus, Grafana, Docker, Terraform, MySQL, Python, AWS, Java, Cloudformation, Gcp, Linux, Incident Management, Azure, Kubernetes, AWS CDK, Go, Production Engineering, anomaly detection, OpenTelemetry, Site Reliability Engineering, LLMs, Platform Engineering
