Search Jobs

Search by job, company or skills

Manager, Cloud Operations Engineering

Manager, Cloud Operations Engineering

MongoDB
6-11 Years
Quick Apply
  • Posted a month ago
  • Over 100 applicants have applied

Job Description

Cloud Operations Engineers are responsible for building internal tools and process automation. Day-to-day duties are creating and monitoring systems alert dashboards, reviewing critical event and system logs, accessing customer instances that underpin their production databases, and performing server administration duties including performance troubleshooting. Applicants must be critical thinkers who are quick to detect, resolve, or escalate issues that are sometimes broad in scope and difficult to trace.

We are looking to speak to candidates who are based in Bengaluru for our hybrid working model.

Responsibilities

  • Help scale the Cloud Operations Engineering team with the strategic implementation and refinement of processes and tools
  • Provide career development feedback and advice to direct reports
  • Identify and measure team health indicators and performance metrics
  • Ensure proper team focus on priorities, objectives, and related deliverables
  • Collaborate with technical and non-technical teams across the company
  • Balance your time between leading your team, working on customer incidents and being involved in projects
  • Be a source of guidance and advice to your own team members and other teams within MongoDB
  • Build a relationship with your team around trust
  • Successfully coordinate with a global team of Cloud Operations Engineers who are tasked with ensuring our uptime guarantees to the MongoDB Atlas customer base
  • Participate in designing and building internal tools
  • Assist in scoping, designing and deploying systems that reduce Mean Time to Resolve for customer incidents
  • Monitor and detect emerging customer-facing incidents on the Atlas platform; assist in their proactive resolution
  • Automate internal processes, routine monitoring and troubleshooting tasks
  • Diagnose live incidents, differentiate between platform issues versus usage issues, and take the next steps toward resolution
  • Cooperate with our Product Management and Cloud Engineering organizations by identifying areas for improvement in the management applications powering the Atlas infrastructure
  • Coordinate and participate in a weekly on-call rotation, where you will handle short term customer incidents (from direct surveillance or through alerts via our Technical Services Engineers)

Requirements

  • Management skills, with hands-on experience running small to mid-sized Engineering Teams in a rapid-growth environment
  • Strong diagnostic/troubleshooting process, with significant experience troubleshooting end-to-end technical issues in production environments
  • Experience supervising, leading and monitoring progress of Software Development projects
  • Patience, empathy, and a genuine desire to help others
  • Excellent communication skills, both written and verbal
  • Ability to think on your feet, remain calm under pressure, and find solutions to challenges in real-time
  • Experience with being an on-call DevOps, SRE, or Cloud Operations engineer
  • Expertise with Linux system administration and networking technologies
  • Knowledge of database and distributed system operations and concepts
  • Knowledgeable about a wide range of web and internet technologies
  • Familiarity with Amazon Web Services and other Cloud infrastructure platforms (e.g. GCP, Azure)
  • Experience in monitoring, system performance data collection and analysis, and reporting
  • Capability to write programs/scripts to solve both short-term systems problems and long-term strategic objectives for the Atlas product
  • A CS/CE degree or equivalent experience
  • At least 2 of the following programming languages: Java, Go, Python, Typescript
  • A keen interest in learning new skills and competencies

More Info

Key Skills

cloud infrastructure (AWS/GCP/Azure)

programming (java/go/python/typescript)

About Company

MongoDB is the next-generation database that helps businesses transform their industries by harnessing the power of data. The world’s most sophisticated organizations, from cutting-edge startups to the largest companies, use MongoDB to create applications never before possible at a fraction of the cost of legacy databases. MongoDB is the fastest-growing database ecosystem, with over 40 million downloads, 6,600+ customers, and over 1,000 technology and service partners. Learn more at www.mongodb.com.

Similar Jobs

8-10 yrs
Bengaluru, India
Skills:
Bdd, Performance Tuning, Maven, Testing Process, Apache Airflow, JUnit, Cicd, Selenium Webdriver With Java, Agile, Selenium, Python, Java, Appium, Anaplan, Etl Testing, TestNG, Jenkins, Istqb Certification, Advanced Sql, Airflow, EC2 based testing, Snowflakes, SDET concepts, Test automation scripts
8-12 yrs
Bengaluru, India
Skills:
Powershell, Bash, Jenkins, Docker, Terraform, Azure, Helm, Python, AWS, Azure DevOps, Github Actions, Gitlab CI, ArgoCD
8-10 yrs
Bengaluru, India
Skills:
Ml, SAP, Alteryx, Power Bi, Rpa, Data Analytics, Certification in internal controls, Internal control design, Accounting Standards, Local GAAP, Ai, Automation programs, ERP systems, It Audit, Compliance frameworks, Celonis, IFRS, Risk management
5-7 yrs
Bengaluru, India
Skills:
Sap Co, SAP Controlling, COPA, Product Costing, CCA, AI Automation, Actual Costing
8-10 yrs
Bengaluru, India
Skills:
Power Bi, Advanced Excel, Automation, Power Query, AI-driven logic