Manager- Infrastructure Operations
Manager - Service Engineering_IT Operations
What will you be responsible for?
Lead end-to-end Incident Management and ensure timely restoration of services within SLA commitments.
Drive and facilitate Major Incident (P1) bridges, coordinating with technical teams, leadership, and customers.
Own Problem Management including RCA reviews, corrective actions, and prevention of recurring incidents.
Oversee infrastructure monitoring and observability across Network, Servers, Cloud, Applications, and Middleware platforms.
Manage and optimize monitoring tools such as SolarWinds and Datadog.
Identify monitoring gaps and drive observability enhancements to improve service visibility and proactive detection.
Lead Change Management activities, ensuring operational readiness and risk mitigation.
Own operational governance, service reviews, SLA/KPI compliance, and executive reporting.
Drive Capacity Management and Availability Management to ensure infrastructure scalability and service resilience.
Lead Continuous Service Improvement (CSI), automation, AIOps, alert rationalization, and operational excellence initiatives.
Manage vendor and partner escalations to ensure timely resolution and support alignment.
Analyze operational trends and performance metrics including MTTR, MTTD, incident volumes, alert effectiveness, and service availability.
Lead, mentor, and develop a team of 10–15 engineers while fostering a high-performance culture.
How would your day look like?
Review service health dashboards and monitoring alerts across infrastructure, cloud, and application environments.
Monitor operational KPIs, SLAs, and service performance metrics.
Lead daily operational reviews and collaborate with infrastructure, cloud, application, and vendor teams.
Drive major incident calls and provide executive and customer communications during service disruptions.
Review recurring incidents and facilitate Problem Management and RCA discussions.
Participate in CAB meetings and assess operational risks related to planned changes.
Analyze monitoring trends, alert patterns, and performance reports to identify improvement opportunities.
Drive automation opportunities, alert tuning, and monitoring optimization initiatives.
Conduct service review meetings with stakeholders and customers to discuss performance and improvement initiatives.
Coach team members, perform performance reviews, and support talent development.
Who are we looking for?
12+ years of experience in IT Operations, Infrastructure Services, Monitoring, or Service Management.
Strong experience in Incident, Problem, Change, and Major Incident Management.
Proven experience managing enterprise-scale production operations and critical business services.
Hands-on knowledge of infrastructure monitoring covering Network, Windows/Linux, Cloud (Azure/AWS/GCP), and Applications.
Strong experience with SolarWinds, Datadog, and enterprise monitoring/observability platforms.
Understanding of observability concepts such as Metrics, Logs, Traces, APM, and Event Management.
Experience with SLI/SLO/SLA management, alert optimization, and monitoring strategy.
Strong knowledge of ITIL Service Management best practices.
Proven experience managing teams of 10–15 members.
Experience driving operational transformation, automation, and continuous improvement programs.
Strong stakeholder, customer, executive, and vendor management skills.
Excellent communication, leadership, analytical, and decision-making capabilities.
Ability to lead high-pressure situations and manage customer-facing escalations effectively.