Lead SRE & Support Engineer

About Providence

Providence, one of the US’s largest not-for-profit healthcare systems, is committed to high quality, compassionate healthcare for all. Driven by the belief that health is a human right and the vision, ‘Health for a better world’, Providence and its 121,000 caregivers strive to provide everyone access to affordable quality care and services.

Providence has a network of 51 hospitals, 1,000+ care clinics, senior services, supportive housing, and other health and educational services in the US.

Providence India is bringing to fruition the transformational shift of the healthcare ecosystem to Health 2.0. The India center will have focused efforts around healthcare technology and innovation, and play a vital role in driving digital transformation of health systems for improved patient outcomes and experiences, caregiver efficiency, and running the business of Providence at scale.


Why Us?

  • Best In-class Benefits
  • Inclusive Leadership
  • Reimagining Healthcare
  • Competitive Pay
  • Supportive Reporting Relation

About the Team

The Cybersecurity Product Engineering Team designs, develops, and operates strategic platforms that support security posture visibility, operational intelligence, enterprise risk management, and secure engineering outcomes.

The team works at the intersection of cybersecurity, cloud engineering, data engineering, site reliability engineering, and artificial intelligence. Reliability and operational excellence are central to how we deliver secure, resilient, scalable services to global stakeholders.

About the Role

We are seeking a highly motivated Lead Site Reliability Engineer (SRE) Support Engineer (3P) to join the Cybersecurity Product (Software) Engineering Team. This role is responsible for maintaining the reliability, availability, and operational performance of critical cloud-native services running in Microsoft Azure currently, with potential to extend to AWS, GCP Cloud Service Providers.

This is a full-time night-shift role (9:00 PM to 6:00 AM IST) supporting US business hours. While office presence requirements can be discussed, candidates must be based in Hyderabad and available to work from the Hyderabad location as needed.
 

The successful candidate will partner with software engineers, platform engineers, data engineers, and cybersecurity stakeholders to support production services, manage incidents, improve observability, automate repetitive work, and drive continuous operational improvement. The ideal candidate combines strong troubleshooting skills with an ownership mindset and practical experience using AI-assisted engineering tools.

TECHNICAL JOB DESCRIPTION 

  • Own and manage DevOps pipelines for the platform across application code, data, RAG, and ML workloads; build and enhance pipelines as needed to improve reliability, automation, and deployment efficiency.
  • Provide hands-on Production Support for product with clear, timely incident updates that summarize impact, investigation status, mitigation, and next actions.
  • Drive incidents through containment, recovery, validation, closure, and handoff when cross-shift follow-up is required.
  • Contribute to root cause analysis and track corrective and preventive actions to completion.
  • Participate in a 24x7 support and on-call model as required by business and service needs.

Site Reliability Engineering 

  • Improve service reliability, resilience, scalability, and operational effectiveness through engineering-led support practices.
  • Identify recurring failure patterns, reliability risks, and opportunities to reduce manual operational toil.
  • Contribute to Service Level Indicators (SLIs), Service Level Objectives (SLOs), availability targets, and service health reviews.
  • Partner with engineering teams on operational readiness, recovery procedures, capacity considerations, and platform hardening.
  • Use incident and service-health learnings to recommend durable technical and process improvements.

Azure-Native Observability

  • Use Azure Monitor, Application Insights, Log Analytics workspaces, Azure Alerts, and Azure Service Health to monitor and diagnose services.
  • Investigate system behavior through metrics, logs, distributed traces, dependency maps, queries, and correlated events.
  • Build and improve actionable dashboards, alert rules, health views, and operational reports.
  • Tune monitoring and alerting to improve signal quality, reduce alert fatigue, and shorten detection and recovery cycles.
  • Contribute to observability standards and consistent telemetry practices across supported services.

Data Platform Operations

  • Support the operational reliability of cloud-based data services and analytics workloads.
  • Troubleshoot Snowflake connectivity, access, workload, query performance, and operational issues within the scope of the support role.
  • Investigate data ingestion, transformation, reporting, and pipeline failures in collaboration with data engineering teams.
  • Validate data availability and operational recovery after incidents, releases, or maintenance activities.
  • Use SQL to investigate data issues, validate processing outcomes, and support incident diagnosis.

Automation and Operational Excellence

  • Develop and maintain Python, PowerShell, or shell-based automation for health checks, evidence gathering, diagnostics, and routine support activities.
  • Create reusable utilities and workflow improvements that reduce manual effort and improve response consistency.
  • Identify opportunities for safe self-service and self-healing capabilities with appropriate controls and auditability.
  • Improve support processes through standardization, measurable outcomes, documentation, and continual learning.
  • Contribute operational feedback to backlog prioritization and engineering improvement plans.

AI-Assisted Operations

  • Use approved enterprise AI assistants and copilots to accelerate troubleshooting, knowledge retrieval, scripting, documentation, and incident summarization.
  • Apply effective prompt engineering techniques to produce clear, context-aware operational outputs and refine results through validation.
  • Use AI assistance to summarize logs and telemetry, detect patterns, propose hypotheses, and organize evidence while independently validating conclusions.
  • Use AI-assisted coding tools to draft or improve scripts, tests, queries, and automation with appropriate review and secure coding practices.
  • Create and improve runbooks, knowledge articles, incident timelines, and post-incident documentation using AI-assisted workflows.
  • Understand foundational AIOps concepts such as anomaly detection, event correlation, alert enrichment, and intelligent triage.
  • Protect confidential, personal, security-sensitive, and regulated information when using AI tools, following organizational data-handling requirements.
  • Recognize AI limitations, including inaccurate or incomplete outputs, and apply human review before operational use.

Required Qualifications

  • Bachelor's degree in Computer Science, Information Technology, Engineering, Cybersecurity, or a related technical discipline, or equivalent practical experience.
  • 6-9 years of experience in Site Reliability Engineering, Production Support, Application Support, Platform Operations, or Cloud Operations.
  • Hands-on experience supporting enterprise-scale production applications and services in Microsoft Azure. With experience/ exposure to other Cloud Service Providers (AWS, GCP).
  • Experience with incident response, service restoration, root cause analysis, change support, and production release validation.
  • Working knowledge of Azure-native monitoring and observability services.
  • Hands-on SQL and Snowflake support or troubleshooting experience.
  • Ability to automate operational work using Python, PowerShell, or shell scripting.
  • Working knowledge of REST APIs, JSON, authentication and authorization concepts, and service connectivity troubleshooting.
  • Experience with Azure DevOps, Git, and CI/CD operational support.
  • Working knowledge of Linux and Windows environments.
  • Strong analytical, problem-solving, written communication, and stakeholder coordination skills.
  • Ability and willingness to work a permanent US shift and participate in a 24x7 or on-call support model as required.

Preferred Qualifications

  • Experience supporting cybersecurity, risk management, governance, compliance, security analytics, or enterprise security platforms.
  • Exposure to cloud-native, microservices-based, and distributed application architectures.
  • Familiarity with containerized environments and Kubernetes or Azure Kubernetes Service concepts.
  • Practical understanding of SRE principles, reliability metrics, SLI/SLO practices, and toil reduction.
  • Experience working in Agile, DevOps, or product engineering delivery models.
  • Experience using approved enterprise AI assistants such as Microsoft Copilot, GitHub Copilot, or comparable tools in engineering workflows.
  • Familiarity with responsible AI usage, secure prompt practices, and validation of AI-generated technical output.
  • Microsoft Azure or related cloud certifications.

Core Technical Capabilities

Capability

Expected Knowledge and Experience

Cloud Platform

Microsoft Azure (M), AWS (O), GCP (O) cloud-native application and platform support 

Observability

Azure Monitor; Application Insights; Log Analytics; Azure Alerts; Azure Service Health

Reliability Operations

Incident and problem management; RCA; change support; release validation; SLI/SLO awareness

Data Technologies

Snowflake; SQL; data investigation; query and pipeline troubleshooting

Automation

Python; PowerShell; shell scripting; operational tooling and runbooks

Application Support

REST APIs; JSON; authentication and authorization; service connectivity

DevOps

Azure DevOps; Git; CI/CD operational support

Operating Systems

Linux; Windows Server fundamentals

AI for Operations

Copilots; prompt engineering; AI-assisted diagnostics, coding, documentation, and AIOps awareness

 

What Success Looks Like

  • Maintain product and platform uptime by ensuring reliable operations, rapid incident resolution, resilient infrastructure, and highly automated deployment and monitoring processes.
  • Independently manages production incidents and operational escalations during assigned support hours.
  • Meets agreed response, restoration, communication, and resolution expectations while minimizing business impact.

CYBERSECURITY PRODUCT ENGINEERING | TECHNICAL JOB DESCRIPTION

  • Improves monitoring coverage, alert quality, operational visibility, and service reliability.
  • Reduces recurring incidents and manual effort through automation, preventive actions, and effective knowledge management.
  • Produces clear incident records, actionable root cause analyses, and well-owned follow-up actions.
  • Uses AI-assisted tools responsibly to improve productivity, troubleshooting quality, and documentation without compromising security or human oversight.
  • Builds trusted working relationships across Cybersecurity, Product Engineering, Cloud, Platform, and Data teams.

Providence’s vision to create ‘Health for a Better World’ aids us to provide a fair and equitable workplace for all in our employment, whether temporary, part-time or full time, and to promote individuality and diversity of thought and background, and acknowledge its role in the organization’s success. This makes us committed towards equal employment opportunities, regardless of race, religion or belief, color, ancestry, disability, marital status, gender, sexual orientation, age, nationality, ethnic origin, pregnancy, or related needs, mental or sensory disability, HIV Status, or any other category protected by applicable law. In furtherance to our mission in building a more inclusive and equitable environment, we shall, from time to time, undertake programs to assist, uplift and empower underrepresented groups including but not limited to Women, PWD (Persons with Disabilities), LGTBQ+ (Lesbian, Gay, Transgender, Bisexual or Queer), Veterans and others. We strive to address all forms of discrimination or harassment and provide a safe and confidential process to report any misconduct.

Contact our Integrity hotline also, read our Code of Conduct.