Principal Cloud Engineer
Organization Background
- Technology Engineering and Operations owns the engineering, governance, reliability, and lifecycle management of the Cloud and OnPrem hosting environments.
- The environment is Azure-first and includes multi-subscription governance, hub-spoke networking, cloud platform services, Infrastructure as Code, monitoring, backup and recovery, security integration, FinOps, ServiceNow/CMDB, and managed operations delivered in partnership vendor partners.
- The Principal Cloud Engineer will serve as the senior technical lead and L4 escalation point, retaining platform engineering ownership while guiding vendor partners responsible for day-to-day operations.
What will you be responsible for?
- Own the technical engineering and operational integrity of the Azure platform.
- Define and maintain Azure Landing Zone standards covering Management Groups, subscriptions, Azure Policy, RBAC, tagging, resource locks, diagnostics, cost controls, and approved service patterns.
- Lead complex Azure operations and engineering across compute, storage, AKS, App Services, Key Vault, databases, backup, monitoring, and resource lifecycle management.
- Provide technical leadership for Azure networking, including hub-spoke connectivity, VNets, NSGs, UDRs, Private Endpoints, DNS, Azure Firewall, and integration with Palo Alto and Infoblox.
- Drive Infrastructure as Code using Terraform or equivalent platforms, including reusable modules, code reviews, testing, deployment pipelines, drift management, and secure state handling.
- Own platform reliability, observability, capacity, patching, backup, restore, BCP/DR readiness, and continual service improvement.
- Act as the L4 escalation point for major incidents, lead technical RCA and problem management, and ensure preventive actions are implemented.
- Guide and govern vendor resources performing L1/L2/L3 monitoring, fulfillment, troubleshooting, patching, backup, and standard changes.
- Partner with Cyber, IAM, Network, M365, EUC, Product Engineering, Compliance, Finance, and Service Management teams to deliver secure and supportable solutions.
- Lead technical reviews through TRB/architecture governance and ensure new workloads meet security, compliance, cost, operational readiness, and service ownership requirements.
- Drive cloud cost optimization, tagging compliance, consumption visibility, and FinOps recommendations.
- Mentor engineers, establish engineering practices, and communicate platform health, risks, decisions, and roadmap updates to senior stakeholders.
What would your day look like?
- Review platform health, availability, alerts, cost, policy compliance, capacity, backup, patching, and operational risks.
- Provide engineering guidance and approval for non-standard Azure requests, network changes, new services, and product onboarding designs.
- Develop or review Terraform modules, automation, pipelines, scripts, and reusable platform patterns.
- Lead complex troubleshooting across Azure infrastructure, AKS, private connectivity, DNS, identity dependencies, and application platform services.
- Participate in Major Incident bridges, approve RCAs, and drive permanent corrective and reliability improvements.
- Review vendor teams operational performance, changes, tickets, runbooks, evidence, SLA trends, and recurring issues.
- Work with Product teams to define monitoring, backup, recovery, RTO/RPO, service ownership, CMDB, and operational acceptance requirements.
- Present technical decisions, risks, investment needs, KPIs, and roadmap progress to leadership and governance forums.
Who are we looking for?
- 10-13 years of experience in cloud, infrastructure, platform engineering, or enterprise operations, including significant hands-on Microsoft Azure experience.
- Proven experience operating and engineering complex, production-grade Azure environments with strong ownership of reliability and service outcomes.
- Deep understanding of Azure Landing Zones, Management Groups, subscriptions, Azure Policy, RBAC/PIM integration, tagging, diagnostics, governance, and cost controls.
- Strong hands-on Azure networking knowledge covering VNets, peering, NSGs, UDRs, Private Endpoints, DNS, load balancing, firewalls, VPN, and hybrid connectivity.
- Hands-on expertise in Infrastructure as Code using Terraform; experience with Bicep, ARM, or other IaC platforms is beneficial.
- Experience with AKS/Kubernetes, Windows and Linux workloads, storage, databases, Key Vault, Azure Monitor, Log Analytics, backup, recovery, and patching.
- Strong understanding of ITSM processes including Incident, Major Incident, Change, Problem, Request, Knowledge, CMDB, and operational acceptance.
- Experience leading technical resolution of P1/P2 incidents, conducting RCA, and implementing preventive engineering actions.
- Working knowledge of cloud security, Zero Trust, encryption, vulnerability remediation, compliance controls, and frameworks such as HIPAA, HITRUST, ISO 27001, SOC 2, CIS, or MCSB.
- Ability to work in a dynamic, ambiguous, and rapidly scaling environment with multiple Product, Enterprise Service, Innovation, vendor, and leadership stakeholders.
- Strong executive communication, stakeholder management, technical documentation, mentoring, and influencing skills.
- Ability to balance hands-on engineering with platform leadership, governance, vendor oversight, and strategic planning.
Good to Have
- Microsoft Certified: Azure Solutions Architect Expert, Azure Administrator Associate, or Azure Network/Security certification.
- HashiCorp Terraform Associate or equivalent IaC/DevOps certification.
- Experience with GitHub, Azure DevOps, CI/CD, policy-as-code, automation using PowerShell/Python/Azure CLI, and self-service provisioning.
- Exposure to Palo Alto, GlobalProtect, Infoblox, Cloudflare, ManageEngine/Site24x7, ServiceNow, CrowdStrike, Rapid7, and Proofpoint integration patterns.
- Experience with healthcare, regulated workloads, customer-hosting platforms, or compliance/audit readiness.
Leadership and Behavioral Expectations
- Demonstrates strong ownership and serves as a trusted technical authority for the platform.
- Leads through influence and can align Product, Platform, Cyber, Operations, and vendor stakeholders.
- Communicates clearly with engineers, service owners, and executive leadership.
- Uses structured decision-making, measurable outcomes, and risk-based prioritization.
- Promotes automation, standardization, operational discipline, knowledge sharing, and continuous improvement.