Lead Site Reliability Engineer, Platforms

Zoom

San Jose, CA · Onsite · Full Time

Posted

Job description

What You Can Expect As a Lead Staff Site Reliability Engineer, you will be one of the technical leads for our DevOps Platforms organization. This group is responsible for DevOps Platforms including cloud infrastructure, physical data center orchestration, critical security services, and our Zoom for Government (ZfG) environment. You will be an uber tech lead working across a broad area, defining projects and guiding work across various teams. Your scope of work is wide and you will have the opportunity to improve our datacenter kubernetes infrastructure, our cloud infrastructure, our security posture, and our operation of ZfG environments. Broadly speaking, you are an exemplary SRE and you will guide all of our teams toward SRE best practices (automation, monitoring, infrastructure as code, etc). About the Team The DevOps Platforms organization owns the full infrastructure stack: cloud infrastructure on AWS and OCI, physical data center orchestration, critical security services (identity, authentication, authorization), and Zoom's FedRAMP-rated federal environment, Zoom for Government (ZfG). The team is currently working on one of the most technically interesting work in the org, hardening our security posture, and building the automation and reliability systems that underpin Zoom's global services. If you want broad visibility, real cross-team influence, and the chance to define how infrastructure gets built, this is the seat. Responsibilities Design and scale DevOps platform services including Kubernetes infrastructure, cloud systems, and compliance-ready environments Define technical roadmaps and architectural direction for infrastructure automation and security Partner with service teams to understand platform needs and deliver solutions that improve reliability and efficiency Establish and advocate for SRE best practices including infrastructure as code, monitoring, and incident management Mentor team members through design, implementation, and production deployment of complex systems What We're Looking For Bring 8+ years of SRE or DevOps experience building and operating production infrastructure at scale Code proficiently in at least one programming language beyond scripting (e.g., Python, Go, Java) Deploy and manage CI/CD pipelines using tools like Git, Jenkins, Argo CD, or JFrog Operate cloud infrastructure on AWS, OCI, or similar platforms using T…

Apply for this job