The posting, in PagerDuty's own words
archived Sep 9, 2026PagerDuty, Inc. (NYSE: PD) is the global leader in AI-first digital operations. By automatically detecting, diagnosing, and remediating issues, the PagerDuty Platform orchestrates AI agents and automated workflows with context from over 750 integrations. Trusted by approximately two-thirds of the Fortune 100 and nearly half of the Fortune 500, PagerDuty is the industry standard for organizations scaling resilient, autonomous operations. Notable customers include Chipotle, Cloudflare, Docusign, Fox, Nvidia, Salesforce, Spotify, Zoom and more. We are growing rapidly and hiring top talent with leading AI skills across engineering, sales, product, marketing, and beyond as we build the leading digital operations platform. As a Site Reliability Engineer I on the Core Infrastructure team in our Atlanta office, you'll help build and operate the foundational infrastructure that powers PagerDuty's real-time digital operations platform. Our systems support millions of events and alerts daily, enabling customers to detect, respond to, and resolve incidents quickly and reliably. You'll work at the intersection of platform evolution and operational excellence, building and evolving foundational network, compute, and ingress infrastructure while scaling and hardening existing systems. Your work will directly impact the reliability, scalability, and security of the services our customers rely on to keep their businesses running as PagerDuty continues to grow across products, regions, and customer use cases. Key Responsibilities ● Support and improve foundational infrastructure, including networking, compute platforms, Kubernetes clusters, and ingress/traffic management systems. ● Contribute to the reliability and scalability of PagerDuty's core platform by hardening existing systems and supporting the rollout of new infrastructure capabilities. ● Participate in agile rituals (standups, planning, retros) and communicate progress/risks early ● You stay current on technical trends to suggest innovative tools and approaches to interesting problems ● Monitor system health using metrics, logs, and alerts, and participate in 24/7 on-call rotations to help detect, respond to, and resolve incidents. Basic Qualifications ● 0 to 1+ years of experience in Site Reliability Engineering, DevOps, or Platform Engineering roles ● Hands-on experience operating Linux-based systems in production environments ● Working knowledge of networking fundamentals, such as load balancing, DNS, TLS, and ingress traffic flow ● Experience with container orchestration (e.g., EKS, Kubernetes) ● Experience working on cloud-native infrastructure (e.g., AWS, GCP, Azure), including networking and compute concepts ● Proficiency in at least one programming language (e.g., Python, Ruby, Go, etc.) ● Experience with Infrastructure as Code (e.g., , CloudFormation) Preferred Qualifications ● Experience with AWS cloud networking concepts such as VPCs, subnets, routing, security groups, and load balancers ● Experience operating or contributing to production Kubernetes platforms (e.g., EKS), including cluster upgrades, networking, or ingress configuration ● Experience with monitoring, observability, and logging platforms (e.g., DataDog, New Relic, SumoLogic, Splunk, Prometheus, Grafana) ● Familiarity with service meshes, ingress controllers, or API gateways (e.g., Envoy, Istio, NGINX) Salary Range: $98,000 to $148,500 Hesitant to apply? We encourage you to submit your resume even if you don't meet every requirement. We value potential and consider each candidate's full professional story. Whether you're exploring a career change or taking your next step, we look forward to reviewing your application. If this just isn’t the right role or time - sign up for job alerts ! Where we work PagerDuty operates a hybrid work model with offices in 8 major cities: Atlanta, Lisbon, London, San Francisco, Santiago, Sydney, Tokyo, and Toronto. While we offer flexibility within our established locations, we cannot employ candidates residing in: