The posting, in Toast's own words
archived Oct 2, 2026Toast creates technology to help restaurants and local businesses succeed in a digital world, helping business owners operate, increase sales, engage customers, and keep employees happy.
About the Role
As a Senior Infrastructure Engineer on the Compute & Storage Platform team in Toast, you will design, build, and operate the foundational distributed systems powering the Toast platform as the company expands into further international markets and broadens its product offerings to encompass retail shopping. A day in the life (Responsibilities) You will bridge the gap between core infrastructure and cutting-edge product development—building low-latency state and storage systems, scalable compute clusters, and high-throughput streaming pipelines. Your work will directly enable Toast product delivery across our Product Development Life Cycle (PDLC). You will bridge the gap between core infrastructure and cutting-edge product development—building low-latency state and storage systems, scalable compute clusters, and high-throughput streaming pipelines. Your work will directly enable Toast product delivery across our Product Development Life Cycle (PDLC). What you'll need to thrive (Requirements)
Read the full posting ↓
Core Infrastructure & Automation
Cloud & Infrastructure as Code: Hands-on experience building and managing scalable cloud infrastructure on AWS using Infrastructure as Code (Terraform, CloudFormation, etc.). Software Development: Fluency in Python , Go , or a similar language for infrastructure automation, tooling, or backend development. Storage & Compute Operations Databases: Deep experience operating and tuning enterprise databases at scale. This should include at least one variety of relational database (such as Postgres, Oracle, or MySQL) and at least one type of nonrelational database (such as MongoDB, DynamoDB, Opensearch). Containerization: Experience working with containerized environments and orchestration tools (e.g., Docker, ECS, Kubernetes).
Systems Reliability & Observability
Solid understanding of Site Reliability Engineering (SRE) principles, including monitoring, alerting, health checks, and performance troubleshooting in high-traffic environments. Toast’s commercial products run on a stack that ranges from guest and restaurant-facing Android tablets to backend microservices, guest- and restaurant-facing web apps, payment systems, event buses, and more. The Infrastructure Engineering group manages Toast's infrastructure via the AWS API, using configuration management tools like Terraform, Ansible, and internally developed clients using Python and Go. We rely heavily on Datadog and Splunk for monitoring and observability and on FireHydrant for alerting and incident response. Infrastructure Engineering provides the products and services to support, deploy, and monitor our product suite, which uses a microservice architecture written using JVM and Node.js apps. AWS products are also in heavy usage, ranging from S3 to RDS, DynamoDB, ECS, Lambda, and many others. We have our own platform for user management, service discovery/elevations, and load balancing. We use Apache Spark for large scale data workloads including query and batch processing. The main Toast POS application is an Android application written in Java and Kotlin. For data synchronization between tablets and our cloud platform we operate RabbitMQ and Apache Pulsar clusters as well as direct tablet communication to the applications' and APIs. ** This is a hybrid role requiring in-office presence two days per week *** #LI-HYBRID #BI-Hybrid