The posting, in Scale AI's own words
archived Oct 2, 2026Scale is building reliable AI systems for the world’s most important decisions. Our products help leading enterprises, governments, and frontier AI labs build, deploy, and evaluate AI systems at scale. The Orchestration Platform is a core technology used across the company to run durable, reliable, and scalable workflows. It powers the complex, long-running processes behind Scale’s products - from data and model-evaluation pipelines to operational systems and customer-facing workflows. As a Senior Software Engineer, you will help build and evolve this foundational orchestration platform. You will design the primitives, services, and developer experience that enable teams across Scale to author, operate, observe, and safely scale distributed workflows. This is a high-leverage role for an engineer excited by distributed systems, reliability, and platform infrastructure used by many internal teams.
You will:
Lead the architecture, design, implementation, and operation of Scale’s core orchestration platform. Build durable workflow infrastructure using technologies such as Temporal, Cadence, Kubernetes, and cloud-native systems. Define platform primitives and APIs for scheduling, retries, state management, task execution, observability, and workflow lifecycle management. Partner with product, infrastructure, data, and application teams to understand workflow needs and turn them into reusable platform capabilities. Improve the reliability, scalability, security, and developer experience of services that run critical company workflows. Establish technical standards and best practices for distributed workflow development, deployment, testing, and incident response. Drive cross-functional technical decisions and communicate platform direction clearly to engineers and stakeholders.
Read the full posting ↓
Ideally, you have:
5+ years of full-time software engineering experience, with a focus on backend, infrastructure, and distributed systems. Experience building and operating production systems with strong requirements for reliability, availability, and scale. Deep familiarity with workflow orchestration platforms such as Temporal, Cadence, AWS Step Functions, Kubernetes, or similar systems. Deep familiarity with Kubernetes and containerized production environments, and familiarity with Terraform, . Strong knowledge of distributed-systems concepts, including asynchronous execution, retries, idempotency, fault tolerance, state management, and observability. A track record of leading technically complex projects from design through rollout and ongoing operation. Excellent communication skills and the ability to collaborate effectively with platform consumers and non-technical stakeholders.