The posting, in Nortal's own words
archived Oct 7, 2026Senior Site Reliability Engineer (SRE) - Talent Connection Real ownership from day one, not check-ins and micromanagement. That's what the Senior Site Reliability Engineer (SRE) role at Nortal looks like. What working with us looks like Remote work with a LATAM team. Coffee breaks, tech talks, and games keep it human, even at a distance. No micromanagement. We hire for autonomy and expect you to use it. Support beyond the job. Our People Care team helps with time off, wellness, and anything in between. Our accounts team handles client relationships, so you focus on the work. What you get Competitive USD salary. Full remote work, with coworking spaces across LATAM if you want to meet the team in person. Paid time off under your country's rules, at full salary. National holidays. Sick leave, no stress attached. A yearly refundable credit. Spend it on anything related to your health and well-being. A day off for your birthday. About this search Great talent doesn't wait for job postings, so we don't either. We're building a network of skilled professionals for roles that come up regularly with our clients. Join our Future Talent network and you'll be one of the first people we reach when the right opening appears. We make around 160 hires a year, so it happens often.
The role
As a Senior Site Reliability Engineer , I help modernize and scale production platforms, most recently a customer-data caching platform that serves 20–25+ consuming applications at about 3,000 requests per second. I work across SRE, DevOps, engineering and architecture to improve resiliency, automation, observability, deployment practices and overall platform reliability. I work independently, set technical direction and do well in small engineering teams.
Read the full posting ↓
Your day-to-day:
Design, build and operate highly available, fault-tolerant distributed systems and caching platforms. Evaluate the existing architecture and identify opportunities to scale the platform toward 6x current traffic while maintaining performance and resiliency. Improve caching strategies, including TTLs, refresh patterns, cache placement, performance and downstream dependency management. Help evolve the platform toward active-active resiliency and validate failure, recovery and capacity scenarios. Design and implement automated pipelines, including rolling, blue/green or canary deployment strategies, automated checks and quality gates. Establish infrastructure, configuration, secrets and application deployment practices using an everything-as-code approach. Build performance and load-testing capabilities and establish meaningful performance gates for releases. Develop monitoring, alerting and observability for cache latency, throughput, availability, errors and other key reliability indicators. Define and improve SLIs, SLOs, error budgets and operational KPIs where appropriate. Automate operational tasks and reduce manual intervention through scripting, tooling and infrastructure automation. Troubleshoot production issues, perform root-cause analysis and drive reliability improvements through blameless post-incident reviews. Provide technical guidance and establish engineering guardrails for SRE and development teams. Collaborate with Java/Spring Boot engineers, architects, client FTEs and other technical teams to deliver platform improvements.