Senior Site Reliability Engineer, Platform Infrastructure

Senior Site Reliability Engineer, Platform Infrastructure

01 Oct 2026
Utah, Southjordan, 84009 Southjordan USA

Senior Site Reliability Engineer, Platform Infrastructure

We're looking for a Senior Site Reliability Engineer, Platform Infrastructure to take hands-on technical ownership of the architecture, reliability, and scalability of our entire AWS infrastructure. Reporting to the Engineering Manager, Platform Infrastructure & SRE, you'll set technical direction, review designs, and raise the bar for reliability engineering across a growing and globally distributed engineering organization.This is a senior individual contributor role. You'll work side by side with our onsite SRE team, Software Engineers, and other Lead Engineers to support seamless 24/7 reliability. It's ideal for an AI-forward engineer with a strong software engineering background who uses AI-assisted development tools to move faster, has a passion for infrastructure-as-code, and a proven track record of mentoring engineers to build highly reliable, scalable, and performant systems.Key ResponsibilitiesOwn the architecture, reliability, and scalability of critical AWS infrastructure, working hands-on across the full stack.Partner with Software Engineers and other Lead Engineers to shape the roadmap and technical strategy for Cricut's platform infrastructure.Take ownership of our AWS environment, driving best practices in security, cost management, and scalability.Champion and expand our "infrastructure-as-code" philosophy across the organization.Use AI-assisted development tools (e.g., Claude Code, GitHub Copilot) to accelerate delivery, applying prompt engineering and context management practices to get reliable results, and evaluate AI/ML infrastructure (e.g., model serving, vector databases, LLM tooling) as it becomes part of the platform.Oversee production monitoring, incident response, and blameless post-mortem processes to continuously improve system reliability.Develop and track Service Level Objectives (SLOs) and Service Level Indicators (SLIs) for critical production systems.Act as a key consultant for our feature-focused pillar and pod teams, ensuring they have the infrastructure resources and support required to deliver their projects successfully.Mentor software engineers who have an affinity for infrastructure, helping them grow their skills in reliability engineering.Collaborate closely with the onsite SRE team, sharing the on-call rotation to ensure seamless 24/7 reliability coverage.

Related jobs

Job Details

Jocancy Online Job Portal by jobSearchi.