
Principal Site Reliability Engineer, Platform Engineering: Dedicated
gitlab • Remote, Canada; Remote, United Kingdom; Remote, United States
Posted: September 23, 2026
Job Description
An overview of this role
We’re looking for a Principal Engineer with deep expertise in Site Reliability, Backend, or Platform Engineering to help shape the next phase of GitLab Dedicated, our fully managed single-tenant SaaS offering. This highly influential technical leadership role will set direction for how we scale a growing fleet of isolated, customer-specific environments while maintaining the reliability, security, and compliance our customers depend on.
You’ll lead platform and operating-model transformation across resilience and failover, tenant orchestration, change management, self-service tooling, and platform integrations. You’ll also help align Dedicated with GitLab’s evolution toward more modular and cell-based architectures, strengthen service ownership, and establish scalable patterns that reduce operational complexity. As a Principal Engineer, you’ll influence across teams, guide complex technical decisions, mentor senior engineers, and help raise the technical maturity of the platform as we scale.
What you will do
- Set technical direction for GitLab Dedicated, shaping architecture and platform strategy as we scale a growing fleet of isolated, single-tenant environments.
- Lead platform transformations across resilience, failover, tenant orchestration, change management, self-service tooling, and platform integrations.
- Drive scalable, modular architecture that aligns Dedicated with GitLab’s broader Cells strategy while preserving its security, isolation, and compliance requirements.
- Strengthen service ownership and operational maturity, helping engineering teams build, operate, and improve the production systems they own.
- Identify and address systemic reliability and scalability risks using production signals, incident patterns, and architectural insight.
- Establish reusable platform patterns and automation that reduce operational toil and allow Dedicated to scale efficiently.
- Lead complex technical decisions across teams, balancing reliability, security, cost, maintainability, and customer needs.
- Advance engineering excellence across the organization through architectural leadership, mentorship, and influence with senior engineers and engineering leaders.
What you will bring
- Deep expertise in Site Reliability, Platform, Infrastructure, or Backend Engineering, with experience designing and operating large-scale production systems.
- Hands-on experience with cloud infrastructure, automation, observability, infrastructure as code, and modern production engineering practices.
- Strong software engineering fundamentals, with experience building production systems or infrastructure tooling in languages such as Go, Ruby, Python, or similar.
- Strong distributed systems and systems-design expertise, with sound judgment around reliability, failure isolation, scalability, and operational complexity.
- A track record of technical leadership across multiple teams, setting direction and driving complex initiatives through influence.
- Experience leading significant platform or infrastructure transformations, including modernization, modularization, or scaling systems through major growth.
- Experience leading changes that improve how engineering teams own and operate production systems, strengthening reliability, operational readiness, and accountability at scale.
- Exceptional technical communication and influence, with the ability to build alignment, mentor senior engineers, and guide complex architectural decisions.