gitlab logo

Principal Site Reliability Engineer, Platform Engineering: Dedicated

gitlab • Remote, Canada; Remote, United Kingdom; Remote, United States


No Relocation

Posted: September 23, 2026

Job Description

An overview of this role

We’re looking for a Principal Engineer with deep expertise in Site Reliability, Backend, or Platform Engineering to help shape the next phase of GitLab Dedicated, our fully managed single-tenant SaaS offering. This highly influential technical leadership role will set direction for how we scale a growing fleet of isolated, customer-specific environments while maintaining the reliability, security, and compliance our customers depend on.

You’ll lead platform and operating-model transformation across resilience and failover, tenant orchestration, change management, self-service tooling, and platform integrations. You’ll also help align Dedicated with GitLab’s evolution toward more modular and cell-based architectures, strengthen service ownership, and establish scalable patterns that reduce operational complexity. As a Principal Engineer, you’ll influence across teams, guide complex technical decisions, mentor senior engineers, and help raise the technical maturity of the platform as we scale.

What you will do

  • Set technical direction for GitLab Dedicated, shaping architecture and platform strategy as we scale a growing fleet of isolated, single-tenant environments.
  • Lead platform transformations across resilience, failover, tenant orchestration, change management, self-service tooling, and platform integrations.
  • Drive scalable, modular architecture that aligns Dedicated with GitLab’s broader Cells strategy while preserving its security, isolation, and compliance requirements.
  • Strengthen service ownership and operational maturity, helping engineering teams build, operate, and improve the production systems they own.
  • Identify and address systemic reliability and scalability risks using production signals, incident patterns, and architectural insight.
  • Establish reusable platform patterns and automation that reduce operational toil and allow Dedicated to scale efficiently.
  • Lead complex technical decisions across teams, balancing reliability, security, cost, maintainability, and customer needs.
  • Advance engineering excellence across the organization through architectural leadership, mentorship, and influence with senior engineers and engineering leaders.

What you will bring

  • Deep expertise in Site Reliability, Platform, Infrastructure, or Backend Engineering, with experience designing and operating large-scale production systems.
  • Hands-on experience with cloud infrastructure, automation, observability, infrastructure as code, and modern production engineering practices.
  • Strong software engineering fundamentals, with experience building production systems or infrastructure tooling in languages such as Go, Ruby, Python, or similar.
  • Strong distributed systems and systems-design expertise, with sound judgment around reliability, failure isolation, scalability, and operational complexity.
  • A track record of technical leadership across multiple teams, setting direction and driving complex initiatives through influence.
  • Experience leading significant platform or infrastructure transformations, including modernization, modularization, or scaling systems through major growth.
  • Experience leading changes that improve how engineering teams own and operate production systems, strengthening reliability, operational readiness, and accountability at scale.
  • Exceptional technical communication and influence, with the ability to build alignment, mentor senior engineers, and guide complex architectural decisions.