Tyk Technologies logo

Site Reliability Engineer - AMER

Tyk Technologies • Canada • Colombia


No Relocation

Posted: August 28, 2026

Job Description

Who are Tyk, and what do we do? 

The Tyk API Management platform is helping to drive the connected world and power new products and services. We’re changing the way that organisations connect any number of their systems and services.Whether internal, external, public or highly encrypted systems, Tyk helps businesses drive value across the retail, finance, telecoms, healthcare, or media industries (to name just a few!)

If you’ve banked online, used an app to check the news, or perhaps even driven a connected car, API’s, and by extension, Tyk, make that possible. Founded in 2015 with offices in London – UK, London – Ontario, Atlanta and Singapore, we have many thousands of users of our B2B platform across the globe. Brands using Tyk range from Lotte, Bell, T Mobile, to RBS, Capital One and Vinci. We have a varied user base hailing from every continent – even Antarctica.

Our Mission

Tyk is on a mission to connect every system in the world. We’ve started by building an API Management platform.

Total flexibility, default remote, radical responsibility

We offer unlimited paid holidays and remote working from anywhere in the world, for everyone, Why? Tyk was founded on the principle of offering flexibility and autonomy to our employees, we believe this allows our employees to achieve their best results. It also means we can build the best possible team, location and working hours are no barrier. 

If this sounds like an environment that you believe could work for you then read on to find out more.

The role:

Tyk Cloud is our managed API management platform, running on multi-region Kubernetes at scale for customers around the world.

We're looking for an SRE who's as comfortable in the code as in the infrastructure. You'll spend most of your time improving and automating the platform, and when you're on call, you'll handle incidents independently. You'll join a small, centralised SRE team that works closely with our product teams.

You don't need to have done everything below. We care most about how you reason through problems and how quickly you learn. You'll have three to four months to get up to speed, shadowing first, before you go on call.

What you'll do:

Most of your time

  • Deliver the team's planned work each quarter, such as optimising the platform, building self-serve tooling for other teams, and rearchitecting parts of the platform as it grows.
  • Help expand Tyk Cloud across regions and clouds, and bring down what it costs to run.
  • Automate operations in Go, including building and maintaining our custom Kubernetes operators.
  • Run the platform's services and databases, including MongoDB and Redis.
  • Improve our observability: find the metrics that matter, and build the dashboards and alerts to act on them.
  • Keep runbooks and documentation current, and support security work such as SOC 2 audits.

When you're on call

You'll be doing one week in three initially (one in four as we grow), Monday-Friday, on a 12-hour shift with secondary backup support. Rotas: 14:00–02:00 UTC

  • Be first line for platform alerts and incidents: restore service, escalate or help fix product bugs, and lead post-incident reviews.
  • Act as second line for our Customer Success team, on requests that come directly from customers.
  • Handle ad hoc requests from other teams across the organisation regarding Tyk Cloud.
Who are Tyk, and what do we do? The Tyk API Management platform is helping to drive the connected world and power new products and services. We’re changing the way that organisations connect any number of their systems and services.Whether int...

What you'll need:

  • 3+ years in SRE, platform or infrastructure roles, across more than one company or production platform.
  • Experience owning on-call and leading incidents yourself.
  • Hands-on experience running production Kubernetes at scale, ideally EKS: operating, upgrading and debugging large, multi-tenant clusters.
  • Experience designing and operating infrastructure on AWS, with Terraform or similar.
  • The ability to write, test and ship Go tooling or services.
  • Experience with Prometheus and Grafana, and with logging systems.
  • Solid Linux and networking fundamentals (DNS, TCP/IP, HTTP, TLS, load balancing).
  • Clear communication across time zones and teams.

Our stack

EKS, Terraform/Terragrunt, Helm, GitHub Actions, Argo CD, MongoDB, Redis, Prometheus and Grafana.