invisibletech logo

Site Reliability Engineer (Contract, Rotation-Based)

invisibletech • Estonia - Remote; India - Remote; South Africa - Remote


No Relocation

Posted: September 2, 2026

Job Description

About The Role

We're looking for Site Reliability Engineers to join a 24/7 incident response rotation supporting a production platform used by one of our key clients. You'll be a first responder for critical incidents — triaging and stabilizing issues at the infrastructure level, working from logs to place a failure to its component, and escalating to the platform team when an issue goes beyond infrastructure-level diagnosis.

We're hiring across a level range. At minimum, you'll own first response: fast, calm triage under time pressure, and good judgment about when to escalate versus resolve. For candidates operating at a more senior level, the role grows to include hardening the systems you're responding to, driving scaling and capacity work, and proposing structural improvements — so the same incidents happen less often over time.

What You’ll Do

  • Serve as first responder for production incidents, triaging and stabilizing at the infrastructure level within defined response SLAs (P1/P2 severity)
  • Diagnose issues primarily from system logs — Kubernetes, RabbitMQ, and Postgres — to place a failure to the right component before escalating
  • Distinguish infrastructure-level failures from application/business-logic failures — for infrastructure-level issues, identify the fix and submit the change yourself; escalate application/business-logic issues to the platform team with clear context
  • Participate in an on-call rotation, including off-hours coverage
  • Communicate incident status clearly to stakeholders during active incidents, and hand off cleanly to the team that owns resolution

What We Need

  • Solid hands-on experience with Kubernetes, RabbitMQ, and PostgreSQL in an enterprise setting, ideally within Financial Services
  • Strong working knowledge of Azure; familiarity with GCP or AWS is a plus
  • Comfort diagnosing an unfamiliar system primarily from its logs rather than its source — this matters more than deep expertise in any one tool
  • Experience troubleshooting production systems under time pressure, with sound judgment about severity and escalation
  • Clear, calm communication during live incidents

Engagement Type & Schedule

This is a contractor role structured around a coverage rotation rather than a standard full-time schedule. The commitment breaks down as:

  • 10+ hours per week during regular business hours
  • Rotational weekend coverage: 8 hours on Saturday and 8 hours on Sunday, every other weekend
  • ~72+ hours per month total, combining weekday and rotational weekend coverage

*Please note: this role is an hourly, contract-based position and is not eligible for bonus or equity compensation*