
Site Reliability Engineer (Contract, Rotation-Based)
invisibletech • Estonia - Remote; India - Remote; South Africa - Remote
Posted: September 2, 2026
Job Description
About The Role
We're looking for Site Reliability Engineers to join a 24/7 incident response rotation supporting a production platform used by one of our key clients. You'll be a first responder for critical incidents — triaging and stabilizing issues at the infrastructure level, working from logs to place a failure to its component, and escalating to the platform team when an issue goes beyond infrastructure-level diagnosis.
We're hiring across a level range. At minimum, you'll own first response: fast, calm triage under time pressure, and good judgment about when to escalate versus resolve. For candidates operating at a more senior level, the role grows to include hardening the systems you're responding to, driving scaling and capacity work, and proposing structural improvements — so the same incidents happen less often over time.
What You’ll Do
- Serve as first responder for production incidents, triaging and stabilizing at the infrastructure level within defined response SLAs (P1/P2 severity)
- Diagnose issues primarily from system logs — Kubernetes, RabbitMQ, and Postgres — to place a failure to the right component before escalating
- Distinguish infrastructure-level failures from application/business-logic failures — for infrastructure-level issues, identify the fix and submit the change yourself; escalate application/business-logic issues to the platform team with clear context
- Participate in an on-call rotation, including off-hours coverage
- Communicate incident status clearly to stakeholders during active incidents, and hand off cleanly to the team that owns resolution
What We Need
- Solid hands-on experience with Kubernetes, RabbitMQ, and PostgreSQL in an enterprise setting, ideally within Financial Services
- Strong working knowledge of Azure; familiarity with GCP or AWS is a plus
- Comfort diagnosing an unfamiliar system primarily from its logs rather than its source — this matters more than deep expertise in any one tool
- Experience troubleshooting production systems under time pressure, with sound judgment about severity and escalation
- Clear, calm communication during live incidents
Engagement Type & Schedule
This is a contractor role structured around a coverage rotation rather than a standard full-time schedule. The commitment breaks down as:
- 10+ hours per week during regular business hours
- Rotational weekend coverage: 8 hours on Saturday and 8 hours on Sunday, every other weekend
- ~72+ hours per month total, combining weekday and rotational weekend coverage
*Please note: this role is an hourly, contract-based position and is not eligible for bonus or equity compensation*