onboardmeetings logo

Site Reliability Engineer

onboardmeetings • United States


No Relocation

Posted: September 1, 2026

Job Description

Cloud Operations Engineer III

Function:  Engineering 

Reports to:  Manager, Cloud Operations 

Location: Remote - United States 

Position Summary 

The Cloud Operations Engineer III is a senior member of the cloud operations team, responsible for the reliability, observability, performance, and operational security of our multi-product SaaS platform. This role owns our Datadog observability practice — instrumentation standards, dashboards, SLOs, monitors, and alert routing — and leads the migration off our legacy monitoring stack. 

It is an engineering role, not a ticket-queue role: the expectation is that recurring operational work gets replaced with code. The Cloud Operations Engineer III participates in on-call, incident response and is measured on fewer customer-impacting incidents, faster detection and recovery, and less manual work year over year. 
 
The ideal candidate is a proactive problem-solver who thrives in dynamic, evolving environments and works effectively across departments to address complex challenges.  They have experience partnering with cross-functional teams to understand and document requirements, then translating those needs into meaningful dashboards that improve service visibility (Observability) and support informed decision-making.  They are passionate about automation, process improvement, and eliminating unnecessary manual effort.  They confidently propose better approaches when opportunities for improvement arise. 

Key Responsibilities 

Observability and Datadog Ownership 

  • Own the Datadog platform across all products and environments, including agent lifecycle, instrumentation standards, unified service tagging, and per-cluster configuration. 
  • Instrument services for APM and distributed tracing, log collection, and synthetic monitoring; partner with engineering teams to close instrumentation gaps in both legacy and modern codebases. 
  • Build and maintain the dashboard, monitor, and SLO catalog; define SLIs and error budgets for critical user journeys and use them to drive prioritization with engineering and product. 
  • Design high-signal alerting: reduce noise and duplicate alerts, tune thresholds, and ensure every alert has an owner and a runbook. 

Automation and Toil Elimination 

  • Develop and maintain automation in PowerShell, Python, and Bash for provisioning, configuration, diagnostics, remediation, and reporting. 
  • Extend our infrastructure-as-code estate — Bicep modules, Kubernetes manifests, Helm releases, and Azure DevOps pipeline templates — so environments and regions are reproducible and drift-free. 
  • Convert manual runbooks into automated or self-service workflows: cluster upgrades, secret and certificate rotation, tenant provisioning, data retention purges, and access provisioning. 

Security, Documentation, and Mentorship 

  • Implement and maintain platform security controls and audit-ready operational evidence: managed identities, secret and key rotation, least-privilege access, and image and dependency scanning. 
  • Author and maintain runbooks, on-call guides, and architecture documentation, and provide technical leadership and mentorship to junior engineers on observability, automation, and incident response. 

Skills and Experience Needed 

  • Bachelor's degree in Computer Science, Information Technology, or a related field, or equivalent practical experience. 
  • 5-7 years of professional experience in cloud operations, site reliability, platform, or DevOps engineering for production SaaS systems. 
  • Demonstrated hands-on depth with a modern observability platform — Datadog strongly preferred — including APM and distributed tracing, log pipelines and indexing controls, dashboards, monitors, and SLOs. 
  • Strong scripting and automation ability in PowerShell, with the judgment to write tooling that other engineers can safely operate. 
  • Strong knowledge of containers, container orchestration, and the Kubernetes ecosystem, including autoscaling, cluster upgrades, and diagnosing pod-level failures. 
  • Production experience with Azure — Kubernetes Service, Azure SQL, Cosmos DB, Redis, Service Bus, Key Vault, and Entra ID — or equivalent depth in another major cloud. 
  • Experience with infrastructure-as-code and CI/CD pipeline authoring (Bicep or Terraform; Helm or Kustomize; Azure DevOps preferred). 
  • Proven incident response experience in a customer-facing production environment, including on-call participation and leading post-incident reviews. 
  • Experience operating multi-region, multi-tenant systems. 
  • Strong knowledge of platform security and operational best practices: secret and key rotation, least-privilege access, and vulnerability remediation. 
  • Excellent problem-solving and analytical abilities, with strong written communication for runbooks, incident updates, and technical proposals. 
  • Strong communication, and teamwork skills, including the ability to work effectively with legacy systems and their constraints. 
  • Nice to have: experience migrating from a legacy monitoring stack to a consolidated observability platform; relevant Azure, Kubernetes, or Datadog certifications. 

Competencies 

Accountability 

Adaptability 

AI Curiosity/Innovation 

Applied Learning 

Business Acumen 

Collaboration 

Customer Focus 

Dealing w/Ambiguity 

Decision Making 

Driving for Results 

Initiating Action 

Planning and Organizing 

Technical/Professional Knowledge