
Site Reliability Engineer (Baremetal)
npv labs • Worldwide
Posted: August 18, 2026
Job Description
tl;dr: SRE; post-acquisition profitable health/adtech; baremetal 10M+ rps k8s production; kubernetes internals, expansion and tuning plus some baremetal infra experience are required; remote, up to 150k USD base, we can talk higher figures and EU/UK/US employment
PulsePoint sits at the intersection of healthcare and adtech. We help brands and agencies interpret the hard-to-read signals across the health journey and unify these digital determinants of health with real-world data to produce the most dimensional view of the customer. We are 300+ and growing, post-acquisition business and one of the leading players in the US healthcare ad market.
We're looking for a Site Reliability Engineer to join PulsePoint and ensure the reliability of the platform behind large-scale data and advertising systems that power business-critical services used every day across the company. The Platform Engineering team owns the foundation that lets engineering teams move quickly and safely – infrastructure supporting Kubernetes workloads, data platforms, developer tooling, observability, and production operations at scale. Unlike many cloud-only environments, the team owns the full lifecycle, from bare-metal hardware and networking to Kubernetes, observability, and developer experience, across multi-petabyte data systems and business-critical services used by multiple engineering organizations.
You will
Help design, build, and operate the Kubernetes platform used across PulsePoint – architecture and lifecycle management.
Own reliability, observability, and incident response across platform services.
Build infrastructure automation and GitOps workflows to reduce operational toil.
Work on networking, service connectivity, and platform security.
Improve developer experience through self-service, reliable platform capabilities.
Operate large-scale distributed systems running on bare-metal infrastructure.
Stack
Kubernetes, ArgoCD, Puppet, Terraform, OpenTelemetry, Prometheus, Alertmanager, Kafka, Redis, Ceph. Bare-metal and hybrid cloud/on-prem infrastructure. Experience with every technology is not required.
Requirements
Experience operating production infrastructure at meaningful scale.
Understanding of how distributed systems fail and recover.
A preference for automation over repetitive operational work, and for simplifying systems rather than adding complexity.
Ownership extending beyond the boundaries of a single component.
Willingness to work 9am–6pm ET US hours (fully remote).
We offer
Remote work, high engineering bar and comfortable culture.
Flat hierarchy with easy access to business, product, and operations.
Enormous scale (10M+ peak rps) with real growth potential.
Ownership and direct impact, you have room to shift focus as your interests evolve.
Up to 150k USD salary, higher figures and EU/UK/US employment are negotiable.
Additional Content
tl;dr: SRE; post-acquisition profitable health/adtech; baremetal 10M+ rps k8s production; kubernetes internals, expansion and tuning plus some baremetal infra experience are required; remote, up to 150k USD base, we can talk higher figures and EU/UK/US employment
PulsePoint sits at the intersection of healthcare and adtech. We help brands and agencies interpret the hard-to-read signals across the health journey and unify these digital determinants of health with real-world data to produce the most dimensional view of the customer. We are 300+ and growing, post-acquisition business and one of the leading players in the US healthcare ad market.
We're looking for a Site Reliability Engineer to join PulsePoint and ensure the reliability of the platform behind large-scale data and advertising systems that power business-critical services used every day across the company. The Platform Engineering team owns the foundation that lets engineering teams move quickly and safely – infrastructure supporting Kubernetes workloads, data platforms, developer tooling, observability, and production operations at scale. Unlike many cloud-only environments, the team owns the full lifecycle, from bare-metal hardware and networking to Kubernetes, observability, and developer experience, across multi-petabyte data systems and business-critical services used by multiple engineering organizations.
You will
Help design, build, and operate the Kubernetes platform used across PulsePoint – architecture and lifecycle management.
Own reliability, observability, and incident response across platform services.
Build infrastructure automation and GitOps workflows to reduce operational toil.
Work on networking, service connectivity, and platform security.
Improve developer experience through self-service, reliable platform capabilities.
Operate large-scale distributed systems running on bare-metal infrastructure.
Stack
Kubernetes, ArgoCD, Puppet, Terraform, OpenTelemetry, Prometheus, Alertmanager, Kafka, Redis, Ceph. Bare-metal and hybrid cloud/on-prem infrastructure. Experience with every technology is not required.
Requirements
Experience operating production infrastructure at meaningful scale.
Understanding of how distributed systems fail and recover.
A preference for automation over repetitive operational work, and for simplifying systems rather than adding complexity.
Ownership extending beyond the boundaries of a single component.
Willingness to work 9am–6pm ET US hours (fully remote).
We offer
Remote work, high engineering bar and comfortable culture.
Flat hierarchy with easy access to business, product, and operations.
Enormous scale (10M+ peak rps) with real growth potential.
Ownership and direct impact, you have room to shift focus as your interests evolve.
Up to 150k USD salary, higher figures and EU/UK/US employment are negotiable.