npv labs logo

Site Reliability Engineer (Hadoop)

npv labs Worldwide


No Relocation

Posted: August 18, 2026

Job Description

tl;dr: SRE (Hadoop, Kafka); post-acquisition profitable health/adtech; 20PB+ scale; remote, up to 150k USD base, we can talk higher figures and EU/UK/US employment

PulsePoint sits at the intersection of healthcare and adtech. We help brands and agencies interpret the hard-to-read signals across the health journey and unify these digital determinants of health with real-world data to produce the most dimensional view of the customer. We are 300+ and growing, post-acquisition business and one of the leading players in the US healthcare ad market.

We're looking for a Site Reliability Engineer to help us ensure the reliability of streaming and storage systems processing billions of events daily through Kafka, Hadoop HDFS, and Ceph. The Data Platform team maintains these systems across hybrid infrastructure (bare-metal on-prem, cloud, and the integration between them) and you'll own reliability across the full lifecycle: architecture, deployment, capacity planning, and incident response.

You will

  • Ensure Kafka reliability – architecture, topic design, governance, partition strategy, throughput and latency optimization.

  • Own Ceph reliability – operations, pool design, placement optimization, capacity planning.

  • Build operational automation to reduce manual toil, speed up incident response, and prevent failures before they happen.

  • Support SQL Server backup and recovery pipelines and basic cluster reliability.

  • Build observability and self-service tooling for the data team.

Stack

  • Apache Kafka, Hadoop, Ceph, SQL Server backup/recovery, Terraform, Ansible, Puppet, ArgoCD, Prometheus, Grafana, Icinga, PagerDuty. Bare-metal and hybrid cloud/on-prem infrastructure.

Requirements

  • 5+ years running distributed systems reliability at scale in production.

  • Deep expertise in Kafka, Ceph, or similar distributed infrastructure.

  • Proven ability to design for scale, reliability, and failure recovery.

  • Experience mentoring engineers and making technical decisions.

  • Willingness to work 9am–6pm ET US hours.

Nice to have

  • Multi-region replication and disaster recovery experience.

  • Cost optimization at infrastructure scale.

  • Hybrid on-prem/cloud operations experience.

We offer

  • Remote work, high engineering bar and comfortable culture.

  • Flat hierarchy with easy access to business, product, and operations.

  • Enormous scale (20PB+ cluster) with real growth potential.

  • Ownership and direct impact, you have room to shift focus as your interests evolve.

  • Up to 150k USD salary, higher figures and EU/UK/US employment are negotiable.

Additional Content

tl;dr: SRE (Hadoop, Kafka); post-acquisition profitable health/adtech; 20PB+ scale; remote, up to 150k USD base, we can talk higher figures and EU/UK/US employment

PulsePoint sits at the intersection of healthcare and adtech. We help brands and agencies interpret the hard-to-read signals across the health journey and unify these digital determinants of health with real-world data to produce the most dimensional view of the customer. We are 300+ and growing, post-acquisition business and one of the leading players in the US healthcare ad market.

We're looking for a Site Reliability Engineer to help us ensure the reliability of streaming and storage systems processing billions of events daily through Kafka, Hadoop HDFS, and Ceph. The Data Platform team maintains these systems across hybrid infrastructure (bare-metal on-prem, cloud, and the integration between them) and you'll own reliability across the full lifecycle: architecture, deployment, capacity planning, and incident response.

You will

  • Ensure Kafka reliability – architecture, topic design, governance, partition strategy, throughput and latency optimization.

  • Own Ceph reliability – operations, pool design, placement optimization, capacity planning.

  • Build operational automation to reduce manual toil, speed up incident response, and prevent failures before they happen.

  • Support SQL Server backup and recovery pipelines and basic cluster reliability.

  • Build observability and self-service tooling for the data team.

Stack

  • Apache Kafka, Hadoop, Ceph, SQL Server backup/recovery, Terraform, Ansible, Puppet, ArgoCD, Prometheus, Grafana, Icinga, PagerDuty. Bare-metal and hybrid cloud/on-prem infrastructure.

Requirements

  • 5+ years running distributed systems reliability at scale in production.

  • Deep expertise in Kafka, Ceph, or similar distributed infrastructure.

  • Proven ability to design for scale, reliability, and failure recovery.

  • Experience mentoring engineers and making technical decisions.

  • Willingness to work 9am–6pm ET US hours.

Nice to have

  • Multi-region replication and disaster recovery experience.

  • Cost optimization at infrastructure scale.

  • Hybrid on-prem/cloud operations experience.

We offer

  • Remote work, high engineering bar and comfortable culture.

  • Flat hierarchy with easy access to business, product, and operations.

  • Enormous scale (20PB+ cluster) with real growth potential.

  • Ownership and direct impact, you have room to shift focus as your interests evolve.

  • Up to 150k USD salary, higher figures and EU/UK/US employment are negotiable.