
Site Reliability Engineer (Hadoop)
npv labs • Worldwide
Posted: August 18, 2026
Job Description
tl;dr: SRE (Hadoop, Kafka); post-acquisition profitable health/adtech; 20PB+ scale; remote, up to 150k USD base, we can talk higher figures and EU/UK/US employment
PulsePoint sits at the intersection of healthcare and adtech. We help brands and agencies interpret the hard-to-read signals across the health journey and unify these digital determinants of health with real-world data to produce the most dimensional view of the customer. We are 300+ and growing, post-acquisition business and one of the leading players in the US healthcare ad market.
We're looking for a Site Reliability Engineer to help us ensure the reliability of streaming and storage systems processing billions of events daily through Kafka, Hadoop HDFS, and Ceph. The Data Platform team maintains these systems across hybrid infrastructure (bare-metal on-prem, cloud, and the integration between them) and you'll own reliability across the full lifecycle: architecture, deployment, capacity planning, and incident response.
You will
Ensure Kafka reliability – architecture, topic design, governance, partition strategy, throughput and latency optimization.
Own Ceph reliability – operations, pool design, placement optimization, capacity planning.
Build operational automation to reduce manual toil, speed up incident response, and prevent failures before they happen.
Support SQL Server backup and recovery pipelines and basic cluster reliability.
Build observability and self-service tooling for the data team.
Stack
Apache Kafka, Hadoop, Ceph, SQL Server backup/recovery, Terraform, Ansible, Puppet, ArgoCD, Prometheus, Grafana, Icinga, PagerDuty. Bare-metal and hybrid cloud/on-prem infrastructure.
Requirements
5+ years running distributed systems reliability at scale in production.
Deep expertise in Kafka, Ceph, or similar distributed infrastructure.
Proven ability to design for scale, reliability, and failure recovery.
Experience mentoring engineers and making technical decisions.
Willingness to work 9am–6pm ET US hours.
Nice to have
Multi-region replication and disaster recovery experience.
Cost optimization at infrastructure scale.
Hybrid on-prem/cloud operations experience.
We offer
Remote work, high engineering bar and comfortable culture.
Flat hierarchy with easy access to business, product, and operations.
Enormous scale (20PB+ cluster) with real growth potential.
Ownership and direct impact, you have room to shift focus as your interests evolve.
Up to 150k USD salary, higher figures and EU/UK/US employment are negotiable.
Additional Content
tl;dr: SRE (Hadoop, Kafka); post-acquisition profitable health/adtech; 20PB+ scale; remote, up to 150k USD base, we can talk higher figures and EU/UK/US employment
PulsePoint sits at the intersection of healthcare and adtech. We help brands and agencies interpret the hard-to-read signals across the health journey and unify these digital determinants of health with real-world data to produce the most dimensional view of the customer. We are 300+ and growing, post-acquisition business and one of the leading players in the US healthcare ad market.
We're looking for a Site Reliability Engineer to help us ensure the reliability of streaming and storage systems processing billions of events daily through Kafka, Hadoop HDFS, and Ceph. The Data Platform team maintains these systems across hybrid infrastructure (bare-metal on-prem, cloud, and the integration between them) and you'll own reliability across the full lifecycle: architecture, deployment, capacity planning, and incident response.
You will
Ensure Kafka reliability – architecture, topic design, governance, partition strategy, throughput and latency optimization.
Own Ceph reliability – operations, pool design, placement optimization, capacity planning.
Build operational automation to reduce manual toil, speed up incident response, and prevent failures before they happen.
Support SQL Server backup and recovery pipelines and basic cluster reliability.
Build observability and self-service tooling for the data team.
Stack
Apache Kafka, Hadoop, Ceph, SQL Server backup/recovery, Terraform, Ansible, Puppet, ArgoCD, Prometheus, Grafana, Icinga, PagerDuty. Bare-metal and hybrid cloud/on-prem infrastructure.
Requirements
5+ years running distributed systems reliability at scale in production.
Deep expertise in Kafka, Ceph, or similar distributed infrastructure.
Proven ability to design for scale, reliability, and failure recovery.
Experience mentoring engineers and making technical decisions.
Willingness to work 9am–6pm ET US hours.
Nice to have
Multi-region replication and disaster recovery experience.
Cost optimization at infrastructure scale.
Hybrid on-prem/cloud operations experience.
We offer
Remote work, high engineering bar and comfortable culture.
Flat hierarchy with easy access to business, product, and operations.
Enormous scale (20PB+ cluster) with real growth potential.
Ownership and direct impact, you have room to shift focus as your interests evolve.
Up to 150k USD salary, higher figures and EU/UK/US employment are negotiable.