
Data Engineer
npv labs • Worldwide
Posted: August 18, 2026
Job Description
tl;dr: Data Engineer; post-acquisition profitable health/adtech; processing 20TB+ daily at near-realtime speed; Python, Spark; remote, up to 150k USD base, we can talk higher figures and EU/UK/US employment
PulsePoint sits at the intersection of healthcare and adtech. We help brands and agencies interpret the hard-to-read signals across the health journey and unify these digital determinants of health with real-world data to produce the most dimensional view of the customer. We are 300+ and growing, post-acquisition business and one of the leading players in the US healthcare ad market.
We're looking for a Data Engineer to join a team, playing a key role at a technology company experiencing exponential growth. The data pipeline processes over 80 billion impressions a day (20TB+ of data, 220TB uncompressed), powering reports, budget updates, and optimization engines against extremely tight SLAs, with stats and reports delivered as close to real-time as possible.
Team responsibilities
Install, maintain, and monitor Kafka, Hadoop, Presto, and RDBMS systems.
Ingest, validate, and process internal and third-party data.
Create, maintain, and monitor data flows in Hive, SQL, and Presto for consistency, accuracy, and lag time.
Maintain and enhance the framework for jobs (primarily aggregate jobs in Hive).
Build Kafka consumers using Spark Streaming for near-real-time aggregation.
Train developers and analysts on data tools.
Evaluate, select, and implement new tools.
Handle backups, retention, high availability, and capacity planning.
Review and approve DDL for databases, Hive framework jobs, and Spark Streaming to ensure standards are met.
Participate in 24×7 on-call rotation for production support.
Stack
Airflow, Docker, Graphite/Beacon, Hive, Impala, Kafka, Kubernetes, Presto, Spark Streaming, SQL Server, Sqoop.
Requirements
BA/BS degree in Computer Science or a related field.
5+ years of software engineering experience.
Strong Spark expertise, Spark Streaming is highly desirable.
Proficiency in Linux.
Fluency in Python; experience with Scala/Java is a strong plus.
Strong understanding of RDBMS and SQL.
Willingness to participate in 24×7 on-call rotation.
Nice to have
Knowledge and exposure to distributed production systems like Hadoop.
Knowledge and exposure to cloud migration.
We offer
Remote work, high engineering bar and comfortable culture.
Flat hierarchy with easy access to business, product, and operations.
Enormous scale (80B+ impressions/day) with real growth potential.
Ownership and direct impact, you have room to shift focus as your interests evolve.
Up to 150k USD salary, higher figures and EU/UK/US employment are negotiable.
Additional Content
tl;dr: Data Engineer; post-acquisition profitable health/adtech; processing 20TB+ daily at near-realtime speed; Python, Spark; remote, up to 150k USD base, we can talk higher figures and EU/UK/US employment
PulsePoint sits at the intersection of healthcare and adtech. We help brands and agencies interpret the hard-to-read signals across the health journey and unify these digital determinants of health with real-world data to produce the most dimensional view of the customer. We are 300+ and growing, post-acquisition business and one of the leading players in the US healthcare ad market.
We're looking for a Data Engineer to join a team, playing a key role at a technology company experiencing exponential growth. The data pipeline processes over 80 billion impressions a day (20TB+ of data, 220TB uncompressed), powering reports, budget updates, and optimization engines against extremely tight SLAs, with stats and reports delivered as close to real-time as possible.
Team responsibilities
Install, maintain, and monitor Kafka, Hadoop, Presto, and RDBMS systems.
Ingest, validate, and process internal and third-party data.
Create, maintain, and monitor data flows in Hive, SQL, and Presto for consistency, accuracy, and lag time.
Maintain and enhance the framework for jobs (primarily aggregate jobs in Hive).
Build Kafka consumers using Spark Streaming for near-real-time aggregation.
Train developers and analysts on data tools.
Evaluate, select, and implement new tools.
Handle backups, retention, high availability, and capacity planning.
Review and approve DDL for databases, Hive framework jobs, and Spark Streaming to ensure standards are met.
Participate in 24×7 on-call rotation for production support.
Stack
Airflow, Docker, Graphite/Beacon, Hive, Impala, Kafka, Kubernetes, Presto, Spark Streaming, SQL Server, Sqoop.
Requirements
BA/BS degree in Computer Science or a related field.
5+ years of software engineering experience.
Strong Spark expertise, Spark Streaming is highly desirable.
Proficiency in Linux.
Fluency in Python; experience with Scala/Java is a strong plus.
Strong understanding of RDBMS and SQL.
Willingness to participate in 24×7 on-call rotation.
Nice to have
Knowledge and exposure to distributed production systems like Hadoop.
Knowledge and exposure to cloud migration.
We offer
Remote work, high engineering bar and comfortable culture.
Flat hierarchy with easy access to business, product, and operations.
Enormous scale (80B+ impressions/day) with real growth potential.
Ownership and direct impact, you have room to shift focus as your interests evolve.
Up to 150k USD salary, higher figures and EU/UK/US employment are negotiable.