npv labs logo

Data Engineer

npv labs Worldwide


No Relocation

Posted: August 18, 2026

Job Description

tl;dr: Data Engineer; post-acquisition profitable health/adtech; processing 20TB+ daily at near-realtime speed; Python, Spark; remote, up to 150k USD base, we can talk higher figures and EU/UK/US employment

PulsePoint sits at the intersection of healthcare and adtech. We help brands and agencies interpret the hard-to-read signals across the health journey and unify these digital determinants of health with real-world data to produce the most dimensional view of the customer. We are 300+ and growing, post-acquisition business and one of the leading players in the US healthcare ad market.

We're looking for a Data Engineer to join a team, playing a key role at a technology company experiencing exponential growth. The data pipeline processes over 80 billion impressions a day (20TB+ of data, 220TB uncompressed), powering reports, budget updates, and optimization engines against extremely tight SLAs, with stats and reports delivered as close to real-time as possible.

Team responsibilities

  • Install, maintain, and monitor Kafka, Hadoop, Presto, and RDBMS systems.

  • Ingest, validate, and process internal and third-party data.

  • Create, maintain, and monitor data flows in Hive, SQL, and Presto for consistency, accuracy, and lag time.

  • Maintain and enhance the framework for jobs (primarily aggregate jobs in Hive).

  • Build Kafka consumers using Spark Streaming for near-real-time aggregation.

  • Train developers and analysts on data tools.

  • Evaluate, select, and implement new tools.

  • Handle backups, retention, high availability, and capacity planning.

  • Review and approve DDL for databases, Hive framework jobs, and Spark Streaming to ensure standards are met.

  • Participate in 24×7 on-call rotation for production support.

Stack

  • Airflow, Docker, Graphite/Beacon, Hive, Impala, Kafka, Kubernetes, Presto, Spark Streaming, SQL Server, Sqoop.

Requirements

  • BA/BS degree in Computer Science or a related field.

  • 5+ years of software engineering experience.

  • Strong Spark expertise, Spark Streaming is highly desirable.

  • Proficiency in Linux.

  • Fluency in Python; experience with Scala/Java is a strong plus.

  • Strong understanding of RDBMS and SQL.

  • Willingness to participate in 24×7 on-call rotation.

Nice to have

  • Knowledge and exposure to distributed production systems like Hadoop.

  • Knowledge and exposure to cloud migration.

We offer

  • Remote work, high engineering bar and comfortable culture.

  • Flat hierarchy with easy access to business, product, and operations.

  • Enormous scale (80B+ impressions/day) with real growth potential.

  • Ownership and direct impact, you have room to shift focus as your interests evolve.

  • Up to 150k USD salary, higher figures and EU/UK/US employment are negotiable.

Additional Content

tl;dr: Data Engineer; post-acquisition profitable health/adtech; processing 20TB+ daily at near-realtime speed; Python, Spark; remote, up to 150k USD base, we can talk higher figures and EU/UK/US employment

PulsePoint sits at the intersection of healthcare and adtech. We help brands and agencies interpret the hard-to-read signals across the health journey and unify these digital determinants of health with real-world data to produce the most dimensional view of the customer. We are 300+ and growing, post-acquisition business and one of the leading players in the US healthcare ad market.

We're looking for a Data Engineer to join a team, playing a key role at a technology company experiencing exponential growth. The data pipeline processes over 80 billion impressions a day (20TB+ of data, 220TB uncompressed), powering reports, budget updates, and optimization engines against extremely tight SLAs, with stats and reports delivered as close to real-time as possible.

Team responsibilities

  • Install, maintain, and monitor Kafka, Hadoop, Presto, and RDBMS systems.

  • Ingest, validate, and process internal and third-party data.

  • Create, maintain, and monitor data flows in Hive, SQL, and Presto for consistency, accuracy, and lag time.

  • Maintain and enhance the framework for jobs (primarily aggregate jobs in Hive).

  • Build Kafka consumers using Spark Streaming for near-real-time aggregation.

  • Train developers and analysts on data tools.

  • Evaluate, select, and implement new tools.

  • Handle backups, retention, high availability, and capacity planning.

  • Review and approve DDL for databases, Hive framework jobs, and Spark Streaming to ensure standards are met.

  • Participate in 24×7 on-call rotation for production support.

Stack

  • Airflow, Docker, Graphite/Beacon, Hive, Impala, Kafka, Kubernetes, Presto, Spark Streaming, SQL Server, Sqoop.

Requirements

  • BA/BS degree in Computer Science or a related field.

  • 5+ years of software engineering experience.

  • Strong Spark expertise, Spark Streaming is highly desirable.

  • Proficiency in Linux.

  • Fluency in Python; experience with Scala/Java is a strong plus.

  • Strong understanding of RDBMS and SQL.

  • Willingness to participate in 24×7 on-call rotation.

Nice to have

  • Knowledge and exposure to distributed production systems like Hadoop.

  • Knowledge and exposure to cloud migration.

We offer

  • Remote work, high engineering bar and comfortable culture.

  • Flat hierarchy with easy access to business, product, and operations.

  • Enormous scale (80B+ impressions/day) with real growth potential.

  • Ownership and direct impact, you have room to shift focus as your interests evolve.

  • Up to 150k USD salary, higher figures and EU/UK/US employment are negotiable.