npv labs logo

Multimodal ML Engineer

npv labs • Paris


No Relocation

Posted: August 17, 2026

Job Description

We're looking for a Multimodal ML Engineer to join White Circle, an AI Safety company building the safety, reliability, and optimization layer for AI systems through natural-language policies it automatically tests, enforces, and improves at scale. Backed by $70M (Series A) from top funds and senior leaders at OpenAI, Anthropic, HuggingFace, Mistral, DeepMind, and others, White Circle processes 100M+ API calls monthly and fine-tunes and trains its own LLMs to run faster and cheaper than open or proprietary models.

You will

  • Train and fine-tune large-scale multimodal models (vision-language, audio, speech, video) from scratch and from pretrained checkpoints.

  • Design experiments, build multimodal data pipelines, and train MoE architectures.

  • Build alignment pipelines (SFT, DPO, GRPO), optimize for production (quantization, distillation, streaming), and deploy end-to-end.

  • Define evaluation metrics that actually matter for the product.

Requirements

  • 3+ years training large-scale multimodal models.

  • Strong PyTorch and distributed training experience (DeepSpeed, FSDP).

  • Deep familiarity with multimodal architectures – LLaVA, Qwen-VL, InternVL, Audio Flamingo, Whisper, HuBERT, Conformer or similar.

  • Hands-on RLHF/alignment across modalities (GRPO, DPO, reward modeling).

  • Both audio and video experience required – sequence modeling for each, plus large-scale dataset curation and production inference optimization.

  • Relocation to Paris or London (hybrid) required.

Bonus

  • Audio signal processing fundamentals – spectrograms, mel features, noise reduction.

  • MoE architecture experience.

We offer

  • Competitive salary + equity.

  • Official employment, visa and relocation help.