Research Scientist

Apply now

Research Scientist

Full-time · Paris

About Gradium

Gradium is a voice AI company building the Text-to-Speech (TTS) engine for the next generation of voice agents.

We believe the best voice agents will win on quality, and that quality starts with the voice itself. We're purpose-built for product engineers and AI-native companies who need voices that sound genuinely human, stay reliable in production, and don't break on real-world content like acronyms, alphanumerics, or cross-lingual names.

We are a small team, early but moving fast, with real customers and a clear thesis: voice agents are the next major frontier of applied AI, and the TTS engine underneath will matter enormously. Our goal is to become the default voice layer for voice agent applications, starting by matching the quality bar, then outrunning the competition on speed of iteration and developer trust.

What you're stepping into: small team, unreasonable ambition, high ownership, fast decisions, and very little process unless it earns its place.

The Role

We're looking for a Research Scientist to help build the next generation of voice AI models, across voice synthesis, transcription and recognition. You'll design and train the models our customers build on, and push the boundaries of what's possible in real-time voice.

You'll work at the core of the company, partnering closely with the founders, research, and product to take models from idea to production. This role exists because our edge is the quality and reliability of our models, and that starts with the science underneath. Expect to own hard problems end to end, from architecture to inference at scale.

What You'll Do

  • Build state-of-the-art speech models: Design and implement models for voice synthesis and recognition that set the quality bar for the industry. You'll own architecture decisions and take models from research idea to something that holds up on real customer content.

  • Make it fast enough for production: Optimize models for real-time inference at scale, so quality never comes at the cost of latency. Reason about the trade-offs between accuracy, speed, and cost, and get the most out of the hardware.

  • Own training end to end: Build and maintain the training pipelines for large-scale model training, from data to distributed runs. Keep experiments fast and reproducible so the team can iterate quickly.

  • Turn research into product: Stay on top of the latest advances in speech AI and bring the ones that matter into our models. Work closely with the product team to translate customer requirements into technical solutions that ship.

Who You Are

  • Deep speech and AI expertise: You have an MS or PhD in Computer Science, Machine Learning, AI or a related field, and 3+ years in machine learning with a focus on speech or audio. You know deep learning architectures (Transformers, CNNs, RNNs) cold and understand how to make them work in practice.

  • Strong engineer, not just a researcher: You write excellent Python and are fluent with PyTorch or TensorFlow. You've run large-scale distributed training and know what breaks at scale.

  • Founder mindset: You act with urgency, take full ownership, and don't wait for permission or perfect information. You are comfortable making high-stakes decisions in ambiguous environments and see the founding team as partners, not hierarchy.

  • Obsessed with real-world quality: You care about how a model behaves on messy, real customer content, not just benchmark scores. You've felt the gap between a research demo and a production system, and you close it.

  • Nice to have: published research at top-tier ML venues (NeurIPS, ICML, ICLR); hands-on experience with speech synthesis models (Tacotron, FastSpeech, VITS); knowledge of audio signal processing and acoustic modeling; or experience with model optimization and quantization.

Still interested?

We'd love to hear about you!