Gradium fournit aux développeurs et aux entreprises des API vocales en temps réel pour la transcription, la synthèse et les agents conversationnels. Sa technologie permet de créer des interactions vocales naturelles avec une faible latence.
Research Scientist
Am I a fit — voir ma compatibilitéGradium recherche un Research Scientist à Paris pour concevoir et entraîner des modèles d’IA vocale dédiés à la synthèse, la transcription et la reconnaissance. Le poste couvre l’optimisation pour l’inférence temps réel, les pipelines d’entraînement distribués et le passage de la recherche à la production. Une expertise en speech AI, Python, PyTorch ou TensorFlow, ainsi qu’un diplôme de niveau master ou doctorat sont demandés.
Repères sur Gradium
- Secteur
- Logiciels et Internet
- Modèle
- B2B
- Type d’entreprise
- Financée par du capital-risque
- Stade
- Startup early-stage
- Siège
- Paris
- Domaine officiel
- gradium.ai
- Création
- 2025 · 1 an
- Financement total publié
- 100 M$Source ↗
- Dernière levée
- Seed · 2026-07-08
- Montant annoncé
- 30 000 000 $
- Offres ouvertes
- 6
Investisseurs mentionnés
Détails de l’offre
La description complète publiée par Gradium.
Description de l’offre
About Gradium Gradium is a voice AI company building the Text-to-Speech (TTS) engine for the next generation of voice agents. We believe the best voice agents will win on quality, and that quality starts with the voice itself. We're purpose-built for product engineers and AI-native companies who need voices that sound genuinely human, stay reliable in production, and don't break on real-world content like acronyms, alphanumerics, or cross-lingual names.
We are a small team , early but moving fast, with real customers and a clear thesis: voice agents are the next major frontier of applied AI , and the TTS engine underneath will matter enormously. Our goal is to become the default voice layer for voice agent applications, starting by matching the quality bar, then outrunning the competition on speed of iteration and developer trust. What you're stepping into: small team, unreasonable ambition, high ownership, fast decisions, and very little process unless it earns its place.
The Role
We're looking for a Research Scientist to help build the next generation of voice AI models, across voice synthesis, transcription and recognition. You'll design and train the models our customers build on, and push the boundaries of what's possible in real-time voice.
You'll work at the core of the company, partnering closely with the founders, research, and product to take models from idea to production. This role exists because our edge is the quality and reliability of our models, and that starts with the science underneath. Expect to own hard problems end to end, from architecture to inference at scale.
What You'll Do
- Build state-of-the-art speech models: Design and implement models for voice synthesis and recognition that set the quality bar for the industry. You'll own architecture decisions and take models from research idea to something that holds up on real customer content.
- Make it fast enough for production: Optimize models for real-time inference at scale, so quality never comes at the cost of latency. Reason about the trade-offs between accuracy, speed, and cost, and get the most out of the hardware.
- Own training end to end: Build and maintain the training pipelines for large-scale model training, from data to distributed runs. Keep experiments fast and reproducible so the team can iterate quickly.
- Turn research into product: Stay on top of the latest advances in speech AI and bring the ones that matter into our models. Work closely with the product team to translate customer requirements into technical solutions that ship.
Who You Are
- Deep speech and AI expertise: You have an MS or PhD in Computer Science, Machine Learning, AI or a related field, and 3+ years in machine learning with a focus on speech or audio. You know deep learning architectures (Transformers, CNNs, RNNs) cold and understand how to make them work in practice.
- Strong engineer, not just a researcher: You write excellent Python and are fluent with PyTorch or TensorFlow. You've run large-scale distributed training and know what breaks at scale.
- Founder mindset: You act with urgency, take full ownership, and don't wait for permission or perfect information. You are comfortable making high-stakes decisions in ambiguous environments and see the founding team as partners, not hierarchy.
- Obsessed with real-world quality: You care about how a model behaves on messy, real customer content, not just benchmark scores. You've felt the gap between a research demo and a production system, and you close it.
-
Nice to have
published research at top-tier ML venues (NeurIPS, ICML, ICLR); hands-on experience with speech synthesis models (Tacotron, FastSpeech, VITS); knowledge of audio signal processing and acoustic modeling; or experience with model optimization and quantization.
Still interested?
We'd love to hear about you!
Prérequis
- Expertise en speech et IA
- Entraînement distribué à grande échelle
- Audio signal processing
- Acoustic modeling
- Model optimization and quantization