USA | Remote•San Francisco, CA•Ann Arbor, MI
Remote
Director
Full Time
💰$ 213,000 - $ 328,300
AIText-to-SpeechResearchLeadershipNeural NetworksSpeech GenerationMachine Learning
Requirements
- •Deep expertise in modern TTS, speech generation, or audio generative modeling with experience training large-scale neural models
- •Command of the modern speech-generation stack and understanding of challenges like naturalness, expressiveness, controllability, robustness, voice consistency, and inference cost
- •Experience setting research direction under uncertainty, prioritizing experiments, allocating resources, and terminating ineffective approaches
- •Experience leading researchers and research engineers through technical leaders, developing tech lead managers, and maintaining technical influence
- •AI as a default mode of work with a clear view of its capabilities and limitations in speech research
- •Ability to communicate complex technical tradeoffs to product, engineering, and executive audiences
What You'll Do
- •Own the TTS research and model roadmap, deciding technical directions to improve speech-generation quality
- •Drive advances in neural audio modeling, prosody, expressiveness, controllability, multilingual speech, voice identity and consistency, data and training strategy, post-training, and inference performance
- •Stay deeply technical by reviewing research, designing experiments, diagnosing model failures, and solving high-leverage problems
- •Build evaluation and benchmarking systems combining automated metrics and human perceptual assessment
- •Lead a team of individual contributors and tech lead managers, hiring and developing talent, setting direction across sub-teams
- •Partner with engineering and product leadership on ship-readiness and represent TTS research internally and externally
Nice to Have
- •TTS or generative-audio models deployed at meaningful production scale
- •Built or scaled a high-performing AI research organization
- •Sophisticated evaluation systems for generative speech including expressive or multilingual generation, voice cloning and adaptation, or controllable generation
- •Recognized external contributions such as publications, open source, patents, or invited talks in relevant fields
- •Experience in fast-moving startup or research environments taking models from idea to production
