Deepgram logo
    D

    Director, Text-to-Speech Synthesis Research

    Deepgram
    USA | Remote•San Francisco, CA•Ann Arbor, MI
    Remote
    Director
    Full Time
    💰$ 213,000 - $ 328,300
    AIText-to-SpeechResearchLeadershipNeural NetworksSpeech GenerationMachine Learning

    Requirements

    • •Deep expertise in modern TTS, speech generation, or audio generative modeling with experience training large-scale neural models
    • •Command of the modern speech-generation stack and understanding of challenges like naturalness, expressiveness, controllability, robustness, voice consistency, and inference cost
    • •Experience setting research direction under uncertainty, prioritizing experiments, allocating resources, and terminating ineffective approaches
    • •Experience leading researchers and research engineers through technical leaders, developing tech lead managers, and maintaining technical influence
    • •AI as a default mode of work with a clear view of its capabilities and limitations in speech research
    • •Ability to communicate complex technical tradeoffs to product, engineering, and executive audiences

    What You'll Do

    • •Own the TTS research and model roadmap, deciding technical directions to improve speech-generation quality
    • •Drive advances in neural audio modeling, prosody, expressiveness, controllability, multilingual speech, voice identity and consistency, data and training strategy, post-training, and inference performance
    • •Stay deeply technical by reviewing research, designing experiments, diagnosing model failures, and solving high-leverage problems
    • •Build evaluation and benchmarking systems combining automated metrics and human perceptual assessment
    • •Lead a team of individual contributors and tech lead managers, hiring and developing talent, setting direction across sub-teams
    • •Partner with engineering and product leadership on ship-readiness and represent TTS research internally and externally

    Nice to Have

    • •TTS or generative-audio models deployed at meaningful production scale
    • •Built or scaled a high-performing AI research organization
    • •Sophisticated evaluation systems for generative speech including expressive or multilingual generation, voice cloning and adaptation, or controllable generation
    • •Recognized external contributions such as publications, open source, patents, or invited talks in relevant fields
    • •Experience in fast-moving startup or research environments taking models from idea to production

    About Deepgram

    Deepgram specializes in providing AI-powered speech-to-text technology that offers audio intelligence, text-to-speech, and voice agent API.

    San Francisco, CA
    100 - 250
    AI & Machine Learning