Deepgram logo
    D

    Senior Software Engineer - Model Evaluation & AI Systems

    Deepgram
    USA | Remote
    Remote
    Senior
    Full Time
    16 days ago
    💰$ 180,000 - $ 240,000
    AIModel EvaluationSpeech-to-TextText-to-SpeechLLMMultimodal SystemsSoftware EngineeringSenior

    Requirements

    • BS, MS, or PhD in Computer Science, AI, Applied Math, or a related field, or equivalent experience
    • 5+ years of professional software or QA engineering experience, with a track record of shipping test infrastructure or evaluation systems
    • Solid backend/scripting experience in a language such as Python, Rust, Go, or similar
    • Experience designing and building automated test pipelines, evaluation frameworks, or data-processing systems
    • Strong analytical skills and comfort reasoning about metrics, thresholds, and statistical variation in results
    • Ability to take charge of ambiguous technical challenges and communicate effectively across research, engineering, and product teams

    What You'll Do

    • Define and build evaluation methodologies for Deepgram's models, spanning speech-to-text, text-to-speech, and emerging LLM, RAG, agent, and multimodal systems
    • Design, build, and maintain automated evaluation pipelines across batch and streaming (e.g. WER, runaway/hallucination detection, latency and time-to-first-byte), with a focus on correctness, reproducibility, and ease of adoption
    • Build scalable, reproducible evaluation infrastructure — harnesses, orchestration, and result-aggregation pipelines — running against production models and, where needed, large GPU clusters
    • Translate Research benchmarks and expected model metrics into automated, enforceable pass/fail gates
    • Build and operate canaries and continuous-monitoring systems that detect quality regressions in production before they reach customers
    • Partner with DevOps/Infra to stand up ephemeral test environments and results-aggregation infrastructure
    • Work alongside Research, model training, inference, and product teams to provide trusted evaluation signals that inform release and optimization decisions
    • Integrate evaluation and quality gates into CI/CD so quality is verified continuously, not manually
    • Help raise the bar through code reviews, technical design discussions, and strong engineering and QA practices

    Nice to Have

    • Hands-on experience evaluating modern AI systems such as LLMs, RAG pipelines, agents, or multimodal models, including model behavior analysis
    • Experience with React Native or other cross-platform mobile frameworks for building tooling that's accessible beyond the desktop
    • Experience building or improving evaluation frameworks, benchmarks, or ML infrastructure used by other teams or external users
    • A strong appreciation for evaluation quality — correctness, reproducibility, and consistency across environments
    • Experience with voice, audio, speech recognition, or real-time systems, and familiarity with metrics like WER, MOS, or latency/TTFB
    • Prior involvement in open-source projects, through contributions, reviews, maintenance, or community engagement
    • Experience acting as a technical bridge across teams or platforms (evaluation, training, inference, agent frameworks), combining architectural understanding with clear communication and influence
    • Familiarity with cloud infrastructure, containerized/ephemeral environments, and monitoring tooling (e.g. Grafana, canaries, anomaly detection)

    About Deepgram

    Deepgram specializes in providing AI-powered speech-to-text technology that offers audio intelligence, text-to-speech, and voice agent API.

    San Francisco, CA
    100 - 250
    AI & Machine Learning