USA | Remote
Remote
Senior
Full Time
16 days ago
💰$ 180,000 - $ 240,000
AIModel EvaluationSpeech-to-TextText-to-SpeechLLMMultimodal SystemsSoftware EngineeringSenior
Requirements
- •BS, MS, or PhD in Computer Science, AI, Applied Math, or a related field, or equivalent experience
- •5+ years of professional software or QA engineering experience, with a track record of shipping test infrastructure or evaluation systems
- •Solid backend/scripting experience in a language such as Python, Rust, Go, or similar
- •Experience designing and building automated test pipelines, evaluation frameworks, or data-processing systems
- •Strong analytical skills and comfort reasoning about metrics, thresholds, and statistical variation in results
- •Ability to take charge of ambiguous technical challenges and communicate effectively across research, engineering, and product teams
What You'll Do
- •Define and build evaluation methodologies for Deepgram's models, spanning speech-to-text, text-to-speech, and emerging LLM, RAG, agent, and multimodal systems
- •Design, build, and maintain automated evaluation pipelines across batch and streaming (e.g. WER, runaway/hallucination detection, latency and time-to-first-byte), with a focus on correctness, reproducibility, and ease of adoption
- •Build scalable, reproducible evaluation infrastructure — harnesses, orchestration, and result-aggregation pipelines — running against production models and, where needed, large GPU clusters
- •Translate Research benchmarks and expected model metrics into automated, enforceable pass/fail gates
- •Build and operate canaries and continuous-monitoring systems that detect quality regressions in production before they reach customers
- •Partner with DevOps/Infra to stand up ephemeral test environments and results-aggregation infrastructure
- •Work alongside Research, model training, inference, and product teams to provide trusted evaluation signals that inform release and optimization decisions
- •Integrate evaluation and quality gates into CI/CD so quality is verified continuously, not manually
- •Help raise the bar through code reviews, technical design discussions, and strong engineering and QA practices
Nice to Have
- •Hands-on experience evaluating modern AI systems such as LLMs, RAG pipelines, agents, or multimodal models, including model behavior analysis
- •Experience with React Native or other cross-platform mobile frameworks for building tooling that's accessible beyond the desktop
- •Experience building or improving evaluation frameworks, benchmarks, or ML infrastructure used by other teams or external users
- •A strong appreciation for evaluation quality — correctness, reproducibility, and consistency across environments
- •Experience with voice, audio, speech recognition, or real-time systems, and familiarity with metrics like WER, MOS, or latency/TTFB
- •Prior involvement in open-source projects, through contributions, reviews, maintenance, or community engagement
- •Experience acting as a technical bridge across teams or platforms (evaluation, training, inference, agent frameworks), combining architectural understanding with clear communication and influence
- •Familiarity with cloud infrastructure, containerized/ephemeral environments, and monitoring tooling (e.g. Grafana, canaries, anomaly detection)
