USA | Remote
Remote
Senior
Full Time
16 days ago
AI-first mindsetvoice AIdata operationsprogram managementspeech AIremote
Requirements
- •Experience owning data, ML, or operations programs end-to-end in program/project or product management
- •Fluency working directly with technical teams; ability to discuss data quality, evaluation, and model impact
- •Systems thinking and ability to design scalable solutions
- •Track record of prioritization under constraint
- •Strong operating instincts to scope, sequence, assign, and ship tasks
- •Demonstrated ability to design and build scalable processes
What You'll Do
- •Design, launch, and own end-to-end data workflows from raw audio ingestion to production-ready datasets
- •Build and evolve labeling specs, style guides, and instructional documentation for global annotation teams
- •Identify opportunities for better tooling, automation, and workflow optimization, and lead their implementation
- •Translate product goals and model requirements into data creation strategies
- •Own the full lifecycle for your domains including customer expectations, data acquisition, preparation, scaling, provenance, and evaluation
- •Prototype and deploy data tools and infrastructure
- •Collaborate with Research and Engineering to align data collection with model training architecture and downstream product impact
- •Track advancements in speech AI research and evolving market use cases to inform labeling approaches and data design priorities
- •Partner with QA and Evaluation leads to deliver high-quality, human-in-the-loop datasets and benchmarks
- •Manage and mentor data vendors, freelancers, and potentially internal ICs
- •Track throughput, data quality, and vendor performance
- •Drive continuous improvement in speed, cost-efficiency, and quality across all data operations
- •Curate and refine datasets to align with specific product goals, linguistic coverage, or research hypotheses
Nice to Have
- •Direct exposure to speech/audio, ASR, or TTS data and nuances of multilingual, code-switched, low-resource, or domain-specific data
- •Experience with active-learning or data-selection approaches
- •Startup or high-ambiguity experience
