Data Ops Lead
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
Job description
Your mission is to turn Neonâs raw consumer audio streams into the cleanest, most reliable training data on the market, and to build the commercial and operational engine that gets it into the hands of the worldâs leading AI labs.
As a Data Ops Lead, youâll own the end-to-end journey that takes raw recordings from our growing community of 500,000+ mobile users and delivers production-ready datasets to frontier labs. In practice, that means three things above all:
- Structuring and managing the data deals that turn our recordings into revenue
- Holding every dataset to a quality bar that keeps buyers coming back
- Standing up human transcription, annotation and other operations, largely overseas, that make it all possible
Youâll work directly with our CEO on commercial priorities and help shape each deal, interface with buyer-side engineering and research teams at frontier labs to translate their exact specifications into deliverable dataset plans, and partner with internal engineering and external vendors to make sure the pipeline supports what weâve sold. This is a foundational role: the datasets and processes you build are the product we sell.
Requirements
- Authorization to work in the US.
- 5+ years of experience building and scaling data pipelines for AI/ML applications, with significant time spent on audio, speech, or multimodal data.
- A track record of structuring and delivering against data or dataset agreements with external partners: taking their requirements, turning them into clear specifications, and owning delivery end to end.
- Experience building and managing overseas or outsourced teams for data tagging, annotation, and QA, with a track record of maintaining quality and throughput across time zones.
- Deep ownership of data quality: designing QA processes, defining acceptance criteria, and catching problems before a customer ever sees them.
- Enough technical fluency to be credible on both sides of a deal. You understand digital audio fundamentals (sample rates, VAD, multichannel formats), can reason about how pipelines are built, and know what âgoodâ looks like, even if youâre not writing every line of code yourself.
- A âFounderâs Mentality.â Youâre comfortable building from zero and making high-stakes calls with incomplete information.
Bonus points
- A background working with audio data in some capacity.
- Direct experience with training data for TTS, ASR, speaker ID, or full-duplex conversational models.
- Familiarity with the modern audio stack (Librosa, FFmpeg, SoX, torchaudio) and cloud data infrastructure (S3, Redshift, BigQuery, or equivalent).
- An understanding of how high-quality, speaker-separated audio gets captured (for example, via WebRTC-based recording tools).
- Experience with active learning loops, human-in-the-loop QA systems, or corpus stratification for balanced dataset design.
- Prior experience leading a data or infrastructure team, including hiring and mentoring engineers.
Benefits & conditions
Compensation Range: $150K - $190K
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.indeed.comGood distractions
Talks and stories from around this role â technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Highest Paying Tech Companies for Developers
Dev Digest 120 - Apple and peers
Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production
Dev Digest 129 - Now that's what I call private data!