Machine Learning Operations (MLOps) Engineer

AGZEN INC.
Somerville, MA, United States
1 day ago
Apply on startup.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$150,000.0 - $200,000.0
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Artificial Neural Networks Cloud Computing Data Visualization Distributed Systems Python (Programming Language) Machine Learning Metadata NumPy Recommender Systems Tensorflow Standard Sql
+9 more
Management of Software Versions Pytorch Deep Learning Pandas Scikit Learn Information Technology Data Management Machine Learning Operations Data Pipelines

Job description

We are looking for a sharp, tenacious, and thorough Senior Machine Learning Operations (MLOps) Engineer to join our team. As part of the the perception team, you’ll own the operational layer around of machine learning models. This role will be responsible for the intake and leveraging crop protection data collected from RealCoverage units installed on sprayers all around the world which is then used improve our CV pipeline and Recommendation Engine. This role will be an essential component of AgZen’s measurement focus group. Strong communication, flexibility, teamwork, the desire to take on different responsibilities and own them will all be essential skills for a successful applicant., * Own the architecture, execution, and operational excellence of large-scale, cloud-native pipelines for multimodal sensor data ingestion, processing, labeling, and validation.

  • Champion model traceability by building a clear lineage for every production model. Track what data trained it, what code produced it, what validation it passed, and how it’s performing. Evaluate and recommend tooling for versioning, metadata, and model registry
  • Partner with data scientists to detect data quality issues, detect drift in upstream sources, and ensure features stay fresh and reliable
  • Track model drift over weeks, flag slow degradation before it crosses a threshold, surface feature freshness problems before they cascade
  • Build diagnostic tooling to root cause pipeline and recommendation issues quickly. Ensure the right context is logged at each stage, candidates, features, serving context, and building the dashboards to tie it collectively
  • Own automated gates that block bad deployments and assist in running model issue retrospectives
  • Work with ML engineers, data engineers, and stakeholders to coordinate on post-deployment metrics, defining what metrics to collect after deployment and why they matter
  • Build tooling and support non-technical domain experts in understanding perception system performance and identifying opportunities for pipeline improvement
  • Collaborate closely with cross-functional teams of software engineers, machine learning scientists, product specialists, and researchers to design, build, and maintain robust data pipelines grounded in sound data organization, domain knowledge, and careful analysis
  • Communicate technical findings, data characteristics, and limitations clearly and effectively to both internal partners and external collaborators

Requirements

  • Bachelor’s or graduate degree in Computer Science, Electrical Engineering, or a closely related field
  • 5+ years of experience building large-scale distributed systems, applications, or advanced ML systems-scale distributed systems, applications, or advanced ML systems
  • Experience with MLOps, data pipelines, and cloud distributed systems
  • Proficiency in Python for system-level and performance-critical implementation
  • Experience operating end-to-end data or ML pipelines for reliability, scale, and observability
  • Communication skills that align collaborators and drive execution across functions
  • Familiarity with deep learning frameworks (e.g., PyTorch, TensorFlow)
  • A record of ownership, accountability, and customer-focused engineering
  • Proven track record of designing robust frameworks with high-quality, durable APIs
  • Deep understanding of machine learning algorithms with hands-on application
  • Expertise in building reliable, high-performance, and cost-efficient systems on modern cloud infrastructure-performance
  • Robust SQL skills and comfort digging into data distributions, feature health, and model behavior

Preferred:

  • Experience with the field of agriculture or related fields such as environmental or life sciences
  • Experience with data science based on real-world physical sensors data
  • Experience with vision-based ML
  • Experience creating intuitive data visualization tools that make complex data approachable for non-technical users
  • Prior experience in developing machine-learning models relevant to biological or crop protection outcomes
  • Advanced scientific Python (NumPy, Pandas, scikit-learn) and hands-on experience with PyTorch and/or TensorFlow, including training and deploying neural networks
  • Experience operating recommendation systems at scale

Benefits & conditions

  • Early-employee equity
  • 401(k) with employer matching at 6 months of employment
  • 6 weeks of PTO per calendar year
  • 12 paid holidays
  • Medical, Dental and Vision insurance, The salary range for this position is $150,000 - $200,000 depending on skills and qualifications evaluated on a per candidate basis.

About the company

AgZen is a fast-growing precision agriculture company headquartered in Somerville, MA, built on MIT research and focused on one problem: making crop spraying more efficient. Our flagship product, RealCoverage, is the world’s first system that measures and controls droplet coverage at the leaf level, giving growers real-time visibility into spray performance and cutting chemical and water use by up to 50% without sacrificing yield.

We are a small, technically deep team working at the intersection of fluid mechanics, computer vision, AI, and real agricultural environments. If you want to build technology with measurable impact on how the world grows food, this is the place to do it.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on startup.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

5:28 min

Defining MLOps and its role in production systems

Hauke Brammer · World Congress 2023

2:34 min

Maximizing execution memory effectively via python numpy broadcasting

Jodie Burchell · LIVE

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · World Congress 2023

2:03 min

Accelerating pandas dataframes using cudf module plugins

Ankit Patel Ankit Patel · World Congress 2024

2:08 min

Essential engineering roles in the generative AI space

Mary Grygleski Mary Grygleski · LIVE

1:25 min

Replacing NumPy with cuPy for straightforward GPU acceleration

Paul Graham Paul Graham · World Congress 2025

Videos

See all

Related articles

See all