Machine Learning Operations Engineer

Nous Infosystems
United States
6 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Amazon Web Services Amazon S3 Data Analysis Computer Vision Batch Processing Cloud Engineering Information Systems Continuous Integration Data Validation Information Engineering Extract Transform Load (ETL)
+21 more
DevOps Monitoring of Systems Identity and Access Management Python (Programming Language) Machine Learning Operational Databases Release Management Cloud Services Workflow Management Systems AWS Cdk Cloud Platform System Delivery Pipeline State Machines Cloudformation Infrastructure Automation Frameworks Information Technology Machine Learning Operations Cloudwatch Terraform Software Version Control Docker

Job description

PG&E is seeking a Machine Learning Operations Engineer with practical AWS experience to help deploy, monitor, and support machine learning and computer vision solutions in production. In this role, you will work with data scientists, machine learning engineers, product teams, and business stakeholders to turn approved models into reliable, repeatable, and well-documented production workflows. The ideal candidate understands the basics of machine learning, enjoys building dependable cloud-based processes, and can communicate clearly with both technical and non-technical teams. What You’ll Do Support the deployment and day-to-day operation of machine learning and computer vision models used for inspection and asset intelligence use cases, including overhead equipment inspection and unauthorized attachment detection. Partner with data scientists and machine learning engineers to package approved models for production use and make sure model handoffs are clear, tested, and documented. Build and maintain practical AWS-based workflows for data movement, model execution, batch inference, and output delivery using services such as Amazon S3, SageMaker, Lambda, Step Functions, CloudWatch, and related AWS tools. Help create repeatable deployment processes so models can move from development to testing to production in a controlled and consistent way. Support CI/CD practices for machine learning workflows, including code versioning, automated checks, deployment readiness steps, and release coordination. Monitor production model runs for job completion, data issues, system errors, performance changes, and operational readiness. Assist with troubleshooting production inference issues by reviewing logs, validating inputs and outputs, coordinating fixes, and communicating status to stakeholders. Maintain clear runbooks, deployment notes, monitoring summaries, and support documentation so production workflows can be operated consistently by the broader team. Work with product managers, SMEs, data teams, cloud platform teams, and business stakeholders to align on production requirements, release timing, support needs, and success measures. Help improve reliability, scalability, security, and cost awareness for machine learning workloads without over-engineering the solution. What You Bring

Requirements

Bachelor s degree in computer science, engineering, data science, information systems, or a related technical field, or equivalent combination of education and relevant experience. 3+ years of experience in machine learning engineering, MLOps, cloud engineering, data engineering, DevOps, or production analytics support. Practical experience working with AWS services used for machine learning or data workflows, such as Amazon S3, SageMaker, Lambda, Step Functions, CloudWatch, IAM, ECR, ECS, or related services. Strong Python skills and comfort working with scripts, APIs, logs, configuration files, and version-controlled repositories. Understanding of how machine learning models move from development into production, including model packaging, testing, deployment, monitoring, and support. Experience supporting batch processing, inference pipelines, data validation, or production data workflows. Familiarity with CI/CD concepts, source control, deployment coordination, and basic release management practices. Ability to troubleshoot issues across data, code, cloud services, permissions, and operational workflows. Ability to work across cross-functional teams and explain technical issues clearly to technical and business stakeholders. Strong analytical, problem-solving, documentation, and communication skills., Experience with computer vision, image-based analytics, inspection workflows, or large-scale image datasets. Experience with Docker, container-based deployments, or model packaging for production use. Exposure to infrastructure-as-code tools such as Terraform, CloudFormation, or AWS CDK. Experience with model monitoring, data quality checks, operational dashboards, or alerting workflows. Familiarity with ML lifecycle tools such as model registries, experiment tracking, or workflow orchestration. Experience in utility, infrastructure, industrial inspection, or similar analytics environments using image-based data for decision-making is a strong advantage.

About the company

The computer vision team develops machine learning solutions that convert aerial inspection imagery into actionable intelligence for PG&E. The team works cross-functionally across product, inspection, data science, machine learning engineering, cloud platform, and business stakeholders to deliver scalable analytics products that support safer operations, better asset visibility, and more informed decisions. The team combines practical model development, AWS-based deployment, structured change management, and production support discipline to move models from concept to operational use.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

3:35 min

Defining a serverless architecture using AWS CDK

Raphael Manke Raphael Manke · WWC 2023

5:28 min

Defining MLOps and its role in production systems

Hauke Brammer · WWC 2023

4:19 min

Introduction to DevOps for AI and MLOps

Aarno Aukia · LIVE

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · WWC Europe 2026

Videos

See all

Related articles

See all