Machine Learning Engineer (SRE Focus)

ConnectedX, Inc.
Plano, TX, United States
9 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Microsoft Windows Amazon Web Services Software Debugging Linux DevOps Document Management Systems Distributed Systems Monitoring of Systems Python (Programming Language) Linux System Administration Linux Servers Machine Learning
+6 more
Reliability Engineering Datadog Reliability of Systems Kubernetes Network Server Docker

Job description

Experienced Machine Learning Engineer with a strong Site Reliability Engineering (SRE) mindset to join our team. The candidate will have hands-on experience maintaining applications on both Windows and Linux environments, managing on-premises servers, and working with Kubernetes clusters. This role requires solid Python programming skills, a good understanding of machine learning concepts, and practical knowledge of ML model deployment, monitoring, and debugging., * Maintain and support machine learning applications running on Windows and Linux servers in on-premises environments.

  • Manage and troubleshoot Kubernetes clusters hosting ML workloads.
  • Collaborate with data scientists and engineers to deploy machine learning models reliably and efficiently.
  • Implement and maintain monitoring and alerting solutions using DataDog to ensure system health and performance.
  • Debug and resolve issues in production environments using Python and monitoring tools.
  • Automate operational tasks to improve system reliability and scalability.
  • Ensure best practices in security, performance, and availability for ML applications.
  • Document system architecture, deployment processes, and troubleshooting guides.

Requirements

  • Proven experience working with Windows and Linux operating systems in production environments.
  • Hands-on experience managing on-premises servers and Kubernetes clusters and Docker containers
  • Strong proficiency in Python programming.
  • Solid understanding of machine learning concepts and workflows.
  • Experience with machine learning model deployment and lifecycle management.
  • Familiarity with monitoring and debugging tools, e.g. DataDog.
  • Ability to troubleshoot complex issues in distributed systems.
  • Experience with CI/CD pipelines for ML applications.
  • Familiarity with AWS cloud platforms
  • Background in Site Reliability Engineering or DevOps practices.
  • Strong problem-solving skills and attention to detail.
  • Excellent communication and collaboration skills.
  • We need an engineer who is also familiar with model development

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:01 min

Executing remote data exploration and model training

Mingshen Sun Mingshen Sun · World Congress 2024

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

Videos

See all

Related articles

See all