Machine Learning Engineer/SRE

Georgia Tek Systems
Chicago, IL, United States
4 days ago

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours

Tech stack

Artificial Intelligence Microsoft Azure Cloud Computing Computer Programming Continuous Delivery Continuous Integration Data Cleansing DevOps Monitoring of Systems Python (Programming Language) Machine Learning Reliability Engineering
+16 more
Tensorflow Azure Machine Learning Virtual Machines Scripting Feature Engineering Pytorch Model Validation Containerization Infrastructure Automation Frameworks Network Support Data Management Machine Learning Operations Software Version Control Data Pipelines Docker Programming Languages

Job description

We are seeking a highly skilled and motivated Machine Learning Engineer who possesses expertise in developing, deploying, and managing machine learning models. In this role, you will be an integral part of our AI Engineering and Site Reliability Engineering (SRE) teams, responsible for managing Azure infrastructure for AI model development and deployment, monitoring and reporting model performance, and responding to outages/incidents related to model operations., Manage Azure Infrastructure: Configure, maintain, and optimize Azure infrastructure for AI model development and deployment, ensuring scalability and performance. Model Performance Monitoring: Implement and maintain monitoring systems to track model performance, proactively identifying and addressing issues as they arise. Incident Response: Collaborate with the SRE team to respond promptly to outages and incidents related to model operations, ensuring minimal downtime and rapid issue resolution.

Requirements

Azure Infrastructure Experience: Proficiency in managing Azure infrastructure components, including virtual machines, storage, and networking, to support AI model development and deployment. CI/CD Pipeline Experience: Experience with Continuous Integration/Continuous Deployment (CI/CD) pipelines, including the automation of model deployment processes. Containerization in the Cloud: Strong knowledge of containerization technologies in the cloud, such as Docker and Kubernetes, for efficient deployment and scaling of machine learning models. Machine Learning Expertise: Proficient in building and optimizing machine learning models, with a deep understanding of various Client algorithms and frameworks. Programming Skills: Proficiency in programming languages commonly used in machine learning, such as Python and libraries like TensorFlow and PyTorch. Data Management: Experience in data preprocessing, feature engineering, and data pipeline development for machine learning. Collaborative Team Player: Excellent communication skills and the ability to work collaboratively with cross-functional teams, including AI engineers and SREs. Documentation: Effective documentation skills to maintain clear and organized records of models, infrastructure configurations, and incident responses. Preferred Qualifications :

Experience with cloud-based machine learning platforms (e.g., Azure Machine Learning). Experience with CI/ CD tools to deploying Client services and applications specific to Azure cloud platform Familiarity with DevOps practices and tools for automating infrastructure and deployments. Knowledge of model versioning and model management tools. Understanding of security best practices in AI model deployment. Certifications in relevant areas, such as Azure certifications or machine learning certifications.

Job titles of folks with these skills may vary - e.g. MLOps Lead, MLOps Solution/Delivery Architect or Senior Client Engineer, Algorithms, Artificial Intelligence (AI), Automation, Best Practices, Cloud Computing, Communication Skills, Computer Programming, Continuous Deployment/Delivery, Continuous Integration, Cross-Functional, Customer Support/Service, Data Management, DevOps, Docker, Documentation, EAD, Identify Issues, Incident Response, Machine Learning, Microsoft Windows Azure, Network Support, Organizational Skills, Performance Analysis, Performance Modeling, Problem Solving Skills, Process Modeling, Programming Languages, Python Programming/Scripting Language, Reliability Engineering, Systems Maintenance, Team Player, Virtual Machine (VM)

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerbuilder.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · WWC 2023

2:08 min

Essential engineering roles in the generative AI space

Mary Grygleski Mary Grygleski · LIVE

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · WWC Europe 2026

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

Videos

See all

Related articles

See all