Site Reliability engineer MLops

Delviom LLC
Austin, TX, United States
6 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
6 years minimum
Working hours
Regular working hours
Job source

Tech stack

Testing (Software) Application Programming Interfaces (APIs) Airflow Amazon Web Services Automation of Tests Cloud Computing Databases Continuous Integration Github Python (Programming Language) Linux System Administration Machine Learning
+18 more
MongoDB Open Source Technology Cloud Services DataOps Software Engineering Apache Solr Management of Software Versions Circleci Google Cloud Large Language Models Containerization Gitlab-ci Kustomize Configuration Management Kubernetes Enterprise Integration Machine Learning Operations Code Restructuring Docker

Job description

  • Design and implement cloud solutions, build MLOps on cloud (AWS or Google Cloud Platform)
  • Build CI/CD pipelines orchestration by GitLab CI, GitHub Actions, Flux, Kustomize, Circle CI, Airflow or similar tools
  • Data science model containerization, deployment using docker, VLLM, Kubernetes
  • Data science model review, run the code refactoring and optimization, containerization, deployment, versioning, and monitoring of its quality
  • Data science models testing, validation and tests automation
  • Communicate with a team of data scientists, data engineers and architects, document the processes
  • Develop and deploy scalable tools and services for our clients to handle machine learning training and inference

Requirements

  • 6+ years of experience in ML Ops with strong knowledge in Kubernetes, Python, MongoDB and AWS.
  • Good understanding of Apache SOLR.
  • Proficient with Linux administration.
  • Knowledge of ML models and LLM.
  • Ability to understand tools used by data scientists and experience with software development and test automation
  • Ability to design and implement cloud solutions and ability to build MLOps pipelines on cloud solutions (AWS or Google Cloud Platform)
  • Experience working with cloud computing and database systems
  • Experience building custom integrations between cloud-based systems using APIs
  • Experience developing and maintaining ML systems built with open-source tools
  • Experience with MLOps Frameworks like Kubeflow, MLFlow, DataRobot, Airflow etc., experience with Docker and Kubernetes
  • Experience developing containers and Kubernetes in cloud computing environments
  • Familiarity with one or more data-oriented workflow orchestration frameworks (Kubeflow, Airflow, Argo, etc.)
  • Ability to translate business needs to technical requirements
  • Strong understanding of software testing, benchmarking, and continuous integration
  • Exposure to machine learning methodology and best practices
  • Good communication skills and ability to work in a team

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com

Good distractions

Talks and stories from around this role β€” technically off-topic, practically not.

2:00 min

Separating dataset creation from low-level software implementation steps

Jan Zawadzki Β· World Congress 2022

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters Β· World Congress 2023

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz Β· World Congress 2025

5:28 min

Defining MLOps and its role in production systems

Hauke Brammer Β· World Congress 2023

2:40 min

Using GitHub primitives for internal documentation and corporate operations

Kyle Daigle Β· Coffee With Developers

2:44 min

Defining core roles and responsibilities in MLOps teams

Bas Geerdink Β· LIVE

Videos

See all

Related articles

See all