DevOps / Site Reliability Engineer (SRE)

SilverSearch, Inc.
United States
3 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$135,200.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Amazon Web Services Application Release Automation Microsoft Azure Cloud Computing DevOps Distributed Systems Github Monitoring of Systems Python (Programming Language) Linux System Administration Machine Learning
+29 more
Performance Tuning Reliability Engineering Prometheus Software Engineering Software Vulnerability Management Datadog Pulumi Scripting Google Cloud Cloud Platform System Spring Cloud Istio Delivery Pipeline Large Language Models Grafana Cloudformation Containerization AI Platforms Gitlab-ci Kubernetes Infrastructure Automation Frameworks Deployment Automation Machine Learning Operations Terraform Splunk Devsecops Docker Jenkins Vulnerability Analysis

Job description

We are seeking a highly skilled DevOps / Site Reliability Engineer (SRE) with experience supporting modern AI platforms and cloud-native infrastructure. This role will focus on building scalable, reliable infrastructure for AI workloads while partnering closely with security and engineering teams to operationalize findings from an emerging AI-driven security platform used to identify code vulnerabilities and infrastructure risks.

This position is ideal for an engineer who enjoys automating infrastructure, improving software delivery pipelines, and supporting the rapid adoption of AI technologies in enterprise environments., * Design, build, and maintain highly available infrastructure supporting AI and machine learning platforms.

  • Develop scalable platform engineering solutions that enable reliable deployment and operation of AI services.
  • Partner with development and security teams to remediate vulnerabilities and infrastructure issues identified by the AI driven security platform.
  • Improve platform reliability through automation, monitoring, observability, and proactive performance tuning.
  • Build and maintain robust CI/CD pipelines for application and infrastructure deployments.
  • Automate operational workflows using Python and Infrastructure-as-Code practices.
  • Implement DevSecOps best practices throughout the software development lifecycle.
  • Support containerized workloads and cloud-native applications.
  • Troubleshoot production issues, perform root cause analysis, and implement long-term reliability improvements.
  • Optimize deployment strategies, release automation, and infrastructure scalability.
  • Collaborate with AI engineering teams to ensure AI services are secure, resilient, and production-ready.

Requirements

  • 5+ years of experience in DevOps, Site Reliability Engineering, or Platform Engineering
  • Strong experience designing and maintaining CI/CD pipelines (GitHub Actions, GitLab CI, Jenkins, Azure DevOps, or similar)
  • Strong Python scripting and automation skills
  • Experience supporting cloud infrastructure (AWS, Azure, or Google Cloud Platform)
  • Experience with Infrastructure as Code (Terraform, CloudFormation, or Pulumi)
  • Hands-on experience with Docker and Kubernetes
  • Strong understanding of Linux systems administration
  • Experience implementing monitoring and observability solutions (Prometheus, Grafana, Datadog, Splunk, etc.)
  • Experience working with security scanning tools and vulnerability remediation
  • Familiarity with DevSecOps principles and secure software delivery
  • Experience supporting AI platform engineering or machine learning infrastructure, * Understanding of AI model deployment, inference infrastructure, and scalability considerations
  • Familiarity with GPU-enabled infrastructure and AI compute environments
  • Knowledge of vector databases, LLM deployment, or MLOps concepts
  • Experience with Kubernetes operators, service mesh, or distributed systems
  • Exposure to AI security tooling
  • Experience integrating automated security scanning into CI/CD pipelines

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

1:55 min

Contrasting Terraform with Pulumi and cloud-specific tools

Devlin Duldulao · LIVE

2:53 min

Configuring dynamic proxy updates with Istio Pilot

Jan Mensch Jan Mensch · WWC Europe 2026

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · WWC 2023

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

Videos

See all

Related articles

See all