Site Reliability Engineer

Optomi LLC
Orlando, FL, United States
10 days ago
Apply on www.thejobnetwork.com
Prepare application

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours

Tech stack

Java (Programming Language) Amazon Web Services Google App Engines Cloud Computing Cloud Computing Security Cloud Engineering Configuration Management Continuous Integration DevOps Elasticsearch Github Identity and Access Management
+25 more
Python (Programming Language) Key Management Network Load Balancing Node.Js Systems Development Life Cycle RabbitMQ Reliability Engineering Ansible Software Engineering Management of Software Versions Data Logging Google Cloud System Availability Grafana Firebase Infrastructure as Code (IaC) Gitlab-ci Kubernetes Information Technology Cloudwatch Terraform Splunk Appdynamics Legacy Systems Jenkins

Job description

Optomi, in partnership with our client, are seeking a Site Reliability Engineer for a 20+ month contract (W2 only no C2C/sponsorship)., We are seeking a highly motivated Site Reliability Engineer (SRE) to join a fast-paced engineering team focused on automation, platform reliability, infrastructure modernization, and operational excellence. This role is responsible for reducing operational toil, improving deployment processes, automating infrastructure management, and supporting the migration of legacy systems to modern cloud-native environments., * Design, implement, and maintain scalable, reliable cloud infrastructure and platform services.

  • Develop automation solutions to reduce manual operational work and improve engineering efficiency.
  • Support infrastructure modernization initiatives, including version upgrades and migration of legacy systems.
  • Build and maintain Infrastructure as Code (IaC) solutions using Terraform.
  • Manage and secure cloud environments, with a strong emphasis on identity, access management, and secrets management.
  • Collaborate with engineering teams to improve software delivery, deployment processes, and platform reliability.
  • Participate in troubleshooting, root cause analysis, and incident resolution efforts.
  • Ensure adherence to security, compliance, and operational best practices.
  • Document solutions, processes, and operational procedures to support knowledge sharing and scalability.

Requirements

The ideal candidate possesses strong technical expertise across cloud platforms, infrastructure as code, CI/CD pipelines, and application reliability, combined with exceptional ownership, communication, and problem-solving skills., * 5+ years of experience in Site Reliability Engineering, DevOps, Cloud Engineering, or a related field.

  • Strong Python development and automation skills.
  • Hands-on experience with Terraform, including version upgrades and infrastructure lifecycle management.

Experience working within AWS environments, specifically:

  • IAM (Identity and Access Management)
  • KMS (Key Management Service)
  • Secrets Manager
  • Understanding of cloud security policies and access controls
  • Experience building and supporting CI/CD pipelines using tools such as:
  • Harness
  • Jenkins
  • GitHub Actions
  • GitLab CI/CD
  • Strong understanding of Software Development Lifecycle (SDLC) practices.
  • Working knowledge of Node.js and Java.
  • Experience with one or more Google Cloud Platform services, including:
  • App Engine
  • Kubernetes
  • Cloud Functions
  • Firebase
  • IAM
  • Experience supporting application and network load balancers.
  • Excellent communication, documentation, and stakeholder collaboration skills.
  • Strong ownership mindset with proven problem-solving abilities.

Preferred Qualifications:

  • Experience with both Chef and Ansible, including configuration management and automation.
  • Experience supporting migrations from Chef to Ansible.
  • Advanced knowledge of Harness CI/CD implementations.
  • Experience with container orchestration technologies, including:
  • Kubernetes
  • Amazon ECS
  • Google App Engine

Experience with monitoring, logging, and observability tools such as:

  • CloudWatch
  • Splunk
  • AppDynamics
  • Grafana
  • Elasticsearch
  • Experience with Atlantis for Terraform automation.
  • Experience with messaging and event-driven technologies, including:
  • RabbitMQ
  • Google Pub/Sub

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.thejobnetwork.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · World Congress 2023

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

2:40 min

Using GitHub primitives for internal documentation and corporate operations

Kyle Daigle · Coffee With Developers

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

Videos

See all

Related articles

See all