Site Reliability Engineer

Intertech, Inc.
United States
10 days ago
Apply on www.intertech.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
4 years minimum
Working hours
Regular working hours

Tech stack

Algorithm Design Microsoft Azure Cloud Computing Data Structures DevOps Elasticsearch Monitoring of Systems Identity and Access Management Internet Protocol Microsoft Visual Studio Open Web Application Security Public Key Infrastructure
+22 more
Reliability Engineering Cloud Services Logstash Prometheus Software Engineering Systems Integration Data Logging Pulumi Cloud Platform System Cloud Monitoring Git Kubernetes Infrastructure Automation Frameworks Information Technology Bare Metal Kibana Restful APIs Terraform Splunk Software Version Control Dynatrace Elk Stack

Job description

The Sr Site Reliability Engineer will architect, develop, and maintain cloud environment in both the commercial and government cloud. The role will work closely with software engineers, architects, and DevOps engineers to architect and maintain a secure, resilient and high performance cloud infrastructure., * Build, maintain, and operate IaaS and PaaS infrastructure in Azure commercial and government clouds

  • Work closely with dev teams to identify and measure SLOs, SLAs and SLIs
  • Act a strong contributor to development of platform services including architecture, provisioning, configuration, deployment, and support
  • Perform integrations with central logging, metrics dashboards, instrumentation, incident monitoring and management
  • Build/integrate/administer systems and tools that enable engineering teams to observe their applications in production with autonomy (Dashboards, APMs).
  • Support software and/or cloud-infrastructure in an on-call rotation basis
  • Assist with identification and remediation of technical problems at the root cause by continuously implementing automation, self-healing, and real-time monitoring to production systems
  • Maintain and improve operational tooling, frameworks, build frameworks that test the performance and resiliency of our platform services/tools
  • Automate alerts for metrics on performance, cost, vulnerabilities, risk, compliance violations
  • Improve processes and champion automation of any manual items around support.

Requirements

  • 4 + years of experience working within a SRE engineer/cloud platform role
  • Experience leveraging AI tools in the software development (or product) lifecycle in order to improve quality and efficiency
  • Expert knowledge of a cloud service provider
  • Expert knowledge and hands on production experience in Kubernetes (bare metal or managed) cluster setup and management required.
  • Experience with infrastructure as code (IaC) tools like Terraform, Pulumi.
  • Experience with Kubernetes deployment tools like Helm, ArgoCD, Flux
  • Experience with monitoring tools (Dynatrace, Azure Monitor, Splunk, Graphana, Prometheus).
  • Experience with the Elastic Stack or ELK stack (Elasticsearch, Logstash, and Kibana)
  • Strong awareness of networking and internet protocols.
  • Understanding of identity and access management (IAM)
  • Experience supporting infrastructure in production cloud environments.
  • Knowledge of Encryption, Public Key Infrastructure (PKI), understanding of OWASP
  • Experience working with RESTful services
  • Familiarity with IDEs and Source Control tools like Visual Studio Code and Git.

Preferences:

  • Bachelor’s Degree in Computer Science, Information Technology, Software Engineering, Math, Physics
  • Master’s Degree with coursework focused on advanced algorithms, mathematics in computing, data structures or related field
  • Expert knowledge of Azure
  • Demonstrate passion about infrastructure automation
  • Ability to prioritize work in a fast-paced environment.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.intertech.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

1:55 min

Contrasting Terraform with Pulumi and cloud-specific tools

Devlin Duldulao · LIVE

3:21 min

Deploying a primary Elasticsearch and Kibana cluster configuration

Philipp Krenn · World Congress 2022

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

56 sec

Favorite git commands and the importance of patch commits

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

Videos

See all

Related articles

See all