SRE Engineer

Randstad
Washington, United States of America
6 days ago

Role details

Contract type
Contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English
Experience level
Intermediate
Compensation
$ 177K

Job location

Washington, United States of America

Tech stack

Amazon Web Services (AWS)
Systems Engineering
Azure
Bioinformatics
Cloud Computing
Cloud Engineering
Linux
DevOps
Github
Monitoring of Systems
Python
Metadata Standards
Network Architecture
NoSQL
Reliability Engineering
Ansible
Amazon Web Services (AWS)
Scripting (Bash/Python/Go/Ruby)
Performance Testing
Computer Network Technologies
System Availability
Delivery Pipeline
Cloudformation
Containerization
Kubernetes
Infrastructure Automation Frameworks
Information Technology
Deployment Automation
Terraform
Dynatrace
Docker
Jenkins
ServiceNow

Job description

job summary: Randstad is partnering with a premier client in the Washington, D.C. area to find a talented Site Reliability Engineer (SRE) to champion system availability, performance, and automation across their enterprise cloud infrastructure. In this role, you will bridge the gap between development and operations by implementing robust CI/CD pipelines and Infrastructure-as-Code (IaC), while heavily leveraging the Dynatrace observability platform to drive deep-dive distributed tracing, build intelligent dashboards, and tune anomaly detection. As a core member of the reliability team, you will apply formal SRE principles-such as defining SLIs/SLOs and managing error budgets-to optimize capacity and resiliency, ensure strict security compliance, and participate in an on-call rotation using ITIL frameworks to minimize incident response times.

location: Washington, Washington, D.C. job type: Contract salary: $75 - 85 per hour work hours: 9am to 5pm education: Bachelors

responsibilities: Observability & Monitoring: Standardize and automate Dynatrace installations, integrate telemetry collection into CI/CD pipelines, enforce tagging/metadata standards, configure distributed tracing with context propagation, and optimize custom dashboards and anomaly alerts.

Deployment & Automation: Design, implement, and maintain CI/CD pipelines using GitHub Actions, AWS CodePipeline, or Jenkins; provision scalable cloud infrastructure using Terraform, CloudFormation, or AWS CDK.

Incident & Problem Management: Serve as a production on-call responder using ITIL frameworks and ServiceNow; lead troubleshooting efforts, conduct deep root-cause analysis (RCA), and author comprehensive knowledge base articles.

Reliability Engineering: Champion SRE metrics including Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets; design and execute resiliency test plans and support performance testing.

Performance & Capacity Optimization: Drive operational cost optimization initiatives across cloud environments and configure robust auto-scaling policies and thresholds.

Security & Compliance: Manage service accounts, access permissions, and digital certificates; respond rapidly to security incidents and execute remediation protocols.

qualifications: Education & Experience: Bachelor's degree in Computer Science, Engineering, or a related technical field, paired with 2 to 4 years of hands-on experience in SRE, DevOps, or infrastructure-focused roles.

Cloud & Containerization: Practical, hands-on experience managing multi-tenant environments within AWS and Azure, alongside a solid understanding of container technologies like Docker, Kubernetes, and Amazon ECS.

Automation & Scripting: Mid-level proficiency in Python (or similar scripting languages) and practical experience with configuration management tools like Ansible to build automated self-service tools.

Systems & Networking Architecture: Strong foundational knowledge of Linux systems engineering, core networking concepts, and navigating relational, cloud-native, and NoSQL databases.

Professional Competencies: Excellent written and verbal communication skills for cross-functional collaboration, a proven ability to work independently, and the flexibility to participate in an on-call rotation outside standard business hours.

Equal Opportunity Employer: Race, Color, Religion, Sex, Sexual Orientation, Gender Identity, National Origin, Age, Genetic Information, Disability, Protected Veteran Status, or any other legally protected group status.

At Randstad Digital, we welcome people of all abilities and want to ensure that our hiring and interview process meets the needs of all applicants. If you require a reasonable accommodation to make your application or interview experience a great one, please contact HRsupport@randstadusa.com.

Pay offered to a successful candidate will be based on several factors including the candidate's education, work experience, work location, specific job duties, certifications, etc. In addition, Randstad Digital offers a comprehensive benefits package, including: medical, prescription, dental, vision, AD&D, and life insurance offerings, short-term disability, and a 401K plan (all benefits are based on eligibility).

This posting is open for thirty (30) days.

,

Observability & Monitoring: Standardize and automate Dynatrace installations, integrate telemetry collection into CI/CD pipelines, enforce tagging/metadata standards, configure distributed tracing with context propagation, and optimize custom dashboards and anomaly alerts.

Deployment & Automation: Design, implement, and maintain CI/CD pipelines using GitHub Actions, AWS CodePipeline, or Jenkins; provision scalable cloud infrastructure using Terraform, CloudFormation, or AWS CDK.

Incident & Problem Management: Serve as a production on-call responder using ITIL frameworks and ServiceNow; lead troubleshooting efforts, conduct deep root-cause analysis (RCA), and author comprehensive knowledge base articles.

Reliability Engineering: Champion SRE metrics including Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets; design and execute resiliency test plans and support performance testing.

Performance & Capacity Optimization: Drive operational cost optimization initiatives across cloud environments and configure robust auto-scaling policies and thresholds.

Security & Compliance: Manage service accounts, access permissions, and digital certificates; respond rapidly to security incidents and execute remediation protocols.

Requirements

Education & Experience: Bachelor's degree in Computer Science, Engineering, or a related technical field, paired with 2 to 4 years of hands-on experience in SRE, DevOps, or infrastructure-focused roles. Cloud & Containerization: Practical, hands-on experience managing multi-tenant environments within AWS and Azure, alongside a solid understanding of container technologies like Docker, Kubernetes, and Amazon ECS. Automation & Scripting: Mid-level proficiency in Python (or similar scripting languages) and practical experience with configuration management tools like Ansible to build automated self-service tools. Systems & Networking Architecture: Strong foundational knowledge of Linux systems engineering, core networking concepts, and navigating relational, cloud-native, and NoSQL databases. Professional Competencies: Excellent written and verbal communication skills for cross-functional collaboration, a proven ability to work independently, and the flexibility to participate in an on-call rotation outside standard business hours.

Apply for this position