Cloud Engineer/Developer

Everforth Apex
Richmond, VA, United States
3 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
3 years minimum
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Agile Methodology Amazon Web Services Amazon Elastic Compute Cloud Amazon S3 Build Automation Automation of Tests Software Quality Code Review Information Systems Continuous Integration DevOps
+21 more
Identity and Access Management Python (Programming Language) Pattern Recognition Reliability Engineering Site Reliability Engineering Practices Software Engineering Software Modules Cloud Platform System Test-Driven Development (TDD) Grafana State Machines Amazon Virtual Private Cloud (VPC) Cloudformation Servicebus Information Technology Functional Programming Cloudwatch Terraform Software Version Control Golang Programming Languages

Job description

  • Design, develop, and maintain reliability solutions and SRE utilities using Python in AWS environments to reduce toil, improve cloud platform reliability, and industrialize SRE practices across the system
  • Build automation scripts, APIs, and utilities in Python to reduce toil and improve platform reliability.
  • Implement observability and monitoring solutions (Grafana, AWS CloudWatch) leveraging Python for custom metrics and dashboards.
  • Build and optimize Infrastructure as Code (IaC) using Terraform to manage AWS resources related to SRE solutions, incorporating cost-efficient design principles

o Optimize Infrastructure as Code (IaC) with Terraform for AWS resources, integrating Python-based workflows.

  • Develop CI/CD pipelines and automated testing to ensure code quality, reliability, and rapid delivery of the solutions
  • Define SRE standards, best practices, and guidelines for adoption across teams; establish SRE metrics like SLI, SLOs, etc.
  • Apply software engineering best practices including version control, code reviews, test-driven development, and documentation to all development

Participate in incident management and on-call rotation, providing technical support for SRE tools, troubleshooting production issues, and collaborating with teams to reduce incident recurrence through proactive detection and pattern analysis

Stay current with emerging AWS services, SRE methodologies, and cloud-native development technologies, and drive adoption of innovative solutions

  • Collaborate within Agile and Scaled Agile frameworks with cross-functional teams to deliver integrated cloud automation solutions
  • Produce clear, blameless postmortems with actionable items and documented failure scenarios

Requirements

  • Must have 5+ years of advanced Python development experience, building enterprise-grade, highly available tools, APIs, and utilities for AWS.
  • Bachelor’s degree in computer science, Information Systems, or equivalent background or equivalent experience
  • 7+ years of extensive experience in software development with focus on reliability and platform engineering
  • 3+ years of hands-on experience developing solutions in AWS environments with deep understanding of core services (EC2, VPC, S3, Lambda, IAM, CloudFormation, EventBridge, Step Functions etc.) and resource cost optimization
  • 3+ years of experience applying SRE principles including observability, toil automation, SLIs/SLOs and reliability engineering
  • Expert-level proficiency with Infrastructure as Code (IaC) using Terraform, including module development and state management
  • Strong experience with CI/CD pipelines, automated testing frameworks, and DevOps practices
  • Experience with observability tools and practices including Grafana, AWS CloudWatch, AWS Canary

Experience defining, implementing, and managing SLOs/SLIs and error budgets; familiarity with conducting RCAs and producing postmortem documentation

  • Working experience in Agile and Scaled Agile environments and familiarity with ITSM processes (incident, change, and problem management), resilience testing and chaos engineering practices

Experience: A minimum of 7 years of experience in software development with a focus on reliability and platform engineering is required. This includes over 3 years of hands-on experience developing solutions in AWS environments and applying SRE principles.

Technical Skills:

  • Advanced Python development skills (5+ years) are required.
  • Expert-level proficiency with Infrastructure as Code (IaC) using Terraform.
  • Strong experience with CI/CD pipelines, automated testing frameworks, and DevOps practices.
  • Experience with observability tools such as Grafana and AWS CloudWatch.
  • Deep understanding of core AWS services (e.g., EC2, VPC, S3, Lambda, IAM, CloudFormation, EventBridge, Step Functions).
  • Experience defining and managing SLOs/SLIs and familiarity with ITSM processes.

Preferred Qualifications

  • Experience with GoLang or additional programming languages is a plus.
  • Working experience in Agile and Scaled Agile environments.

About the company

Everforth Apex is a world-class IT services company that serves thousands of clients across the globe. When you join Everforth Apex, you become part of a team that values innovation, collaboration, and continuous learning. We offer quality career resources, training, certifications, development opportunities, and a comprehensive benefits package. Our commitment to excellence is reflected in many awards, including ClearlyRateds Best of Staffing in Talent Satisfaction in the United States and Great Place to Work in the United Kingdom and Mexico.

Everforth Apex uses a virtual recruiter as part of the application process. Click for more details. By applying for this job, you agree to receive calls, AI-generated calls, text messages, or emails from Everforth Apex and its affiliates, and contracted partners. Frequency varies for text messages. Message and data rates may apply. Carriers are not liable for delayed or undelivered messages. You can reply STOP to cancel and HELP for help. You can access our privacy policy at

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

6:16 min

Event-driven Golang backend architecture and cloud deployment

Irina Branovic Irina Branovic · World Congress 2026 Europe

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

Videos

See all

Related articles

See all