TELECOMMUTE Capacity Engineer with Data Engineer

Mpower Plus Rezolve Ai Group Ltd
United States
2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Information Engineering Extract Transform Load (ETL) Load Testing Reliability Engineering Site Reliability Engineering Practices Ansible Prometheus Service Design Datadog Grafana Infrastructure Automation Frameworks Data Management
+2 more
Terraform Data Pipelines

Job description

· Data Pipeline Development: Design and maintain ETL/ELT pipelines to collect, transform, and store infrastructure usage data.

· Data Modeling: Build models to analyze system metrics and predict future resource needs.

· Demand Forecasting: Analyze historical usage patterns to predict CPU, memory, and storage requirements.

· Load Testing & Scaling: Simulate traffic spikes to identify bottlenecks and ensure systems scale linearly.

· Cost Efficiency: Optimize resource allocation to avoid unnecessary costs while maintaining service availability.

· Automation: Use Infrastructure as Code (IaC) tools like Terraform to automate scaling and provisioning.

· Architecture Review: Collaborate with software teams to flag single points of failure and ensure resilient service design.

Requirements

Data Engineer with strong Site Reliability Engineering (SRE) expertise in capacity planning. This role ensures our infrastructure scales efficiently to meet user demand, balancing performance with cost. The engineer will forecast growth, analyze usage trends, and automate resource provisioning to prevent outages, over-provisioning, or under-provisioning. In addition, the role requires building robust data pipelines and analytical models to support forecasting and decision-making., · Strong background in data engineering and SRE practices.

· Hands-on experience with capacity planning, forecasting, and scaling.

· Proficiency in IaC tools (Terraform, Ansible, Harness).

· Experience with data pipelines, ETL/ELT frameworks, and big data tools.

· Familiarity with monitoring/observability platforms (Prometheus, Grafana, Datadog).

· Knowledge of chaos engineering and resilience testing.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:42 min

Automating Skupper deployments using Ansible

Alex Soto Alex Soto · WWC 2024

1:36 min

Visualizing memory limits and isolating suspicious endpoints

Dina Matveev Dina Matveev · Europe 2026 Virtual

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

3:19 min

Executing complex workflows using Ansible Automation Platform

Goetz Rieger Goetz Rieger · WWC 2025

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

Videos

See all

Related articles

See all