Cloud Reliability Solutions Engineer

The MathWorks, Inc.
United States
10 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours
Job source

Tech stack

Microsoft Windows Amazon Web Services Amazon Elastic Compute Cloud Amazon S3 Application Performance Management Microsoft Azure Backup Devices Ubuntu (Operating System) Cloud Computing Cloud Storage Continuous Integration Data Centers
+49 more
Linux Disaster Recovery Domain Name System (DNS) VMware ESX Servers Github Identity and Access Management Python (Programming Language) Key Management Linux Distribution Log Analysis Windows Servers Network Segmentation Network Virtualization Windows PowerShell Role-Based Access Control Red Hat Enterprise Linux Reliability Engineering Migration Manager Ansible Prometheus VMware VSphere Software Vulnerability Management YAML Datadog SSL Certificate Management Data Logging Curam Configuration Tools Google Cloud Cloud Platform System Cloud Monitoring System Availability Grafana Software Troubleshooting Multi-Cloud Mathworks HybridCloud Amazon Virtual Private Cloud (VPC) Gitlab Git Flow Kubernetes Vcenter Azure AKS Route53 Cloudwatch Terraform Splunk Docker Key Vault Vmware

Job description

MathWorks has a hybrid work model that enables staff members to split their time between office and home. The hybrid model provides the advantage of having both in-person time with colleagues and flexible at-home life optimizations. Learn More: ;br> MathWorks is a company built around problem solving, collaboration, and helping customers do their best technical work. The SSG Hosting team provides reliable, secure, and scalable infrastructure services that support teams across the company. We are looking for a friendly, curious, and technically strong Senior Cloud Reliability Solutions Engineer who enjoys designing practical solutions, automating repeatable work, and partnering with application, security, finance, and infrastructure teams to make our platforms easier to use and operate. In this role, you will help shape Azure-centered solutions that support workloads across Azure, AWS, Google Cloud Platform, and on-premises VMware environments, while advancing reliability engineering, automation, observability, and operational excellence practices.

MathWorks nurtures growth, appreciates inclusivity, encourages initiative, values teamwork, shares success, and rewards excellence.

Responsibilities

  • Design and deliver resilient hybrid-cloud solutions across Azure, AWS, Google Cloud Platform, and on-premises VMware environments, with deep emphasis on Azure architecture and operations.
  • Partner with internal customers to understand requirements, evaluate tradeoffs, estimate cost and effort, and recommend solutions that are reliable, secure, supportable, and aligned with business needs.
  • Build and improve infrastructure automation using Terraform, Packer, PowerShell, Python, Git-based workflows, and CI/CD pipelines to reduce manual work and improve consistency.
  • Apply reliability engineering practices including service ownership, operational readiness, incident response, root cause analysis, capacity planning, disaster recovery, and continuous improvement.
  • Create and maintain observability through useful dashboards, actionable alerts, telemetry standards, health checks, and reliability reporting across cloud and data center platforms.
  • Support secure and governed platform operations through RBAC, least privilege access, secrets management, policy enforcement, patching practices, vulnerability remediation, tagging, and cost visibility.
  • Contribute to Kubernetes and platform services including AKS, EKS, GKE, container networking, ingress, persistent storage, Helm, and GitOps-oriented operating models.
  • Bridge traditional and cloud-native infrastructure by supporting VMware, Windows, Linux, storage, networking, DNS, backup, and hybrid connectivity patterns.
  • Collaborate openly and constructively with peers and partner teams, sharing knowledge, documenting designs, mentoring others, and learning from the team.
  • Participate in production support including a rotating on-call schedule for infrastructure services and critical operational issues., * Azure platform expertise: Azure networking, compute, storage, identity, Azure Arc, Azure Monitor, Log Analytics, Azure Policy, RBAC, Azure Update Manager, Key Vault, Backup, and Site Recovery.

Requirements

  • A bachelor’s degree and 6 years of professional work experience (or a master’s degree and 3 years of professional work experience, or a PhD degree, or equivalent experience) is required., * Multi-cloud literacy: working knowledge of AWS services such as EC2, VPC, IAM, S3, EKS, CloudWatch, Systems Manager, Route 53, and Organizations; familiarity with Google Cloud Platform Compute Engine, VPC, IAM, Cloud Storage, GKE, and Cloud Logging/Monitoring.
  • Infrastructure as Code and automation: Terraform, Packer, Ansible or similar configuration tools, PowerShell, Python, YAML, GitHub or GitLab, and CI/CD pipeline practices.
  • Reliability engineering: SLA/SLO, error budgets, incident management, RCA, disaster recovery, high availability, capacity planning, and operational readiness reviews.
  • Containers and Kubernetes: Docker, Kubernetes, AKS, EKS, GKE, Helm, ingress controllers, container networking, persistent storage, secrets management, and cluster lifecycle practices.
  • VMware and enterprise infrastructure: vSphere, ESXi, vCenter, virtual networking, storage, backup, migration planning, and hybrid integration patterns.
  • Security and governance: identity and access management, least privilege, network segmentation, certificate management, vulnerability remediation, cloud policy, compliance reporting, and cost/tagging governance.
  • Observability tools: Azure Monitor, Log Analytics, Application Insights, Prometheus, Grafana, Splunk, Datadog, or similar monitoring and logging platforms.
  • Systems administration: Windows Server, Linux distributions such as Ubuntu, RHEL, or Rocky Linux, DNS, networking, patching, performance troubleshooting, and service management.
  • Communication and consulting skills: clear written documentation, design reviews, stakeholder engagement, influence without authority, and the ability to explain technical options to both engineering and business audiences.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

6:13 min

Defining cloud proficiency by technical role

Piet Van Dongen · LIVE

1:40 min

Managing containerized infrastructure with Podman Desktop

Cedric Clyburn Cedric Clyburn +1 · World Congress 2025

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

1:35 min

Centralizing configuration logic with native YAML block references

Matthieu Vincent Matthieu Vincent · Europe 2026 Virtual

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

1:41 min

Parallels between cloud and legacy infrastructure lock-ins

Björn Stahl Björn Stahl · World Congress 2024

Videos

See all

Related articles

See all