Senior Platform Engineer (Cloud Workloads)

Veeam Software Corporation
San Jose, CA, United States
about 2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$172,800.0 - $320,900.0
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services JIRA Microsoft Azure Bash Shell Software as a Service Continuous Integration Data Infrastructure Elasticsearch Github Python (Programming Language) Pattern Recognition Windows PowerShell
+16 more
Reliability Engineering Cloud Services Kusto Query Language Salesforce.Com Pulumi Scripting Cloud Platform System Infrastructure Automation Frameworks Bicep Cosmos DB Azure AKS Kibana Veeam Terraform Key Vault Servicenow

Job description

We are looking for a Senior Platform Engineer to join the Workload team within the Veeam R&D Department. You will own critical observability infrastructure, drive incident response maturity, and help scale proactive support capabilities as operational accountability., * Design, build, and maintain observability pipelines using the Elastic Stack (Elasticsearch, Kibana, Fleet) across Azure and AWS workloads

  • Develop and own SLO/SLI dashboards and error budget reporting for BaaS platform services
  • Respond to and lead incident response for distributed, multi-tenant cloud workloads; own runbook creation, maintenance, and continuous improvement
  • Build and refine proactive support tooling, including pattern analysis, tenant correlation dashboards, and baseline deviation alerting, to reduce reactive support burden
  • Manage and maintain Elastic Fleet agent policies, enrollment health, and log streaming pipelines across Azure and AWS worker fleets
  • Partner with SRE, R&D, and Proactive Support teams to close observability gaps, including tenant identification workflows and admin portal integrations

Technologies we work with

  • Elastic Stack - Elasticsearch, Kibana, Elastic Fleet, KQL, Query DSL
  • Azure Kubernetes Service (AKS), Azure Container Apps, VMs
  • Azure Security - Entra ID, Managed Identities (user/system assigned), App Registrations, Key Vault
  • Infrastructure as Code - Azure Bicep, Terraform, or Pulumi
  • CI/CD - Azure DevOps, GitHub Actions
  • ITSM tooling - ServiceNow, Salesforce, Jira, Incident.io (for tenant and incident workflows)

Requirements

  • 5+ years of experience in cloud platform engineering, SRE, or infrastructure roles supporting commercial SaaS products
  • Deep hands-on experience with Elastic Stack: Building dashboards, writing KQL/Query DSL, managing Fleet
  • Proven experience operating and troubleshooting distributed, multi-tenant workloads on Azure and/or AWS
  • Strong understanding of Azure cloud services: AKS, Entra ID, Key Vault, Service Bus, Cosmos DB, Private Endpoints, etc.
  • Experience with incident response in production cloud environments, including runbook development and post-incident review
  • Experience with IaC tools (Azure Bicep, Terraform) and CI/CD pipelines (Azure DevOps, GitHub Actions)
  • Strong scripting skills in Bash, Python, or PowerShell
  • Ability to work cross-functionally with SRE, product, and customer-facing support teams

About the company

Veeam is the Data and AI Trust Company, specializing in helping organizations ensure their data and AI are fully understood, secured, and resilient to enable the acceleration of safe AI at scale. As the market leader in both data resilience and data security posture management, Veeam is built for the convergence of identity, data, security, and AI risk. Headquartered in Seattle with offices in more than 30 countries, Veeam protects over 550,000 customers worldwide, who trust Veeam to keep their businesses running. Join us as we go fearlessly forward together, growing, learning, and making a real impact for some of the world’s biggest brands.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

2:56 min

Provisioning a secure container infrastructure with Bicep

Matthias Falkenberg +1 · WWC 2022

3:05 min

Integrating an assistant application with Jira software

Felix Augenstein · LIVE

4:09 min

Selecting infrastructure tools and determining proper abstraction layers

Alayshia Knighten Alayshia Knighten · WWC 2024

5:47 min

Integrating user stories and test automation via Jira tools

Christoph Ruggenthaler · LIVE

2:14 min

Solving complex platform architecture challenges at an enterprise scale

Maria Apazoglou · Coffee With Developers

Videos

See all

Related articles

See all