Platform Infrastructure SRE / Software Engineer

Bayside Solutions
Cupertino, CA, United States
21 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Compensation
$124,800.0 - $145,600.0
Working hours
Regular working hours
Job source

Tech stack

User Authentication Cloud Computing Cloud Computing Security Digital Architecture Distributed Data Store Distributed Systems Domain Name System (DNS) Routing Reliability Engineering Data Logging Pulumi Google Cloud
+11 more
Load Balancing Istio Apache Spark Multi-Cloud Amazon Virtual Private Cloud (VPC) Cloudformation Kubernetes Infrastructure Automation Frameworks Apache Flink Deployment Automation Terraform

Job description

We are looking for a strong Senior Platform Infrastructure SRE / Software Engineer to support the productionization and operation of large-scale Kubernetes-based platform services across multiple cloud environments. This role is best suited for someone with deep Kubernetes and infrastructure experience who can take capabilities developed by a platform engineering team and make them repeatable, scalable, observable, reliable, and production-ready across many environments., * Productionize Kubernetes-based platform services developed by platform engineering teams.

  • Deploy and operate platform infrastructure across multiple cloud and production environments.
  • Build repeatable environment provisioning and deployment automation.
  • Develop reusable infrastructure templates, blueprints, and deployment patterns.
  • Provision Kubernetes clusters, cloud infrastructure, networking, and service dependencies.
  • Configure environment-specific infrastructure, connectivity, and platform services.
  • Deploy, validate, upgrade, and maintain platform services throughout their lifecycle.
  • Build monitoring, alerting, dashboards, logging, health checks, and operational controls.
  • Establish and validate production-readiness standards for new platform capabilities.
  • Troubleshoot complex failures across applications, Kubernetes, infrastructure, networking, and distributed systems.
  • Work closely with platform developers to understand application behavior and identify operational gaps before production rollout.
  • Implement and maintain Infrastructure as Code using technologies such as Crossplane, Terraform, Pulumi, or CloudFormation.
  • Build and maintain Helm-based Kubernetes packaging and deployment patterns.
  • Design safe rollout, rollback, upgrade, recovery, and lifecycle-management processes.
  • Validate platform capacity, availability, scalability, and reliability.
  • Configure and troubleshoot DNS, load balancing, VPC networking, routing, service connectivity, security policies, and certificates.
  • Improve operational automation and reduce manual environment-specific work.
  • Document architecture, deployment patterns, operational procedures, and troubleshooting guidance.
  • Participate in production support, incident response, and root-cause analysis as appropriate.
  • Independently own technical work and drive complex problems through resolution with limited supervision.

Requirements

  • Deep experience with Crossplane for infrastructure provisioning and platform automation.
  • Experience with Alibaba Cloud and its Kubernetes, networking, and infrastructure services.
  • Experience with AWS EKS and/or Google Cloud Platform.
  • Experience designing or operating multi-cloud platforms.
  • Experience with service mesh technologies and Kubernetes service networking.
  • Experience with distributed data technologies such as Apache Spark, Apache Flink, or Trino.
  • Experience implementing authentication, authorization, cloud security, and governance controls.
  • Experience building and maintaining CI/CD pipelines for Kubernetes-based platforms.
  • Experience designing highly available and resilient platform architectures.
  • Experience automating provisioning and lifecycle management across a large number of environments.
  • Strong production SRE, incident response, reliability engineering, and operational automation experience.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

2:04 min

Enhancing network privacy with routing fees and onion routing

Andreas M Antonopoulos · LIVE

1:55 min

Contrasting Terraform with Pulumi and cloud-specific tools

Devlin Duldulao · LIVE

2:53 min

Configuring dynamic proxy updates with Istio Pilot

Jan Mensch Jan Mensch · World Congress 2026 Europe

2:14 min

Solving complex platform architecture challenges at an enterprise scale

Maria Apazoglou · Coffee With Developers

1:51 min

Overview of the three Google Maps routing applications

Germán Álvarez · LIVE

Videos

See all

Related articles

See all