Staff Software Engineer, Devops

TEAM Inc.
United States
2 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
10 years minimum
Compensation
$155,854.0 - $233,781.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Amazon Web Services Cloud Computing Continuous Integration Data as a Services Information Engineering Software Design Documents DevOps Domain Name System (DNS) Github Identity and Access Management Virtual Private Networks (VPN)
+28 more
Python (Programming Language) Key Management Machine Learning Routing Octopus Deploy Role-Based Access Control Reliability Engineering Prometheus TypeScript Virtualization Technology Policy as Code AWS Cdk Enterprise Software Applications Load Balancing Istio Delivery Pipeline Grafana Amazon Virtual Private Cloud (VPC) Backend Kubernetes Information Technology Rancher Github Enterprise Bare Metal Cloudwatch Data Pipelines Serverless Computing Vmware

Job description

  • Set the Direction for Our AWS Foundation: Evolve our multi-account AWS environment, including account governance and guardrails, identity and access management, and the networking and DNS that connect the cloud to our sites. Decide where new workloads live and how they connect.
  • Build the Reliability Practice: Establish our on-call rotation, SLOs and error budgets, alerting standards, incident response and blameless postmortems. Stand up the observability (metrics, logs, traces and synthetic checks) that tells us something is wrong before our users do.
  • Enable the Teams We Serve: Partner with our Digital teams (customer-facing website and backend services), Data Engineering, Manufacturing Systems and enterprise application teams. Give them well-supported paths: CI/CD templates, deployment patterns, reusable infrastructure libraries and self-service environments.
  • Plan for What’s Next: Shape how the platform supports manufacturing and plant-floor workloads, data pipelines, and emerging AI and agentic workloads, whether each one runs on-premises or in the cloud.
  • Build In Security and Cost Discipline: Make least-privilege access, policy guardrails, secrets management and software supply chain controls the default. Keep cloud spend visible and deliberate through tagging, right-sizing and commitment planning.
  • Ship and Mentor: Write production infrastructure and code every week. Lead design reviews, mentor engineers on the team and across the organization, and write the documentation and runbooks that let others move without waiting on you., At Slate, we’re fueled by grit, determination, and attention to detail. The start-up spirit of ingenuity and resourcefulness move our business forward. Team Slate fosters a culture of excellence, innovation, and mutual respect, and is motivated by shared principles.
  • Safety First
  • Delight Customers
  • One Team
  • Relentless Improvement
  • Fast, Frugal, and Scrappy
  • Respectful Collaboration
  • Positive Legacy

Requirements

  • 10+ years in DevOps, site reliability, platform or infrastructure engineering, including several years setting technical direction for production platforms.
  • Deep AWS Expertise: Extensive hands-on experience designing and operating production AWS environments across many accounts, including multi-account governance, identity and access management, VPC networking and hybrid connectivity, DNS, EKS, containers, serverless and managed data services. An AWS Professional or Specialty certification is a plus.
  • Kubernetes in Production, in the Cloud and On-Premises: Proven experience running Kubernetes in production on EKS and on self-managed clusters built on VMs or bare metal, including cluster lifecycle, CNI networking, ingress, storage, RBAC and upgrades. Experience with Rancher or a comparable multi-cluster management platform is a strong plus.
  • Python and Infrastructure as Code: Strong experience with infrastructure as code; AWS CDK is strongly preferred. Our platform code is written in Python, so strong Python skills are required. You can read, understand and make targeted changes to TypeScript, which some of our application teams use.
  • CI/CD and GitOps: Experience building delivery pipelines with GitHub Actions or similar tools, and deploying to Kubernetes with GitOps tooling (Argo CD or Flux), Helm and Kustomize.
  • Reliability Engineering: A track record of establishing SLOs, observability (Prometheus, Grafana, OpenTelemetry, CloudWatch or similar), incident management and on-call practices, and of participating in on-call yourself.
  • Hybrid Networking: A solid grasp of networking across cloud and on-premises environments: routing, VPN, DNS, load balancing, certificates and firewalls, and how each of them fails.
  • Technical Leadership Without Authority: Ability to drive architectural decisions across teams you don’t manage, write clear design documents, and bring skeptical stakeholders along.
  • Startup Orientation: Comfort in a fast-moving, resource-constrained environment. You know how to ship, how to cut scope without cutting corners, and how to improve a system while it’s running.
  • Bachelor’s degree in Computer Science, Engineering or a related field, or equivalent practical experience.

Nice to have

  • Exposure to manufacturing, industrial or OT environments and the constraints of running software near the plant floor.
  • Experience running GPU, machine learning or AI and agentic workloads on Kubernetes.
  • Familiarity with enterprise virtualization platforms such as VMware.
  • Policy as code (OPA Gatekeeper or Kyverno), service mesh, or GitHub Enterprise administration.

WORK AUTHORIZATION REQUIREMENTApplicants must be authorized to work in the United States on a permanent basis. We are unable to offer visa sponsorship at this time.

About the company

At Slate, we’re building safe, reliable vehicles that people can afford, personalize and love-and doing it here in the USA as part of our commitment to reindustrialization. The spirit of DIY and customization runs throughout every element of a Slate, because people should have control over how their trucks look, feel, and represent them., Slate is looking for a Staff DevOps Engineer to set the technical direction for how we build, run and operate our infrastructure, and to build a large share of it yourself. We don’t separate DevOps from SRE. The team that builds the platform also runs it, carries the pager, and is accountable for its reliability.

Our AWS foundation is in place: a governed multi-account environment, private connectivity between the cloud and our sites, and infrastructure as code throughout. Next is a hybrid Kubernetes platform that spans Amazon EKS and self-managed clusters in our own facilities, serving our digital, data, manufacturing and enterprise teams. You will design that platform and help ship it.

This is a hands-on role. You’ll lead through architecture, code and example rather than through direct reports, and you’ll take your turn on call like everyone else on the team. You will report to the leader of the DevOps team and work closely with Digital, Data Engineering, IT, Manufacturing Systems and Security.

  • The Hands-On Architect: You design systems and then build them. You’re as comfortable writing a design doc as you are writing a CDK construct in Python or a Helm chart, and you can tell a platform people choose to use from one they are forced to use.
  • The Reliability Owner: You think in SLOs and failure modes. You’ve been paged at 2 a.m., you’ve written the postmortem, and you’ve made sure the same thing didn’t page anyone again.
  • The Hybrid Thinker: You know when a workload belongs in the cloud and when it belongs next to the machines it talks to. You weigh latency, cost, data gravity and operational burden, and you can explain the trade-off to people who don’t work in infrastructure.
  • The Force Multiplier: You’d rather build a paved road than approve every request by hand. Other teams ship faster because of what you build and document.

What you get to do

  • Architect the Hybrid Kubernetes Platform: Define and build Slate’s Kubernetes ecosystem across Amazon EKS and self-managed clusters running on virtualized infrastructure provided by our IT team. Own the reference architecture for cluster provisioning, multi-cluster management (such as Rancher), networking, identity, secrets, storage and upgrades.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Loading talks and stories from around this role…