Sr. Manufacturing Site Reliability Engineer, Windows Platform

Tesla Motors
Austin, TX, United States
about 1 month ago
Apply on diversityjobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Working hours
Regular working hours

Tech stack

Microsoft Windows Active Directory JIRA Build Automation C Sharp (Programming Language) Configuration Management Cyber Security Linux Firmware Python (Programming Language) System Center Configuration Manager Windows API
+18 more
Windows Servers NT File System (NTFS) Windows PowerShell Reliability Engineering Ansible Prometheus OPC Unified Architecture Windows Software Trace Preprocessor Datadog Microsoft Deployment Toolkit Grafana Microsoft InTune Puppet Splunk Pagerduty Jenkins Artifactory Golang

Job description

Our Manufacturing SRE team owns the underlying platforms that keep production lines moving across all manufacturing sites. The fleet spans tens of thousands of Linux and Windows hosts driving station controllers, Industrial PCs (IPCs), camera systems, robot controllers, and label printers across every shop on every line.

The Controls organization is shifting more workloads back to Windows, and a major new robotics program is on track to run more Windows computers than Linux in production. Today, one engineer carries the Windows specialty for the entire team. This role exists to change that - to scale the depth of Windows expertise on the team so the platform can absorb the robotics rollout, IPC fleet growth, and a steady increase in Windows software complexity per host without losing reliability.

You partner with Controls Engineering, IT Manufacturing Operations, and the robotics team to make sure Windows hosts boot, image, monitor, recover, and patch the same way Linux hosts already do, and you contribute back to the broader SRE platform (imaging pipeline, host inventory, observability stack) so the wins compound. What You’ll Do

  • Own the Windows production fleet end to end: Industrial PCs, test benches, station HMIs, robotics dyno PCs, and the long tail of factory Windows hosts across all sites
  • Extend GPOSentry, our internal GPO drift-detection service, to catch and auto-remediate policy drift across the Windows fleet before it reaches the floor
  • Drive Windows imaging through a custom PXE pipeline (Ansible, Jenkins, Artifactory) so a new IPC can boot, join the domain, and report green telemetry without manual intervention
  • Extend the Windows host inventory and management service to cover new platforms and new sites; co-maintain Windows LAPS rotation, AD lifecycle, and the endpoint management agent runtime
  • Push the Windows fleet onto the same observability surface as Linux: Grafana Alloy and Prometheus collecting metrics, structured logs into Splunk, alerting in Opsgenie or Jira Service Management
  • Build automation for the Windows-specific operational pain that does not exist on Linux: GPO drift, driver and firmware management, Windows Update windows, NTFS permission audits, time zone enforcement across sites, certificate rotation for OPC-UA
  • Roll out and sustain the SentinelOne agent across factory Windows endpoints, partnering with Infosec on detections and exclusions tuned for production hardware
  • Own the deployment story for Windows-resident in-house factory applications (print management, label printing services, camera/footage systems, station controllers) when they touch Windows boxes
  • Carry production on-call rotation for Windows incidents: triage P1 line-down events, write the runbook, file the Jira, drive the post-mortem, and turn the fix into automation
  • Contribute to the cross-platform fleet tooling the team owns (imaging, inventory, observability, Windows configuration management); submit PRs in Go, Python, PowerShell, or Ansible as the work demands

Requirements

  • 5+ years operating Windows in production at scale, including Active Directory, GPO, Windows Server, and Windows 10 or 11 LTSC on industrial hardware
  • Strong PowerShell skills: scripts that hit the Win32 API, parse event logs, drive WMI, and integrate with REST endpoints, not one-liners
  • Hands-on experience with at least one configuration management or imaging platform: Ansible, Intune, SCCM, MDT, Puppet, or equivalent custom PXE work
  • SRE practice fundamentals: SLO design, alert hygiene, runbook discipline, blameless post-mortem authorship, error-budget thinking
  • Working knowledge of Linux as a peer platform; you do not need to be a kernel hacker, but you can read a systemd unit, write a bash one-liner, and submit a clean Ansible PR
  • Comfort writing application code in at least one of Python, Go, or C#, enough to ship a small service or extend an existing one
  • Production experience with observability tooling: Prometheus or Grafana, Splunk or equivalent log platform, OpenTelemetry concepts
  • A bias for automating away repeated work, even at the cost of more upfront engineering effort
  • Direct, low-ego communication style; comfort working asynchronously across multiple sites and time zones

Benefits & conditions

Along with competitive pay, as a full-time Tesla employee, you are eligible for the following benefits at day 1 of hire:

  • Medical plans > plan options with $0 payroll deduction
  • Family-building, fertility, adoption and surrogacy benefits
  • Dental (including orthodontic coverage) and vision plans, both have options with a $0 paycheck contribution
  • Company Paid (Health Savings Accounts) HSA Contribution when enrolled in the High-Deductible medical plan with HSA
  • Healthcare and Dependent Care Flexible Spending Accounts (FSA)
  • 401(k) with employer match, Employee Stock Purchase Plans, and other financial benefits
  • Company paid Basic Life, AD&D
  • Short-term and long-term disability insurance (90 day waiting period)
  • Employee Assistance Program
  • Sick and Vacation time (Flex time for salary positions, Accrued hours for Hourly positions), and Paid Holidays
  • Back-up childcare and parenting support resources
  • Voluntary benefits to include: critical illness, hospital indemnity, accident insurance, theft & legal services, and pet insurance
  • Weight Loss and Tobacco Cessation Programs
  • Tesla Babies program
  • Commuter benefits
  • Employee discounts and perks program

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on diversityjobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:46 min

Introduction to the speaker and engineering background

Llywelyn Griffith-Swain · World Congress 2023

3:05 min

Integrating an assistant application with Jira software

Felix Augenstein · LIVE

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

43 sec

Software engineering journey and local Manchester roots

Jonathan Tang · Coffee With Developers

2:11 min

Updating the delivery architecture with Jira and Tekton pipelines

Lian Li · World Congress 2022

Videos

See all

Related articles

See all