> Markdown version of [/jobs/ext/2756421-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2756421-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** Tesla, Inc. - **Location:** Fremont, CA, United States - **Experience:** Experienced - **Salary:** $140,000.0 - $210,000.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Applications Architecture, Software Applications, Databases, Continuous Delivery, Continuous Integration, DevOps, Middleware, Github, Python (Programming Language), Linux System Administration, Networking Basics, Routing, Reliability Engineering, Prometheus, Software Engineering, Systems Architecture, Virtual Local Area Networks, Virtual Machines, Virtualization Technology, Data Logging, Network Switches, Network Routing, Scripting, Load Balancing, Grafana, Firewalls (Computer Science), Git, Data Layers, Containerization, Kubernetes, Production Code, Splunk, Dynatrace, Docker - **Published:** September 6, 2026 - **Apply:** https://www.careerbuilder.com/job-details/site-reliability-engineer-factory-software-fremont-ca--9e49bb4d-6ceb-453c-900a-a73a0641c89f ## About the Role * 4+ years of experience in Site Reliability Engineering, Platform Engineering, DevOps, or a closely related systems role * Working knowledge of Kubernetes and hands-on experience with Docker/containerization or virtualization * Strong understanding of Observability concepts with practical experience using Prometheus, Grafana, Tempo, and/or Splunk * Expert-level Linux administration skills * Solid understanding of networking fundamentals (routing, switching, VLANs, firewalls, and load balancers) * Experience with Git and CI/CD pipelines (GitHub Actions is a plus) * Proficiency in at least one high-level language (Go, Python, or Java) with demonstrable experience writing production-grade code * Comfortable with on-call rotations and performing live troubleshooting during outages * Strong documentation habits and a track record of effectively sharing knowledge across teams * Strong bias for action - comfortable getting hands dirty, shipping solutions quickly, and learning from mistakes, Accidental Death and Dismemberment (AD&D), Administrative Skills, Best Practices, Budgeting, Compensation and Benefits, Consulting, Continuous Deployment/Delivery, Continuous Integration, Control Engineering, DevOps, Disability Insurance, Docker, Documentation, Employee Assistance Plan, Firewalls, Git, GitHub, Health Plan, Identify Issues, Incident Response, Infrastructure Software, Instrumentation, Java, Linux Administration, Load Balancing, Machine Tool, Manufacturing Execution Systems (MES), Manufacturing Management, Manufacturing/Industrial Processes, Mentoring, Metrics, Middleware, Network Routing, Network Switching, On Call, Orthodontics, Payroll Tax, Product Lifecycle, Programmable Logic Controller (PLC), Python Programming/Scripting Language, Reliability Engineering, Software Engineering, Splunk, Stock Purchase Plans, System Architecture, Technical Writing, Telemetry, VLAN (Virtual Local Area Network), VMS Operating System, Virtual Machine (VM), Virtualization, Vision Plan, Warehousing ## Description This role sits at the intersection of infrastructure (Kubernetes clusters, VMs, servers, and databases) and the software applications running on top of them. As an SRE on the Factory Software team, you will own the reliability of the full stack - from compute and data layers to the middleware connecting factory equipment (including PLCs) with MES systems and other services. You will implement advanced monitoring and observability to detect issues before they impact production. Your mission is to make both infrastructure and applications more reliable, observable, and standardized by catching speed bottlenecks, database contention, and infrastructure problems early while driving best practices and tooling consistency across teams. * Design, implement, and evolve end-to-end observability and telemetry across services and infrastructure, including OTEL instrumentation, logging, metrics, and distributed tracing * Build and maintain robust monitoring using Prometheus, Grafana, Tempo, and related tools to proactively detect speed bottlenecks, database contention, resource exhaustion, and infrastructure issues * Define and track SLIs, SLOs, and error budgets; standardize observability practices, golden signals, and tooling across engineering teams * Implement effective, low-noise alerting systems and drive strong incident response processes * Collaborate closely with Platform, Infrastructure, and Software Engineering teams to embed reliability and observability into the development lifecycle * Write production-grade code to reduce toil, automate operations, manage deployments, and treat infrastructure and reliability as a software engineering problem * Consult on infrastructure, systems, and application architecture with a reliability-first mindset * Participate in on-call rotations, live troubleshooting on NOC bridges/outage calls, and blameless post-mortems * Document solutions, create and maintain technical documentation, and actively mentor engineers across the organization ## Related Videos - [Remote Driving on Plant Grounds with State-of-the-Art Cloud Technologies](https://www.wearedevelopers.com/videos/251-remote-driving-on-plant-grounds-with-state-of-the-art-cloud-technologies) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Handling incidents collaboratively is like solving a rubix cube](https://www.wearedevelopers.com/videos/680-handling-incidents-collaboratively-is-like-solving-a-rubix-cube) - [Bringing AI Model Testing and Prompt Management to Your Codebase with GitHub Models](https://www.wearedevelopers.com/videos/1536-bringing-ai-model-testing-and-prompt-management-to-your-codebase-with-github-models) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [What is Software Engineering?](https://www.wearedevelopers.com/magazine/289-what-is-software-engineering) - [Highest Paying Tech Companies in Europe](https://www.wearedevelopers.com/magazine/162-highest-paying-tech-companies-in-europe)