Principal Site Reliability Engineer, Platform

Blue River Technology
Santa Clara, CA, United States
12 days ago
Apply on www.bluerivertechnology.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
6 years minimum
Compensation
$174,000.0 - $305,000.0
Working hours
Regular working hours

Tech stack

JavaScript (Programming Language) Artificial Intelligence Amazon Web Services Data Analysis Computer Vision Software as a Service Cloud Computing Continuous Integration Software Design Patterns Github Intrusion Detection and Prevention Python (Programming Language)
+16 more
Machine Learning Object-Oriented Software Development Reliability Engineering Robotic Automation Software Software Engineering Web Services Workflow Management Systems Rust (Programming Language) System Availability IT Architecture Backend Kubernetes SailPoint Terraform Artifactory Golang

Job description

Architects and owns highly available infrastructure and Kubernetes-based platforms supporting autonomous systems. Builds Golang backend services, platform tooling, observability systems, dashboards, alerts, and log aggregation. Partners with product teams to launch services, performs performance analysis, manages cloud upgrades, and participates in incident response and postmortems. Collaborates on cloud security risk assessments, intrusion detection, threat-feed systems, risk mitigation, and SaaS payment processes. Provides architectural leadership and mentorship across engineering teams. The summary above was generated by AI

We’re Blue River, a team of innovators driven to create intelligent machinery that solves monumental problems for our customers. We empower our customers - farmers, construction crews, and foresters - to implement safer and more sustainable solutions, driving increased profitability with less reliance on scarce labor. We believe that focusing on the small stuff - pixel-by-pixel and task-by-task - leads to big gains.

Blue River Technology aligns with John Deere’s vision to “innovate on behalf of humanity” by quickly identifying and solving high-value, high-uncertainty challenges in AI, machine learning, computer vision, and robotics. BRT acts as a research and development flywheel, building not only new products but also new platforms that reliably create value for both Deere and its customers. From fully autonomous machines to highly precise farming equipment, BRT and Deere are partnering to create technical breakthroughs in industries like agriculture and construction.

Our people are at the heart of what we do. Through cross-disciplinary collaboration, this mission-driven team is eager to define the new frontier of robotics. We are always asking hard questions, rapidly iterating, and getting our boots in the field to figure it out. We won’t give up until we’ve made a tangible and positive impact on the planet!, We are seeking a Principal Site Reliability Engineer to join the Platform organization, which accelerates company-wide adoption and scaling of automation and robotics. The Platform’s product is a set of API services and infrastructure designed to overcome scaling hurdles, such as operational complexity and system exceptions, thereby enabling the rapid launch and scaling of new autonomy innovations and products at Blue River. In this role, you will join a fun, fast-moving engineering team to drive architectural decisions, mentor engineers across teams, and shape our platform’s direction., A combination, not necessarily all-inclusive, of the following:

  • Architect, scale, and own essential infrastructure.
  • Build and maintain a Kubernetes-based platform supporting multiple teams and services.
  • Build backend services (Golang) to support autonomous systems.
  • Partner with product teams to launch new products on the platform.
  • Grow our high availability infrastructure while maintaining key metrics such as uptime.
  • Build tooling to support our platform and development teams.
  • Perform end-to-end performance analysis, identify areas for improvement, and implement robust solutions.
  • Work with cloud vendors and external technical support for upgrades and rapid problem resolution.
  • Participate in on-call rotation, triaging and resolving production incidents with thorough root cause analysis and postmortem documentation.
  • Design and maintain observability infrastructure, dashboards, alerts, and log aggregation to ensure visibility into platform health and service performance.
  • Collaborate with the security team to conduct regular risk assessments.
  • Maintain the risk register and develop and implement mitigation plans.
  • Assess intrusion detection alerts. Improve systems and services that digest threat feeds.
  • In collaboration with IT and purchasing teams, establish and maintain the payment process for each SaaS service., Artificial Intelligence * Big Data * Cloud * Machine Learning * Software * Business Intelligence * Data Privacy The role involves designing cloud infrastructure, managing production Kubernetes clusters, optimizing CI/CD pipelines, enhancing developer experience, and ensuring reliable AI workloads. Candidates should have extensive experience in infrastructure and distributed systems engineering with strong coding skills and cloud expertise. Top Skills: AWSAzureDatadogDockerElkGCPGoGrafanaJavaKubernetesPrometheusPythonTerraform Civic Roundtable

Requirements

  • Min. of 8 years of deep experience in building and maintaining infrastructure for data-intensive, high-availability applications, including six years building and maintaining public cloud solutions.
  • Deep understanding of cloud orchestration tools such as Kubernetes and Terraform.
  • Deep understanding of software design methodologies, information systems architecture, object-oriented design, and software design patterns.
  • Deep understanding of securing cloud infrastructure (preferably AWS and Kubernetes).
  • Deep experience in one or more of the following languages: Golang (preference), Python, JavaScript, Rust.
  • Deep experience in CI/CD tooling (GitHub Actions, ArgoCD, ArgoCD Image Updater, Artifactory).

Preferred Experience and Skills

  • You are interested in robotic applications and developing software that assists robots.
  • You are excited about robotics and the future of automation.
  • You are a self-starter with infectious enthusiasm, energy, and problem-solving abilities.

Benefits & conditions

At Blue River, your base pay is one part of your total compensation package. For this position, the reasonably expected pay range is between $174,000 - $305,000/year for the level at which this job has been scoped. Your base pay will depend on several factors, including your experience, qualifications, education, location, and skills. This position is also eligible for an annual performance bonus and a competitive benefit package. During the recruitment process, we may identify an alternative role or level to which you are more suited. If your ideal role at Blue River differs from the advertised position, we will provide an updated pay range as soon as possible during the hiring process., An Hour Ago In-Office or Remote 140K-170K Annually Senior level 140K-170K Annually Senior level Social Impact * Software Owns core product areas across the full lifecycle, from discovery and strategy through execution, launch, and iteration. Conducts customer research, workflow analysis, usability testing, competitive research, and data analysis. Partners closely with engineering and design to deliver scalable user experiences, defines success metrics and instrumentation, and supports launches through messaging and cross-functional readiness. The role also requires hands-on AI experience and operates in a high-ownership, low-structure B2B SaaS environment. Top Skills: AICRM SailPoint

Customer Success Manager

An Hour Ago Remote or Hybrid United States 60K-101K Annually Mid level 60K-101K Annually Mid level Artificial Intelligence * Cloud * Sales * Security * Software * Cybersecurity * Data Privacy Manage assigned client accounts to ensure satisfaction and renewals. Coach clients on SailPoint/IdentityIQ identity and access solutions, monitor usage and risks, provide strategic updates, identify expansion opportunities, and drive resolutions to customer issues. Top Skills: IdentityiqSailpoint

What you need to know about the Colorado Tech Scene

With a business-friendly climate and research universities like CU Boulder and Colorado State, Colorado has made a name for itself as a startup ecosystem. The state boasts a skilled workforce and high quality of life thanks to its affordable housing, vibrant cultural scene and unparalleled opportunities for outdoor recreation. Colorado is also home to the National Renewable Energy Laboratory, helping cement its status as a hub for renewable energy innovation.

Key Facts About Colorado Tech

  • Number of Tech Workers: 260,000; 8.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Lockheed Martin, Century Link, Comcast, BAE Systems, Level 3
  • Key Industries: Software, artificial intelligence, aerospace, e-commerce, fintech, healthtech
  • Funding Landscape: $4.9 billion in VC funding in 2024 (Pitchbook)
  • Notable Investors: Access Venture Partners, Ridgeline Ventures, Techstars, Blackhorn Ventures
  • Research Centers and Universities: Colorado School of Mines, University of Colorado Boulder, University of Denver, Colorado State University, Mesa Laboratory, Space Science Institute, National Center for Atmospheric Research, National Renewable Energy Laboratory, Gottlieb Institute

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.bluerivertechnology.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · World Congress 2023

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

1:52 min

Structuring and scaling the backend engineering team

Stefan Lingler Stefan Lingler +1 · Coffee With Developers

4:18 min

Prioritizing communication and structural awareness over strict tool mastery

Liam Hurrel +1 · World Congress 2021

2:40 min

Using GitHub primitives for internal documentation and corporate operations

Kyle Daigle · Coffee With Developers

Videos

See all

Related articles

See all