> Markdown version of [/jobs/ext/2603264-senior-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2603264-senior-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Site Reliability Engineer - **Company:** Ford Motor Company - **Location:** Palo Alto, CA, United States - **Experience:** Expert - **Salary:** $150,200.0 - $283,500.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Agile Methodology, Amazon Web Services, Application Performance Management, Application Services, BigQuery, Software as a Service, Cloud Computing, Computer Engineering, Continuous Integration, DevOps, Middleware, Github, PostgreSQL, Enterprise Messaging Systems, Octopus Deploy, Reliability Engineering, Prometheus, Software Engineering, SonarQube, Alwayson, Datadog, Performance Testing, Delivery Pipeline, Grafana, Spring-boot, Build Server, Backend, Kubernetes, Information Technology, Deployment Automation, Data Analytics, Apache Kafka, Data Management, Database Replication, Restful APIs, Terraform, Grpc, Dynatrace, Serverless Computing, Jenkins, Microservices - **Published:** August 11, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=e9002ba35285ed13 ## About the Role * Bachelor's degree in Computer Science, Electrical/Computer Engineering, or related field (or equivalent experience) * 5+ years of progressive experience in cloud-based software development. * 5+ years of experience designing, deploying, and supporting cloud-based solutions in production environments. * 5+ years of experience supporting mission-critical, always-on applications with high reliability, availability, and performance requirements. Even better, you will have... * 5+ years of hands-on expertise with GCP or AWS and cloud-native services including Pub/Sub, MSK, GCS, BigQuery, and container orchestration platforms such as GKE, EKS, or Kubernetes. * Strong expertise in messaging technologies: Kafka, Kafka Connect, Pub/Sub, and related data replication/management tooling. * Strong expertise in observability technologies: OpenTelemetry, Vector, Prometheus, VictoriaMetrics, Grafana, Dynatrace. * Experience with infrastructure as code a d deployment automation using Terraform, as well as CI/CD tools such as Cloud Build, ArgoCD, and Tekton. * Experience operating and supporting critical, shared middleware infrastructure in 24x7, "always-on" production environments used by multiple downstream teams. * Experience developing backend applications/services (beyond middleware/infra) is a plus, given occasional need to support broader app development efforts. ## Description We are seeking a highly skilled Senior Software Engineer to join our Messaging & Observability team, responsible for the mission-critical middleware and tooling that every engineering team at Ford relies on. You will design, build, and operate the messaging middleware platform, data management and replication tooling, and observability infrastructure that gives visibility into services running across our Kubernetes (EKS) environments - while also contributing to and supporting application development efforts as needed. You will balance strong data-driven decision-making with the ability to move forward effectively when information is incomplete. What you'll do... Engineering & System Design * Design, develop, maintain, and support the messaging middleware platform used by all engineering teams, along with associated data management and replication tooling. * Build and evolve observability tooling and pipelines (metrics, logs, traces) that provide visibility into services running in Kubernetes/EKS across the organization. * Own end-to-end delivery of middleware and observability services, including the platform infrastructure that supports them. * Design, develop, and operate high-performance, cloud-based and microservices-driven platforms at scale. * Build resilient backend services and tooling using technologies such as Java, Spring Boot, Kafka, PostgreSQL, gRPC, REST, and Kubernetes. * Develop and deploy services and tooling on AWS and/or GCP, leveraging managed services (e.g., MSK, Pub/Sub, EKS, GKE). * Contribute to application-level development efforts when needed, applying the same engineering rigor used for middleware and platform work. Cloud, DevOps & Delivery Excellence * Implement and maintain robust CI/CD pipelines with a strong focus on security, reliability, and efficiency. * Apply TDD and DevOps best practices using tools such as Jenkins, ArgoCD, SonarQube, Fossa, and GitHub. * Champion automation and operational consistency across environments for messaging, observability, and application infrastructure. Quality, Performance & Reliability * Write clean, maintainable, and well-tested code that meets high quality standards. * Perform load, stress, and performance testing on messaging middleware, data replication tooling, and application services to ensure scalability and reliability. * Own service health for the messaging and observability platforms by proactively identifying performance bottlenecks, throughput limits, and system risks. Operational Excellence * Implement comprehensive monitoring, alerting, and performance management strategies for the messaging middleware and dependent services. * Ensure messaging and observability platforms consistently meet SLA and reliability targets, minimizing impact to all downstream teams. * Collaborate with team members and downstream consumers to establish and evolve best practices that reduce operational risk across shared infrastructure. Innovation & Continuous Improvement * Proactively identify opportunities to adopt emerging messaging, observability, and cloud-native technologies to improve system efficiency and reliability. * Lead or contribute to refactoring initiatives to improve middleware and application performance, scalability, and maintainability. * Anticipate future challenges through data-informed decision-making and pragmatic engineering judgment. Agile Collaboration * Work within Agile development environments, partnering closely with product managers and cross-functional teams that depend on the messaging and observability platforms, as well as teams building applications on top of them. * Translate business and platform requirements into incremental, production-ready solutions. Technical Leadership & Communication * Participate in and lead design discussions, contributing to architectural decisions and technical standards for shared middleware, observability infrastructure, and consuming applications. * Influence technical direction while fostering collaboration across teams that consume the platform. * Communicate complex technical concepts clearly to both * technical and non-technical stakeholders. * Drive to sound conclusions even when data is incomplete or ambiguous. Customer-Centric Mindset * Stay ahead of emerging industry trends in messaging middleware, observability tooling, and cloud-native application development. * Align technical solutions with the needs of internal engineering teams and business outcomes, ensuring long-term value creation. ## Related Videos - [Remote Driving on Plant Grounds with State-of-the-Art Cloud Technologies](https://www.wearedevelopers.com/videos/251-remote-driving-on-plant-grounds-with-state-of-the-art-cloud-technologies) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Exploring the Power of gRPC-Gateway for Writing RESTful Services](https://www.wearedevelopers.com/videos/2072-exploring-the-power-of-grpc-gateway-for-writing-restful-services) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Agile work at CARIAD – Creating a customer web application for controlling the vehicle ](https://www.wearedevelopers.com/videos/200-agile-work-at-cariad-creating-a-customer-web-application-for-controlling-the-vehicle) - [Bringing AI Model Testing and Prompt Management to Your Codebase with GitHub Models](https://www.wearedevelopers.com/videos/1536-bringing-ai-model-testing-and-prompt-management-to-your-codebase-with-github-models) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [The Best Software Developer Blogs to Read](https://www.wearedevelopers.com/magazine/156-the-best-software-developer-blogs-to-read) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers)