> Markdown version of [/jobs/ext/1503140-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/1503140-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** Spectraforce - **Location:** Scottsdale, AZ, United States - **Experience:** Expert - **Contract:** Temporary contract - **Skills:** Java (Programming Language), Artificial Intelligence, Application Performance Management, Confluence, BigQuery, Cloud Computing, Databases, Continuous Integration, Linux, Distributed Systems, Domain Name System (DNS), Github, Monitoring of Systems, Hypertext Transfer Protocols (HTTP), Python (Programming Language), PostgreSQL, Microsoft SQL Server, MongoDB, Network Protocols, Node.Js, Oracle Databases, Redis, Reliability Engineering, Ansible, Prometheus, TCP/IP, Rust (Programming Language), Google Cloud, Load Balancing, Istio, System Availability, Grafana, Containerization, Kubernetes, Rancher, Hashicorp, Graphql, Vertica, Api Gateway, Terraform, Splunk, Appdynamics, Dynatrace, Golang, Programming Languages - **Published:** July 30, 2026 - **Apply:** http://leoforce.us/Careers/Spectraforce/JobDetails.html?jobid=1a64e20d-8844-49f7-8cbd-f81a2b788876&OrgId=1&UserId=47 ## About the Role * Service reliability/operation experience running large-scale, high-performance applications in a hybrid environment (on-prem and cloud). * Experience in writing automation scripts and building dashboards for Application Performance management to manage Transaction journeys. * Experience working with Programming languages such as Go, Python, Java, Rust etc. * Working knowledge on with one or more databases- Oracle, SQL Server, Redis, Clickhouse, postgres, Mongo or any time-series databases · Experience in transitioning platforms to the cloud and Containerization - GCPand Rancher * Experience maintaining containerized app in GKE/RKE/AKE environments. * Experience Implementing Cloud observability using OTEL to enable real-time monitoring, distributed tracing and incident resolution. · Experience working with specific GraphQL Framework (Apollo, Prisma, Hasura etc.). * Experience using knowledge of networking protocols such as TCP/IP, HTTP, DNS, Load balancing and service mesh to troubleshoot issues in high pressure situations. Preferred Skills: · * Proven experience managing Application availability, building creative solutions to manage repetitive activities, improving gating and detect for applications at every touchpoint for a 24 x 7 High availability platform exposed to critical clients and customers. * Working knowledge of Monitoring tools - Splunk, App-dynamics, grafana/Prometheus and Dynatrace. * Experience with tools like Rally, Confluence and other CI/CD extenders. * Hands-on experience with implementing in-memory caching solutions. * Experience on Redis DB is a plus. * Excellent debugging skills across variety of integrated technical platforms on API gateway. * Hands-on with GCS, Cloud SQL, Spanner and Firestore. * Extensive experience in Enterprise level Infrastructure and Operations. * Experience in High Availability and distributed systems, Linux and Windows administration, troubleshooting and support. * Monitor and troubleshoot HashiCorp Vault environments, ensuring minimal downtime and rapid recovery from incidents. * Working knowledge on Vertex AI, Gen AI and Bigquery · Google Cloud Platform (GCP) Containerization, Kubernetes · Infrastructure as Code (Terraform), CI/CD (GitHub Actions), and Helm · Automation and scripting using Python, Ansible, and Node.js · ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Reducing LLM Calls with Vector Search Patterns - Raphael De Lio (Redis)](https://www.wearedevelopers.com/videos/1714-reducing-llm-calls-with-vector-search-patterns-raphael-de-lio-redis) - [Rate-limiting using eBPF and Istio: How to protect your SaaS customers from themselves](https://www.wearedevelopers.com/videos/100220-rate-limiting-using-ebpf-and-istio-how-to-protect-your-saas-customers-from-themselves) - [Remote Driving on Plant Grounds with State-of-the-Art Cloud Technologies](https://www.wearedevelopers.com/videos/251-remote-driving-on-plant-grounds-with-state-of-the-art-cloud-technologies) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers)