> Markdown version of [/jobs/ext/3531857-senior-deployment-and-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/3531857-senior-deployment-and-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Deployment and Site Reliability Engineer - **Company:** GE Vernova - **Location:** Cambridge, MA, United States (Remote available) - **Experience:** Expert - **Salary:** $113,200.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Microsoft Azure, Bash Shell, Databases, Relational Databases, Disaster Recovery, Embedded Software, Virtual Private Networks (VPN), Python (Programming Language), PostgreSQL, Octopus Deploy, Role-Based Access Control, Reliability Engineering, Robotic Automation Software, Software Engineering, Software Systems, Data Logging, Transport Layer Security, Load Balancing, Istio, System Availability, Delivery Pipeline, Firewalls (Computer Science), Containerization, Kubernetes, Information Technology, Deployment Automation, Performance Monitor, Operational Systems, Docker, Servicenow, Artifactory - **Published:** October 1, 2026 - **Apply:** https://www.techcareers.com/job.asp?id=3407945054&tx=JT9688TYV&pt=1&aff=0B19D771-A501-4A5E-8338-2A822B784D54&utm_source=Job%20Feed&utm_medium=textkernel&utm_campaign=DE&utm_term=0B19D771-A501-4A5E-8338-2A822B784D54 ## About the Role * Education: Bachelor's or Master's degree in Software Engineering, Computer Science, or a related technical discipline. * Systems Deployment: 5-8 years of experience deploying, commissioning, and supporting complex software systems in production or customer environments across cloud (Azure/AWS) and on-premise architectures. * Kubernetes & Containerization: 5-8 years of deep expertise in full-stack Kubernetes operations (Helm, RBAC, Networking) and container/Docker proficiency (building, inspecting, and troubleshooting runtime behavior). * Incident Management: 5-8 years of experience managing critical production issues end-to-end, including triage, service restoration, root cause analysis, and corrective action under customer-facing pressure. * Technical Infrastructure: 5-8 years of experience with relational databases (e.g., PostgreSQL), networking/security configurations (TLS, firewalls, load balancing, VPNs), and scripting automation (Python, Bash)., * Operational Technology (OT): Experience supporting software in industrial, substation, or utility environments, including familiarity with IEC 62443 standards and hardened or air-gapped systems. * Travel & Availability: Ability to travel to customer sites (approx. 30% of time) and flexibility to support maintenance windows or critical issues outside normal business hours. * Communication: Exceptional ability to communicate technical findings clearly to both technical and non-technical stakeholders, maintaining customer confidence during high-pressure incidents. * Environment Management: Proven experience maintaining staging/pre-production environments integrated with release pipelines and artifact repositories (e.g., Artifactory). * Advanced Orchestration: Familiarity with GitOps-based delivery (Argo CD or Flux CD) and service mesh technologies such as Istio. * Service Management: Experience with structured incident/problem management practices (e.g., ITIL) and service management tools like Jira or ServiceNow. * High Availability: Experience with disaster recovery architectures, database clustering, and connection pooling (e.g., pgpool). * Continuous Improvement: A proactive mindset focused on turning field experiences into permanent improvements for products, deployment procedures, and automation. * Global Collaboration: Ability to interface effectively within international, multi-cultural, and matrixed organizations to deliver cohesive technical solutions. ## Description Grid Automation offers its customers a complete range of innovative products, systems and services covering design and manufacture, as well as commissioning and long-term maintenance of Substation Protection and Automation solutions. This includes, amongst others, Protection relays, control systems, cyber security solutions, device management, asset performance, and many more. Grid Automation Software operates at the boundaries of Grid technologies, embedded software and advanced analytics solutions. If you're passionate about Software and want to work on improving Grid resilience through Software, this is the job for you. We are looking for a Deployment and Site Reliability Engineer to join the GridBeats team. The GridBeats Software Portfolio aggregates all Grid Automation applications which monitor, optimize and manage grid connected devices. It is a very complex and exciting portfolio which is setting the basis for grid connected devices management for the future. The Deployment and Site Reliability Engineer takes GridBeats software into customer operation and keeps it running there. The role deploys, commissions, and upgrades our software at customer sites and remotely, and is the technical escalation point when something critical breaks in a live installation. It is the engineer customers and Regional Operations rely on when a deployment has to succeed inside a maintenance window, and when a production issue has to be understood and resolved rather than merely logged. The portfolio runs on Kubernetes, so deep Kubernetes skill is the central technical requirement for this position. This is a customer-facing, hands-on role, and it is not a product engineering role. The software is built by the product teams; this role gets it installed, proven, and supported in real substation and utility environments, and feeds what it learns in the field back into the products. Travel to customer sites, work inside planned maintenance windows, and availability for critical issues outside normal hours are an integral part of the job., * Deployment & Commissioning: Lead end-to-end installation, configuration, and commissioning of GridBeats software across diverse architectures (cloud, on-premise, hybrid, and air-gapped) while ensuring seamless technical handover to customers. * System Upgrades & Recovery: Execute software upgrades, patching, and data migrations; maintain rigorous backup/restore protocols and rehearsed disaster recovery plans to ensure system integrity. * Critical Issue Resolution: Serve as the technical escalation point for high-severity incidents, conducting deep-dive analysis of logs and network behavior to restore service and drive permanent resolutions with product teams. * Staging & Validation: Prepare and manage staging and pre-production environments that mirror customer architectures to validate release packages, configurations, and upgrade paths prior to field deployment. * Reliability & Monitoring: Implement proactive monitoring, logging, and alerting systems; analyze system health and resource consumption to recommend preventive maintenance and capacity planning. * Automation & Efficiency: Utilize infrastructure-as-code and deployment automation to minimize manual effort, advocating for platform improvements that enhance field deployability and repeatability. * Operational Documentation: Author and maintain critical technical documentation, including runbooks, installation guides, and troubleshooting procedures to support consistent service delivery. * Knowledge Transfer: Provide training and technical support to Regional Operations, service teams, and customer personnel to ensure operational excellence across the installed base. * Cross-Functional Advocacy: Champion the "field perspective" within product development, contributing insights on deployability and supportability to influence product design and future release quality. * The successful candidate will sit onsite at GEV Electrification Lab ## Related Videos - [Enabling automated 1-click customer deployments with built-in quality and security](https://www.wearedevelopers.com/videos/83-enabling-automated-1-click-customer-deployments-with-built-in-quality-and-security) - [Kubernetes and Microservices with Multi-Model Databases](https://www.wearedevelopers.com/videos/382-kubernetes-and-microservices-with-multi-model-databases) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Rate-limiting using eBPF and Istio: How to protect your SaaS customers from themselves](https://www.wearedevelopers.com/videos/100220-rate-limiting-using-ebpf-and-istio-how-to-protect-your-saas-customers-from-themselves) - [Next-gen CI/CD with Gitops and Progressive Delivery](https://www.wearedevelopers.com/videos/1603-next-gen-ci-cd-with-gitops-and-progressive-delivery) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Where to Find Entry-Level Software Engineering Jobs](https://www.wearedevelopers.com/magazine/397-where-to-find-entry-level-software-engineering-jobs) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)