> Markdown version of [/jobs/ext/2801125-site-reliability-engineer-service-assurance-systems](https://www.wearedevelopers.com/jobs/ext/2801125-site-reliability-engineer-service-assurance-systems). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - Service Assurance Systems - **Company:** Viasat, Inc. - **Location:** Greater London, UK - **Contract:** Permanent contract - **Skills:** Microsoft Windows, Agile Methodology, Airflow, Amazon Web Services, Amazon Elastic Compute Cloud, Amazon S3, Command-Line Interface, Computer Networks, Continuous Integration, Extract Transform Load (ETL), Data Transformation, Linux, DevOps, Data Flow Control, Github, Monitoring of Systems, Issue Tracking Systems, Python (Programming Language), Operational Data Store, Queueing Systems, Raw Data, Reliability Engineering, Ansible, Prometheus, SAS (Software), Software Engineering, Software Systems, SQL Databases, Data Streaming, Transmission Control Protocol (TCP), Scripting, Google Cloud, ServiceNow IT Service Management, Containerization, Gitlab-ci, Kubernetes, Infrastructure Automation Frameworks, Apache Flink, Deployment Automation, Integration Frameworks, Apache Kafka, Cloudwatch, Terraform, Data Pipelines, Docker, Jenkins, Servicenow - **Published:** September 9, 2026 - **Apply:** https://www.collegerecruiter.com/job/2850543356-site-reliability-engineer-service-assurance-systems ## About the Role * Familiarity with both Windows and Linux operating systems, including command-line administration and troubleshooting. * Solid experience with SQL or data pipeline technologies, * Experience with data processing frameworks such as Apache Airflow, Apache Flink or Google Dataflow, with a practical understanding of ETL pipeline design and the ability to diagnose issues across data transformation and scheduling workflows. * Demonstrable CI/CD experience - building, maintaining and improving pipelines using tools such as Jenkins, GitLab CI, GitHub Actions or equivalent. * Hands-on experience with AWS services (e.g. EC2, S3, ECS, Lambda, CloudWatch) and containerisation using Docker. * Experience with monitoring and observability tooling, including Prometheus and AWS CloudWatch, for metrics collection, alerting and dashboarding. * Proficiency in scripting with Python for automation, tooling and operational support tasks. * Strong communication skills - able to clearly articulate technical issues and their impact to both technical and non-technical audiences, and to provide timely updates on incident progress. * Experience using an ITSM/ticketing system (e.g. ServiceNow) for incident management, request fulfilment and problem tracking. * A proactive, solution-oriented approach with strong attention to detail and a commitment to operational excellence. * A reasonable understanding and appreciation of IT and security best practices in an operational environment., * Understanding of TCP/IP networking principles, including the ability to diagnose connectivity issues and interpret network traffic. * Familiarity with event streaming platforms such as Apache Kafka, including an understanding of producer/consumer patterns and how message queues are used to move data reliably between systems at scale. * Familiarity with infrastructure-as-code tools such as Terraform or Ansible. * Experience with log aggregation and analysis platforms such as the OTEL Stack or AWS CloudWatch Logs Insights. * Exposure to Kubernetes or other container orchestration platforms. * Experience working in an Agile or DevOps team environment. ## Description The Service Assurance Systems (SAS) Group, part of Global Operations, develop and maintain many software systems and applications which support the operation of Viasat services. The team collects assurance and operational data from a wide range of systems across the Viasat estate. This data is shared between internal platforms and distributed through Google Cloud Platform (GCP), where it underpins observability and monitoring capabilities that give operations teams real-time insight into service health and performance. To ensure the data is accurate and fit for purpose, the team builds and maintains data pipelines that cleanse, transform and enrich raw data, making it readily consumable by our stakeholders across operations, engineering and management. As a Site Reliability Engineer (SRE), you will play a key role in bridging the gap between software development and operational reliability. You will be responsible for supporting the applications built and maintained by the SAS group, managing deployments, investigating operational issues, and driving the team's transition toward a modern DevOps culture. You will act as a first point of contact for service issues raised through the ServiceNow ticketing system, working to diagnose, triage and resolve incidents efficiently while collaborating closely with developers and operations teams. The day-to-day You will be working as part of a small team of developers and operations engineers, supporting the evolution of Viasat's network and service monitoring capabilities, ensuring it remains world class in support of existing and future services. Day-to-day the role will involve: * Investigate and resolve incidents, service requests and problems raised through the ServiceNow ticketing system, ensuring timely and thorough resolution within agreed SLAs. * Manage and oversee application deployments across development, staging and production environments, ensuring smooth and reliable release processes. * Monitor application and infrastructure health using observability tools such as Prometheus and AWS CloudWatch; proactively identify and respond to anomalies and performance degradation. * Own and maintain CI/CD pipelines, working to improve build, test and deployment automation to reduce manual effort and increase release confidence. * Champion and drive the adoption of DevOps practices and culture within the SAS group, working towards greater automation, infrastructure-as-code, and operational maturity. * Manage and maintain containerised workloads using Docker, and support the operation of services hosted on AWS cloud infrastructure. * Write and maintain operational scripts (primarily Python) to automate routine tasks, support incident investigations and improve team efficiency. * Collaborate with software developers to identify recurring operational issues and feed findings back into the development process to improve application resilience. * Maintain clear and up-to-date operational documentation including runbooks, deployment guides and incident post-mortems. * Provide written and verbal progress updates on open incidents, deployments and operational improvements to the SAS group and wider stakeholders. * Support on-call and out-of-hours incident response as required. * Liaise with engineering and infrastructure teams to ensure system changes are communicated and operationally risk-assessed before deployment. ## Related Videos - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Unlocking the potential of Digital & IT at Vodafone](https://www.wearedevelopers.com/videos/602-unlocking-the-potential-of-digital-it-at-vodafone) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [The Most Popular IT Jobs on the Market](https://www.wearedevelopers.com/magazine/376-the-most-popular-it-jobs-on-the-market) - [The 12 Best Jobs for Software Engineers](https://www.wearedevelopers.com/magazine/401-the-12-best-jobs-for-software-engineers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [IT Salaries in UK](https://www.wearedevelopers.com/magazine/288-it-salaries-in-uk)