Principal Site Reliability Engineer

Veson Nautical
London, UK
17 days ago
Apply on www.adzuna.co.uk
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
£78,994.0
Working hours
Regular working hours

Tech stack

Artificial Intelligence Amazon Web Services BigTable BigQuery Software as a Service Cloud Computing Cloud Database Configuration Management Computer Programming Continuous Integration Data as a Services Extract Transform Load (ETL)
+26 more
Data Stores DevOps Elasticsearch Data Flow Control Identity and Access Management Python (Programming Language) PostgreSQL Microsoft SQL Server Octopus Deploy Reliability Engineering TypeScript Google Cloud Load Balancing Grafana Multi-Cloud Amazon Virtual Private Cloud (VPC) Gitlab Gitlab-ci Kubernetes Information Technology Sentry Dynamic Data Elastic Beanstalk Terraform Splunk Pagerduty

Job description

As a Principal Site Reliability Engineer at Veson Nautical, you will design, build, monitor, and support the cloud infrastructure that underpins our rapidly growing SaaS platform. This is a multi-cloud role spanning both AWS and Google Cloud Platform, with an immediate focus on growing our GCP footprint and the systems that connect our environments across regions, accounts, and clouds.

The work is a mix of greenfield and stewardship. You’ll stand up infrastructure for new applications from scratch, and you’ll take on the harder problem of making our existing estate more consistent, more observable, and easier to operate. You will have significant influence over the architectural direction of the platform.

Our suite of products includes:

· Veson Platform - the core commercial maritime platform used by the world’s leading shipping organizations to manage vessel communication, operations, and trade decisions

· Oceanbolt - a dynamic data intelligence platform, tracking over 23,000 vessels in real time to deliver accurate, timely market intelligence to drive decision making

· Shipfix - using proprietary AI-driven tools to infer cargo and vessel information, extracting, anonymizing, and aggregating billions of data points with near real-time processing of email exchanges in the shipping market

The Team

You’ll join a global Site Reliability Engineering team with members in the United States and the United Kingdom. This is a senior individual contributor role without direct reports, but with real leadership expectations. If you are the kind of engineer who measures success partly by what your teammates can do without you, this role will suit you.

This position is based in London in a hybrid model, with an expectation of 2-3 days per week in our office. We think those days are worth showing up for: our beautiful office is stocked with snacks, there’s a weekly team lunch on us, and the atmosphere is friendly, informal, and genuinely collaborative.

The team participates in an on-call rotation -more details will be provided in the interview process.

Our Stack

· Google Cloud Platform - primarily PaaS services (Bigtable, Cloud SQL, Dataflow, Datastore, GKE, GCS, KMS, Pub/Sub)

· Amazon Web Services - multi-region, multi-account, with a broad range of managed services

· Containers and orchestration - Kubernetes (GKE and EKS)

· Infrastructure-as-Code - Terraform, Terragrunt, and Atlantis

· CI/CD - GitLab Pipelines, ArgoCD, Octopus Deploy

· Data - ElasticSearch hosted with Kubernetes Operator, PostgreSQL, SQL Server, BigQuery

· Monitoring and Security - Splunk, Grafana / Grafana Tempo, OpenTelemetry, Cloud Armor Enterprise, OpsGenie, Renovate, Sentry

· AI Tools - Claude, Amazon Bedrock, Gemini, Vertex AI, · Design, implement, and operate scalable, reliable, and secure infrastructure across Google Cloud Platform and AWS

· Lead greenfield infrastructure builds for new applications, and modernize existing infrastructure toward common patterns

· Solve cross-region, cross-account, and cross-cloud problems - networking, identity, data movement, and the operational patterns that hold across environments

· Drive automation of infrastructure provisioning and configuration management using Terraform and related IaC tooling

· Establish and maintain comprehensive monitoring, alerting, and observability practices

· Cross-train and mentor other engineers, with the explicit goal of broadening GCP and multi-cloud capability across the team

· Partner closely with development teams to ensure the reliability, performance, and scalability of our platforms

· Set technical direction through design reviews, architecture proposals, and clear written documentation

· Participate in and improve incident response, and drive the follow-through that keeps the same incident from recurring

· Improve the cost effectiveness of our cloud footprint through visibility, analysis, and sound architectural choices

Requirements

· Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience

· Previous experience working on a large-scale Software-as-a-Service (SaaS) platform supporting thousands of global users in a 24x7x365 environment

· 5+ years of hands-on experience with Google Cloud Platform services and architecture, including running production workloads at scale

· Production Kubernetes experience, particularly Google Kubernetes Engine (GKE)

· Proficiency with Infrastructure-as-Code, preferably Terraform with the GCP provider

· Experience with Google Cloud networking, including VPC, Cloud Load Balancing, and Cloud CDN

· Strong programming skills in Python, Go, or TypeScript for automation and tool development

· Demonstrated ability to raise the technical capability of a team through mentorship, documentation, and knowledge sharing

· A track record of building consensus around technical decisions that span multiple teams

Highly Desirable:

· Hands-on experience with both GCP and AWS, and a clear point of view on where multi-cloud helps and where it hurts

· Working knowledge of Google Cloud security best practices and IAM implementation

· Experience with Google Cloud data services including BigQuery and Dataflow

· Experience with cloud cost management (budgeting, anomaly detection, cost analysis and reporting)

· Experience working on a geographically distributed team across time zones

Nice to have:

· Google Cloud Professional certifications (Cloud Architect, Data Engineer, or DevOps Engineer)

· Experience with GitLab CI or Octopus Deploy

About the company

Veson Nautical empowers the global maritime industry to navigate complexity on all sides of the trade. Veson’s platform combines AI-driven workflows, trusted data, and seamless collaboration, to deliver the insight and context needed for confident, competitive decision-making., Veson Nautical is a successful, rapidly growing global software company. Our clients are the world’s leading commercial maritime owners, operators, and commodity trading companies. Veson’s solutions enable our clients to identify new opportunities and proactively manage their business to make more profitable decisions. With offices in Singapore, Tokyo, London, Houston and headquarters in Boston, USA, Veson Nautical is a dynamic organization with a committed team of professionals. Dedicated to ensuring the highest levels of client satisfaction, Veson Nautical brings decades of experience, technical knowledge, enthusiasm, and commitment to clients around the world.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.adzuna.co.uk
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

6:14 min

Structuring CI/CD pipelines with integrated security and quality checks

Christoph Ruggenthaler · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

1:22 min

Overview of the Sentry error and performance monitoring platform

Priscila Oliveira · World Congress 2023

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

4:54 min

Implementing geographic salary tiers for compensation equity and fairness

Rudi Bauer Rudi Bauer +1 · Cappuccino with HR

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

Videos

See all

Related articles

See all