Senior Devops Engineer
Role details
Job location
Tech stack
Job description
This Financial Services Client is seeking a hands-on DevOps Engineer to lead and support the tooling and platform infrastructure that powers our software development and delivery lifecycle within our retirement business. This role is responsible for managing and optimizing the full DevOps toolchain, including CI/CD platforms, source control, orchestration, operational monitoring, work management, and knowledge management tools, while enabling development teams to build, test, deploy, and run applications efficiently. The environment is primarily on-premises with some AWS integration, requiring a strong foundation in both modern DevOps practices and traditional infrastructure management. We are looking for a well-rounded engineer who can independently drive initiatives, take ownership of solutions, collaborate across teams and business units, and help standardize tooling across the organization. While strong technical skills are important, the ideal candidate is a, Platform Implementation & Rollout
-
Implement and roll out core platform capabilities (Kubernetes-based runtime, build/deploy tooling, observability) across teams and environments.
-
Build reusable templates, Helm charts, pipeline definitions, and IaC modules that teams can adopt with minimal friction.
-
Onboard scrum teams onto the platform - migrating workloads, standardizing configuration, and documenting self-service paths.
-
Drive consistent adoption of platform standards while accommodating legitimate team-specific needs.
Platform Maintenance & Support
-
Own day-to-day health, configuration, and lifecycle of the platform and its supporting infrastructure.
-
Plan and execute infrastructure and platform upgrades (Kubernetes versions, node pools, runtimes, agents, tooling) with minimal disruption.
-
Manage platform configuration changes through controlled, auditable processes.
-
Provide responsive support to engineering teams using the platform, acting as the escalation point for platform-related issues.
Production Maintenance & Operations
-
SSL/TLS certificate renewals and rotation, secret/credential rotation, patching, and capacity adjustments performed on a recurring, reliable cadence.
-
Plan and carry out infrastructure upgrades and maintenance windows, including rollback planning and stakeholder communication.
-
Maintain and tune platform configuration for scalability, resiliency, performance, and cost efficiency.
-
Lead incident response for production issues - triage, mitigation, coordination, and resolution - and run blameless post-incident reviews with concrete follow-up actions.
-
Improve observability, monitoring, and alerting so issues are caught proactively rather than reported by users.
CI/CD Enablement
-
Design, implement, and maintain automated CI/CD pipelines covering build, test, release, and deployment across Windows and Linux workloads.
-
Enable scrum teams to own their pipelines through shared templates, reusable steps, and clear documentation.
-
Promote sound source control, branching, and release management practices.
-
Embed security and compliance controls directly into pipelines (DevSecOps) - scanning, policy gates, and secrets handling.
-
Continuously reduce lead time, deployment friction, and manual steps in the delivery lifecycle.
Reliability & Continuous Improvement
-
Provide Level-3 support for complex production and platform incidents.
-
Identify and act on opportunities to improve system stability, operational maturity, and self-service.
-
Optimize platforms and processes based on measurable outcomes (deployment frequency, change failure rate, MTTR, cost).
-
Maintain strong, current technical documentation and runbooks.
Collaboration, Mentorship & Innovation
-
Partner with Application Architects and engineering teams to align platform and DevOps solutions with product needs.
-
Mentor and coach engineers on DevOps practices, pipeline ownership, and operational discipline.
-
Promote a culture of automation, ownership, and continuous improvement.
-
Research emerging tools and practices that improve the delivery lifecycle, and introduce cost-effective solutions that increase speed, quality, and reliability.
-
Active AI practitioner- leverages AI-assisted tooling (Claude, Cursor, Copilot) to accelerate engineering and operational work.
Requirements
self-starter with excellent communication skills, a continuous learning mindset, and the ability to see the bigger picture while mentoring and supporting others., EKS, kubernetes, Docker, Helm, Policy as a code, Secrets Management, DevSecOps tooling, Dora, AWS
Top Skills Details
EKS,kubernetes,Docker,Helm
Additional Skills & Qualifications
Additional Skills
-
Log collection and dashboarding (ELK, New Relic).
-
Practical experience with production operations: SSL/TLS certificate management, networking, load balancing, caching, high availability, and disaster recovery.
-
Demonstrated experience operating production systems and leading incident response.