> Markdown version of [/jobs/ext/2869849-principal-site-reliability-architect-engineer](https://www.wearedevelopers.com/jobs/ext/2869849-principal-site-reliability-architect-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal Site Reliability Architect / Engineer - **Company:** Propertyvalue Northern Trust - **Location:** United States - **Experience:** Expert - **Salary:** $137,400.0 - $233,600.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Application Lifecycle Management, Systems Engineering, C Sharp (Programming Language), Configuration Management, Computer Programming, Distributed Systems, Fault Tolerance, Python (Programming Language), Reliability Engineering, Ansible, Software Deployment, Data Logging, Scripting, Bicep, Puppet, Terraform - **Published:** September 12, 2026 - **Apply:** https://www.dice.com/job-detail/ef604e16-e9b8-4073-80e2-789b24eb98e6 ## About the Role * Bachelor's degree or equivalent practical experience, with 5+ years in site reliability, systems engineering, or infrastructure engineering role. * Experience contributing to and evolving reliable, scalable system architectures, including availability, resiliency, and fault-tolerance patterns. * Proven ability to translate architectural standards into practical implementations in partnership with engineering and operations teams. * Experience across cloud and on-prem environments, with an architectural focus on consistency, observability, and operational maturity. * Solid understanding of distributed systems, networking, and modern application architectures. * Hands-on experience with observability solutions, including contributing to and applying standards for metrics, logging, tracing, dashboards, and alerting. * Proficiency in at least one programming or scripting language (e.g., Python, C#, Java), with the ability to contribute code and guide implementation. * Proficiency with at least one IaC or Configuration Management language (e.g., Terraform, Bicep / ARM, Chef, Puppet, Ansible), with the ability to contribute code and guide implementation. * Experience delivering Infrastructure as Code and application deployment automation through CI/CD pipelines. * Strong communication skills, with the ability to articulate architectural trade-offs and reliability strategies to diverse stakeholders. ## Description NT is seeking a Principal Site Reliability Architect / Engineer who sets the architectural vision and resiliency strategy for critical platforms across Northern Trust. At Northern Trust, reliability is a design-time concern. This role partners with senior technology leaders, architects, and engineering organizations to define how reliability is designed, measured, and sustained at enterprise scale. It is accountable for reliability architecture outcomes, ensuring that resiliency principles are embedded into system design decisions, lifecycle governance, and engineering standards - rather than relying solely on operational practices after deployment. You will influence long-term platform evolution by defining reference architectures, architectural guardrails, and maturity models that guide multiple teams and domains toward highly resilient, observable, and operable systems. Architectural leadership in this role does not end at the whiteboard. The Principal Site Reliability Architect is expected to leave the design session able to actively participate in solution creation-validating architectural decisions through targeted implementation, pairing with engineers to shape critical components, or contributing focused code and automation where needed. This hands-on engagement ensures architectural decisions remain actionable, technically credible, and grounded in real-world operational realities. This role will be responsible for several key functions that both support and drive improvements to the reliability of Northern Trust's IT Landscape. What you will do: * Create and own resiliency guidelines + Create, maintain, and continuously improve one or more resiliency guidelines working directly with technology management teams. + Translate resiliency and reliability principles into practical, measurable guidance application teams can implement. + Guide adoption and maturity of resilient software and infrastructure architectures to proactively improve our product and service landscape. * Provide guidance in incident response and root cause analysis + Support root cause analysis sessions by helping teams identify systemic failure modes, resiliency gaps, and contributing factors. + Partner with teams to define preventive measures and guide implementation plans that reduce recurrence. + Feed learnings from incidents back into standards, guidelines, and automation patterns. * Build infrastructure configuration and application management automation + Develop, maintain, and expand automation logic to streamline repetitive operational tasks and improve resiliency. + Build and standardize infrastructure configuration management and application lifecycle management patterns to increase consistency across environments. + Implement and improve monitoring and observability capabilities that help detect and address issues before they impact the business. * Foster a broader community of practice + Help build and sustain an SRE community of practice by collaborating to identify common improvement areas and define standards and governance. + Communicate effectively with stakeholders across different teams and levels within the organization. + Collaborate with product, development, and operations teams to align reliability and resiliency efforts with overarching business goals. ## Related Videos - [Back(end) to the Future: Embracing the continuous Evolution of Infrastructure and Code](https://www.wearedevelopers.com/videos/440-back-end-to-the-future-embracing-the-continuous-evolution-of-infrastructure-and-code) - [Automate everything via NodeJS and Puppeteer](https://www.wearedevelopers.com/videos/322-automate-everything-via-nodejs-and-puppeteer) - [Dev & Test in the Cloud? Deploy your cloud environments with Ansible & Terraform](https://www.wearedevelopers.com/videos/1607-dev-test-in-the-cloud-deploy-your-cloud-environments-with-ansible-terraform) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) - [The Memory Leak That Ate Our Cluster: A Postmortem](https://www.wearedevelopers.com/videos/2057-the-memory-leak-that-ate-our-cluster-a-postmortem) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [The 12 Best Jobs for Software Engineers](https://www.wearedevelopers.com/magazine/401-the-12-best-jobs-for-software-engineers) - [Trustworthy AI Starts at Deployment: 5 Checks Before You Ship](https://www.wearedevelopers.com/magazine/753-trustworthy-ai-starts-at-deployment-5-checks-before-you-ship) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers) - [Top-Paying Tech Jobs (with Salaries)](https://www.wearedevelopers.com/magazine/372-top-paying-tech-jobs-with-salaries)