Observability SRE (Site Reliability Engineer)

NEOMORPH INVESTMENTS LLC
Philadelphia, PA, United States
15 days ago
Apply on www.careerjet.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
2 years minimum
Compensation
$115,000.0 - $135,000.0
Working hours
Regular working hours

Tech stack

Apache ActiveMQ Artificial Intelligence Data Analysis Confluence JIRA Collaborative Software Continuous Integration Relational Databases Linux Middleware Monitoring of Systems Web Servers
+21 more
Python (Programming Language) Microsoft SQL Server MySQL Cisco Nexus Switches Open Source Technology Scrum Methodology Release Management Reliability Engineering Ansible Standard Sql Load Balancing Grafana Sybase Reliability of Systems Gitlab Containerization Kubernetes Information Technology TIBCO (Software) Docker Jenkins

Job description

SRE within the Group Platform Services & Engineering division which provides the Nomura group common services to Development, Infrastructure and Production Services. This is an SRE/support position responsible for administering and supporting Production environment as well as engineering reliability into the products / services we support i.e. monitoring & observability platform. The successful candidate will have a vital role in shaping future monitoring strategy and direction within the Nomura Group. A fantastic opportunity for somebody with 3+ years IT experience to work with state-of-the-art technologies to deliver industry leading solutions in the Telemetry, Observability and Monitoring space. The successful candidate would join a team of enthusiastic, creative and forward-thinking SRE in the US who are working in tandem with the engineers to radically transform how the Nomura Group manages the operation of its estate. The position is within a global team consisting of 20 team members, across Engineering and SRE, bringing change across the organisation. The candidate will work closely with their peers in other regions as well as other teams to facilitate the strategic objectives of the team. The challenges we strive to solve include availability, scalability and performance related to delivering a platform used by the entire Nomura Group. The observability platform consists of a combination of platforms and frameworks from in-house, vendors, and open source. These include:

  • Grafana LGTM stack (open source)
  • RightITNow, EverBridge, Sentinel (3rd Party tools)
  • AMBER, Bing, MCM, CMS (homegrown)

  • Cross functional engagement to champion and provide necessary support for the adoption of TOM platform across the Nomura group of companies.
  • Gain understanding of the various tools and frameworks that together provide observability and notification service to the organization and assist development and production support teams with queries / issues related to their usage of our platform.
  • Act as custodian of production environment and engage within the team and outside, if need be, towards building and maintaining robust, scalable, highly available production systems in accordance with our service level objectives
  • Preventing production incidents but when they do occur, performing effective incident and problem management and RCA to minimize downtime as well as possibility of recurrence.
  • Pushing out changes and releases to production environment reliably via effective change and release management
  • Quick and effective response to alerts before they become incidents, with an approach to prevent them from occurring ever again
  • Effectively triaging alerts, requests, emails such that things that needs attention get addressed first and in a timely manner in the order of their priority, the drivers for which should be production stability and user satisfaction
  • Continuous and effective engagement with users, with the required empathy, providing the right guidance so as to provide a good customer experience
  • Collaborate in a global agile team environment using established support practices, participating in sprint planning, reviews, and continuous improvement initiatives
  • Build and maintain scalable, reliable monitoring solutions that support Nomura’s global infrastructure
  • Engage with engineers, architect towards contributing to architectural decisions that influence the future direction of Nomura’s observability platform
  • Champion observability best practices across the organization, helping teams leverage data-driven insights to improve system reliability and performance
  • Partner with engineers as needed to optimize operational efficiency and enhance system resilience
  • Effectively leveraging AI tools such as Claude, CoPilot etc. with adequate guardrails to bring efficiencies into operational processes in a consistent, repeatable and risk averse manner.
  • Mentor and guide other SREs, sharing your knowledge and expertise across other team members for the benefit of the team., * Explore Insights & Vision: Identify the underlying causes of problems faced by you or your team and define a clear vision and direction for the future.
  • Making Strategic Decisions: Evaluate all the options for resolving the problems and effectively prioritize actions or recommendations.
  • Inspire Entrepreneurship in People: Inspire team members through effective communication of ideas and motivate them to actively enhance productivity.
  • Elevate Organizational Capability: Engage proactively in professional development and enhance team productivity through the promotion of knowledge sharing.
  • Inclusion: Foster a culture of inclusion and psychological safety in the workplace and cultivate a “Risk Culture” (Challenge, Escalate and Respect).

Requirements

  • Minimum 2 years’ experience with Grafana or any other modern observability tools in an administrative capacity for a medium/large scale enterprise.
  • At least 2 years’ exposure to Linux OS with a decent hold on general purpose troubleshooting and day to day commands
  • Exposure to one or more of following - Python / Ansible
  • Production support experience - Request handling, incident management, problem management, change management, release management, on-call handling, user engagement, responding to alerts etc.
  • Good communication and interpersonal skills
  • Strong analytical and trouble-shooting skills, with the ability to exercise mature judgement
  • Basic understanding of cloud platforms
  • Basic understanding of CI/CD tools such as GitLab, Jenkins, Ansible, Nexus etc.
  • Good Team player

Preffered:

  • Understanding of Open Telemetry standards
  • Understanding of containerization technologies such as Kubernetes, EKS, Docker etc.
  • Supporting a medium / large scale production environment
  • Knowledge of ITIL
  • Decent understanding of DB Platforms - Sybase / MySQL / MSSQL - general RDBMS concepts, SQL
  • Collaboration Tools - Confluence / JIRA
  • Basic knowledge of / familiarity with other infrastructure technologies such as Middleware (ActiveMQ / Solace / EMS / Tibco etc.), Web servers, Load balancers, Directory Services etc.
  • Experience working with a globally dispersed team

Benefits & conditions

  • base pay offered may vary depending on multiple individualized factors, including market location, corporate and functional title and duties, job-related knowledge and advanced degrees, skills, and experience. The total compensation package for this position may also include other elements, including a sign-on bonus, restricted stock units, and discretionary awards in addition to a full range of medical, financial, and/or other benefits (including 401(k) eligibility and various paid time off benefits, such as vacation, sick time, and parental leave), dependent on the position offered. Details of participation in these benefit plans will be provided if an employee receives an offer of employment.

If hired in the U.S., employee will be in an “at-will position” and the Company reserves the right to modify base salary (as well as any other discretionary payment or compensation program) at any time, including for reasons related to individual performance, Company or individual department/team performance, and market factors”. Nomura is an Equal Opportunity Employer

About the company

Nomura is a financial services group with an integrated global network. By connecting markets East & West, we service the needs of individuals, institutions, corporates and governments through our four business divisions: Wealth Management, Investment Management, Wholesale (Global Markets and Investment Banking) and Banking. Driven by the insights of some 28,000 people worldwide, we put our clients at the center of everything we do, delivering unparalleled access to, from and within Asia. For further information about Nomura, visit Aon’s Benefit Index®, Nomura’s benefits rank #1 amongst our competitors Department Overview: The Information Technology department at Nomura is at the forefront of innovation, driving technology solutions that empower our business and enhance client experiences. We leverage cutting-edge technologies to develop and maintain robust systems and infrastructure, ensuring the security, reliability, and efficiency of our operations. Join our team and be part of a dynamic and collaborative environment that embraces technological advancements to deliver value and drive our digital transformation journey., + Philadelphia, PA

  • $103,200-141,900 per year Who We Are: We’re powering a cleaner, brighter future. Exelon is leading the energy transformation, and we’re calling all problem solvers, innovators, community builders and chan…

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerjet.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:05 min

Integrating an assistant application with Jira software

Felix Augenstein · LIVE

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

2:18 min

Scaling MySQL databases for massive user growth

Johannes Nicolai Johannes Nicolai +1 · LIVE

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

8:22 min

Simulating a Linux terminal and running Spring Boot

Jakov Semenski · LIVE

2:11 min

Updating the delivery architecture with Jira and Tekton pipelines

Lian Li · World Congress 2022

Videos

See all

Related articles

See all