> Markdown version of [/jobs/ext/2045035-senior-software-engineer-app-reliability](https://www.wearedevelopers.com/jobs/ext/2045035-senior-software-engineer-app-reliability). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Software Engineer - App Reliability - **Company:** MetLife - **Location:** Clarks Summit, PA, United States - **Experience:** Expert - **Salary:** $90,000.0 - $120,000.0 - **Contract:** Temporary to permanent - **Skills:** Clean Code Principles, Java (Programming Language), JavaScript (Programming Language), Application Programming Interfaces (APIs), Agile Methodology, Artificial Intelligence, Software Applications, User Authentication, Automation of Tests, Microsoft Azure, Batch Processing, Bioinformatics, C Sharp (Programming Language), Cloud Computing, Code Review, Information Systems, Continuous Integration, Software Debugging, DevOps, HP Systems Insight Manager, Python (Programming Language), MongoDB, Operational Data Store, Object-Oriented Software Development, Reliability Engineering, Software Tools, Site Reliability Engineering Practices, Secure Coding, Software Engineering, Data Streaming, Systems Integration, TypeScript, Web Platforms, Application Enhancement Tool, Large Language Models, Grafana, Prompt Engineering, Technical Debt, Containerization, Kubernetes, Information Technology, Enterprise Integration, Cosmos DB, Splunk, Appdynamics, Software Version Control, Docker, Servicenow, Programming Languages - **Published:** August 13, 2026 - **Apply:** https://www.careerjet.com/job/us7699f4385225ea4aecde61ae638a4fb7/eaa ## About the Role * Bachelor's degree in computer science, Information Systems, Engineering, or a related field (or equivalent experience), with 7+ years of experience building and supporting full-stack applications, APIs, integrations, batch processes, data flows, and customer-facing digital platforms. * Proficiency in one or more programming languages, including Java, Python, C#, JavaScript/TypeScript, or similar technologies. Strong knowledge of software engineering practices, including object-oriented design, secure coding, automated testing, code reviews, source control, CI/CD, cloud platforms, DevOps methodologies, and ITSM tools such as ServiceNow. * Experience applying Site Reliability Engineering (SRE) principles and capabilities, including observability, monitoring, alerting, performance analysis, incident response, root cause analysis, reliability improvement, and automation to enhance application stability and operational effectiveness. * Experience with cloud platforms, DevOps practices, container technologies, and application monitoring/observability tools (ex: Splunk, Grafana, Ealstic, Azure DevOps, Docker, Kubernetes, MongoDB, CosmosDB & AppDynamics etc.,) * Experience with AI-enabled engineering, AI SRE, or AIOps platforms, including building dashboards, agents, automation tools, AI-assisted development practices (LLM), prompt engineering., * Hands-on experience with ServiceNow ticket management, building SRE level logs & operational dashboards, and production support reporting. * Demonstrated ability to engineer solutions for complex applications, infrastructure, API, batch, data, authentication, and integration challenges in highly available production environments. * Strong knowledge of production reliability practices, including incident management, triage facilitation, root cause analysis, post-incident reviews, and corrective action tracking. * Ability to analyze technical and operational data from multiple sources, identify patterns, draw conclusions, and recommend engineering improvements. * Strong understanding of service level objectives, service level agreements, customer-facing metrics, and production health indicators. Location Expectation: This is a hybrid role requiring a minimum of 3 days per week in office. ## Description When you join MetLife's Global Technology team, you'll be part of a forward-thinking group dedicated to shaping the future of digital solutions for customers worldwide. You'll develop, maintain and support technology applications and delivery, leveraging AI, automation, and contemporary ways of working to enhance experiences and drive business outcomes. Your work will simplify complex processes, improve tech resiliency, and ensure high-performing, seamless solutions that power life's most important moments. In this dynamic environment, you'll collaborate with talented peers across teams and functions, expanding your skills in impactful ways. Ready to push boundaries and set new industry standards? Join us and help drive the future of technology forward. The Opportunity At MetLife, we seek to make a meaningful impact on the lives of our customers and our communities. Global Technology & Operations group (GTO) is a diverse team of Agile practitioners comprised of engineers, developers, and technology leaders with the freedom to create innovative solutions that address core business challenges. This role is for a software engineer who will design, build, enhance, and support software applications and platforms for US Digital applications. You'll work in a collaborative, agile environment and be hands-on with full stack engineering, API and integration development, cloud and DevOps practices, production reliability, observability, and AI-enabled engineering tools., * Design, develop, test, deploy, and maintain full stack software applications across UI, API, data, and integration layers. * Deliver high-quality, secure, scalable, and maintainable code using modern software engineering practices, including code reviews, automated testing, version control, and CI/CD pipelines. * Analyze business and technical requirements, translate them into engineering solutions, and partner with product and platform teams to deliver application enhancements. * Troubleshoot, debug, and resolve complex software defects across applications, API, data, batch, authentication, and integration components. * Apply site reliability engineering practices to improve application availability, performance, resiliency, observability, monitoring, alerting, incident response, automation, and problem management. * Participate in production support and incident response activities as an engineering owner, including root cause analysis, corrective actions, and preventive improvements. * Use AI-enabled engineering and AIOps capabilities to accelerate development, identify incident patterns, analyze trends, automate repetitive tasks, support root cause analysis, and improve operational efficiency. * Contribute to the integration of AI capabilities into applications, including GenAI APIs, LLM-based tools, agents, and intelligent self-service or support features. * Collaborate with cross-functional teams to improve system design, reduce technical debt, strengthen operational readiness, and deliver reliable customer-facing digital experiences. * Perform related duties as assigned or requested., MetLife is an Equal Opportunity Employer. All employment decisions are made without regards to race, color, national origin, religion, creed, sex (including pregnancy, childbirth, or related medical conditions), sexual orientation, gender identity or expression, age, disability, marital or domestic/civil partnership status, genetic information, citizenship status (although applicants and employees must be legally authorized to work in the United States), uniformed service member or veteran status, or any other characteristic protected by applicable federal, state, or local law ("protected characteristics"). If you need an accommodation due to a disability, please email us at accommodations@metlife.com. This information will be held in confidence and used only to determine an appropriate accommodation for the application process. MetLife maintains a drug-free workplace. This posting is for a current vacancy and is anticipated to remain open for at least 90 days from the listed posting date. ## Related Videos - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [The 12 Best Jobs for Software Engineers](https://www.wearedevelopers.com/magazine/401-the-12-best-jobs-for-software-engineers) - [How Much Does a Software Engineer Make? Realistic Software Engineering Salaries](https://www.wearedevelopers.com/magazine/425-how-much-does-a-software-engineer-make-realistic-software-engineering-salaries) - [Best Countries for Software Engineers](https://www.wearedevelopers.com/magazine/267-best-countries-for-software-engineers)