> Markdown version of [/jobs/ext/2312096-site-reliability-engineer-lead-sre-internal-kubernetes-container-platform-ikcp](https://www.wearedevelopers.com/jobs/ext/2312096-site-reliability-engineer-lead-sre-internal-kubernetes-container-platform-ikcp). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer Lead (SRE) - Internal Kubernetes Container Platform (IKCP) - **Company:** Bank of America - **Location:** Jersey City, NJ, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Kubernetes Security, Bash Shell, Cloud Computing, Cloud Computing Security, Information Systems, Computer Networks, Continuous Integration, DevOps, Disaster Recovery, Distributed Systems, Domain Name System (DNS), Github, Monitoring of Systems, Python (Programming Language), Linux System Administration, Networking Basics, Octopus Deploy, OpenShift, Performance Tuning, Role-Based Access Control, Reliability Engineering, Ansible, Prometheus, Runbook, Software Engineering, Virtualization Technology, Software Vulnerability Management, Backup and Restore, Data Logging, Scripting, Load Balancing, Cloud Platform System, Istio, Grafana, Reliability of Systems, Gitlab, Containerization, Git Flow, Kubernetes, Information Technology, Rancher, Bitbucket, Terraform, Splunk, Dynatrace, Jenkins, Golang, Vmware - **Published:** August 30, 2026 - **Apply:** https://dejobs.org/x/x/389B71DCDE2E4E04B5BD369880DA983A/job/ ## About the Role * 8+ years of infrastructure, cloud, platform engineering, or SRE experience. * 5+ years managing Kubernetes and/or OpenShift production environments. * Experience operating large-scale mission-critical distributed systems. * Experience supporting enterprise production environments with 24x7 operational responsibilities. Technical Skills * Kubernetes, OpenShift, Rancher, VKS container orchestration platforms. * Linux administration and troubleshooting. * Terraform, Ansible, GitOps, ArgoCD, Helm * CI/CD platforms such as Jenkins, GitHub, GitLab, Bitbucket, or equivalent. * Monitoring and observability tools such as Dynatrace, Prometheus, Grafana, Splunk, ELK, OpenTelemetry. * Infrastructure as Code and automation frameworks. * Networking fundamentals, load balancing, ingress, DNS, and service mesh concepts. * Storage platforms, backup technologies, and disaster recovery solutions. * Scripting in Python, Go, Bash, or similar languages., * BS /MS degree in Computer Science, Engineering, Information Systems, or related technical discipline, or equivalent experience. * OpenShift Administration or Kubernetes certifications. * Experience running large-scale enterprise container platforms. * Experience with virtualization technologies including VMware, VCF, and OpenShift Virtualization. * Experience implementing cloud-native security controls and platform governance. * Knowledge of platform engineering, developer experience, and platform-as-a-product operating models. * Experience with vulnerability management and container security scanning solutions. * Drives operational excellence and continuous improvement. * Demonstrates strong ownership and accountability. * Influences cross-functional teams without direct authority. * Communicate effectively with senior technical and business leaders. * Champions automation-first and reliability-first engineering culture., * Automation * Collaboration * Influence * Production Support * Result Orientation * Analytical Thinking * Application Development * Architecture * Solution Design * Stakeholder Management * Adaptability * DevOps Practices * Project Management * Risk Management * Solution Delivery Process ## Description The IKCP Site Reliability Engineer Lead is responsible for ensuring the reliability, scalability, performance, security, and operational excellence of the enterprise Internal Kubernetes Container Platform (IKCP). This role serves as a technical lead within the platform organization, driving automation, observability, incident management, capacity planning, platform resilience, and continuous improvement across OpenShift, Kubernetes, Rancher, VKS and emerging container platform services. The role partners closely with Engineering, Architecture, Product Management, Security, Infrastructure, and Central Operations teams to deliver a highly available platform-as-a-product experience for application teams. Responsibilities are aligned with IKCP's focus on SLOs, error budgets, observability, runbooks, L3 operations, upgrade orchestration, and platform governance., Reliability & Operations * Own platform reliability objectives, including service availability, resiliency, recoverability, and operational health. * Lead critical incident response, root cause analysis, and problem management activities. * Serve as a senior escalation point for L3 platform support and on-call operations. * Develop and maintain operational runbooks, recovery procedures, and standard operating practices. * Drive production readiness reviews for new platform capabilities and services. * Ensure platforms meet enterprise resiliency and availability objectives. * Conduct resilience exercises and continuous improvement activities following recovery testing. Kubernetes & OpenShift Platform Engineering * Execute platform upgrades, patching strategies, cluster modernization, and release orchestration. * Improve platform scalability, performance, and resource utilization across production and non-production environments. * Support platform modernization initiatives including OpenShift virtualization, VKS, and cloud-native technologies * Collaborate with Product, Architecture, Engineering, and Operations teams to improve developer experience and platform adoption. Observability & Automation * Design and implement enterprise observability solutions leveraging monitoring, logging, tracing, and alerting platforms. * Automate operational processes using Infrastructure-as-Code, GitOps, CI/CD, and scripting frameworks. * Reduce operational toil through self-healing, intelligent automation, and proactive remediation capabilities. * Drive operational efficiency through automation of cluster provisioning, upgrades, compliance, and day-2 operations. Capacity & Performance Engineering * Perform platform capacity planning and trend analysis. * Forecast infrastructure growth requirements and optimize platform resource consumption. * Conduct performance tuning for clusters, workloads, networking, and storage services. * Support enterprise-scale growth while maintaining platform stability and customer experience. Security & Compliance * Partner with security teams to implement platform security controls and governance requirements. * Support vulnerability remediation, image compliance, platform hardening, and policy enforcement. * Implement and maintain RBAC, Network Policies, and container security controls. * Drive compliance with enterprise standards, vulnerability management processes, and audit requirements., This job is responsible for partnering with engineering and technology teams to implement measures prescribed by the Site Reliability Engineer teams it leads. Key responsibilities include ensuring appropriate instrumentation, tooling, ticketing, alerting and on call routines are in place for key services, demonstrating technical expertise within domains, and decomposing objectives into work units. Job expectations include advancing efficient solution delivery practices and promoting exceptional design, engineering, and organizational practices., * Collaborates with Development and Infrastructure teams to understand technical solutions and implement monitoring capabilities outlined in the application and system monitoring designs put forward by the Senior Site Reliability Engineer (SRE) * Develops and maintains reliability scripts, tools and libraries and leverages them for common instrumentation, automation, and operational needs, and when mentoring SRE resources on reliability practices and established tools/capabilities * Partners to implement code changes to make use of common reliability libraries and tools and helps Application Production Services and Application Development teammates understand how to use them * Participates regularly in architecture community of practice meetings and communication via other channels * Identifies vulnerabilities and opportunities for reliability improvement, such as investigating low level error rates and 'noise' in monitoring, and defines solutions to reduce manual support effort and/or improve system reliability * Engages as a subject matter expert in major incident triage efforts and failure scenario modelling and diagnosis with Problem Manager root causes for major incident/problem management investigations, Bank of America and its affiliates consider for employment and hire qualified candidates without regard to race, religious creed, religion, color, sex, sexual orientation, genetic information, gender, gender identity, gender expression, age, national origin, ancestry, citizenship, protected veteran or disability status or any factor prohibited by law, and as such affirms in policy and practice to support and promote the concept of equal employment opportunity, in accordance with all applicable federal, state, provincial and municipal laws. The company also prohibits discrimination on other bases such as medical condition, marital status or any other factor that is irrelevant to the performance of our teammates. View your "Know your Rights (https://www.eeoc.gov/sites/default/files/2023-06/22-088_EEOC_KnowYourRights6.12.pdf) " poster. View the LA County Fair Chance Ordinance (https://dcba.lacounty.gov/wp-content/uploads/2024/08/FCOE-Official-Notice-Eng-Final-8.30.2024.pdf) . ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Rate-limiting using eBPF and Istio: How to protect your SaaS customers from themselves](https://www.wearedevelopers.com/videos/100220-rate-limiting-using-ebpf-and-istio-how-to-protect-your-saas-customers-from-themselves) - [WeAreDevelopers LIVE - Modern DevOps for IoT Devices and More](https://www.wearedevelopers.com/videos/1805-wearedevelopers-live-modern-devops-for-iot-devices-and-more) - [Enabling automated 1-click customer deployments with built-in quality and security](https://www.wearedevelopers.com/videos/83-enabling-automated-1-click-customer-deployments-with-built-in-quality-and-security) - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Learning Kubernetes made easy with KubeCampus](https://www.wearedevelopers.com/magazine/348-learning-kubernetes-made-easy-with-kubecampus) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers)