Senior Software Engineer - SRE

T-Mobile Us, Inc.
Atlanta, GA, United States
about 1 month ago
Apply on us.experteer.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
3 years minimum
Working hours
Regular working hours

Tech stack

JavaScript (Programming Language) Artificial Intelligence Amazon Web Services Automation of Tests Microsoft Azure Mobile Application Development Cloud Engineering Software Quality Computer Programming Continuous Integration Identity and Access Management Python (Programming Language)
+14 more
Key Management Node.Js Windows PowerShell Runbook Software Engineering Vault (Revision Control System) Software Vulnerability Management Data Logging Enterprise Software Applications Cloud Platform System Large Language Models Kubernetes Terraform Docker

Job description

Experteer Overview In this role you lead the design, development, testing, and operation of secure, scalable software platforms using SRE and AI-native practices. You collaborate across teams to support cloud-native, containerized systems with strong production reliability and automation. You design and implement infrastructure-as-code and standardize capabilities like certificate lifecycle and secrets management. You apply AI-assisted development to improve software quality and delivery, while mentoring the SRE team and strengthening regional capabilities. This is a hands-on leadership role focused on resilience and scalable, secure software delivery. Compensation / Benefits * Lead SRE architecture and standardization across cloud platforms, containers, deployment patterns, certificate lifecycle management, secrets management, and Vault solutions * Integrate AI-native development practices and automation to improve scalability and delivery performance, mentoring the SRE team and strengthening GCC capabilities * Maintain technical documentation, runbooks, architecture decisions, and reusable patterns to improve supportability * Own production reliability for enterprise applications, including monitoring, incident response, root-cause analysis, and service readiness * Design, develop, automate, test, and optimize software using Terraform-based infrastructure-as-code and modern testing frameworks * Contribute to design innovations and adopt new frameworks and best practices * Collaborate with technical teams to deliver solutions and mentor others through knowledge sharing and training * Support technology strategy by evaluating current technologies aligned with business goals * Create clear documentation for code, designs, and requirements to support knowledge sharing * Undertake other duties/projects as assigned by management Tasks * Bachelor’s degree with 5 years of related work experience or advanced degree with 3 years * 4-7 years technical engineering experience * Experience owning reliability, monitoring, incident response, and operational readiness for production services * Hands-on design and implementation using Terraform; scripting in Python, PowerShell, or JavaScript/Node.js * Cloud-native platforms with AWS and/or Azure; Docker, Kubernetes; performance and cost awareness * CI/CD and deployment engineering with secure, repeatable pipelines; automated testing; release validation; rollback controls * Security-focused platform operations including vulnerability remediation, IAM, certificate lifecycle and Vault * Observability practices: logging, metrics, tracing, alerting, SLOs/SLIs, dashboards * Experience with AI-enabled services, AI-assisted development, LLM, RAG, embeddings, or anomaly detection * Strong analytical thinking, leadership and communication for cross-team collaboration * Programming, software design and documentation skills * Legally authorized to work in the United States * Travel willingness Key requirements * medical, dental and vision insurance * 401(k) and stock plans * paid time off and holidays * paid parental and family leave * tuition assistance and college coaching * commuter and transit programs

Requirements

response, owning reliability, monitoring, incident response, and operational readiness for production services * Hands-on design and implementation using Terraform; scripting in Python, PowerShell, or JavaScript/Node.js * Cloud-native platforms with AWS and/or Azure; Docker, Kubernetes; performance and cost awareness * CI/CD and deployment engineering with secure, repeatable pipelines; automated testing; release validation; rollback controls * Security-focused platform operations including vulnerability remediation, IAM, certificate lifecycle and Vault * Observability practices: logging, metrics, tracing, alerting, SLOs/SLIs, dashboards * Experience with AI-enabled services, AI-assisted development, LLM, RAG, embeddings, or anomaly detection * Strong analytical thinking, leadership and communication for cross-team collaboration * Programming, software design and documentation skills * Legally authorized to work in the United States * Travel willingness Key requirements * medical, dental a

  • vision insurance * 401(k) and stock plans * paid time off and holidays * paid parental and family leave * tuition assistance and college coaching * commuter and transit programs

About the company

Experteer Overview In this role you lead the design, development, testing, and operation of secure, scalable software platforms using SRE and AI-native practices. You collaborate across teams to support cloud-native, containerized systems with strong production reliability and automation. You design and implement infrastructure-as-code and standardize capabilities like certificate lifecycle and secrets management. You apply AI-assisted development to improve software quality and delivery, while mentoring the SRE team and strengthening regional capabilities. This is a hands-on leadership role focused on resilience and scalable, secure software delivery. Compensation / Benefits * Lead SRE architecture and standardization across cloud platforms, containers, deployment patterns, certificate lifecycle management, secrets management, and Vault solutions * Integrate AI-native development practices and automation to improve scalability and delivery performance, mentoring the SRE team and aaaaaaa

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on us.experteer.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

7:01 min

Career progression from backend programming to professional mobile engineering

Edoardo Dusi · LIVE

45 sec

Working securely with Node.js path application programming interfaces

Sonya Moisset · World Congress 2023

2:50 min

Introduction and the value of runbooks

Hila Fish · World Congress 2023

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

2:14 min

Exploring internal AI product initiatives and global engineering roles

Maria Apazoglou · Coffee With Developers

3:55 min

Identifying underlying Node.js runtime vulnerabilities using fuzzing tools

Sonya Moisset · World Congress 2023

Videos

See all

Related articles

See all