Associate-Technology Operations Engineering

American Express Company
Fort Lauderdale, FL, United States
3 days ago
Apply on www.miamigigs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$89,250.0 - $150,250.0
Working hours
Regular working hours

Tech stack

C (Programming Language) Java (Programming Language) Application Programming Interfaces (APIs) Artificial Intelligence Amazon Web Services Business Analytics Applications Computing Platforms JIRA Microsoft Azure Bash Shell BigTable BigQuery
+55 more
Business Software Cloud Computing Cloud Computing Security Computer Engineering Couchbase Servers IBM DB2 Linux DevOps Disaster Recovery Distributed Data Store Distributed Systems Endevor Identity and Access Management Spring Framework Python (Programming Language) PostgreSQL Machine Learning Mainframes Apache Maven Model View Controller (MVC) Node.Js NoSQL Object-Oriented Software Development Windows PowerShell IBM Resource Access Control Facility Redis Reliability Engineering Ansible Software Engineering SQL Databases Data Logging Scripting Google Cloud Enterprise Software Applications System Availability Grafana Mttr Git Event Driven Architecture Containerization Kubernetes Information Technology Enterprise Integration Omegamon Real Time Data Apache Kafka Graphql Software Coding Restful APIs Splunk Grpc Appdynamics Dynatrace Docker Jenkins

Job description

This role is responsible for driving the reliability, resiliency, performance, and modernization of critical American Express platforms across Distributed environments. You will leverage deep technical expertise in software engineering, runtime engineering, production support, and platform operations to quickly assess and remediate complex availability, performance, and operational issues. As part of our technology team, you will partner with engineering, product, infrastructure, and operations teams to design, build, automate, and support highly available enterprise platforms. You will help accelerate modernization initiatives, improve operational excellence, and deliver secure, scalable, and resilient solutions that power critical customer and business capabilities., + Serve as a hands-on engineer with experience in supporting complex enterprise applications, platforms, and operational tooling across Distributed environments.

  • This role requires and must have 5+ years of distributed experience and knowledge.
  • Design, develop, prototype, code, test, and implement scalable software solutions using technologies such as Java, Python, SQL, and related frameworks.
  • Act as a technical contributor in change management reviews, root cause analysis, and troubleshooting of complex technical issues.
  • Design and implement automation solutions, and engineering practices that improve platform resiliency, operational efficiency, and security.
  • Use best practices in incident management, problem management and change management as this role focuses on application production support. Runtime Engineering, Reliability & Operations

  • Contribute to the technical roadmap for runtime systems, ensuring platform reliability, scalability, availability, recoverability, and performance.
  • Establish, monitor, and continuously improve key performance indicators (KPIs), service level objectives (SLOs), and operational metrics (MTTR, MTBF) for platform health and resiliency.
  • Perform diagnosis and resolution of production incidents, batch failures, application outages, performance bottlenecks, and infrastructure issues across Mainframe and Distributed platforms.
  • Apply Site Reliability Engineering (SRE) principles and operational excellence practices to improve system stability and reduce operational risk.
  • Support disaster recovery, high availability, workload management, capacity planning, and business continuity initiatives. Distributed Platform Engineering

  • Must have experience in Support and optimize enterprise platforms across distributed technologies including Java, Python, C, SQL, NodeJS, Bash, JS/HTML/CSS, Spanner, BigTable, BigQuery, Spring, GraphQL, OOP, MVC, Algos, Git, CoPilot, Jenkins, XLR, Elastic, Jira, AI, ML, Docker, Kubernetes, Kafka, Rest API, GRPC, Grafana.
  • Desirable Support and optimize Distributed Platform technologies including cloud infrastructure, Linux/Unix, containers, APIs, Java-based services, distributed databases, and modern application platforms.
  • Implement and support Cloud/Distributed architectures, modernization initiatives, API enablement, and enterprise integration capabilities.
  • Knowledge of cloud platforms AWS, GCP or general Cloud fundamentals. Data, Integration & Automation

  • Develop and support enterprise integration solutions utilizing APIs, MQ, Connect:Direct, event-driven architectures, and batch and real-time data integration patterns.
  • Utilize relational and NoSQL databases including DB2, PostgreSQL, Redis, and Couchbase to support critical business applications.
  • Automate operational processes, deployments, monitoring, reporting, and remediation activities using Python, Bash, Ansible, Jenkins, and related technologies.
  • Collaborate with engineering teams to adopt scalable automation and self-service capabilities for deployment, monitoring, and operational support. Observability & Continuous Improvement

  • Implement and utilize monitoring, observability, logging, and analytics solutions using tools such as Splunk, OMEGAMON, RMF/SMF, Sysview, MainView, Dynatrace, AppDynamics, ELK, or equivalent technologies.
  • Analyze operational trends, identify opportunities for optimization, and formulate strategic recommendations to improve platform health and engineering effectiveness.
  • Contribute to continuous improvement initiatives focused on reliability, performance, security, operational maturity, and customer experience. DevOps, Security & Governance

  • Understand CI/CD pipelines and DevOps practices using tools such as Git, Jenkins, Maven, DBB, Endevor, Changeman, ISPW, UrbanCode Deploy, or equivalent platforms.
  • Apply enterprise security controls, compliance requirements, audit standards, and access management practices, including RACF, ACF2, Top Secret, and cloud security principles.
  • Ensure solutions meet non-functional requirements (NFRs) including availability, scalability, performance, security, recoverability, and maintainability.

Requirements

  • Bachelor’s degree in Computer Science, Computer Engineering, and/or comparable experience
  • Work experience in software engineering, app support or infrastructure operations or runtime engineering.
  • A working understanding of cloud infrastructure, distributed systems, and containerization technologies, with experience in supporting critical business applications being a plus.
  • Familiarity with monitoring and logging tools, and incident management best practices, to ensure reliability and performance of applications in a production environment.
  • Solid programming and scripting skills, with hands on experience to automate operational tasks using tools such as Python
  • Knowledge of scripting languages (e.g., PowerShell, Python) for automation tasks
  • Experience in technology operations work
  • Hands on experience with relational and NoSQL databases such as DB2, Redis, Postgres, Couchbase etc.
  • Experience in cloud platforms such as AWS, Azure, or Google Cloud, Public Cloud certification is a plus Employment eligibility to work with American Express in the United States is required as the company will not pursue visa sponsorship for these positions.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.miamigigs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:06 min

Developer experience and project variety at scale

Alexandra Petri · World Congress 2023

3:08 min

Aligning engineering processes with core business impact metrics

Chris Riley · World Congress 2021

4:56 min

Establishing internal service communication with gRPC

Florian Bader Florian Bader · World Congress 2026 Europe

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

3:07 min

Establishing service level agreements directly for internal platforms

Pawel Piwosz · LIVE

1:11 min

Evaluating architectural trade-offs between REST and gRPC

Sakshi Nasha Sakshi Nasha · Europe 2026 Virtual

Videos

See all

Related articles

See all