> Markdown version of [/jobs/ext/2773823-infrastructure-engineer-iii](https://www.wearedevelopers.com/jobs/ext/2773823-infrastructure-engineer-iii). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Infrastructure Engineer III - **Company:** American Express Company - **Location:** Phoenix, AZ, United States - **Experience:** Experienced - **Salary:** $103,750.0 - $174,750.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Amazon Web Services, Application Performance Management, Audit Trail, Microsoft Azure, Bash Shell, Cloud Computing, Cloud Engineering, Cluster Analysis, Data Masking, Elasticsearch, Monitoring of Systems, Python (Programming Language), OpenShift, Performance Tuning, Role-Based Access Control, Cloud Services, Ansible, Prometheus, Security Assertion Markup Language (SAML), Working Model 2D, Scripting, Google Cloud, System Availability, Grafana, Containerization, Kubernetes, Information Technology, Influxdb, Restful APIs, Terraform, Splunk, Data Pipelines, Dynatrace, Api Management, Microservices - **Published:** September 7, 2026 - **Apply:** https://dejobs.org/x/x/987D583CD8314C6E91BA72CEFBD09F37/job/ ## About the Role * Overall 5+ years of IT experience, with 3+ years of hands-on experience administering Splunk, Grafana, and/or Dynatrace in large-scale enterprise environments. * Strong expertise in observability concepts, including logs, metrics, traces, APM, infrastructure monitoring, and synthetic monitoring. * Proven experience with Splunk indexers, forwarders, search heads, clustering, and performance optimization. * Hands-on experience with Grafana dashboard development and data sources such as Prometheus, Loki, Elasticsearch, and cloud-native monitoring tools. * Strong Dynatrace experience, including OneAgent deployment, service monitoring, custom metrics, and AI-driven root cause analysis. * Experience integrating observability platforms with cloud infrastructure, Kubernetes, and microservices-based architectures. * Proficiency in scripting, automation, and API integrations (Python, Shell, REST APIs, Terraform, Ansible). * Knowledge of security, compliance, and governance practices in monitoring platforms (SOC2, ISO, NIST, audit requirements). * Relevant certifications (Splunk Certified Admin/Architect, Dynatrace Associate/Professional, Grafana certifications) preferred. * Strong communication skills with the ability to collaborate across engineering, SRE, and leadership teams. Depending on factors such as business unit requirements, the nature of the position, cost and applicable laws, American Express may provide visa sponsorship for certain positions ## Description The Technology organization enables and accelerates the company's growth strategies, delivering global capabilities and services in support of Amex's customers and colleagues, while maintaining 24/7 servicing and availability to ensure an uninterrupted, high-quality customer experience. Technology provides the foundation for everything we do in the company while driving differentiation through building and leveraging innovative technology and data insights., * Administer, configure, upgrade, and optimize Splunk, Grafana, and Dynatrace platforms, ensuring high availability, security, scalability, and performance. * Manage end-to-end observability solutions, including log aggregation (Splunk), metrics and visualization (Grafana), and application performance monitoring (Dynatrace). * Design, implement, and maintain monitoring, alerting, and dashboarding standards across infrastructure, applications, and cloud platforms. * Perform platform upgrades, patching, and lifecycle management with minimal downtime and adherence to enterprise change management processes. * Integrate Splunk, Grafana, and Dynatrace with cloud services (AWS, Azure, GCP), container platforms (Kubernetes, OpenShift), and CI/CD pipelines. * Configure and manage data ingestion pipelines, indexes, retention policies, parsing rules, and performance tuning for large-scale log and metric volumes. * Implement Dynatrace OneAgent deployments, SmartScape topology mapping, service flow monitoring, and root cause analysis (Davis AI). * Develop and maintain Grafana dashboards using Prometheus, Loki, InfluxDB, and other supported data sources. * Enforce security, access controls, and compliance requirements (SSO/SAML, RBAC, audit logging, data masking, retention policies). * Automate operational tasks using scripting and APIs (Python, Shell, REST APIs, Terraform, Ansible). * Troubleshoot platform issues, performance bottlenecks, ingestion delays, alert noise, and data gaps across observability tools. * Collaborate with application, infrastructure, and SRE teams to improve monitoring coverage, incident response, and operational resilience. * Document platform standards, best practices, and provide guidance or training to engineering and operations team., We back our colleagues with the support they need to thrive, professionally and personally. That's why we have Amex Flex, our enterprise working model that provides greater flexibility to colleagues while ensuring we preserve the important aspects of our unique in-person culture. Depending on role and business needs, colleagues will either work onsite, in a hybrid model (combination of in-office and virtual days) or fully virtually. ## Related Videos - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Dev & Test in the Cloud? Deploy your cloud environments with Ansible & Terraform](https://www.wearedevelopers.com/videos/1607-dev-test-in-the-cloud-deploy-your-cloud-environments-with-ansible-terraform) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Keycloak case study: Making users happy with service level indicators and observability](https://www.wearedevelopers.com/videos/1599-keycloak-case-study-making-users-happy-with-service-level-indicators-and-observability) - [Easy Mode Monitoring and Logging with Shiftmon](https://www.wearedevelopers.com/videos/2114-easy-mode-monitoring-and-logging-with-shiftmon) ## Related Articles - [Top-Paying Tech Jobs (with Salaries)](https://www.wearedevelopers.com/magazine/372-top-paying-tech-jobs-with-salaries) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Best Paying Jobs in Technology](https://www.wearedevelopers.com/magazine/256-best-paying-jobs-in-technology) - [The Most Popular IT Jobs on the Market](https://www.wearedevelopers.com/magazine/376-the-most-popular-it-jobs-on-the-market) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Highest Paying Tech Companies in Europe](https://www.wearedevelopers.com/magazine/162-highest-paying-tech-companies-in-europe)