> Markdown version of [/jobs/ext/3006769-senior-software-engineer-data-and-ai-infrastructure](https://www.wearedevelopers.com/jobs/ext/3006769-senior-software-engineer-data-and-ai-infrastructure). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Software Engineer, Data and AI Infrastructure - **Company:** Airwallex US, LLC - **Location:** Seattle, WA, United States - **Experience:** Expert - **Salary:** $150,000.0 - $245,000.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Artificial Intelligence, Amazon Web Services, User Authentication, Microsoft Azure, Batch Processing, Big Data, Cloud Computing, Software Debugging, DevOps, Distributed Data Store, Distributed Systems, Domain Name System (DNS), Memory Management, Fault Tolerance, Python (Programming Language), Network Security, Machine Learning, Routing, Open Source Technology, Performance Tuning, Service Discovery, Data Streaming, AI Infrastructure, Data Processing, Transport Layer Security, Google Cloud, Load Balancing, Cloud Platform System, Autoscaling, Apache Spark, Parallel Computation, Rate Limiting, Kubernetes, Apache Flink, Real Time Data, Apache Kafka, Machine Learning Operations - **Published:** September 19, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=828b16062fb735f2 ## About the Role * 5+ years in DevOps, SRE, or platform engineering, owning production systems end to end * Strong experience designing, operating, and troubleshooting production Kubernetes environments. * Experience building or operating distributed data infrastructure with technologies such as Kafka, Spark, or Flink OR experience developing AI infrastructure such as AI gateways or model-routing platforms. * Hands-on experience with at least one major public cloud platform, such as AWS, Google Cloud, or Microsoft Azure. * Strong knowledge of cloud and container networking, including DNS, load balancing, ingress, service discovery, TLS, routing, and network security. * Proficiency in one or more of Go, Python, or Java, with experience writing maintainable production software. * A solid understanding of distributed-systems concepts, including availability, consistency, fault tolerance, backpressure, and horizontal scalability. * Experience operating critical infrastructure using infrastructure-as-code, automated delivery, and modern observability practices. * Strong debugging skills and the ability to work methodically across multiple layers of a complex system. * Clear communication skills and a track record of collaborating effectively across engineering disciplines. * An ownership mindset: you identify important problems, drive them to resolution, and improve the underlying system rather than treating symptoms., * Platform engineering experience, particularly building internal developer platforms or paved-road workflows used by multiple engineering teams. * Hands-on experience with serving technologies such as SGLang, vLLM, or NVIDIA Triton Inference Server. * Knowledge of GPU scheduling, batching, model parallelism, memory management, autoscaling, and inference-performance optimization. * Experience improving the cost efficiency of large-scale data processing or AI inference workloads. * Contributions to infrastructure, data-platform, Kubernetes, or AI-serving open-source projects. ## Description We are looking for a Senior Software Engineer to build the infrastructure that powers our data and AI platforms. You will design and operate distributed systems that support high-throughput data processing, real-time workloads, and production AI applications. This includes evolving our Kubernetes and cloud foundations, improving the reliability and scalability of platforms such as Kafka, Spark, and Flink, and building infrastructure for AI traffic management and model serving. This is a high-impact role for an engineer who enjoys solving complex infrastructure problems, writing production software, and giving other engineering teams reliable self-service platforms. You will work across application, data, machine learning, security, and infrastructure teams to establish the technical foundations for the company's next stage of growth., * Design, build, and operate highly available data and AI infrastructure on Kubernetes and public cloud platforms. * Develop scalable platforms for streaming, batch processing, and real-time data workloads using technologies such as Kafka, Spark, and Flink. * Build and evolve AI infrastructure, including AI gateways, model-routing layers, traffic management, rate limiting, authentication, observability, and usage controls. * Develop self-service capabilities that enable data, AI, and application teams to deploy and operate workloads safely and independently. * Partner with engineering teams to translate emerging data and AI requirements into durable platform capabilities., * Delivered meaningful improvements to the scalability, reliability, or efficiency of our data and AI infrastructure. * Reduced the operational effort required to deploy and manage data or AI workloads. * Improved visibility into system performance, reliability, capacity, and cost. * Established reusable platform capabilities adopted by engineering teams. * Helped define the technical direction for our next generation of data and AI infrastructure. ## Related Videos - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Creating a routing app with Google Maps API from scratch](https://www.wearedevelopers.com/videos/831-creating-a-routing-app-with-google-maps-api-from-scratch) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) - [A Technical Introduction to Bitcoin's 2nd Layer- The Lightning Network](https://www.wearedevelopers.com/videos/15-a-technical-introduction-to-bitcoin-s-2nd-layer-the-lightning-network) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)