> Markdown version of [/jobs/ext/2533423-director-ai-platform-reliability](https://www.wearedevelopers.com/jobs/ext/2533423-director-ai-platform-reliability). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Director, AI Platform Reliability - **Company:** LogicMonitor, Inc. - **Location:** San Francisco, CA, United States - **Experience:** Expert - **Salary:** $247,500.0 - $275,000.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Application Programming Interfaces (APIs), Amazon Web Services, Computing Platforms, Microsoft Azure, Batch Processing, Cloud Computing, Profiling, Continuous Integration, Distributed Systems, Memory Management, Fault Tolerance, Java Virtual Machine (JVM), Load Testing, NoSQL, Reliability Engineering, Data Streaming, Multithreading, Google Cloud, Cloud Platform System, Data Ingestion, Spring Cloud, Concurrency, Reliability of Systems, Technical Debt, Backend, Event Driven Architecture, Data Lakes, AI Platforms, Kubernetes, Storage Technologies, Low Latency, Apache Kafka, Data Management, Data Lakehouse, Data Pipelines, Serverless Computing, Microservices - **Published:** August 19, 2026 - **Apply:** https://dejobs.org/x/x/8A6306EE81E549E1A589153F9027356A/job/ ## About the Role * 10+ years of professional software-engineering experience, including significant experience building large-scale distributed systems. * Experience leading engineering teams, architects, and staff engineers. * Demonstrated success delivering and operating platforms that process hundreds of millions of transactions, requests, or events. * Deep technical expertise in Java, JVM performance, concurrency, multithreading, memory management, and application profiling. * Strong experience designing and operating microservice-based and event-driven architectures. * Extensive production experience with Apache Kafka or a comparable distributed streaming platform. * Strong understanding of Kafka partitioning, replication, consumer groups, offset management, ordering, delivery semantics, schema evolution, and reprocessing. * Experience designing low-latency, highly available APIs and backend services. * Experience managing terabyte- or petabyte-scale datasets across relational, NoSQL, streaming, and object-storage technologies. * Strong understanding of distributed-systems concepts, including consensus, replication, partitioning, consistency models, idempotency, backpressure, and fault tolerance. * Experience operating cloud-native applications using Kubernetes, containers, infrastructure as code, and automated CI/CD pipelines. * Experience with at least one major cloud platform, such as AWS, Google Cloud, or Microsoft Azure. * Strong knowledge of observability practices involving metrics, logs, traces, profiling, alerting, dashboards, and service-level objectives. * Demonstrated experience improving system reliability, scalability, latency, cost efficiency, and engineering productivity. * Strong written and verbal communication skills, including the ability to explain complex technical decisions to engineering teams, executives, and business stakeholders. * Proven ability to build inclusive, accountable, and high-performing engineering organizations., At this time, we are able to consider candidates who are authorized to work in the United States on a full-time, permanent basis without requiring new or initial employer-sponsored work authorization. Candidates who currently hold valid U.S. work authorization that can be transferred to a new employer (such as certain H-1B statuses) may be considered on a case-by-case basis. ## Description We are looking for an accomplished and hands-on Director of AI Platform Reliability to lead the architecture, development, and operation of highly scalable, distributed software platforms. This leader will be responsible for systems that process hundreds of millions/billions of transactions and events , manage terabytes to petabytes of data , and deliver reliable, low-latency services to enterprise customers. The ideal candidate combines strong engineering depth in Java, Kafka, distributed systems, and cloud-native microservices with a demonstrated ability to build and lead high-performing engineering organizations. This is a strategic leadership role, but it requires a leader who can remain close to the technology, participate in architecture reviews, challenge design decisions, guide teams through complex production problems, and establish the engineering practices required to operate mission-critical platforms at scale. Here's a closer look at this key role: * Lead and scale multiple engineering teams responsible for high-volume, business-critical distributed systems and data platforms. * Define the technical strategy and architecture for platforms processing hundreds of millions of transactions and terabytes of data. * Guide the development of Java-based microservices, APIs, Kafka streaming pipelines, batch-processing workflows, and cloud-native services. * Build and evolve scalable data lake and Data Lakehouse platforms supporting real-time, near-real-time, and batch analytics workloads. * Establish reliable data ingestion, transformation, storage, governance, lineage, retention, and data-quality practices across streaming and batch pipelines. * Build low-latency, highly available, fault-tolerant systems with strong scalability, resiliency, and disaster-recovery capabilities. * Define and own operational SLAs, SLOs, availability targets, recovery objectives, and performance metrics for critical services and data pipelines. * Drive capacity planning, load testing, throughput optimization, and improvements to p95 and p99 latency. * Ensure effective Kafka design, including partitioning, consumer groups, ordering, schema evolution, replay, and lag management. * Establish engineering standards for architecture, coding, testing, security, observability, and production readiness. * Partner with Product, Architecture, SRE, Security, Data, and Infrastructure teams to deliver strategic platform initiatives. * Strengthen operational excellence through monitoring, incident management, on-call practices, root-cause analysis, and continuous reliability improvements. * Recruit, mentor, and develop engineering managers, architects, and senior technical leaders. * Improve developer productivity, CI/CD automation, deployment safety, and release predictability. * Manage technical debt, platform modernization, cloud costs, and long-term scalability investments. ## Related Videos - [Leveraging Real time data in FSIs](https://www.wearedevelopers.com/videos/806-leveraging-real-time-data-in-fsis) - [Developing the Backend with Stefan Lingler, CTO at Shpock](https://www.wearedevelopers.com/videos/100360-developing-the-backend-with-stefan-lingler-cto-at-shpock) - [Profiling Symfony & PHP apps with Blackfire](https://www.wearedevelopers.com/videos/265-profiling-symfony-php-apps-with-blackfire) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [NoSQL Data Modeling for Front-end Developers](https://www.wearedevelopers.com/videos/297-nosql-data-modeling-for-front-end-developers) - [Inside Bitpanda's Tech Stack: Scaling a European Fintech Leader - Markus Dorner](https://www.wearedevelopers.com/videos/1979-inside-bitpanda-s-tech-stack-scaling-a-european-fintech-leader-markus-dorner) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers)