> Markdown version of [/jobs/ext/2506354-sr-platform-engineer](https://www.wearedevelopers.com/jobs/ext/2506354-sr-platform-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Sr. Platform Engineer - **Company:** Alkami - **Location:** United States (Remote available) - **Experience:** Expert - **Salary:** $145,000.0 - $165,000.0 - **Contract:** Permanent contract - **Skills:** Query Performance, Component-Based Software Engineering, Application Performance Management, Cloud Computing, Continuous Integration, Relational Databases, Software Design Patterns, Distributed Systems, Github, Message Broker, Node.Js, Reliability Engineering, Software Engineering, TypeScript, Datadog, Caching, Containerization, Kubernetes, Information Technology, Apache Kafka, Terraform, Dynatrace, Microservices - **Published:** August 5, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=e5db70238baf3319 ## About the Role 4 to 7 years of experience in software engineering, platform engineering, or a hybrid development/reliability engineering role, Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent work experience, * Strong proficiency in TypeScript, with production experience in a Node.js/TypeScript service environment * Hands-on experience deploying and troubleshooting containerized workloads in Kubernetes * Experience building and maintaining CI/CD pipelines, GitHub Actions preferred * Experience configuring monitoring, dashboards, and alerting in an APM/observability tool, Datadog preferred * Experience troubleshooting distributed system and microservice communication failures across a variety of integration points and protocols * Experience with distributed tracing tools and practices * Working familiarity with relational databases, sufficient to spot and tune slow or problematic queries * Experience diagnosing and resolving application performance issues (caching, inefficient code paths, slow queries) * Experience designing for and testing failure modes (fault injection, resilience/chaos-style testing) and implementing defensive patterns such as idempotency, retries, and circuit breakers * Strong analytical and troubleshooting skills; ability to work independently on ambiguous, reactive reliability problems * Ability to communicate technical root cause and remediation plans clearly to both engineering and non-technical stakeholders Preferred * Experience with message brokers or event-streaming platforms such as Kafka * Experience with OpenTelemetry or comparable distributed tracing frameworks * Experience in a regulated or compliance-driven environment (fintech, banking, or similar) * Familiarity with infrastructure-as-code tooling (Terraform or similar) * Experience partnering with infrastructure/SRE teams on issues that cross application and infrastructure boundaries ## Description The Sr Platform Engineer is responsible for locating, making visible, and remediating sources of unreliability in the MANTL platform, including correctness problems that surface under failure conditions. This role works directly in the platform's application codebase (TypeScript) and its container-native deployment environment, combining application-engineering skill with reliability-engineering practice. Much of the work is reactive, investigating and resolving issues as they surface, balanced against a standing roadmap of known reliability risks the team has identified and prioritized ahead of time, independent of feature-delivery timelines. This role partners with, but is organizationally and functionally distinct from, both Cloud Infrastructure Engineering and product application engineering teams, focusing specifically on reliability concerns that span or fall between those domains., * Investigate, troubleshoot, and resolve reliability issues within MANTL platform application code, including microservice communication failures and correctness issues that emerge under failure conditions * Identify and address failure modes across the platform's third-party and internal system integrations, developing resilience strategies suited to each integration's specific behavior * Design, configure, and maintain monitoring, dashboards, and alerting (Datadog preferred) to increase visibility into platform health and surface emerging issues before they become incidents * Implement and extend distributed tracing across microservices to accelerate root-cause identification for cross-service failures * Diagnose and remediate application performance issues, including caching strategy, inefficient code paths, and query performance * Harden platform and application components against known failure modes through fault-injection and resilience testing, implementing defensive design patterns to prevent recurrence * Build, maintain, and troubleshoot CI/CD build pipelines (GitHub Actions) supporting deployment of the platform * Deploy and troubleshoot container-native (Kubernetes) workloads as part of diagnosing and resolving platform reliability issues * Maintain and execute against a roadmap of known reliability risks, independent of feature-delivery timelines * Create and maintain documentation and runbooks covering platform reliability issues, root causes, and remediations * Act as an escalation point for complex platform reliability issues, partnering with Cloud Infrastructure Engineering and application engineering teams on issues that cross domain boundaries * Contribute to defining reliability targets (SLOs/SLIs) for key platform services ## Related Videos - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [HTTP headers that make your website go faster](https://www.wearedevelopers.com/videos/1676-http-headers-that-make-your-website-go-faster) - [Stop using Node.js like in 2020! What changed and what you can do today with Node.js](https://www.wearedevelopers.com/videos/100011-stop-using-node-js-like-in-2020-what-changed-and-what-you-can-do-today-with-node-js) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Bringing AI Model Testing and Prompt Management to Your Codebase with GitHub Models](https://www.wearedevelopers.com/videos/1536-bringing-ai-model-testing-and-prompt-management-to-your-codebase-with-github-models) - [Forget Developer Platforms, Think Developer Productivity!](https://www.wearedevelopers.com/videos/1121-forget-developer-platforms-think-developer-productivity) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [The Best X (Twitter) Accounts for Developers](https://www.wearedevelopers.com/magazine/294-the-best-x-twitter-accounts-for-developers) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs)