> Markdown version of [/jobs/ext/2623314-senior-platform-software-engineer-cache-query-and-compute-platforms](https://www.wearedevelopers.com/jobs/ext/2623314-senior-platform-software-engineer-cache-query-and-compute-platforms). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Platform Software Engineer (Cache, Query and Compute Platforms) - **Company:** Medallia - **Location:** McLean, VA, United States - **Experience:** Expert - **Salary:** $138,000.0 - $205,000.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Amazon S3, Cloud Computing, DevOps, Failover, Fault Tolerance, Apache Hive, Python (Programming Language), Metadata Repositories, Performance Tuning, Redis, Reliability Engineering, System Software, System Availability, Apache Spark, Kubernetes, Infrastructure Automation Frameworks, Golang, Programming Languages - **Published:** August 14, 2026 - **Apply:** https://jobs.localjobnetwork.com/apply/add/88030183/1 ## About the Role We are seeking a hands-on Senior Platform Software Engineer with experience operating cache, distributed query, and compute platforms at scale. As a key member of the Platform Services team, you will help ensure the availability, reliability, performance, and operational readiness of Redis, Kvrocks, Trino, and Spark., * 5 years of experience in software, systems, platform, DevOps, or reliability engineering. * 3 or more years of experience operating distributed platforms in production. * Cache & Query Platforms: Hands-on operational experience managing and scaling at least two of the following platforms in production: Redis, Kvrocks, Trino, or Spark. * Experience leading production incident triage, root-cause analysis (RCA), and performance optimization for memory fragmentation, query execution bottlenecks, cluster failovers, and CPU/memory resource utilization. * High Availability & Fault Tolerance: Experience designing, testing and executing disaster recovery plans, automated node failovers, cluster re-sharding, and zero-downtime upgrades for high-throughput platforms * Demonstrated experience in Java, Go, Python, or a similar programming language. * Technical Collaboration & Review: Experience conducting formal architecture reviews, authoring technical design documents (e.g., RFCs/ADRs), and establishing operational runbooks across engineering teams. Preferred Qualifications * Experience with additional platforms among Redis, Kvrocks, Trino, and Spark. * Experience operating distributed platforms on Kubernetes or cloud infrastructure. * Experience integrating Trino with object storage, Hive Metastore, or external data catalogs. * Experience validating topology, failover, recovery, and application-health behavior. * Experience with platform automation, observability, upgrades, capacity planning, and failure-mode testing. * Experience supporting large-scale, business-critical cache, query, or compute platforms. ## Description * Operate and maintain Redis, Kvrocks, Trino, and Spark platforms. * Manage Redis Cluster and Sentinel deployments, replication, failover, persistence, upgrades, and performance. * Validate Kvrocks topology, replication, failover, recovery, and operational readiness. * Troubleshoot Trino coordinators, workers, catalogs, query execution, S3 access, and Hive Metastore integrations. * Operate Spark clusters and troubleshoot drivers, executors, scheduling, failover, resource usage, and application health. * Lead production incident triage, root-cause analysis, and performance tuning. * Plan and validate upgrades, capacity, failover scenarios, and recovery procedures. * Review application architectures and production-readiness requirements. * Build operational automation, monitoring, alerts, documentation, and troubleshooting playbooks. * Participate in a periodic on-call rotation to maintain 24/7 reliability and performance of production services. Candidates based in the Tysons vicinity will be prioritized as this role is Hybrid, 3 days per week onsite. ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Reducing LLM Calls with Vector Search Patterns - Raphael De Lio (Redis)](https://www.wearedevelopers.com/videos/1714-reducing-llm-calls-with-vector-search-patterns-raphael-de-lio-redis) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) - [Accelerating Authentication Architecture: Taking Passwordless to the Next Level](https://www.wearedevelopers.com/videos/733-accelerating-authentication-architecture-taking-passwordless-to-the-next-level) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline)