Senior Platform Software Engineer (Cache, Query and Compute Platforms)

Medallia
McLean, VA, United States
22 days ago
Apply on jobs.localjobnetwork.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$138,000.0 - $205,000.0
Working hours
Regular working hours

Tech stack

Java (Programming Language) Amazon S3 Cloud Computing DevOps Failover Fault Tolerance Apache Hive Python (Programming Language) Metadata Repositories Performance Tuning Redis Reliability Engineering
+7 more
System Software System Availability Apache Spark Kubernetes Infrastructure Automation Frameworks Golang Programming Languages

Job description

  • Operate and maintain Redis, Kvrocks, Trino, and Spark platforms.
  • Manage Redis Cluster and Sentinel deployments, replication, failover, persistence, upgrades, and performance.
  • Validate Kvrocks topology, replication, failover, recovery, and operational readiness.
  • Troubleshoot Trino coordinators, workers, catalogs, query execution, S3 access, and Hive Metastore integrations.
  • Operate Spark clusters and troubleshoot drivers, executors, scheduling, failover, resource usage, and application health.
  • Lead production incident triage, root-cause analysis, and performance tuning.
  • Plan and validate upgrades, capacity, failover scenarios, and recovery procedures.
  • Review application architectures and production-readiness requirements.
  • Build operational automation, monitoring, alerts, documentation, and troubleshooting playbooks.
  • Participate in a periodic on-call rotation to maintain 24/7 reliability and performance of production services.

Candidates based in the Tysons vicinity will be prioritized as this role is Hybrid, 3 days per week onsite.

Requirements

We are seeking a hands-on Senior Platform Software Engineer with experience operating cache, distributed query, and compute platforms at scale. As a key member of the Platform Services team, you will help ensure the availability, reliability, performance, and operational readiness of Redis, Kvrocks, Trino, and Spark., * 5 years of experience in software, systems, platform, DevOps, or reliability engineering.

  • 3 or more years of experience operating distributed platforms in production.
  • Cache & Query Platforms: Hands-on operational experience managing and scaling at least two of the following platforms in production: Redis, Kvrocks, Trino, or Spark.
  • Experience leading production incident triage, root-cause analysis (RCA), and performance optimization for memory fragmentation, query execution bottlenecks, cluster failovers, and CPU/memory resource utilization.
  • High Availability & Fault Tolerance: Experience designing, testing and executing disaster recovery plans, automated node failovers, cluster re-sharding, and zero-downtime upgrades for high-throughput platforms
  • Demonstrated experience in Java, Go, Python, or a similar programming language.
  • Technical Collaboration & Review: Experience conducting formal architecture reviews, authoring technical design documents (e.g., RFCs/ADRs), and establishing operational runbooks across engineering teams.

Preferred Qualifications

  • Experience with additional platforms among Redis, Kvrocks, Trino, and Spark.
  • Experience operating distributed platforms on Kubernetes or cloud infrastructure.
  • Experience integrating Trino with object storage, Hive Metastore, or external data catalogs.
  • Experience validating topology, failover, recovery, and application-health behavior.
  • Experience with platform automation, observability, upgrades, capacity planning, and failure-mode testing.
  • Experience supporting large-scale, business-critical cache, query, or compute platforms.

Benefits & conditions

Medallia is committed to equal pay and transparency. The annual base salary range for this position is $138,000 - $205,000. Please note that the salary range information provided is a general guideline and combines all of the distinct labor markets within the US. It is uncommon for an individual to be hired at or near the top of the range for their role and compensation decisions are dependent on a variety of factors. Medallia considers factors such as (but not limited to) scope and responsibilities of the position, candidate’s work experience, candidate’s work location, education/training, key skills, internal peer equity, external market data, as well as, market and business considerations when making compensation decisions.

Medallia also offers competitive health and wellness benefits, including but not limited to medical, dental, vision, 401(k), short-term and long-term disability, life and AD&D insurance, statutory leaves, paid parental leave, and paid holidays. Benefits and eligibility may vary by location and role.

About the company

Medallia is the pioneer and market leader in Experience Management. Our award-winning SaaS platform, Medallia Experience Cloud, leads the market in the management of experiences, insights, and actions for candidates, customers, employees, patients, and residents alike.

We believe that every experience is a memory that can last a lifetime. Experiences shape the way people feel about a company. And they greatly influence how likely people are to advocate, contribute, and stay. At Medallia, we are committed to creating a world where organizations are loved by their customers and their employees.

We empower exceptional people to create extraordinary experiences together.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on jobs.localjobnetwork.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:55 min

Demonstrating semantic routing thresholds with the Redis vector library

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

4:36 min

Hiring passionate software engineers to tackle unprecedented scaling challenges

Dana Lawson Dana Lawson +1 · World Congress 2026 Europe

3:42 min

Comparing in-memory and Redis storage for cache scalability

Simone Sanfratello · World Congress 2022

2:14 min

Solving complex platform architecture challenges at an enterprise scale

Maria Apazoglou · Coffee With Developers

Videos

See all

Related articles

See all