Principal Software Engineer

TALENT Software Services
United States
9 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Automation of Tests Backup Devices Bash Shell Business Software Disaster Recovery Distributed Data Store Monitoring of Systems Python (Programming Language) Linux System Administration Performance Tuning Redis Ansible
+13 more
Distributed Caching Runbook Software Requirements Analysis Software Systems Test Data Scripting Test Scripts Caching Scalability Testing Containerization Kubernetes Infrastructure Automation Frameworks Docker

Job description

Join a Platform Engineering team responsible for operating and supporting distributed caching services used by critical business applications. This role focuses on the day-to-day administration, monitoring, troubleshooting, and maintenance of Redis-based caching platforms running in enterprise environments. Working alongside senior engineers, you will help ensure platform reliability, performance, and availability while contributing to automation and operational improvements. Experience with container platforms is beneficial but not required., * Support daily operations of Redis and caching platforms in production and non-production environments

  • Monitor platform health, performance, capacity, and availability
  • Investigate and resolve operational incidents, alerts, and service disruptions
  • Execute maintenance activities including upgrades, patching, backups, and recovery procedures
  • Assist with Redis configuration, deployment, and troubleshooting activities
  • Develop and maintain operational automation using Ansible and scripting tools
  • Create and maintain operational documentation, runbooks, and knowledge articles
  • Collaborate with application teams to support caching-related performance and connectivity issues
  • Participate in on-call and operational support rotations as required
  • Support platform modernization initiatives, including containerized deployments when applicable, * Designs, codes, tests, debugs, and documents software according to systems quality standards, policies, and procedures
  • Analyzes business needs and creates software solutions
  • Responsible for preparing design documentation
  • Prepares test data for unit, string, and parallel testing
  • Evaluates and recommends software and hardware solutions to meet user needs
  • Resolves customer issues with software solutions and responds to suggestions for improvements and enhancements
  • Works with business and development teams to clarify requirements to ensure testability
  • Drafts, revises, and maintains test plans, test cases, and automated test scripts
  • Executes test procedures according to software requirements specifications
  • Logs defects and makes recommendations to address defects
  • Retests software corrections to ensure problems are resolved
  • Documents evolution of testing procedures for future replication
  • May conduct performance and scalability testing, * Plans, conducts, and leads assignments generally involving moderate, high-budget projects or more than one project
  • Manages user expectations regarding appropriate milestones and deadlines
  • Assists in training, work assignment, and checking of less experienced developers
  • Serves as technical consultant to leaders in the IT organization and functional user groups
  • Subject matter expert in one or more technical programming specialties; employs expertise as a generalist or a specialist
  • Performs estimation efforts on complex projects and tracks progress
  • Works on the highest level of problems where analysis of situations or data requires an in-depth evaluation of various factors
  • Documents, evaluates, and researches test results; documents evolution of testing scripts for future replication
  • Identifies, recommends, and implements changes to enhance the effectiveness of quality assurance strategies

Primary responsibility will be for the day-to-day operation, maintenance, and reliability of enterprise caching services, including Redis and related distributed caching technologies, deployed across diverse infrastructure environments. Key duties will include monitoring platform health and performance, executing planned maintenance activities such as patching, upgrades, and capacity adjustments, troubleshooting incidents, and driving root cause analysis.

You will contribute to the development and enhancement of automation frameworks to streamline operational workflows, reduce manual intervention, and improve service consistency. Collaboration with engineering, application, and infrastructure teams will be essential to ensure caching services meet performance, availability, and scalability requirements across the enterprise.

Requirements

  • Experience supporting Linux-based systems in enterprise environments
  • Understanding of system administration fundamentals including networking, processes, storage, and troubleshooting
  • Experience supporting production applications or infrastructure platforms
  • Exposure to Redis, caching technologies, or similar distributed data platforms
  • Experience with Ansible or similar automation and configuration management tools
  • Familiarity with Bash, Python, or similar scripting languages
  • Strong analytical and problem-solving skills
  • Ability to document procedures and work effectively within operational support processes
  • Good communication and teamwork skills, * Experience with Redis administration or support
  • Understanding of caching concepts such as key-value stores, data expiration, replication, and performance tuning
  • Exposure to Kubernetes, containers, Docker, or Helm
  • Experience with monitoring and observability tools
  • Experience supporting high-availability systems
  • Relevant certifications, coursework, or equivalent practical experience

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

47 sec

Establishing robust testing lifecycles to offset specialized talent shortages

Ilie-Daniel Gheorghe-Pop Ilie-Daniel Gheorghe-Pop · World Congress 2026 Europe

2:50 min

Introduction and the value of runbooks

Hila Fish · World Congress 2023

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

3:55 min

Demonstrating semantic routing thresholds with the Redis vector library

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

1:32 min

Structuring automated incident workflows between runbooks and raw models

Aram Hakobyan Aram Hakobyan +1 · World Congress 2026 Europe

Videos

See all

Related articles

See all