Senior Technical Lead - DevOps, Python, Kubernetes

HCL America Inc.
Santa Clara, CA, United States
about 2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
7 years minimum
Compensation
$134,000.0
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Bash Shell Configuration Management Data as a Services DevOps Disaster Recovery Distributed Data Store Distributed Systems Python (Programming Language) Lightweight Directory Access Protocols (LDAP) PostgreSQL Linux System Administration
+14 more
Performance Tuning Prometheus Service Discovery Software Vulnerability Management Apache Zookeeper Scripting Cloud Monitoring Grafana Apigee Data Layers Kubernetes Infrastructure Automation Frameworks Cassandra Terraform

Job description

We are seeking an experienced Data Services Lead Engineer to own the technical direction, architecture, and operational excellence of our data platform. This role requires deep expertise in Cassandra, ZooKeeper, and Consul operations, strong leadership skills, and a passion for building robust, scalable distributed data systems. You will guide the team on best practices, lead complex technical projects, and act as the primary escalation point for data-platform-related issues. The team is also responsible for ZooKeeper, Consul, LDAP, PostgreSQL, and Qpid., Lead the design, architecture, and implementation of highly available, scalable, and performant distributed data stores (including Cassandra and PostgreSQL) across cloud and OnPrem environments. Define and drive the technical roadmap and strategy for the persistence services layer within Apigee Edge Data Services. Lead incident response and management with clear communication. Lead comprehensive post-mortem analyses for production incidents to identify root causes, document findings, and drive the implementation of preventative measures across the data platform. Lead vulnerability management initiatives, including the execution of regular version and security upgrades for all supported data services. Establish and enforce best practices for distributed systems data modeling, capacity planning, performance tuning, security, and disaster recovery. Develop and improve automation for cluster provisioning, configuration management, and upgrades. Serve as the primary technical escalation point for complex production issues, including root cause analysis. Mentor and provide technical guidance to other engineers across the organization. Collaborate with Engineering, SRE, and Support teams to align the data layer with platform requirements. Drive continuous improvement initiatives to enhance reliability and maintainability. Participate in the team’s on-call rotation for production support.

Requirements

7+ years of experience managing large-scale, mission-critical distributed data systems (e.g., Cassandra, ZooKeeper) in a production environment. Understanding of Consul for service discovery and configuration management. Deep understanding of distributed system architectures, data modeling, internals, and performance tuning. Proficiency in Linux environments and scripting languages (e.g., Python, Bash). Experience with infrastructure-as-code tools (e.g., Terraform). Experience with monitoring and alerting systems (e.g., Prometheus, Grafana, Cloud Monitoring). Experience working in cloud environments (GCP, AWS, etc.).

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

4:36 min

Hiring passionate software engineers to tackle unprecedented scaling challenges

Dana Lawson Dana Lawson +1 · WWC Europe 2026

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

1:44 min

Career transition into cloud native and data management

Michael Cade · LIVE

1:04 min

Visualizing Keycloak performance via standard Grafana troubleshooting dashboards

Alexander Schwartz Alexander Schwartz · WWC 2025

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

Videos

See all

Related articles

See all