Data Platform SRE / Kafka Platform Engineer

VDart, Inc.
Frisco, TX, United States
2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Working hours
Regular working hours
Job source

Tech stack

Java (Programming Language) Application Programming Interfaces (APIs) Amazon Web Services Microsoft Azure Computer Programming Continuous Integration Data Infrastructure Distributed Systems Monitoring of Systems Identity and Access Management Issue Tracking Systems JSON
+26 more
Python (Programming Language) Metadata Cisco Nexus Switches Queue Management Systems Reliability Engineering Prometheus Cloudera Software Vulnerability Management Parquet File Transfer Protocol (FTP) Data Ingestion System Availability Grafana Reliability of Systems Git Gitlab-ci Kubernetes Apache Kafka Data Management Splunk Software Version Control Data Pipelines Docker Jenkins Servicenow Microservices

Job description

  • Support, monitor, and maintain microservices-based data platforms in production environments.
  • Proactively monitor, troubleshoot, and resolve incidents across distributed systems and streaming platforms.
  • Ensure high availability, reliability, and performance of Kafka-based ingestion and processing pipelines.
  • Own operational readiness across non-production and production environments, including: Start/stop procedures, Dependency validation, Release support
  • Drive automation initiatives across: Deployments, Vulnerability remediation, Monitoring and alerting, Operational recovery workflows

Requirements

  • Strong programming expertise in Java and/or Python.
  • Hands-on experience with Apache Kafka, including: Topics, partitions, brokers, Consumer groups, Kafka Connect, Lag monitoring and alert handling
  • Proven experience in microservices architecture (build, support, and troubleshooting).
  • Solid understanding of SRE principles, including: SLAs, SLOs, Incident response, Threshold-based alerting, Observability and system resilience
  • Experience with CI/CD pipelines and tools such as: Jenkins, GitLab CI/CD, Nexus, Source control platforms (Git-based)
  • Strong troubleshooting capability across logs, metrics, traces, infrastructure dependencies, and application failures.
  • Kafka, Platform Operations & Environment Readiness
  • Experience with data ingestion mechanisms including Kafka, SFTP, and APIs.
  • Knowledge of data formats such as JSON and Parquet.
  • Ability to manage and operate: Kafka topic inventory, Metadata/entity mapping, Topic classifications (K0/K1/K2), Consumer group operations
  • Experience with: Kafka backlog monitoring Alert triage and incident response
  • Exposure to Cloudera-managed Kafka environments (preferred).
  • Ability to validate end-to-end operational readiness of non-production environments for releases.
  • Platform, Infrastructure & Observability
  • Experience with cloud platforms (AWS or Azure).
  • Hands-on knowledge of: Docker, Kubernetes (AKS preferred)
  • Familiarity with observability and monitoring tools: Grafana, Prometheus, Pushgateway, Splunk
  • Ability to validate: Dashboard ownership, Log routing, Reconciliation metrics, Monitoring coverage
  • Understanding of platform dependencies across: Jenkins, Nexus, GOGS, Supporting infrastructure services
  • Experience handling production incidents (P1/P2) and escalations.
  • Strong exposure to: Root Cause Analysis (RCA) preparation, Support runbook execution
  • Familiarity with ServiceNow (or similar ticketing systems) and queue management workflows.

Good understanding of:

  • Application runbooks (start/stop order, escalation paths, thresholds)
  • Access governance (AD/IDM, privileged access management)
  • Change management processes and deployment ownership
  • Exit procedures and access revocation controls

Key Skills: Kafka, SRE, Microservices, Platform Ops, (AWS/Azure), Kubernetes, CI/CD tools

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on careerjet.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

3:47 min

Exploring JSON, CBOR, and JOSE for data serialization

Aaron Russell · LIVE

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

2:03 min

Distinguishing type definition constructs from data validation routines

Clemens Vasters Clemens Vasters · WWC 2025

Videos

See all

Related articles

See all