Kafka Platform Engineer

Bright Vision Technologies
United States
about 1 month ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$130,000.0 - $180,000.0
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Amazon Web Services Authentication Protocols Bash Shell Code Review Data Governance DevOps Disaster Recovery Distributed Systems Failover Monitoring of Systems Python (Programming Language)
+27 more
Meta-Data Management Performance Tuning Role-Based Access Control Ansible Prometheus Data Streaming Data Logging Scripting Transport Layer Security Google Cloud Cloud Platform System System Availability Grafana Apache Spark Event Driven Architecture Kubernetes Infrastructure Automation Frameworks Information Technology Data Lineage Apache Flink Real Time Data Apache Kafka Terraform Stream Processing Dynatrace Apache Beam Confluent

Job description

Bright Vision Technologies is seeking a highly experienced Kafka Platform Engineer with 10+ years of experience in distributed systems, event streaming, and platform engineering, including extensive expertise with Apache Kafka and the Confluent Platform. The ideal candidate will architect, deploy, and manage enterprise-scale streaming platforms that power mission-critical, real-time data processing across the organization. This role requires deep technical expertise in Kafka internals, platform automation, security, observability, and cloud-native infrastructure, along with the ability to mentor engineering teams and define enterprise event-streaming standards., * Architect, deploy, and manage highly available Apache Kafka and Confluent Platform environments supporting enterprise-scale event-driven applications.

  • Design Kafka cluster topology, broker configuration, partitioning strategies, replication, capacity planning, and performance optimization for high-throughput workloads.
  • Implement enterprise-grade security using SASL, SSL/mTLS, ACLs, RBAC, encryption, and authentication mechanisms.
  • Design and manage Kafka Connect, Schema Registry, Kafka Streams, ksqlDB, and event streaming pipelines for real-time data integration.
  • Develop and maintain Infrastructure as Code using Terraform, Ansible, or similar automation frameworks to provision and manage Kafka infrastructure.
  • Design and implement High Availability (HA), Disaster Recovery (DR), backup, failover, and cross-region replication strategies.
  • Build comprehensive monitoring, alerting, logging, and observability solutions using Prometheus, Grafana, OpenTelemetry, ELK, and Confluent monitoring tools.
  • Optimize cluster performance, troubleshoot production issues, conduct root cause analysis, and implement long-term platform improvements.
  • Support Kafka deployments on Kubernetes using Strimzi, Confluent Operator, or managed cloud services including AWS MSK, Azure Event Hubs (Kafka API), and Confluent Cloud.
  • Collaborate with application developers, data engineers, DevOps, SRE, and enterprise architects to establish event-driven architecture standards and best practices.
  • Conduct architecture reviews, establish platform governance, perform code reviews, and mentor engineers on Kafka development and operations.

Requirements

Sponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position., * Bachelor’s or Master’s degree in Computer Science, Engineering, Information Technology, or a related technical discipline.

  • 10+ years of professional experience in distributed systems, platform engineering, or infrastructure engineering, including 5+ years of hands-on Apache Kafka or Confluent Platform experience.
  • Expert-level knowledge of Kafka internals, including partitions, replication, brokers, ISR, consumer groups, offset management, and performance tuning.
  • Extensive experience implementing Kafka security using SASL, SSL/mTLS, ACLs, RBAC, and enterprise authentication mechanisms.
  • Strong hands-on experience with Kafka Connect, Schema Registry, Kafka Streams, ksqlDB, and real-time event streaming architectures.
  • Experience implementing High Availability (HA), Disaster Recovery (DR), multi-cluster replication, and cross-region failover strategies.
  • Advanced scripting skills using Python, Bash, or Go, along with Infrastructure as Code experience using Terraform, Ansible, or similar tools.
  • Strong experience with observability platforms including Prometheus, Grafana, OpenTelemetry, ELK, and distributed tracing solutions.
  • Excellent troubleshooting, communication, documentation, stakeholder management, and technical leadership skills., * Confluent Certified Administrator or Confluent Certified Developer certification.
  • Experience operating Kafka on Kubernetes using Strimzi, Confluent Operator, or similar operators.
  • Hands-on experience with managed Kafka services including AWS MSK, Confluent Cloud, Azure Event Hubs (Kafka API), or Google Cloud Managed Kafka.
  • Experience with stream processing technologies such as Apache Flink, Apache Spark Structured Streaming, or Apache Beam.
  • Knowledge of data governance, schema evolution, metadata management, and data lineage for streaming platforms.

Benefits & conditions

4.24.2 out of 5 stars Remote $130,000 - $180,000 a year - Full-time

About the company

Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States. This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:55 min

Infrastructure challenges when combining Kafka with Apache Flink

Bobur Umurzokov · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

1:42 min

Automating Skupper deployments using Ansible

Alex Soto Alex Soto · World Congress 2024

1:47 min

Advantages of adding Kafka to streaming architecture

Developersteve · LIVE

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

Videos

See all

Related articles

See all