Big Data DevOps Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+36 more
Job description
Experienced Big Data Administrator responsible for managing, supporting, automating, and optimizing enterprise-scale Hadoop, Kafka, AWS cloud, and DevOps platforms. The role focuses on ensuring high availability, performance, security, scalability, and operational excellence across big data ecosystems while collaborating with development, architecture, and business teams., Kafka Administration
-
Deploy, configure, and manage Apache Kafka clusters and AWS MSK environments.
-
Monitor broker health, partitions, replication factors, and consumer lag.
-
Perform capacity planning and cluster scaling activities.
-
Manage Kafka security using SSL, SASL, ACLs, and encryption standards.
-
Troubleshoot producer, consumer, and broker performance issues.
-
Support Kafka Connect, Schema Registry, Cruise Control, and MirrorMaker implementations.
AWS Cloud Administration
-
Manage cloud infrastructure services including EC2, S3, IAM, VPC, EBS, CloudWatch, CloudTrail and AWS Glue.
-
Support AWS Managed Streaming for Kafka (MSK), EMR, Lambda, and Airflow environments.
-
Implement cloud security best practices and governance controls.
-
Perform infrastructure provisioning and automation using Infrastructure as Code (IaC).
-
Monitor cloud resource utilization and optimize operational costs.
DevOps & Automation
-
Design and maintain CI/CD pipelines using Jenkins, GitHub Actions, GitLab CI, or similar tools.
-
Automate infrastructure deployment using Terraform, CloudFormation, and Ansible.
-
Manage source control repositories and release processes.
-
Implement monitoring and alerting solutions using Prometheus, Grafana, Splunk, ELK, or CloudWatch.
-
Support containerization technologies such as Docker and Kubernetes.
-
Develop automation scripts using Python, Shell, or Bash.
Operations & Support
-
Provide Level 2 and Level 3 production support.
-
Participate in on-call support rotations and incident management activities.
-
Perform root cause analysis (RCA) and implement preventive measures.
-
Create and maintain operational documentation and standard operating procedures.
-
Ensure compliance with security, audit, and regulatory requirements., * Experience supporting large-scale production environments handling petabyte-scale data workloads.
Key Achievements Expected
-
Maintain platform availability above 99.9%.
-
Adopt AI-assisted engineering practices to improve operational efficiency, reduce manual effort, and accelerate troubleshooting and documentation.
-
Automate repetitive operational tasks.
-
Improve cluster performance and resource utilization.
-
Ensure secure, scalable, and reliable data platform operations.
Requirements
-
Apache Kafka Administration
-
AWS Cloud Services
-
Linux (RHEL/Rocky Linux)
-
Shell Scripting and Python
-
Jenkins, Git, Ansible, Terraform
-
Docker and Kubernetes
-
Monitoring Tools (Grafana, Prometheus, Splunk)
-
Networking, Security, and High Availability Concepts
-
Performance Tuning and Capacity Planning
Preferred Qualifications
-
Bachelor’s degree in Computer Science, Information Technology, or related field.
-
Experience with Cloudera CDP, AWS MSK, Airflow, and Spark.
-
AWS, Google Cloud Platform, Kafka, or Kubernetes certifications.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.dice.comGood distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Highest Paying Tech Companies for Developers
Making Data Warehouses Fast: A Developer’s Story
DevOps Engineer Salary [2023]
Data Engineer Salary UK