> Markdown version of [/jobs/ext/2625195-cloudera-public-cloud-platform-engineer-cdp](https://www.wearedevelopers.com/jobs/ext/2625195-cloudera-public-cloud-platform-engineer-cdp). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Cloudera Public Cloud Platform Engineer (CDP) - **Company:** Ark Infotech Spectrum - **Location:** United States (Remote available) - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Amazon S3, Data Analysis, Microsoft Azure, Bash Shell, Command-Line Interface, Cloud Computing, Cloudera Impala, Computer Networks, System Configuration, Data as a Services, Data Validation, Information Engineering, Data Infrastructure, Data Integration, Data Security, Data Visualization, Disaster Recovery, Apache Hive, Identity and Access Management, Subnetting, Python (Programming Language), Kerberos (Protocol), Log Analysis, Metadata, Routing, Performance Tuning, Reliability Engineering, Azure Active Directory, Cloud Services, Prometheus, Software Defined Everything, Cloudera, Azure Machine Learning, Runbook, Service Pack, Data Streaming, Transport Layer Security, Google Cloud, Data Ingestion, Autoscaling, Cloudera Manager, System Availability, Grafana, Apache Spark, Amazon Virtual Private Cloud (VPC), Containerization, Data Lakes, AI Platforms, Kubernetes, Data Lineage, Apache Kafka, Apache Nifi, Data Management, Machine Learning Operations, Terraform, Software Version Control, Data Pipelines, Serverless Computing, Legacy Systems - **Published:** August 31, 2026 - **Apply:** https://www.dice.com/job-detail/2c0929ef-9368-4754-9add-494289dec64d ## About the Role * 4-6 years of experience in Big Data Platform / Cloud Operations / Infrastructure Support roles * 3+ years of experience with Cloudera ecosystem (CDH/CDP) * Hands-on experience with: o CDP services (CDE, CDW, CDF, CAI) o Monitoring, alerting, and operational support * Strong understanding of: o Cloud platforms (AWS/Azure/Google Cloud Platform basics - IAM, storage, networking) o Kubernetes concepts (pod-level troubleshooting, logs, resource usage) * Experience in: o Production support and incident handling (P2/P3) o Platform monitoring, issue triage, and escalation, * Strong experience with Cloudera CDP Public Cloud * Expertise in: o Cloud platforms (AWS/Azure/Google Cloud Platform) o Kubernetes concepts (troubleshooting-focused) * Hands-on with: o CDE, CDW, CDF (NiFi), CAI * Knowledge of: o IAM, networking, observability tools * Platforms operating at multi-terabyte to petabyte scale with high concurrency workloads * Hands-on experience with: o Kafka (or similar streaming platforms) including monitoring, troubleshooting, and performance tuning * Experience with Cloudera CDP CLI (Command Line Interface) for: o Platform operations and administration o Job execution and service management (CDE/CDW/CDL) o Automation of routine operational tasks, o Modernization of legacy data platforms/applications to Cloudera CDP Public Cloud o Migration and onboarding of workloads to CDE, CDW, and CAI environments o Supporting hybrid or multi-environment transitions (on-prem * cloud) * Familiarity with: o Cloud platforms (AWS, Azure, Google Cloud Platform) including storage, IAM, and networking concepts o Kubernetes-based runtime environments (troubleshooting-focused) * Strong scripting and automation skills (Python, Shell, Terraform) for platform operations ## Description * Responsible for: o Executing day-to-day platform operations o Supporting deployments, monitoring, and troubleshooting o Following runbooks and operational procedures, o Cloud IAM (AWS IAM / Azure AD) including roles, policies, and cross-service access o User and group mapping across CDP, cloud IAM, and Ranger policies o Troubleshooting access issues across storage (S3/ADLS), CDP services, and data access layers, * Enterprise-scale Cloudera CDP platform supporting data engineering, analytics, and AI workloads across multiple applications * Modernization of legacy platforms and applications into cloud-native CDP services * Operational support and scaling of: o Data services (CDE, CDW, CDF, CDL) o AI/ML platforms (CAI, inference, workbenches) * Platform performance optimization, observability, and reliability engineering for mission-critical workloads Why This Role Matters * Ensures availability, stability, and performance of the CDP platform supporting all data and AI workloads * Enables successful modernization of legacy applications into scalable, cloud-native services * Maintains high availability, observability, and operational excellence across enterprise platforms * Acts as the backbone for data engineering, analytics, and AI initiatives * This role focuses on platform reliability and infrastructure operations and does not include data-layer ownership (e.g., Iceberg table management or data validation)., We are seeking a highly skilled Cloudera Public Cloud Platform Engineer to operate and manage the end-to-end CDP platform ecosystem, including data services, NiFI, Kafka, AI/ML platforms, and enterprise observability. This role is responsible for ensuring availability, scalability, security, and performance of all platform services supporting data, analytics, and AI workloads across environments. The ideal candidate brings strong expertise in CDP on-prem, public cloud services, cloud infrastructure, Kubernetes-based runtime environments, and platform observability, supporting high-concurrency, mission-critical workloads at multi-terabyte to petabyte scale This role is critical to ensuring uninterrupted operation of data, analytics, and AI platforms-any degradation directly impacts downstream business reporting, data pipelines, and model execution., CDP Platform & Multi-Service Operations * Own end-to-end operational responsibility for Cloudera Public Cloud services across Dev / Stage / UAT / Prod: o CDE, CDW, COD, CDL, CDF (NiFi), CDV, CAI, Kafka * Ensure multi-cluster stability, workload isolation, and SLA adherence * Support onboarding and operations of multiple applications across environments * Manage and support multi-environment, multi-cluster deployments with strict isolation, governance, and release coordination across Dev/UAT/Prod AI/ML Platform Operations * Operate and support Cloudera AI (CAI) environments: o AI Workbenches, AI Studios o Model training and development environments o AI inference endpoints and model serving * Troubleshoot: o Resource contention (CPU/GPU) o Model deployment/runtime failures CDP Runtime & Kubernetes-Aware Operations * Operate CDP services running on Cloudera-managed Kubernetes infrastructure * Apply strong understanding of containerized workloads and Kubernetes concepts for troubleshooting * Diagnose and resolve: o Pod failures, restarts, and resource contention o Spark job failures in containerized environments (CDE) o Service-to-service communication issues * Analyze logs and metrics to identify runtime failures and performance issues * Collaborate with Cloudera support for managed service-level issues Data Integration & Platform Services * Operate and support: o CDF (NiFi) for ingestion pipelines o CDV (Data Visualization) for reporting workloads o Octopai for data lineage and catalog integration * Ensure reliability and performance of data pipelines and integrations * Monitor and troubleshoot Kafka environments: o Topic configurations, partitions, and replication o Consumer lag and throughput issues o Broker connectivity and performance bottlenecks Security, Governance & SDX Administration * Implement and manage: o Kerberos, TLS/SSL, Ranger policies * Administer SDX for: o Centralized security o Metadata and policy enforcement * Support Atlas and Octopai integration * Manage and troubleshoot user access and identity mapping across layers, including: o Cloud IAM roles and permissions o CDP users/groups and identity providers o Ranger policies for fine-grained data access * Resolve access-related issues impacting: o Data access (S3/ADLS) o Query execution (CDW/CDE) o Application and service-level permissions Cloud Infrastructure & Networking * Troubleshoot: o S3 / ADLS storage issues o IAM roles and permissions o VPC, subnets, routing, security groups o Bastion host access and connectivity * Ensure secure and reliable connectivity across services * Understand and troubleshoot S3-based data lake patterns, including: o Bucket structure, prefix design, and access patterns o Performance issues related to small files, request rates, and throughput limits o Encryption (SSE-S3, SSE-KMS) and access policies * Manage and troubleshoot cross-account IAM roles and access patterns for CDP environments * Ensure secure access between: o CDP environments and cloud resources o Multiple AWS accounts (dev/prod separation) Disaster Recovery & Resiliency * Support and validate disaster recovery and failover strategies across CDP environments * Ensure backup, recovery, and environment resiliency for critical workloads * Participate in DR drills and recovery validation Observability, Monitoring & Alerting (Critical) * Implement and manage end-to-end observability: o Metrics, logs, and alerting * Use: o Cloudera observability, Cloudera Manager, Prometheus, Grafana * Monitor: o Cluster health o Workload performance o AI inference endpoints * Enable proactive issue detection and prevention * Define and implement SLIs/SLOs and alerting thresholds to ensure platform reliability and performance * Support high-severity (P1/P2) incident response, triage, and resolution within defined SLAs Operational Support & On-Call * Participate in on-call rotation to support 24/7 platform operations * Respond to production incidents, alerts, and service disruptions within defined SLAs * Handle P1/P2 incidents, including triage, troubleshooting, and resolution * Perform root cause analysis (RCA) and implement preventive measures Upgrades, Patching & Platform Lifecycle * Execute: o CDP upgrades and version management o Security patches and hotfixes * Perform: o Rolling upgrades o Validation and rollback strategies Performance Optimization & Cost Efficiency * Optimize: o Platform-level performance (Spark, Hive, Impala workloads) o Cluster utilization and workload distribution * Drive: o Autoscaling strategies o Cost optimization (FinOps practices) Automation & Operational Excellence * Utilize and support existing automation frameworks for: o Platform provisioning o Monitoring and alerting o Routine operational tasks * Work with infrastructure teams that manage Infrastructure-as-Code (Terraform) for environment setup and changes * Leverage scripting (Python / Shell) for: * Operational support * Task automation * Troubleshooting and diagnostics * Maintain and follow runbooks, SOPs, and operational procedures to ensure consistent platform operations ## Related Videos - [Technical Documentation - How Can I Write Them Better and Why Should I Care?](https://www.wearedevelopers.com/videos/681-technical-documentation-how-can-i-write-them-better-and-why-should-i-care) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Creating a routing app with Google Maps API from scratch](https://www.wearedevelopers.com/videos/831-creating-a-routing-app-with-google-maps-api-from-scratch) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Bridging AI and Nomad: a Go-based MCP Server for Cluster Control](https://www.wearedevelopers.com/videos/2063-bridging-ai-and-nomad-a-go-based-mcp-server-for-cluster-control) - [A Technical Introduction to Bitcoin's 2nd Layer- The Lightning Network](https://www.wearedevelopers.com/videos/15-a-technical-introduction-to-bitcoin-s-2nd-layer-the-lightning-network) ## Related Articles - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story)