> Markdown version of [/jobs/ext/1266181-expert-cloud-engineer](https://www.wearedevelopers.com/jobs/ext/1266181-expert-cloud-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Expert Cloud Engineer - **Company:** Technatomy Corporation - **Location:** United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Amazon Web Services, Audit Trail, Automation of Tests, Cloud Computing Security, Cloud Engineering, Information Systems, Continuous Integration, DevOps, Disaster Recovery, Distributed Systems, Elasticsearch, Identity and Access Management, Interoperability, Key Management, Performance Tuning, Role-Based Access Control, Site Reliability Engineering Practices, Prometheus, Zero Trust Network Access, Software Deployment, Software Engineering, Vault (Revision Control System), Data Streaming, Systems Integration, Software Technical Review, WebSocket, Policy as Code, Planning Software, Fast Healthcare Interoperability Resources, System Availability, Grafana, Software Security, Appian, Event Driven Architecture, Kubernetes, Infrastructure Automation Frameworks, Information Technology, Health Level Seven International, Low-code, Apache Kafka, Bitbucket, Cloud Optimization, Kibana, Terraform, Devsecops, Serverless Computing, Docker, Jenkins, Microservices - **Published:** July 14, 2026 - **Apply:** https://www.clearancejobs.com/jobs/9026985/expert-cloud-engineer ## About the Role · 10-12 years of progressive experience in cloud engineering, platform engineering, software engineering, DevSecOps, or distributed systems roles within enterprise or mission-critical environments. · Expert-level AWS cloud engineering experience, including the design and operation of secure, scalable, highly available, and observable production platforms. · Deep hands-on expertise with Kafka or AWS MSK, including producer/consumer design, topic architecture, schema evolution, resiliency patterns, and production operations. · Demonstrated experience architecting event-driven systems, distributed microservices, asynchronous workflows, and high-throughput integration patterns. · Advanced experience with Kubernetes, Docker, Terraform, CI/CD pipelines, infrastructure automation, automated testing, and production deployment strategies. · Strong background designing real-time streaming, WebSocket-based, or event-notification capabilities for responsive user experiences. · Expert understanding of DevSecOps, secure software delivery, cloud security, secrets management, IAM, encryption, auditability, and Zero Trust architecture principles. · Advanced knowledge of high availability, disaster recovery, observability, incident management, performance engineering, capacity planning, and production readiness. · Proven ability to lead complex technical initiatives across multiple engineering teams without requiring direct management authority. · Experience authoring architecture decision records, technical roadmaps, operational playbooks, engineering standards, and production readiness criteria. · Exceptional communication and influencing skills with the ability to align architects, engineers, security teams, delivery leadership, and customer stakeholders. KNOWLEDGE AND SKILLS DESIRED: · Experience serving as a principal engineer, lead platform engineer, cloud architect, or senior technical authority on Federal, VA, healthcare, or other regulated enterprise programs. · Familiarity with Federal compliance expectations, continuous monitoring, audit evidence, security control implementation, FedRAMP-aligned environments, or VA enterprise standards. · Experience integrating Appian, low-code platforms, or workflow engines with cloud-native services, APIs, event buses, and enterprise data sources. · Advanced experience with Site Reliability Engineering practices, service-level objectives, performance tuning, incident management, and cloud cost optimization. · Experience with healthcare interoperability, scheduling platforms, case management systems, HL7/FHIR-adjacent integrations, or mission-critical public-sector workflows. · Relevant certifications such as AWS Certified Solutions Architect - Professional, AWS Certified DevOps Engineer - Professional, AWS Certified Security - Specialty, or Certified Kubernetes Administrator are preferred. EDUCATION: · Bachelor's degree in Computer Science, Information Technology, Information Systems, Engineering, or a related discipline, or equivalent practical experience; an advanced technical degree is valued. CLEARANCE: · Must be able to obtain and maintain a Public Trust clearance. ## Description We are seeking an Expert Cloud Engineer to serve as a hands-on technical authority and platform engineering leader for a mission-critical, cloud-native, event-driven platform supporting the Department of Veterans Affairs. This role partners with the Solution Architect, security stakeholders, product leadership, and engineering teams to define and execute the technical direction for a Kafka/MSK-based event bus, RPC-driven core services, resilient microservices, and real-time WebSocket experiences. The Expert Cloud Engineer owns complex design decisions, establishes engineering standards, guides implementation across teams, and ensures enterprise-grade scalability, reliability, security, compliance, observability, and operational excellence., · Serve as the principal technical authority for AWS cloud engineering, event-driven architecture, distributed systems, and platform reliability. · Define the target cloud architecture, engineering standards, reference patterns, and implementation roadmap for mission-critical platform capabilities. · Architect and guide delivery of cloud-native microservices supporting RPC-based appointment, scheduling, and workflow functions. · Lead implementation of Kafka/MSK producers, consumers, topic strategies, schema governance, replay patterns, and event lifecycle controls. · Establish enterprise event-processing patterns for idempotency, ordering, backpressure, retry, dead-letter handling, fault isolation, and data consistency. · Design and guide secure WebSocket-based streaming capabilities that provide responsive, real-time user experiences. · Set standards for high availability, disaster recovery, automated rollback, resiliency testing, incident response, and production readiness. · Lead technical design reviews and architecture decision records for integrations with external healthcare, Federal, and VA enterprise systems. · Own and mature CI/CD, DevSecOps, and infrastructure automation practices using Terraform, Jenkins, Bitbucket, Packer, automated testing, policy-as-code, and secure deployment patterns. · Lead deployment and operation of Kubernetes, Docker, and AWS-native workloads with a focus on scalability, security, maintainability, and cost efficiency. · Define observability strategy, dashboards, service-level indicators, and service-level objectives using Prometheus, Grafana, Elasticsearch, Kibana, and AWS-native monitoring services. · Establish secure-by-default platform patterns using Vault, IAM, encryption, secrets management, least privilege, Zero Trust principles, and audit-ready controls. · Partner with cybersecurity, compliance, and operations teams to align implementation with Federal security expectations, enterprise guardrails, and continuous monitoring needs. · Mentor senior engineers, raise the technical bar through design and code reviews, and coach teams on cloud-native engineering, distributed systems, and production operations. · Identify systemic risks, performance bottlenecks, reliability gaps, and cost drivers and influence platform priorities and remediation plans across stakeholders. ## Related Videos - [Debug a Kubernetes Operator](https://www.wearedevelopers.com/videos/487-debug-a-kubernetes-operator) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Remote Driving on Plant Grounds with State-of-the-Art Cloud Technologies](https://www.wearedevelopers.com/videos/251-remote-driving-on-plant-grounds-with-state-of-the-art-cloud-technologies) - [My journey into DevOps world - How it all started!](https://www.wearedevelopers.com/videos/545-my-journey-into-devops-world-how-it-all-started) - [Add Location-based Searching to Site with ElasticSearch](https://www.wearedevelopers.com/videos/77-add-location-based-searching-to-site-with-elasticsearch) ## Related Articles - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Best Paying Jobs in Technology](https://www.wearedevelopers.com/magazine/256-best-paying-jobs-in-technology) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs)