Expert Cloud Engineer

Technatomy Corporation
United States
28 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
10 years minimum
Working hours
Regular working hours

Tech stack

Application Programming Interfaces (APIs) Amazon Web Services Audit Trail Automation of Tests Cloud Computing Security Cloud Engineering Information Systems Continuous Integration DevOps Disaster Recovery Distributed Systems Elasticsearch
+38 more
Identity and Access Management Interoperability Key Management Performance Tuning Role-Based Access Control Site Reliability Engineering Practices Prometheus Zero Trust Network Access Software Deployment Software Engineering Vault (Revision Control System) Data Streaming Systems Integration Software Technical Review WebSocket Policy as Code Planning Software Fast Healthcare Interoperability Resources System Availability Grafana Software Security Appian Event Driven Architecture Kubernetes Infrastructure Automation Frameworks Information Technology Health Level Seven International Low-code Apache Kafka Bitbucket Cloud Optimization Kibana Terraform Devsecops Serverless Computing Docker Jenkins Microservices

Job description

We are seeking an Expert Cloud Engineer to serve as a hands-on technical authority and platform engineering leader for a mission-critical, cloud-native, event-driven platform supporting the Department of Veterans Affairs. This role partners with the Solution Architect, security stakeholders, product leadership, and engineering teams to define and execute the technical direction for a Kafka/MSK-based event bus, RPC-driven core services, resilient microservices, and real-time WebSocket experiences. The Expert Cloud Engineer owns complex design decisions, establishes engineering standards, guides implementation across teams, and ensures enterprise-grade scalability, reliability, security, compliance, observability, and operational excellence., · Serve as the principal technical authority for AWS cloud engineering, event-driven architecture, distributed systems, and platform reliability.

· Define the target cloud architecture, engineering standards, reference patterns, and implementation roadmap for mission-critical platform capabilities.

· Architect and guide delivery of cloud-native microservices supporting RPC-based appointment, scheduling, and workflow functions.

· Lead implementation of Kafka/MSK producers, consumers, topic strategies, schema governance, replay patterns, and event lifecycle controls.

· Establish enterprise event-processing patterns for idempotency, ordering, backpressure, retry, dead-letter handling, fault isolation, and data consistency.

· Design and guide secure WebSocket-based streaming capabilities that provide responsive, real-time user experiences.

· Set standards for high availability, disaster recovery, automated rollback, resiliency testing, incident response, and production readiness.

· Lead technical design reviews and architecture decision records for integrations with external healthcare, Federal, and VA enterprise systems.

· Own and mature CI/CD, DevSecOps, and infrastructure automation practices using Terraform, Jenkins, Bitbucket, Packer, automated testing, policy-as-code, and secure deployment patterns.

· Lead deployment and operation of Kubernetes, Docker, and AWS-native workloads with a focus on scalability, security, maintainability, and cost efficiency.

· Define observability strategy, dashboards, service-level indicators, and service-level objectives using Prometheus, Grafana, Elasticsearch, Kibana, and AWS-native monitoring services.

· Establish secure-by-default platform patterns using Vault, IAM, encryption, secrets management, least privilege, Zero Trust principles, and audit-ready controls.

· Partner with cybersecurity, compliance, and operations teams to align implementation with Federal security expectations, enterprise guardrails, and continuous monitoring needs.

· Mentor senior engineers, raise the technical bar through design and code reviews, and coach teams on cloud-native engineering, distributed systems, and production operations.

· Identify systemic risks, performance bottlenecks, reliability gaps, and cost drivers and influence platform priorities and remediation plans across stakeholders.

Requirements

· 10-12 years of progressive experience in cloud engineering, platform engineering, software engineering, DevSecOps, or distributed systems roles within enterprise or mission-critical environments.

· Expert-level AWS cloud engineering experience, including the design and operation of secure, scalable, highly available, and observable production platforms.

· Deep hands-on expertise with Kafka or AWS MSK, including producer/consumer design, topic architecture, schema evolution, resiliency patterns, and production operations.

· Demonstrated experience architecting event-driven systems, distributed microservices, asynchronous workflows, and high-throughput integration patterns.

· Advanced experience with Kubernetes, Docker, Terraform, CI/CD pipelines, infrastructure automation, automated testing, and production deployment strategies.

· Strong background designing real-time streaming, WebSocket-based, or event-notification capabilities for responsive user experiences.

· Expert understanding of DevSecOps, secure software delivery, cloud security, secrets management, IAM, encryption, auditability, and Zero Trust architecture principles.

· Advanced knowledge of high availability, disaster recovery, observability, incident management, performance engineering, capacity planning, and production readiness.

· Proven ability to lead complex technical initiatives across multiple engineering teams without requiring direct management authority.

· Experience authoring architecture decision records, technical roadmaps, operational playbooks, engineering standards, and production readiness criteria.

· Exceptional communication and influencing skills with the ability to align architects, engineers, security teams, delivery leadership, and customer stakeholders.

KNOWLEDGE AND SKILLS DESIRED:

· Experience serving as a principal engineer, lead platform engineer, cloud architect, or senior technical authority on Federal, VA, healthcare, or other regulated enterprise programs.

· Familiarity with Federal compliance expectations, continuous monitoring, audit evidence, security control implementation, FedRAMP-aligned environments, or VA enterprise standards.

· Experience integrating Appian, low-code platforms, or workflow engines with cloud-native services, APIs, event buses, and enterprise data sources.

· Advanced experience with Site Reliability Engineering practices, service-level objectives, performance tuning, incident management, and cloud cost optimization.

· Experience with healthcare interoperability, scheduling platforms, case management systems, HL7/FHIR-adjacent integrations, or mission-critical public-sector workflows.

· Relevant certifications such as AWS Certified Solutions Architect - Professional, AWS Certified DevOps Engineer - Professional, AWS Certified Security - Specialty, or Certified Kubernetes Administrator are preferred.

EDUCATION:

· Bachelor’s degree in Computer Science, Information Technology, Information Systems, Engineering, or a related discipline, or equivalent practical experience; an advanced technical degree is valued.

CLEARANCE:

· Must be able to obtain and maintain a Public Trust clearance.

About the company

At Technatomy, we deliver innovative solutions through the efforts of our diverse and talented people who are dedicated to our customer’s success. We provide solutions to agencies and entities including the Department of Veterans Affairs, Department of Defense, Defense Logistics Agency, National Institute of Health, and more. Everything we do is built on a commitment to do the right thing for our customers, our people, and our community. Our Mission, Vision, and Values guide the way we do business.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.clearancejobs.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

3:21 min

Deploying a primary Elasticsearch and Kibana cluster configuration

Philipp Krenn · WWC 2022

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

4:51 min

Executing simple full-text search queries using the Kibana interface

Derek Binkley · LIVE

Videos

See all

Related articles

See all