Principal DevSecOps Engineer - AI Infrastructure

Comtech Global
United States
16 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Amazon Web Services Amazon Elastic Compute Cloud Amazon S3 Continuous Integration Distributed Systems Failover Fault Tolerance Github Identity and Access Management Python (Programming Language)
+16 more
Key Management Machine Learning Prometheus TypeScript AI Infrastructure Fluentd Grafana Amazon Virtual Private Cloud (VPC) Backend Cloudformation Amazon Relational Database Service Containerization Kubernetes Software Coding Terraform Devsecops

Job description

  • Serve as the founding infrastructure engineer, building the core platform that scales the company and raises the reliability bar.
  • Establish secure, repeatable IaC/GitOps patterns (Terraform/CloudFormation) and automated delivery (GitHub Actions, ArgoCD).
  • Partner with teams pre-GA on design reviews, capacity planning, and readiness.
  • Define and drive SLIs/SLOs/SLAs and an error-budget culture for services and ops.
  • Eliminate toil with end-to-end automation across provisioning, config, testing, and operations.
  • Co-design platforms with ML, backend, and security to safely power AI/ML workloads.
  • Architect multi-region resilience-backup, DR, and failover-balancing availability, consistency, and cost.
  • Advance observability and incident excellence; make smart bets on emerging infra tools.
  • Codify production engineering standards and coach teams toward operational excellence.

Requirements

This position is for a Principal DevSecOps Engineer who will serve as the founding infrastructure engineer, responsible for building the core platform that scales the company and raises the reliability bar. The ideal candidate will have experience in operating high-availability, fault-tolerant distributed systems with IaC and GitOps, as well as strong coding skills in Go/Python/Rust and solid shell skills., * 8+ years operating high-availability, fault-tolerant distributed systems with IaC and GitOps.

  • Strong coding in Go/Python/Rust plus solid shell skills; comfortable extending Kubernetes via CRDs.
  • Deep Kubernetes/EKS expertise; mastery of containerization and service networking.
  • Hands-on with AWS primitives (VPC, EC2, S3, IAM, RDS) and multi-region traffic/failover.
  • Observability pro (Prometheus, Grafana, OpenTelemetry, Fluentd, Jaeger) with strong RCA/incident chops.
  • Security fundamentals: IAM, secrets management, and compliance guardrails (SOC2/HIPAA/GDPR).
  • Experience building secure, self-service platforms (SDKs/APIs/portals, e.g., Backstage/TypeScript).
  • Proven SRE practice-SLIs/SLOs, error budgets-and strong testing, reviews, and CI/CD habits.
  • Clear communicator and mentor who thrives in fast-moving environments and collaborates across ML.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · World Congress 2023

1:52 min

Structuring and scaling the backend engineering team

Stefan Lingler Stefan Lingler +1 · Coffee With Developers

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

2:40 min

Using GitHub primitives for internal documentation and corporate operations

Kyle Daigle · Coffee With Developers

Videos

See all

Related articles

See all