Software Engineer, Infrastructure

NEXT GENERATION LLC
United States
23 days ago
Apply on www.builtincolorado.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Compensation
$145,000.0 - $195,000.0
Working hours
Regular working hours

Tech stack

Artificial Intelligence Amazon Web Services Audit Trail Microsoft Azure Cloud Computing Computer Networks Continuous Integration Software Design Documents Programming Tools Github Python (Programming Language) PostgreSQL
+12 more
RabbitMQ Reliability Engineering Prometheus Datadog SSL Certificate Management Cloud Platform System Grafana Kubernetes Infrastructure Automation Frameworks Terraform Docker Service Stack

Job description

Epiq AI Labs is the innovation and engineering hub behind Epiq’s next-generation AI platform for corporate legal departments and global law firms. Operating with the speed and autonomy of a startup and the resources of a global alternative legal services provider, the team builds intelligent agents, reasoning engines, knowledge systems, and structured workflows for litigation, investigations, compliance, and corporate knowledge work.

The team is highly collaborative, deeply technical, and focused on rapid iteration, thoughtful design, and end-to-end ownership.

The Opportunity

You will build the cloud infrastructure beneath Epiq AI Labs’ AI platform, spanning infrastructure as code, Kubernetes, CI/CD and release engineering, networking, secrets management, observability, security, and compliance infrastructure. Ownership extends from initial design through production operation.

Infrastructure is treated as a product whose reliability, scalability, and security shape the capabilities of the entire platform. You will partner closely with backend engineering, AI engineering, product management, and security to establish deployment, observability, and security patterns that can scale with the organization. This role will be in the office 3~4 days a week., * Design and implement cloud infrastructure using Terraform, including reusable modules, environment topology, and drift detection and remediation.

  • Operate Kubernetes in production, including orchestration, autoscaling, resource governance, network policy, and cluster lifecycle management.
  • Build and maintain CI/CD and release infrastructure with progressive delivery, rollback mechanisms, and efficient paths from merge to production.
  • Define platform service-level objectives and build the metrics, tracing, alerting, and error-budget practices required to support them.
  • Implement security and compliance infrastructure, including hardening, audit logging, data-residency controls, retention, legal hold, and audit evidence collection.
  • Build and maintain network, secrets, key, credential, TLS, and certificate-lifecycle infrastructure.
  • Contribute to incident response, post-incident review, developer tooling, technical design documentation, architectural review, and platform operational readiness., Builds trust and alignment through open communication, shared goals, and strong partnerships to drive collective success.
  • Build trust-based partnerships
  • Nurture long-term relationships
  • Remove collaboration barriers
  • Celebrate cross-team success

Engages & Influences

Inspires action and alignment through clear communication, purposeful influence, and a compelling vision.

  • Use storytelling to build buy-in
  • Align communication with organizational goals
  • Guild alignment through strong engagement

Maximizes Performance

Sets and reinforces performance standards that drive results, ensure accountability, and align with Epiq’s goals.

  • Use data to identify improvement opportunities
  • Make informed decisions
  • Align team goals with boarder strategy
  • Empower teams to manage their own goals
  • Translate vision into clear priorities
  • Prepare for disruptions with strong change management

Achieves Operational Success

Drives continuous improvement and operational excellence through smart processes, data insights, and quality execution.

  • Improve workflows for team efficiency
  • Use clear documentation and expectations
  • Resolve issues quickly using data and feedback

Requirements

  • 3+ years of experience in infrastructure engineering, platform engineering, or site reliability engineering.
  • Demonstrated experience building and operating production infrastructure, including on-call responsibility for systems of your own design.
  • Hands-on experience with at least one major cloud platform such as AWS, GCP, or Azure.
  • Experience with infrastructure-as-code tools, particularly Terraform, including reusable module design.
  • Production Kubernetes experience, including scaling, upgrades, resource limits, network policy, and troubleshooting.
  • Ownership of CI/CD pipelines using GitHub Actions, Azure DevOps, or a comparable platform.
  • Experience with observability tooling such as Prometheus, Grafana, or OpenTelemetry, including defining and maintaining service-level objectives.
  • Demonstrated incident-command experience in production environments.
  • Experience with secrets and certificate management at organizational scale.
  • Proficiency in Python, Go, or a comparable language sufficient to build tooling and automation.
  • Strong system-design and architecture experience, including production of technical design documents.

Technology Stack

· Azure · Terraform · Kubernetes · Docker · CI/CD platform to be confirmed · Prometheus · Grafana · OpenTelemetry · PostgreSQL · RabbitMQ · Python.

LI-KS1

The Compensation range for this role is $145,000 -$195,000 USD annually and may be eligible for an annual bonus.

In compliance with federal law, all persons hired will be required to verify identity and eligibility to work in the United States and to complete the required employment eligibility verification form upon hire.

Must be authorized to work in the United States for any employer.

Benefits & conditions

Posted Yesterday In-Office or Remote Hiring Remotely in USA 145K-195K Annually Mid level In-Office or Remote Hiring Remotely in USA 145K-195K Annually Mid level Build and operate cloud infrastructure for an AI platform: design Terraform modules, run production Kubernetes, manage CI/CD and release engineering, implement observability and SLOs, enforce security/compliance controls, manage secrets and certificates, and support incident response and platform readiness. The summary above was generated by AI

About the company

At Epiq, your work contributes to complex, global legal outcomes. You’ll join a values-driven community where integrity guides decisions, relentless service sets the bar, and we thrive on big challenges together. We invest in your growth with enterprise-wide learning and mobility. We celebrate who you are, and we respect life beyond work with flexibility that’s recognized externally. Enabled by modern platforms and AI, you’ll do the most meaningful work of your career and see your impact at scale., available upon request. Epiq is pleased to provide such assistance and no applicant will be penalized as a result of such a request. Pursuant to relevant law, where applicable, Epiq will consider for employment qualified applicants with arrest and conviction records.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.builtincolorado.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · World Congress 2023

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

4:36 min

Hiring passionate software engineers to tackle unprecedented scaling challenges

Dana Lawson Dana Lawson +1 · World Congress 2026 Europe

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · World Congress 2026 Europe

Videos

See all

Related articles

See all