Software Engineer, Infrastructure & Reliability
Role details
Job location
Tech stack
Job description
Reposted 12 Hours Ago Remote Hiring Remotely in United States Mid level Remote Hiring Remotely in United States Mid level Build and operate CrewAI's infrastructure across AWS, Azure, and GCP. Own containerized deployments, CI/CD pipelines, observability, secrets, networking, databases, and on-call reliability. Harden security, automate deployments and installs, partner with runtime and product engineers, and reduce operational toil through tooling and automation. The summary above was generated by AI About CrewAI
CrewAI is the leading framework and enterprise platform for building and orchestrating multi-agent AI systems, powering 300M+ agent executions per month across thousands of companies. The Agent Management Platform is our control plane for deploying, monitoring, governing, and scaling agents in production. This role owns the infrastructure foundation that keeps it reliable, secure, and fast. The Role
You'll build and operate the platform infrastructure behind CrewAI's cloud and enterprise deployments. You'll work across multiple hyperscalers - AWS, Azure, and GCP. You'll work on containers, CI/CD, deployment automation, observability, secrets, networking, and runtime reliability. Your job is to make the product and runtime teams faster while making customer's production environments safer.
This is not a pure DevOps support role. You'll write code, improve systems, design deployment paths, harden production, and build the internal platform that lets CrewAI scale and scale our customer deployments. What You'll Do
- Own and improve the infrastructure that runs CrewAI's platform: AWS, ECS/ECR, Docker, Kubernetes/Helm, networking, secrets, databases, Redis, and related services.
- Build and maintain CI/CD pipelines for build, test, image publishing, migrations, environment promotion, rollbacks, and deploy safety.
- Improve reliability across cloud and enterprise deployments: health checks, alerting, incident response, capacity planning, recovery paths, and operational runbooks - and own the front-line on-call rotation and its SLAs.
- Partner with runtime engineers on Celery/FastAPI/Redis workloads and with product engineers on Rails/Solid Queue/Postgres production behavior.
- Manage production observability and telemetry infrastructure: logs, metrics, traces, dashboards, Sentry/OpenTelemetry plumbing, actionable alerts, and telemetry export to customers' own monitoring systems.
- Harden security and compliance posture across IAM, workload identity, secrets management, vulnerability scanning, dependency/image hygiene, and least-privilege access.
- Build the tooling and automation that lets field engineers and customers run self-hosted installs themselves - Helm charts, environment config, release artifacts, pre-flight checks, and install runbooks - so engineering does fewer hands-on installs over time.
- Reduce operational toil by automating recurring workflows and making deployments boring., 10 Minutes Ago Remote or Hybrid United States 74K-98K Annually Mid level 74K-98K Annually Mid level Automotive * Professional Services * Software * Consulting * Energy * Chemical * Renewable Energy Lead and support GHG emissions assurance engagements: perform analytical reviews of emissions and environmental data, manage project activities, prepare client materials, engage clients, support business development, and contribute to thought leadership and industry events. Top Skills: ExcelSustainability Reporting Software Wipfli
Senior Accountant
33 Minutes Ago Remote or Hybrid Denver, CO, USA 83K-113K Annually Senior level 83K-113K Annually Senior level Cloud * Fintech * Software * Business Intelligence * Consulting * Financial Services Provide assistant controller-level support to healthcare clients: manage financial reporting and general ledgers, prepare and finalize month-end close and monthly reporting packages, review balance sheets and account classifications, maintain client procedures, coordinate project workflows, supervise and estimate accountant work, and collaborate across teams to develop best practices. Top Skills: Microsoft Office Suite Wipfli
Senior Accountant
33 Minutes Ago Remote or Hybrid Senior level Senior level Cloud * Fintech * Software * Business Intelligence * Consulting * Financial Services Provide assistant controller-level support to healthcare clients: manage financial reporting accuracy, general ledger and month-end close, prepare and review reporting packages and KPIs, maintain client procedure manuals, oversee accountants' work, coordinate projects, and ensure balance sheet and workpaper accuracy. Top Skills: ExcelMS Office
What you need to know about the Colorado Tech Scene
With a business-friendly climate and research universities like CU Boulder and Colorado State, Colorado has made a name for itself as a startup ecosystem. The state boasts a skilled workforce and high quality of life thanks to its affordable housing, vibrant cultural scene and unparalleled opportunities for outdoor recreation. Colorado is also home to the National Renewable Energy Laboratory, helping cement its status as a hub for renewable energy innovation.
Key Facts About Colorado Tech
- Number of Tech Workers: 260,000; 8.5% of overall workforce (2024 CompTIA survey)
- Major Tech Employers: Lockheed Martin, Century Link, Comcast, BAE Systems, Level 3
- Key Industries: Software, artificial intelligence, aerospace, e-commerce, fintech, healthtech
- Funding Landscape: $4.9 billion in VC funding in 2024 (Pitchbook)
- Notable Investors: Access Venture Partners, Ridgeline Ventures, Techstars, Blackhorn Ventures
- Research Centers and Universities: Colorado School of Mines, University of Colorado Boulder, University of Denver, Colorado State University, Mesa Laboratory, Space Science Institute, National Center for Atmospheric Research, National Renewable Energy Laboratory, Gottlieb Institute
Requirements
- Strong infrastructure/platform engineering experience in production SaaS environments.
- Deep practical experience with AWS, Docker, CI/CD, GitHub Actions, and containerized services.
- Experience with ECS and/or Kubernetes; Helm experience is a strong plus.
- Comfort operating PostgreSQL, Redis, background job systems, queues, and web services in production.
- Strong debugging instincts across app, infra, network, deploy, and dependency layers.
- Security-minded approach to IAM, secrets, workload identity, vulnerability management, and production access.
- Ability to write reliable automation in Python, Ruby, Go, Bash, or similar.
- Calm, rigorous approach to incidents, rollbacks, migrations, and production change management.
Bonus
- Experience with AI/agent platforms, workflow runtimes, or high-volume async execution systems.
- Experience supporting enterprise/self-hosted deployments.
- Terraform or other IaC experience.
- SRE background: SLOs, incident review, capacity planning, load testing.
- Familiarity with Rails, FastAPI, Celery, OpenTelemetry, or multi-service observability.