> Markdown version of [/jobs/ext/2844860-software-engineer](https://www.wearedevelopers.com/jobs/ext/2844860-software-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer - **Company:** Wells Fargo - **Location:** United States (Remote available) - **Experience:** Starter - **Salary:** $47,840.0 - $64,480.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Microsoft Azure, Bash Shell, Software as a Service, Cloud Computing, Databases, Continuous Integration, DevOps, Distributed Systems, Github, Monitoring of Systems, Identity and Access Management, Python (Programming Language), Key Management, PostgreSQL, Load Testing, Redis, Ruby, Software Vulnerability Management, Web Services, Google Cloud, Kubernetes Helm Charts, Fastapi, Kubernetes, Information Technology, Deployment Automation, Sentry, Celery, Terraform, Docker, Vulnerability Analysis, Golang - **Published:** September 11, 2026 - **Apply:** https://www.builtincolorado.com/job/software-engineer-infrastructure-reliability/11125505?handler=ApplyRedirect ## About the Role * Strong infrastructure/platform engineering experience in production SaaS environments. * Deep practical experience with AWS, Docker, CI/CD, GitHub Actions, and containerized services. * Experience with ECS and/or Kubernetes; Helm experience is a strong plus. * Comfort operating PostgreSQL, Redis, background job systems, queues, and web services in production. * Strong debugging instincts across app, infra, network, deploy, and dependency layers. * Security-minded approach to IAM, secrets, workload identity, vulnerability management, and production access. * Ability to write reliable automation in Python, Ruby, Go, Bash, or similar. * Calm, rigorous approach to incidents, rollbacks, migrations, and production change management. Bonus * Experience with AI/agent platforms, workflow runtimes, or high-volume async execution systems. * Experience supporting enterprise/self-hosted deployments. * Terraform or other IaC experience. * SRE background: SLOs, incident review, capacity planning, load testing. * Familiarity with Rails, FastAPI, Celery, OpenTelemetry, or multi-service observability. ## Description Hiring Remotely in United States Entry level Remote Hiring Remotely in United States Entry level Build and operate CrewAI's cloud and enterprise infrastructure across AWS, Azure, and GCP. Develop CI/CD pipelines, deployment automation, observability, security controls, self-hosted installation tooling, and reliability processes. Support containers, Kubernetes, databases, queues, networking, secrets, and production workloads. Participate in incident response and on-call operations while partnering with runtime and product engineering teams to improve deployment safety, scalability, and customer production environments. The summary above was generated by AI About CrewAI CrewAI is the leading framework and enterprise platform for building and orchestrating multi-agent AI systems, powering 300M+ agent executions per month across thousands of companies. The Agent Management Platform is our control plane for deploying, monitoring, governing, and scaling agents in production. This role owns the infrastructure foundation that keeps it reliable, secure, and fast. The Role You'll build and operate the platform infrastructure behind CrewAI's cloud and enterprise deployments. You'll work across multiple hyperscalers - AWS, Azure, and GCP. You'll work on containers, CI/CD, deployment automation, observability, secrets, networking, and runtime reliability. Your job is to make the product and runtime teams faster while making customer's production environments safer. This is not a pure DevOps support role. You'll write code, improve systems, design deployment paths, harden production, and build the internal platform that lets CrewAI scale and scale our customer deployments. What You'll Do * Own and improve the infrastructure that runs CrewAI's platform: AWS, ECS/ECR, Docker, Kubernetes/Helm, networking, secrets, databases, Redis, and related services. * Build and maintain CI/CD pipelines for build, test, image publishing, migrations, environment promotion, rollbacks, and deploy safety. * Improve reliability across cloud and enterprise deployments: health checks, alerting, incident response, capacity planning, recovery paths, and operational runbooks - and own the front-line on-call rotation and its SLAs. * Partner with runtime engineers on Celery/FastAPI/Redis workloads and with product engineers on Rails/Solid Queue/Postgres production behavior. * Manage production observability and telemetry infrastructure: logs, metrics, traces, dashboards, Sentry/OpenTelemetry plumbing, actionable alerts, and telemetry export to customers' own monitoring systems. * Harden security and compliance posture across IAM, workload identity, secrets management, vulnerability scanning, dependency/image hygiene, and least-privilege access. * Build the tooling and automation that lets field engineers and customers run self-hosted installs themselves - Helm charts, environment config, release artifacts, pre-flight checks, and install runbooks - so engineering does fewer hands-on installs over time. * Reduce operational toil by automating recurring workflows and making deployments boring., The Senior Infrastructure Engineer will manage and scale LiveKit's core infrastructure, focusing on reliability and performance, while working on Golang code and automation for distributed systems., Design, implement, and manage LiveKit's infrastructure, focusing on performance, reliability, and automation while ensuring effective incident management and coordination with vendors. ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Coffee with Developers: David Heinemeier Hansson](https://www.wearedevelopers.com/videos/875-coffee-with-developers-david-heinemeier-hansson) - [Reducing LLM Calls with Vector Search Patterns - Raphael De Lio (Redis)](https://www.wearedevelopers.com/videos/1714-reducing-llm-calls-with-vector-search-patterns-raphael-de-lio-redis) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Inside Bitpanda's Tech Stack: Scaling a European Fintech Leader - Markus Dorner](https://www.wearedevelopers.com/videos/1979-inside-bitpanda-s-tech-stack-scaling-a-european-fintech-leader-markus-dorner) - [Coroutine explained yet again 60 years later](https://www.wearedevelopers.com/videos/690-coroutine-explained-yet-again-60-years-later) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [The Best X (Twitter) Accounts for Developers](https://www.wearedevelopers.com/magazine/294-the-best-x-twitter-accounts-for-developers) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers)