> Markdown version of [/jobs/ext/1483458-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/1483458-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** Insight Global - **Location:** Chicago, IL, United States - **Experience:** Expert - **Salary:** $89,440.0 - $112,320.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Microsoft Azure, Bash Shell, Cloud Computing, Continuous Integration, Software Debugging, DevOps, Github, Python (Programming Language), Reliability Engineering, AI Infrastructure, Large Language Models, Gitlab-ci, Kubernetes, Jenkins - **Published:** July 29, 2026 - **Apply:** https://jobs.insightglobal.com/jobs/find_a_job/illinois/chicago/site-reliability-engineer/job-556966/ ## About the Role · 5+ years of hands-on DevOps, SRE, or platform engineering experience in a production environment. · Strong Kubernetes experience - you have run production Kubernetes workloads and debugged real cluster issues. · Solid experience with cloud infrastructure (AWS strongly preferred; Azure also relevant). · Proficiency in Python, Bash, or a similar language for automation and tooling. · Experience with modern CI/CD tooling (GitHub Actions, Jenkins, GitLab CI, or similar) and with GitOps patterns. · Strong debugging, systems-thinking, and root-cause analysis skills. Clear written and verbal communication - comfortable working across geographically distributed teams and with engineers outside your immediate area. Nice to Have Skills & Experience · Direct experience hosting or operating LLM gateways, model gateways, or LLM-adjacent infrastructure (LiteLLM, model routers, inference platforms). · Familiarity with the emerging MCP and agent ecosystem (MCP servers, agent runtimes such as KA Agent or LangSmith Fleet, agent registries). · Experience with observability platforms (i.e. Langsmith) and cost-attribution patterns for shared multi-tenant infrastructure. · Experience operating platforms in a regulated or financial services environment. AWS certifications. ## Description As a Senior DevOps / SRE Engineer on contract, you will be embedded with the Central Technology AI enablement team, working alongside engineers from Direct, PitchBook, Retirement, and other business units. Your initial focus will be on the SRE and hosting side of our growing AI platform footprint. As the platform matures, we expect this role could extend into hands-on contributions to the AI enablement components themselves, including the agent registry, agent runtime platform, and MCP tooling. This is a hands-on individual contributor role. You will have the opportunity to shape how a large financial services organization operates production AI infrastructure at scale, working with a small, senior team that moves quickly and makes decisions in the open. ## Related Videos - [The Road to MLOps: How Verivox Transitioned to AWS](https://www.wearedevelopers.com/videos/1050-the-road-to-mlops-how-verivox-transitioned-to-aws) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [DevOps for AI: running LLMs in production with Kubernetes and KubeFlow](https://www.wearedevelopers.com/videos/1222-devops-for-ai-running-llms-in-production-with-kubernetes-and-kubeflow) - [Our GitOps approach for deploying an Identity Provider and an API Gateway in a SaaS company](https://www.wearedevelopers.com/videos/776-our-gitops-approach-for-deploying-an-identity-provider-and-an-api-gateway-in-a-saas-company) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path) - [Dev Digest 210: AI Agents Are Go! Is MCP Dead? LLMs Crack Anonymity](https://www.wearedevelopers.com/magazine/709-dev-digest-210-ai-agents-are-go-is-mcp-dead-llms-crack-anonymity) - [Everything a Developer Needs to Know About MCP with Neo4j](https://www.wearedevelopers.com/magazine/604-everything-a-developer-needs-to-know-about-mcp-with-neo4j) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline)