Site Reliability Engineer

Insight Global
Chicago, IL, United States
15 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$89,440.0 - $112,320.0
Working hours
Regular working hours

Tech stack

Amazon Web Services Microsoft Azure Bash Shell Cloud Computing Continuous Integration Software Debugging DevOps Github Python (Programming Language) Reliability Engineering AI Infrastructure Large Language Models
+3 more
Gitlab-ci Kubernetes Jenkins

Job description

As a Senior DevOps / SRE Engineer on contract, you will be embedded with the Central Technology AI enablement team, working alongside engineers from Direct, PitchBook, Retirement, and other business units. Your initial focus will be on the SRE and hosting side of our growing AI platform footprint. As the platform matures, we expect this role could extend into hands-on contributions to the AI enablement components themselves, including the agent registry, agent runtime platform, and MCP tooling. This is a hands-on individual contributor role. You will have the opportunity to shape how a large financial services organization operates production AI infrastructure at scale, working with a small, senior team that moves quickly and makes decisions in the open.

Requirements

· 5+ years of hands-on DevOps, SRE, or platform engineering experience in a production environment. · Strong Kubernetes experience - you have run production Kubernetes workloads and debugged real cluster issues. · Solid experience with cloud infrastructure (AWS strongly preferred; Azure also relevant). · Proficiency in Python, Bash, or a similar language for automation and tooling. · Experience with modern CI/CD tooling (GitHub Actions, Jenkins, GitLab CI, or similar) and with GitOps patterns. · Strong debugging, systems-thinking, and root-cause analysis skills. Clear written and verbal communication - comfortable working across geographically distributed teams and with engineers outside your immediate area.

Nice to Have Skills & Experience

· Direct experience hosting or operating LLM gateways, model gateways, or LLM-adjacent infrastructure (LiteLLM, model routers, inference platforms). · Familiarity with the emerging MCP and agent ecosystem (MCP servers, agent runtimes such as KA Agent or LangSmith Fleet, agent registries). · Experience with observability platforms (i.e. Langsmith) and cost-attribution patterns for shared multi-tenant infrastructure. · Experience operating platforms in a regulated or financial services environment. AWS certifications.

Benefits & conditions

Benefit packages for this role will start on the 1st day of employment and include medical, dental, and vision insurance, as well as HSA, FSA, and DCFSA account options, and 401k retirement account access with employer matching. Employees in this role are also entitled to paid sick leave and/or other paid time off as provided by applicable law.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on jobs.insightglobal.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · World Congress 2023

1:02 min

Applying an ETL methodology to infrastructure configuration management

Axel Barbier · World Congress 2023

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

2:08 min

Applying large language models to infrastructure tasks

Alfonso Sandoval Rosas Alfonso Sandoval Rosas · Europe 2026 Virtual

Videos

See all

Related articles

See all