> Markdown version of [/jobs/ext/2994842-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2994842-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** BELONGINGS - **Location:** New York, United States - **Experience:** Expert - **Salary:** $150,000.0 - $225,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon S3, Software Quality, Amazon DynamoDB, Elasticsearch, Github, Python (Programming Language), PostgreSQL, MySQL, NoSQL, Reliability Engineering, Cloud Services, Systems Integration, Software Organization, Datadog, System Availability, Large Language Models, Generative AI, Kubernetes, Virtual Agents - **Published:** September 19, 2026 - **Apply:** https://www.dice.com/job-detail/19936cbd-0ec2-4d1c-b38a-f541e6ab2754 ## About the Role * 5+ years of professional cloud automation or site reliability engineering experience, or similar * Familiarity with software development best practices and experience creating applications in Python * Experience working creating and integrating REST APIs, and event-driven and asynchronous architectures * Experience working with both relational (Postgres, MySQL) and NoSQL databases (DynamoDB, Elastic Search) * Experience with AWS cloud services (S3, OpenSearch, DynamoDB, etc.), Kubernetes-driven deployments, Github Actions, and modern CI/CD pipelines * A passion for AI, creative ideas for how it can be a tool for problem solving and productivity * A thoughtful approach to problem-solving; ability to break down complex issues methodically * Excellent communication skills; Able to translate technical concepts for non-technical stakeholders We'd love if you had: * Experience developing agentic AI applications, harnesses, MCP servers, and workflows * Familiarity with or experience building AI agents, Retrieval-Augmented Generation (RAG) pipelines, or agentic tools * Experience with observability tooling such as DataDog ## Description We're looking for a Site Reliability Engineer to join our AI Technology team and play a critical role in ensuring the stability and scalability of the firm's internal Agentic AI platform. This is a high-impact, hands-on engineering position where you'll contribute code, automation workflows, observability, metrics, and support users as issues arise. You'll work in a dynamic, fast-paced environment where product requirements evolve, and business stakeholders are deeply engaged. What you'll do * You will set the reliability standards for Enterprise AI, defining Service Level Objectives (SLOs), error budgets, and custom incident response runbooks. * You will own the observability, incident response, reliability, and scalability of our AI platform. * Ensure that our agents, gateways, LLM proxies, and RAG pipelines operate with high availability, accuracy, and financial efficiency. * Support users in a dedicated help channel when issues arise - investigating root causes, providing solutions, monitoring the status of upstream dependencies, and communicating updates back to users. * You'll also be involved in the team's code quality and best practices, identify gaps in development lifecycles, and continuously improve both the core platform and how developers build with it. ## Related Videos - [Designing UX for SRE Agents in High-Stakes Incidents](https://www.wearedevelopers.com/videos/100003-designing-ux-for-sre-agents-in-high-stakes-incidents) - [MySQL Protocol Features You Should Be Aware Of](https://www.wearedevelopers.com/videos/100267-mysql-protocol-features-you-should-be-aware-of) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Leveraging Real time data in FSIs](https://www.wearedevelopers.com/videos/806-leveraging-real-time-data-in-fsis) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Coding for Good: Achieving social change with an app](https://www.wearedevelopers.com/videos/1645-coding-for-good-achieving-social-change-with-an-app) ## Related Articles - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Trustworthy AI Starts at Deployment: 5 Checks Before You Ship](https://www.wearedevelopers.com/magazine/753-trustworthy-ai-starts-at-deployment-5-checks-before-you-ship) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline)