> Markdown version of [/jobs/ext/2283145-sre-ai-platforms](https://www.wearedevelopers.com/jobs/ext/2283145-sre-ai-platforms). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # SRE, AI Platforms - **Company:** Brady Mainz Group Inc - **Location:** Portland, OR, United States (Remote available) - **Experience:** Expert - **Contract:** Temporary contract - **Skills:** Bash Shell, Business Systems, Cloud Computing, Data Centers, Python (Programming Language), Machine Learning, Prometheus, Data Logging, High Performance Computing, Grafana, AI Platforms, Information Technology, Pagerduty, Elk Stack, Golang - **Published:** August 28, 2026 - **Apply:** https://mbg.com/jobs/software-asset-manager/ ## About the Role The ideal candidate is a senior, hands-on engineer who can troubleshoot complex production environments while driving long-term improvements in platform availability and resilience., * 7+ years of SRE, production operations, or reliability-focused infrastructure engineering experience. * Strong hands-on expertise with Linux, Kubernetes, networking, and distributed systems. * Experience building observability, monitoring, alerting, and incident-response workflows. * Strong scripting and automation experience with Python, Bash, Go, or similar technologies. * Experience with incident management, root cause analysis, and post-incident remediation. * Ability to automate repetitive operational processes and reduce production support overhead. * Strong troubleshooting skills across complex, highly available production environments. * Excellent communication and ability to work across technical teams. Preferred Experience * AI platforms, machine learning infrastructure, or large-scale HPC environments. * Prometheus, Grafana, ELK Stack, OpenSearch, Loki, and PagerDuty. * Defining and managing SLIs, SLOs, and error budgets. * Supporting both batch and service-based workloads. * Cloud and data center infrastructure. * Logging, tracing, infrastructure monitoring, and service reliability tooling. ## Description * Operate and improve the reliability of AI platform services, cluster dependencies, and shared infrastructure. * Lead and support incident triage involving Kubernetes, Linux, storage, networking, scheduling platforms, job orchestration, and dependency failures. * Define and improve SLIs, SLOs, alerting thresholds, runbooks, escalation paths, and post-incident corrective actions. * Analyze recurring failures and turn manual operational processes into automation and preventative controls. * Build observability across systems and services using metrics, logs, traces, and event correlation. * Troubleshoot performance and availability issues impacting AI/ML training workloads, inference services, and internal platforms. * Partner with infrastructure and validation teams to improve production readiness and change-management safety. * Drive operational reviews, resilience testing, and readiness assessments., Senior Software Asset Management Analyst Join a global enterprise technology team in a two-year hybrid assignment based in Beaverton, Oregon. You'll manage complex software licensing, strengthen compliance, and optimize software assets across multiple business and IT organizations. Top 3..., Senior AI Security Engineer Help secure next-generation AI applications by partnering with development teams to identify vulnerabilities, red-team LLM systems, and build security into products from the start. Top 3 Technical Skills AI/LLM Application Security: Deep understanding of agentic AI, MCP,..., Director of Systems & Enablement About the Role We are seeking an experienced Director of Systems & Enablement to own and evolve the business systems ecosystem supporting a rapidly growing, multi-market organization. This is a highly visible leadership role responsible for the systems,..., Senior Software Engineer - Distributed AI Systems Help build production-grade distributed systems that power cutting-edge applied AI solutions for American manufacturing. You'll lead backend architecture, develop scalable microservices and data pipelines, and mentor engineers while helping... ## Related Videos - [AI in Leadership: How Technology is Reshaping Executive Roles](https://www.wearedevelopers.com/videos/1705-ai-in-leadership-how-technology-is-reshaping-executive-roles) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) - [Designing UX for SRE Agents in High-Stakes Incidents](https://www.wearedevelopers.com/videos/100003-designing-ux-for-sre-agents-in-high-stakes-incidents) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)