> Markdown version of [/jobs/ext/2835985-principal-site-reliability-engineer-in-beverly-hills](https://www.wearedevelopers.com/jobs/ext/2835985-principal-site-reliability-engineer-in-beverly-hills). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal Site Reliability Engineer in Beverly Hills - **Company:** Energy Jobline - **Location:** Beverly Hills, CA, United States - **Experience:** Expert - **Contract:** Temporary contract - **Skills:** Amazon Web Services, Cloud Computing, Code Review, DevOps, Global Distribution Systems, Performance Tuning, Reliability Engineering, Prometheus, Web Services, System Availability, Grafana, Reliability of Systems, Containerization, Kubernetes, Infrastructure Automation Frameworks, Data Analytics, Terraform - **Published:** September 10, 2026 - **Apply:** https://www.energyjobline.com/job/principal-site-reliability-engineer-beverly-hills-31600382 ## About the Role * 7+ years of experience in Site Reliability Engineering, DevOps, or related roles, with a track record of improving system reliability and operational maturity * Strong expertise in cloud platforms and modern infrastructure environments (e.g., AWS, containerized workloads, or similar ecosystems) * Experience with infrastructure automation and container orchestration (e.g., Terraform, Kubernetes or equivalent technologies) * Deep understanding of multi-tenant architecture, security principles, and data protection practices * Hands-on experience with observability tools and monitoring frameworks (e.g., Prometheus, Grafana or similar) * Experience implementing automated compliance and governance practices (e.g., SOC 2, GDPR, ISO 27001 or similar standards) * Strong leadership and mentoring capabilities, with the ability to influence engineering teams and drive adoption of reliability-focused practices ## Description KēSTA I.T. is actively seeking a Principal Engineer for an immediate full-time opportunity with our industry creating client. Are you on the lookout for a unique career opportunity that offers leadership, responsibility, and the chance to make a significant impact? If you're eager to contribute to a thriving and stable organization while maintaining your confidentiality, continue reading. The Opportunity An innovative technology company is seeking experienced Site Reliability Engineers to take ownership of building reliable, scalable platforms that deliver advanced 3D/4D spatial content to global users across AR/VR environments. This is a high-impact role focused on ensuring system reliability at scale, requiring deep expertise in observability, multi-tenant architectures, and data-driven operational decision-making. You will play a key role in designing and maintaining infrastructure that supports high-volume streaming workloads while meeting enterprise-grade security and compliance standards. This role partners closely with web services and platform engineering teams to implement SRE best practices, establish robust monitoring, and build infrastructure capable of supporting rapid growth and global distribution. What You'll Do * Design, configure, and maintain cloud infrastructure using infrastructure-as-code tools (e.g., Terraform), with a focus on optimizing content delivery and CDN performance * Develop and execute capacity planning strategies and performance optimization initiatives for large-scale streaming platforms * Instrument services to monitor system health, building dashboards and alerting systems that provide actionable insights into performance and user experience * Define and implement observability strategies, including SLI/SLO frameworks and error budget management * Establish escalation protocols and participate in on-call rotations to ensure 24/7 system availability * Lead incident response efforts and conduct post-incident reviews to drive continuous improvement * Implement and promote reliability engineering practices, including deployment safety, code review standards, and operational readiness * Mentor engineering teams on best practices for reliability, scalability, and production operations ## Related Videos - [Shifting Stress to Progress— Understanding DevOps to do DevOps Better](https://www.wearedevelopers.com/videos/268-shifting-stress-to-progress-understanding-devops-to-do-devops-better) - [Monitoring as Code - Managing your dashboards at scale](https://www.wearedevelopers.com/videos/753-monitoring-as-code-managing-your-dashboards-at-scale) - [Infrastructure as Code: The Developer's Secret Weapon](https://www.wearedevelopers.com/videos/1221-infrastructure-as-code-the-developer-s-secret-weapon) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Implementing Feature Environments with AWS and Terraform](https://www.wearedevelopers.com/videos/531-implementing-feature-environments-with-aws-and-terraform) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)