Topic mix

Site reliability engineering

10 moments from 9 videos · 41:20 min total

This compilation of conference talks provides site reliability engineers and developers with practical methods for managing system scale, measuring reliability, and automating operations.

Staying Safe in the AI Future
Play section Adopting site reliability engineering practices for machine learning
Adopting site reliability engineering practices for machine learning thumbnail

Adopting site reliability engineering practices for machine learning

Building protective safety nets to mitigate system failures when algorithmic execution inevitably deviates from human expectations.

What Developers Get Wrong About Application Quality
Play section Scaling shift left practices within large engineering organizations
Scaling shift left practices within large engineering organizations thumbnail

Scaling shift left practices within large engineering organizations

Embedded site reliability engineers act as guides to help developers manage service maturity and technical debt.

Beyond Chat: AI Workflows That Actually Investigate Alerts (So You Don't Have To Know Everything)
Play section Encoding senior site reliability engineering methodologies into observable pipelines
Encoding senior site reliability engineering methodologies into observable pipelines thumbnail

Encoding senior site reliability engineering methodologies into observable pipelines

Creating verifiable investigation artifacts relies on programmatically mapping resource topology and formally testing system hypotheses.

Serverless Observability: where SLOs meet transforms
Play section Integrating service level objectives into incident management
Integrating service level objectives into incident management thumbnail

Integrating service level objectives into incident management

How site reliability engineers coordinate with product teams and customer support to align user expectations.

Play section Balancing feature releases with reliable service targets
Balancing feature releases with reliable service targets thumbnail

Balancing feature releases with reliable service targets

Distributing responsibilities between software developers and site reliability teams while keeping runbook documentation updated.

Designing UX for SRE Agents in High-Stakes Incidents
Play section Empowering site reliability engineers with integrated AI agents
Empowering site reliability engineers with integrated AI agents thumbnail

Empowering site reliability engineers with integrated AI agents

AI agents embedded directly within production infrastructure help engineers quickly identify the root cause of late-night system failures.

Data binning and understanding histograms
Play section Key takeaways for monitoring service latency
Key takeaways for monitoring service latency thumbnail

Key takeaways for monitoring service latency

Why site reliability engineers must continuously probe APIs and customize buckets to capture real user experiences.

The Human Side of Software Engineering in the Age of AI - Olena Babenko
Play section Building engineering communities and finding technical inspiration
Building engineering communities and finding technical inspiration thumbnail

Building engineering communities and finding technical inspiration

Engaging with colleagues and developer communities provides valuable insights into complex engineering roles like site reliability and operations.

API‑First: How Twilio Designs for Developers - Justin Kitagawa (Twilio)
Play section Engineering practices for extreme platform reliability
Engineering practices for extreme platform reliability thumbnail

Engineering practices for extreme platform reliability

Adopting an operational mindset through cross-team alignment, incremental deployments, and cellular architecture to isolate hardware failures.

Handling incidents collaboratively is like solving a rubix cube
Play section Navigating specialized roles and toolsets across engineering teams
Navigating specialized roles and toolsets across engineering teams thumbnail

Navigating specialized roles and toolsets across engineering teams

Incident resolution requires cross-functional collaboration because backend developers, SREs, and DevOps professionals primarily operate within domain-specific workflows.

Your mix. Instantly.

More mixes