WeAreDevelopers LIVE Jan 26, 2022

SRE Methods In an Agency Environment

Martin Beránek

Exhausting your error budget every month isn't bad luck—it's a flawed engineering process. Discover how to successfully adapt SRE principles for fast-paced agency environments and seamless client handovers.

Pause
Mute Enter Fullscreen
#1 about 2 min

Defining service level terminology for agencies

The core differences between SLIs, SLOs, and SLAs dictate how availability thresholds and penalties are applied.

#2 about 3 min

Mapping the agency and customer relationship for software delivery

Understanding how customers, agencies, and end-users interact clarifies the context of application delivery and service agreements.

#3 about 1 min

Overcoming agency objections to reliability planning

Project managers and developers often push back against service level agreements when striving for minimum viable products.

#4 about 3 min

Introduction to aki agency and operations context

A fast-paced software agency structure relies heavily on DevOps workflows and standardized application deployments.

#5 about 4 min

Establishing realistic SLO documents before production

Expectations must be managed during pre-production to set rational latency and availability error budgets.

#6 about 4 min

Defining project handover scenarios and knowledge silos

The operational approach shifts depending on whether a customer intends to maintain an application or requires ongoing agency reliability support.

#7 about 2 min

Structuring three core documents for incident support

Service level documents, support playbooks, and postmortems function together to streamline continuous reliability improvements.

#8 about 4 min

Writing blameless and detailed incident postmortems

A constructive framework structures postmortems with root causes, impact assessments, timeline logs, and actionable resolution items.

#9 about 4 min

Assigning strict roles for incident response teams

Clear escalation boundaries isolate the specific duties of an incident commander, communications lead, and operations lead during an active issue.

#10 about 4 min

Accounting for unexpected cloud provider environment events

Proactive communication strategies mitigate the broader stakeholder impact of sudden architecture shifts or pricing changes enforced by cloud vendors.

#11 about 3 min

Aligning automated security issues with error budgets

Predefined error budget policies accurately determine the prioritization response required for automated dependency security scans.

#12 about 3 min

Managing cloud access credentials during project handovers

Routine practices like rotating identities and trimming unused IAM permissions ensure project infrastructure remains secure upon client transfer.

#13 about 2 min

Transferring robust application secrets to external clients

Native cloud secret managers usually provide a safer external handover experience than self-hosted solutions or simple continuous integration variables.

#14 about 3 min

Conducting a transparent support and toolset adoption period

Stripping away agency-specific automation constraints and reviewing past postmortems directly empowers client teams during operational deployment transitions.

#15 about 3 min

Advocating for SRE practices within agency environments

Educating stakeholders early and focusing directly on end-user experience ultimately secures project management buy-in for sustainable reliability work.

#16 about 7 min

Answering inquiries on SLA negotiations and observability tooling

Strategic responses to audience inquiries detail practical approaches for resolving early service agreement disputes and selecting native tracking metrics.

Matching moments

8:02 min

Integrating service level objectives into incident management

Diana Todea · LIVE

3:53 min

Applying software development methodologies to incident response

Tobias Dunn-Krahn · LIVE

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · WWC 2023

3:50 min

Scaling shift left practices within large engineering organizations

Chris Riley · WWC 2021

3:08 min

Shifting software delivery bottlenecks to operations and incident response

Milin Desai Milin Desai +1 · WWC Europe 2026

3:44 min

Balancing feature releases with reliable service targets

Diana Todea · LIVE

Upcoming sessions on this topic

Open session

World Congress 2026 North America

Boring Failover: Predictable Region Recovery Across 5,000 Microservices

Sahil Sabharwal, Garvit Kataria

Sahil Sabharwal
Garvit Kataria
Open session

World Congress 2026 North America

Closing the Visibility Gap: Lessons from Safety Critical Agentic Systems

Vivek Pandit

Principal Engineer at Cadence

Vivek Pandit
Open session

World Congress 2026 North America

Your Evals Passed. Your Agent Just Emptied a Database.

Tejas Pravinbhai Patel

IEEE Award-Winning Researcher | Best Keynote Speaker | Sr. Software Engineer at Amazon | AI Systems & Agent Architect

Tejas Pravinbhai Patel
Open session

World Congress 2026 North America

Reinventing Incident Response with AI Agents and MCP

Jayant Tyagi

Lead Member of Technical Staff @ Salesforce

Jayant Tyagi
Open session

World Congress 2026 North America

Give the Agent a Budget, Not a Token

Sachin Malhotra

MTS @Anthropic

Sachin Malhotra
Open session

World Congress 2026 North America

Agents Can't Iterate Against Tests That Lie

Rocky Warren

Senior Staff Software Engineer at Clipboard

Rocky Warren