AWS Solution Architect
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+26 more
Job description
Lead SRE observability workstream focusing on AWS-native monitoring solutions, migrating from third-party tools (Datadog/Splunk) to AWS CloudWatch ecosystem. Design and implement comprehensive observability architecture using CloudWatch Metrics V2, AWS AppSignals, AWS Distro for OpenTelemetry (ADOT), and AWS X-Ray. Establish enterprise-grade dashboards, alerting frameworks, and CloudFormation-based infrastructure automation for monitoring stack deployment., * Observability Migration Leadership: Lead migration from Datadog/Splunk to AWS CloudWatch Metrics V2, ensuring feature parity, cost optimization, and minimal disruption to monitoring capabilities
-
Deep Metrics Integration: Implement advanced CloudWatch metrics collection from ZeHorizon platforms, custom application metrics, and infrastructure telemetry with enhanced resolution and dimensionality
-
Distributed Tracing & APM: Design and deploy AWS X-Ray distributed tracing architecture integrated with ADOT (AWS Distro for OpenTelemetry) for end-to-end application performance monitoring
-
AppSignals Implementation: Configure AWS AppSignals for automatic service-level objective (SLO) tracking, anomaly detection, and application health monitoring
-
Dashboard & Alerting: Create CloudWatch dashboards with cross-service correlation, CloudWatch Alarms with composite alarms, and EventBridge-based alert routing to incident management systems
-
Infrastructure as Code: Develop CloudFormation templates (CFT) for repeatable observability stack deployment across multi-account/multi-region environments
-
Cost Optimization: Optimize CloudWatch costs through metric filtering, log retention policies, and efficient query patterns while maintaining observability coverage
Key Deliverables:
-
Migration runbooks from Datadog/Splunk to CloudWatch ecosystem
-
ADOT Collector configurations for custom metrics and traces
-
CloudWatch dashboard templates for application and infrastructure monitoring
-
CloudFormation templates for observability stack automation
-
Alert escalation policies and runbook automation
-
Cost optimization recommendations and implementation
AWS Skills & Services (added)
Observability & Monitoring:
-
Amazon CloudWatch: Expert-level proficiency in CloudWatch Metrics V2 (high-resolution metrics, metric math, anomaly detection), CloudWatch Logs (Logs Insights queries, subscription filters, metric filters), CloudWatch Alarms (composite alarms, alarm actions), and CloudWatch Dashboards (cross-account/cross-region dashboards, custom widgets)
-
AWS X-Ray: Distributed tracing architecture, service maps, trace analysis, sampling rules, X-Ray SDK integration, and X-Ray daemon configuration
-
AWS Distro for OpenTelemetry (ADOT): ADOT Collector deployment (ECS/EKS/EC2), OpenTelemetry SDK instrumentation, exporters configuration (CloudWatch, X-Ray, Prometheus), and custom metric pipelines
-
AWS AppSignals: Service-level objective (SLO) configuration, automatic instrumentation, anomaly detection, and application health scoring
-
Amazon Managed Service for Prometheus (AMP): Prometheus-compatible metrics collection, PromQL queries, and Grafana integration (if hybrid monitoring required)
-
Amazon Managed Grafana (AMG): Dashboard creation, data source integration (CloudWatch, AMP, X-Ray), and role-based access control
Logging & Analytics:
-
CloudWatch Logs Insights: Advanced query syntax, saved queries, query visualization, and log pattern analysis
-
Amazon OpenSearch Service: Centralized log aggregation (if replacing Splunk), index management, and Kibana dashboards
-
AWS CloudTrail: API activity monitoring, CloudTrail Insights for anomaly detection, and integration with CloudWatch Logs
Alerting & Incident Management:
-
Amazon EventBridge: Event-driven alerting, custom event patterns, cross-account event routing, and integration with incident management tools (PagerDuty, Opsgenie)
-
Amazon SNS: Multi-channel alert notifications (email, SMS, Lambda, SQS), topic subscriptions, and message filtering
-
AWS Chatbot: Slack/Microsoft Teams integration for alert notifications and interactive runbook execution
Infrastructure as Code & Automation:
-
AWS CloudFormation: StackSets for multi-account deployments, nested stacks, custom resources, and drift detection for observability infrastructure
-
AWS CDK: Programmatic infrastructure definition for complex observability patterns
-
AWS Systems Manager: Parameter Store for configuration management, Automation documents for remediation runbooks, and OpsCenter for incident tracking
Container & Compute Monitoring:
-
Amazon ECS/EKS: Container Insights, ADOT sidecar deployment, service mesh observability (AWS App Mesh), and pod-level metrics
-
AWS Lambda: Lambda Insights, custom CloudWatch metrics from Lambda functions, and X-Ray tracing for serverless applications
-
Amazon EC2: CloudWatch agent configuration, custom metrics publishing, and Systems Manager integration
Networking & Security:
-
Amazon VPC: VPC Flow Logs analysis, network performance monitoring, and Transit Gateway metrics
-
AWS Network Firewall: Firewall logs integration with CloudWatch, alert rules for security events
-
AWS IAM: Fine-grained access control for CloudWatch resources, cross-account monitoring roles, and service-linked roles
-
Cost Management:
-
AWS Cost Explorer: CloudWatch cost analysis, usage patterns, and cost allocation tags
-
AWS Budgets: Cost alerts for observability services, anomaly detection for unexpected spend
Migration & Integration:
- Datadog/Splunk Migration Patterns: Understanding of Datadog APM, metrics, and log forwarding to design equivalent CloudWatch solutions; Splunk HEC (HTTP Event Collector) to CloudWatch Logs migration strategies
Third-Party Integrations: Webhook integrations, API-based metric ingestion, and custom CloudWatch metric publishers
Requirements
-
Large regulated financial-services delivery with formal change-control, audit and risk governance
-
Operational resilience expectations including RTO/RPO, multi-region DR and evidence for audit review
-
Awareness of applicable controls and regulations such as DORA, NIST CSF 2.0, PCI DSS, SEC cyber rules, RegSCI and SIFMU/FMI expectations where relevant
-
Ability to create Tech Risk-ready documentation including ADRs, runbooks, design docs, threat models and validation evidence
-
Clear communication with client engineering, security, SRE, data and platform stakeholders as an embedded SME
Certifications / Qualifications
-
AWS Certified Solutions Architect - Associate / Professional
-
AWS Certified Developer - Associate
-
AWS Certified DevOps Engineer - Professional preferred
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Why Attend a Developer Event in 2026?
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
Events like RSAC Get You CISOs. Developers Decide What Actually Gets Deployed.
Top Must-Visit Developer Conferences in the US in 2026