AWS Solution Architect

FSTONE Technologies
Dallas, TX, United States
17 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Amazon Web Services Amazon Elastic Compute Cloud Data Analysis Application Performance Management Configuration Management Software Design Documents DevOps Monitoring of Systems Identity and Access Management Issue Tracking Systems Network Security
+26 more
Automation of Marketing Pattern Recognition PCI Data Security Standards Role-Based Access Control Reliability Engineering Prometheus Runbook Datadog AWS Cdk Data Logging Chatbots Grafana AWS Lambda Firewalls (Computer Science) Amazon Virtual Private Cloud (VPC) Cloudformation Infrastructure Automation Frameworks Integration Frameworks Api Design Cloudwatch Kibana Amazon Simple Queue Service (SQS) Splunk Dynatrace Serverless Computing Pagerduty

Job description

Lead SRE observability workstream focusing on AWS-native monitoring solutions, migrating from third-party tools (Datadog/Splunk) to AWS CloudWatch ecosystem. Design and implement comprehensive observability architecture using CloudWatch Metrics V2, AWS AppSignals, AWS Distro for OpenTelemetry (ADOT), and AWS X-Ray. Establish enterprise-grade dashboards, alerting frameworks, and CloudFormation-based infrastructure automation for monitoring stack deployment., * Observability Migration Leadership: Lead migration from Datadog/Splunk to AWS CloudWatch Metrics V2, ensuring feature parity, cost optimization, and minimal disruption to monitoring capabilities

  • Deep Metrics Integration: Implement advanced CloudWatch metrics collection from ZeHorizon platforms, custom application metrics, and infrastructure telemetry with enhanced resolution and dimensionality

  • Distributed Tracing & APM: Design and deploy AWS X-Ray distributed tracing architecture integrated with ADOT (AWS Distro for OpenTelemetry) for end-to-end application performance monitoring

  • AppSignals Implementation: Configure AWS AppSignals for automatic service-level objective (SLO) tracking, anomaly detection, and application health monitoring

  • Dashboard & Alerting: Create CloudWatch dashboards with cross-service correlation, CloudWatch Alarms with composite alarms, and EventBridge-based alert routing to incident management systems

  • Infrastructure as Code: Develop CloudFormation templates (CFT) for repeatable observability stack deployment across multi-account/multi-region environments

  • Cost Optimization: Optimize CloudWatch costs through metric filtering, log retention policies, and efficient query patterns while maintaining observability coverage

Key Deliverables:

  • Migration runbooks from Datadog/Splunk to CloudWatch ecosystem

  • ADOT Collector configurations for custom metrics and traces

  • CloudWatch dashboard templates for application and infrastructure monitoring

  • CloudFormation templates for observability stack automation

  • Alert escalation policies and runbook automation

  • Cost optimization recommendations and implementation

AWS Skills & Services (added)

Observability & Monitoring:

  • Amazon CloudWatch: Expert-level proficiency in CloudWatch Metrics V2 (high-resolution metrics, metric math, anomaly detection), CloudWatch Logs (Logs Insights queries, subscription filters, metric filters), CloudWatch Alarms (composite alarms, alarm actions), and CloudWatch Dashboards (cross-account/cross-region dashboards, custom widgets)

  • AWS X-Ray: Distributed tracing architecture, service maps, trace analysis, sampling rules, X-Ray SDK integration, and X-Ray daemon configuration

  • AWS Distro for OpenTelemetry (ADOT): ADOT Collector deployment (ECS/EKS/EC2), OpenTelemetry SDK instrumentation, exporters configuration (CloudWatch, X-Ray, Prometheus), and custom metric pipelines

  • AWS AppSignals: Service-level objective (SLO) configuration, automatic instrumentation, anomaly detection, and application health scoring

  • Amazon Managed Service for Prometheus (AMP): Prometheus-compatible metrics collection, PromQL queries, and Grafana integration (if hybrid monitoring required)

  • Amazon Managed Grafana (AMG): Dashboard creation, data source integration (CloudWatch, AMP, X-Ray), and role-based access control

Logging & Analytics:

  • CloudWatch Logs Insights: Advanced query syntax, saved queries, query visualization, and log pattern analysis

  • Amazon OpenSearch Service: Centralized log aggregation (if replacing Splunk), index management, and Kibana dashboards

  • AWS CloudTrail: API activity monitoring, CloudTrail Insights for anomaly detection, and integration with CloudWatch Logs

Alerting & Incident Management:

  • Amazon EventBridge: Event-driven alerting, custom event patterns, cross-account event routing, and integration with incident management tools (PagerDuty, Opsgenie)

  • Amazon SNS: Multi-channel alert notifications (email, SMS, Lambda, SQS), topic subscriptions, and message filtering

  • AWS Chatbot: Slack/Microsoft Teams integration for alert notifications and interactive runbook execution

Infrastructure as Code & Automation:

  • AWS CloudFormation: StackSets for multi-account deployments, nested stacks, custom resources, and drift detection for observability infrastructure

  • AWS CDK: Programmatic infrastructure definition for complex observability patterns

  • AWS Systems Manager: Parameter Store for configuration management, Automation documents for remediation runbooks, and OpsCenter for incident tracking

Container & Compute Monitoring:

  • Amazon ECS/EKS: Container Insights, ADOT sidecar deployment, service mesh observability (AWS App Mesh), and pod-level metrics

  • AWS Lambda: Lambda Insights, custom CloudWatch metrics from Lambda functions, and X-Ray tracing for serverless applications

  • Amazon EC2: CloudWatch agent configuration, custom metrics publishing, and Systems Manager integration

Networking & Security:

  • Amazon VPC: VPC Flow Logs analysis, network performance monitoring, and Transit Gateway metrics

  • AWS Network Firewall: Firewall logs integration with CloudWatch, alert rules for security events

  • AWS IAM: Fine-grained access control for CloudWatch resources, cross-account monitoring roles, and service-linked roles

  • Cost Management:

  • AWS Cost Explorer: CloudWatch cost analysis, usage patterns, and cost allocation tags

  • AWS Budgets: Cost alerts for observability services, anomaly detection for unexpected spend

Migration & Integration:

  • Datadog/Splunk Migration Patterns: Understanding of Datadog APM, metrics, and log forwarding to design equivalent CloudWatch solutions; Splunk HEC (HTTP Event Collector) to CloudWatch Logs migration strategies

Third-Party Integrations: Webhook integrations, API-based metric ingestion, and custom CloudWatch metric publishers

Requirements

  • Large regulated financial-services delivery with formal change-control, audit and risk governance

  • Operational resilience expectations including RTO/RPO, multi-region DR and evidence for audit review

  • Awareness of applicable controls and regulations such as DORA, NIST CSF 2.0, PCI DSS, SEC cyber rules, RegSCI and SIFMU/FMI expectations where relevant

  • Ability to create Tech Risk-ready documentation including ADRs, runbooks, design docs, threat models and validation evidence

  • Clear communication with client engineering, security, SRE, data and platform stakeholders as an embedded SME

Certifications / Qualifications

  • AWS Certified Solutions Architect - Associate / Professional

  • AWS Certified Developer - Associate

  • AWS Certified DevOps Engineer - Professional preferred

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:24 min

Evaluating formal AWS certifications versus raw practical engineering experience

Jan Giacomelli · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

3:21 min

Deploying a primary Elasticsearch and Kibana cluster configuration

Philipp Krenn · World Congress 2022

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

Videos

See all

Related articles

See all