Solutions Architect - Mythos SRE

Propertyvalue Lighthouse Technology Services
United States
2 months ago
Apply on dice.com
Prepare application

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$135,200.0 - $166,400.0
Working hours
Regular working hours
Job source

Tech stack

Agile Methodology Artificial Intelligence Microsoft Azure Cloud Computing Disaster Recovery Failover Performance Tuning Systems Development Life Cycle Reliability Engineering Site Reliability Engineering Practices Software Engineering System Availability
+3 more
Delivery Pipeline AI Platforms Programming Languages

Job description

Lighthouse Technology Services is partnering with our client to fill their Solution Architect - Mythos SRE position! This is a 12 - 18 month contract that can be remote in the United States. This role will be a W2 employee of Lighthouse Technology Services. No C2C or subcontracting arrangements will be considered., * Lead the design and implementation of comprehensive reliability, scalability, and operational architecture for the AI platform across Azure cloud and co-location environments, ensuring alignment with enterprise SRE principles and infrastructure standards

  • Architect solutions for reliability and resilience patterns including high availability, failover, disaster recovery, and geo-distribution strategies for AI workloads and model-serving infrastructure
  • Design and guide observability and telemetry frameworks that provide visibility into system health, model performance, drift detection, and risk indicators aligned with AI governance requirements
  • Collaborate with domain leadership and enterprise architects to translate conceptual and logical designs into detailed physical solution architecture that meets business capabilities and financial targets
  • Establish and implement SRE practices including SLIs, SLOs, error budgets, and operational readiness frameworks while guiding automation patterns and infrastructure-as-code deployment pipelines
  • Partner with Agile teams throughout the SDLC to validate architecture decisions, ensure compliance with enterprise standards, and support the delivery of scalable, resilient AI platform operations
  • Drive performance optimization initiatives for model-serving and AI workloads while identifying and evaluating emerging technologies and trends that could impact the domain
  • Facilitate governance activities and collaborate with stakeholders across infrastructure, platform engineering, and observability teams to ensure consistent adoption of architecture patterns

Requirements

  • 5+ years of solution architecture or software engineering experience with demonstrated ability to design and integrate applications using modern architecture principles and patterns
  • Proven expertise in Service Reliability Engineering with hands-on experience establishing SLIs, SLOs, error budgets, and operational readiness frameworks
  • Strong technical proficiency across multiple programming languages and cloud technologies (Azure experience highly preferred)
  • Experience architecting for reliability, scalability, and resilience including high availability, disaster recovery, and performance optimization strategies
  • Solid understanding of observability, telemetry, and monitoring frameworks with ability to implement comprehensive system health and performance visibility
  • Demonstrated ability to work in Agile environments and effectively communicate complex architectural concepts to stakeholders at all levels of the organization
  • Experience with infrastructure-as-code, automation patterns, and deployment pipeline design
  • Industry-recognized certifications in cloud technologies or programming languages preferred

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

5:06 min

Modernizing legacy applications through proactive leadership

Babette Wagner Babette Wagner · Coffee With Developers

2:59 min

Scaling clusters and handling automated replica failover

Jürgen Pilz · World Congress 2023

2:36 min

Choosing between managed AI platforms and custom governance

Péter Farkas Péter Farkas · Europe 2026 Virtual

1:06 min

Outline of free tools for Microsoft Azure

Radu Vunvulea Radu Vunvulea · World Congress 2022

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

5:25 min

Implementing redundancy, failover, and architectural load balancing patterns

Mihaela-Roxana Ghidersa · LIVE

Videos

See all

Related articles

See all