Senior AI Platform and Reliability Engineer

ECHOSTAR
Littleton, United States of America
6 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English
Experience level
Senior
Compensation
$ 145K

Job location

Littleton, United States of America

Tech stack

Java
API
Artificial Intelligence
Amazon Web Services (AWS)
Bash
Mobile Application Development
Cloud Computing
Code Review
Data Retrieval
Python
Machine Learning
Open Source Technology
Reliability Engineering
Ansible
Runbook
Software Engineering
Unstructured Data
Large Language Models
Multi-Agent Systems
Prompt Engineering
Reliability of Systems
Generative AI
Backend
AI Platforms
Kubernetes
Information Technology
Virtual Agents
Cloudwatch
Dynatrace
Serverless Computing
Microservices

Job description

Candidates must be willing to participate in at least one in-person interview, which may include a live whiteboarding or technical assessment session.

This role addresses the critical challenge of transitioning from reactive, threshold-based monitoring to a proactive, self-healing infrastructure. The primary focus is bridging the gap between experimental machine learning and production-grade software by architecting resilient systems that orchestrate Large Language Models and agentic workflows. By leveraging predictive intelligence and deep observability, this position ensures the stability and efficiency of high-traffic order provisioning and device activation journeys in complex cloud environments.

What Success Looks Like (Objectives)

  • Architect and deploy self-healing infrastructure that moves operations toward predictive intelligence and autonomous remediation
  • Design high-scale AI orchestration workflows and RAG pipelines to solve real-world business problems while optimizing for token costs and performance
  • Automate the resolution of recurring Tier-1 and Tier-2 incidents using Runbooks-as-Code to reduce manual toil and improve system reliability
  • Lead the fine-tuning and evaluation of open-source models to ensure high-quality data retrieval and minimize hallucination rates in production
  • Establish AI engineering best practices through comprehensive code reviews and mentorship of junior engineers transitioning to AI-native development
  • Implement full-stack observability across the customer order journey to diagnose and resolve complex provisioning issues before they impact the user experience

Requirements

  • Expert proficiency in Python and Java to develop production-grade backend systems and complex AI microservices
  • Deep technical expertise in Generative AI architectures, including Transformer models, prompt engineering, and the integration of Large Language Models
  • Critical experience architecting Retrieval-Augmented Generation (RAG) pipelines and managing unstructured data within vector databases like Milvus
  • Advanced knowledge of observability platforms such as Dynatrace and AWS CloudWatch to drive data-driven decision-making and system performance
  • Demonstrated ability to design and implement automation tools using Python, Ansible, and Bash to achieve autonomous infrastructure remediation
  • AI literacy and the capacity to innovate by applying emerging AI-Ops tools and predictive intelligence to traditional site reliability challenges, * Contributions to open-source AI projects or a documented portfolio of Agentic AI applications
  • Professional certifications in Cloud AI (e.g., AWS Certified Machine Learning - Specialty)
  • Experience building scalable applications leveraging IBM Watsonx.ai and/or AWS Bedrock, ensuring our AI solutions are governed, secure, and performant, * Minimum Education: Bachelor's Degree in Computer Science, Information Technology, or a relevant field (or 5+ years of equivalent software engineering experience)
  • Minimum Experience: 5+ years of experience in backend or full-stack software development, with at least 2 years focused on AI-powered development
  • Required Technical Skills:
  • Critical experience in Python-based API orchestration and custom tool development
  • Hands-on expertise with AWS cloud services, Serverless architectures, and Kubernetes
  • Proven track record in AI-driven software development with 2+ years of specialized experience in architecting and deploying autonomous AI agents and intelligent bot frameworks

Benefits & conditions

We offer versatile health perks, including flexible spending accounts, HSA, a 401(k) Plan with company match, ESPP, career opportunities, and a flexible time away plan; all benefits can be viewed here: EchoStar Benefits .

The base pay range shown is a guideline. Individual total compensation will vary based on factors such as qualifications, skill level, and competencies; compensation is based on the role's location and is subject to change based on work location.

Candidates need to successfully complete a pre-employment screen, which may include a drug test and DMV check. Our company is committed to fostering an inclusive and equitable workplace where every individual has the opportunity to succeed. We are dedicated to providing individuals with criminal or arrest records a fair chance of employment in accordance with local, state, and federal laws.

The posting will be active for a minimum of 3 days. The active posting will continue to extend by 3 days until the position is filled.

About the company

EchoStar is reimagining the future of connectivity. Our business reach spans satellite television service, live-streaming and on-demand programming, smart home installation services, mobile plans and products. Today, our brands include Boost Mobile, DISH TV, Gen Mobile, Hughes and Sling TV. Our Technology teams challenge the status quo and reimagine capabilities across industries. Whether through research and development, technology innovation or solution engineering, our team members play a vital role in connecting consumers with the products and platforms of tomorrow.

Apply for this position