> Markdown version of [/jobs/ext/2709164-sre](https://www.wearedevelopers.com/jobs/ext/2709164-sre). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # SRE - **Company:** Lovelace Ai - **Location:** Pittsburgh, PA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Amazon Elastic Compute Cloud, Amazon S3, Microsoft Azure, Bash Shell, Continuous Integration, Distributed Systems, Domain Name System (DNS), Monitoring of Systems, Hypertext Transfer Protocols (HTTP), Python (Programming Language), Networking Basics, Performance Tuning, Reliability Engineering, Ansible, Prometheus, Software Engineering, TCP/IP, Software Vulnerability Management, Datadog, Circleci, Scripting, Google Cloud, Load Balancing, Grafana, Reliability of Systems, Cloudformation, Gitlab-ci, Kubernetes, Infrastructure Automation Frameworks, Build Tools, Functional Programming, Terraform, Dynatrace, Docker, Elk Stack, Jenkins, Golang, Microservices - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/software-engineer-site-reliability-engineer-sre-lovelace-ai-8290473 ## About the Role * 5+ years of experience in site reliability engineering, DevOps, systems administration, or related roles. * Proven track record of managing complex infrastructure, troubleshooting production issues, and optimizing system performance in high-scale environments. * Strong experience with Linux/Unix administration and proficiency in scripting languages (e.g., Python, Bash, Go). * Deep understanding of cloud platforms (AWS, GCP, Azure) and related services (e.g., EC2, S3, Lambda, Kubernetes). * Experience with containerization and orchestration technologies like Docker and Kubernetes. * Proficiency with monitoring and observability tools (e.g., Prometheus, Grafana, Datadog, Dynatrace, ELK Stack). * Strong understanding of networking fundamentals (DNS, HTTP, TCP/IP), load balancing, and CDNs. * Experience with CI/CD tools (e.g., Jenkins, GitLab CI, CircleCI) and infrastructure automation. * Familiarity with distributed systems and microservices architecture. * Excellent problem-solving and troubleshooting skills. * Strong analytical skills with the ability to identify Service Level Indicators (SLIs) and align efforts to meet availability and latency objectives. * Ability to balance both development and support roles effectively. * Strong interpersonal skills and excellent communication skills, with the ability to collaborate effectively across various teams. * Experience in working on projects that involve business segments. * Must be a US Citizen. ## Description * Lovelace AI is seeking a highly skilled and motivated Site Reliability Engineer (SRE) to join our growing team. As an SRE at Lovelace AI, you will play a critical role in ensuring the availability, scalability, and performance of our cutting-edge AI-powered applications and infrastructure. You will bridge the gap between software development and operations, applying sound engineering principles and automation to maintain and improve our systems., * Design, implement, and maintain robust monitoring, alerting, and observability solutions to proactively detect and resolve issues before they impact end-users. * Lead troubleshooting efforts for complex production issues, providing detailed root cause analysis (RCA) and implementing preventative measures. * Develop and maintain automation scripts, build systems (Bazel) and infrastructure as code (IaC) using tools like Terraform, Ansible, or CloudFormation to eliminate manual tasks and improve system reliability and efficiency. * Collaborate closely with software engineering teams to influence the design of new services and applications, ensuring they are scalable, reliable, and resilient from the outset. * Participate in on-call rotations to respond to platform emergencies, alerts, and escalations, ensuring high service uptime. * Analyze system performance and recommend optimizations for scalability, reliability, and efficiency. * Implement and enforce best practices in deployment, monitoring, and incident management to continuously improve overall system reliability and reduce downtime. * Develop and maintain internal tools that streamline complex operations, track bugs, manage CI/CD pipelines, and facilitate cross-team communication. * Conduct post-incident reviews, documenting software problems and solutions in a shared knowledge base to prevent similar issues in the future. * Assist with vulnerability management, system patching, and implementing security measures to protect the integrity and availability of services. ## Related Videos - [An Applied Introduction to eBPF with Go](https://www.wearedevelopers.com/videos/1075-an-applied-introduction-to-ebpf-with-go) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Turning Container security up to 11 with Capabilities](https://www.wearedevelopers.com/videos/718-turning-container-security-up-to-11-with-capabilities) - [Retooling and refactoring - an investment in people.](https://www.wearedevelopers.com/videos/371-retooling-and-refactoring-an-investment-in-people) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers)