> Markdown version of [/jobs/ext/1219168-senior-software-engineer-observability](https://www.wearedevelopers.com/jobs/ext/1219168-senior-software-engineer-observability). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Software Engineer, Observability - **Company:** Together Ai - **Location:** San Francisco, CA, United States (Remote available) - **Experience:** Expert - **Salary:** $200,000.0 - $280,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Microsoft Azure, Computer Clusters, Computer Programming, Databases, Distributed Systems, HP Systems Insight Manager, Python (Programming Language), PostgreSQL, MongoDB, Open Source Technology, Redis, Reliability Engineering, Software Reliability Testing, Ansible, Prometheus, Google Cloud, Data Ingestion, Istio, Grafana, Model Validation, Build Management, Kubernetes, Infrastructure Automation Frameworks, Machine Learning Operations, Vertica, Terraform, Dynatrace, Docker, Golang, Microservices - **Published:** July 9, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=7798608e6c8e6d0a ## About the Role * Expertise in observability platforms (Prometheus, Grafana, ClickStack, OpenTelemetry) and cloud-native monitoring services (AWS, GCP, Azure). * Strong programming skills in Go, Python, or similar languages, with proficiency in infrastructure-as-code tools (Terraform, Ansible, Helm). * Experience designing, operating, and scaling large-scale distributed systems and pipelines for high-volume data ingestion and real-time querying. * Deep understanding of containerization (Docker) and orchestration (Kubernetes). * Knowledge of microservices architecture, service mesh technologies, CI/CD pipelines, and GitOps workflows. * Expertise in managing databases (PostgreSQL, MongoDB, Redis) and time-series databases with high-cardinality data. Preferred * Experience monitoring AI/ML infrastructure, GPU clusters, and custom metrics for model performance and training pipelines. * Background in high-frequency, low-latency systems monitoring, chaos engineering, and reliability testing. * Contributions to open-source observability projects. * Familiarity with security monitoring and compliance frameworks. ## Description * Design and implement a scalable observability platform (metrics, logs, traces) using tools like Prometheus, Grafana, ClickHouse, ClickStack, and OpenTelemetry, including telemetry data pipelines and log aggregation workflows. * Develop automated monitoring, alerting, and anomaly detection systems, including SLIs/SLOs, runbooks, and predictive analytics for critical services. * Build and deploy custom observability tools and infrastructure-as-code using Go, Python, Terraform, Ansible, and Helm. * Collaborate with engineering teams to enhance distributed tracing and application monitoring, and lead incident response with post-mortem analysis. * Define observability best practices. ## Related Videos - [Unlocking the AI Black Box: Building Trust in the Era of Agentic Production](https://www.wearedevelopers.com/videos/100086-unlocking-the-ai-black-box-building-trust-in-the-era-of-agentic-production) - [Rate-limiting using eBPF and Istio: How to protect your SaaS customers from themselves](https://www.wearedevelopers.com/videos/100220-rate-limiting-using-ebpf-and-istio-how-to-protect-your-saas-customers-from-themselves) - [Reducing LLM Calls with Vector Search Patterns - Raphael De Lio (Redis)](https://www.wearedevelopers.com/videos/1714-reducing-llm-calls-with-vector-search-patterns-raphael-de-lio-redis) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Mastering AI-Driven Problem Solving in Engineering with Observability](https://www.wearedevelopers.com/videos/994-mastering-ai-driven-problem-solving-in-engineering-with-observability) - [Get started with securing your cloud-native Java microservices applications](https://www.wearedevelopers.com/videos/123-get-started-with-securing-your-cloud-native-java-microservices-applications) ## Related Articles - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [What is Software Engineering in the Age of AI?](https://www.wearedevelopers.com/magazine/640-what-is-software-engineering-in-the-age-of-ai)