> Markdown version of [/jobs/ext/1958732-principal-architect-lead-observability](https://www.wearedevelopers.com/jobs/ext/1958732-principal-architect-lead-observability). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Principal Architect/Lead- Observability - **Company:** Nokia - **Location:** Sunnyvale, CA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Cloud Computing, Code Review, Databases, Extract Transform Load (ETL), Distributed Systems, Design of User Interfaces, Monitoring of Systems, Online Analytical Processing, Prometheus, Data Streaming, AI Infrastructure, Datadog, Data Processing, Grafana, Apache Spark, Data Layers, Kubernetes, Information Technology, Apache Flink, Apache Kafka, Codebase - **Published:** August 6, 2026 - **Apply:** https://diversityjobs.com/main/sendform/8/8/28176/1/17836692?backUrl=%2Fcareer%2F17836692%2FPrincipal-Architect-Lead-Observability-California-Sunnyvale ## About the Role * Expert knowledge of observability tooling, including Prometheus, OpenTelemetry, Grafana, Tempo, and Victoria Metrics. * Experience with OLAP databases, Kafka, ETL pipelines, and streaming platforms (Spark, Flink) is essential. * Must be hands-on in designing and implementing code; Experience with Kubernetes; Experience with cloud technologies * Familiarity with graph-databases, ontologies, semantic layers, and knowledgebases is highly advantageous. * Understanding of distributed systems and AI infrastructure at scale. * Ability to work closely with engineering teams and drive adoption of observability practices. * Excellent communication and collaboration skills to work effectively with cross-functional teams. * Experience in leading and mentoring a team of observability experts is desirable. * A proven track record of implementing and optimizing observability stacks in large-scale projects. * Strong problem-solving and analytical skills, with the ability to identify and resolve complex issues. * A passion for staying updated with the latest advancements in observability and monitoring technologies. * Master's degree in Computer Science, Engineering, or a related field; PhD preferred. Advancing connectivity to secure a brighter world. ## Description We are seeking an experienced and visionary Principal Architect/Lead to take ownership of our observability practices and instrumentation standards. This role is crucial in ensuring we have a robust and efficient observability stack, enabling our engineering teams to monitor and optimize our distributed AI infrastructure effectively. The successful candidate will work closely with various teams, including Platform, SRE, UI, and UX, to establish best practices and drive the adoption of industry-leading observability tools and methodologies. * Own and manage the observability stack, including tooling and instrumentation standards, across the entire codebase. * Work closely with SRE, engineering, and platform teams to embed observability practices from day one of a project's lifecycle. * Ensure proper instrumentation of services and push for high-quality observability data collection. * Stay up-to-date with the latest observability tooling and technologies, such as Prometheus, OpenTelemetry, Grafana, Tempo, and Victoria Metrics. * Collaborate with UI/UX designers to create intuitive and informative observability dashboards and visualizations. * Develop and maintain OLAP databases, Kafka, ETL pipelines, and streaming platforms to support observability data processing and analysis. * Explore and implement graph-databases, ontologies, and semantic layers to enhance observability and knowledge-base capabilities. * Provide expertise and guidance to engineering teams on best practices for distributed AI infrastructure observability. * Actively participate in code reviews and provide feedback to ensure proper instrumentation and observability practices are followed. ## Related Videos - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) - [Getting to Know Your Legacy (System) with AI-Driven Software Archeology](https://www.wearedevelopers.com/videos/1437-getting-to-know-your-legacy-system-with-ai-driven-software-archeology) - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Software Engineering Social Connection: Yubo’s lean approach to scaling an 80M-user infrastructure](https://www.wearedevelopers.com/videos/1583-software-engineering-social-connection-yubo-s-lean-approach-to-scaling-an-80m-user-infrastructure) - [Keycloak case study: Making users happy with service level indicators and observability](https://www.wearedevelopers.com/videos/1599-keycloak-case-study-making-users-happy-with-service-level-indicators-and-observability) ## Related Articles - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Top-Paying Tech Jobs (with Salaries)](https://www.wearedevelopers.com/magazine/372-top-paying-tech-jobs-with-salaries)