Principal Architect/Lead- Observability

Nokia
Sunnyvale, CA, United States
6 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours

Tech stack

Cloud Computing Code Review Databases Extract Transform Load (ETL) Distributed Systems Design of User Interfaces Monitoring of Systems Online Analytical Processing Prometheus Data Streaming AI Infrastructure Datadog
+9 more
Data Processing Grafana Apache Spark Data Layers Kubernetes Information Technology Apache Flink Apache Kafka Codebase

Job description

We are seeking an experienced and visionary Principal Architect/Lead to take ownership of our observability practices and instrumentation standards. This role is crucial in ensuring we have a robust and efficient observability stack, enabling our engineering teams to monitor and optimize our distributed AI infrastructure effectively. The successful candidate will work closely with various teams, including Platform, SRE, UI, and UX, to establish best practices and drive the adoption of industry-leading observability tools and methodologies.

  • Own and manage the observability stack, including tooling and instrumentation standards, across the entire codebase.
  • Work closely with SRE, engineering, and platform teams to embed observability practices from day one of a project’s lifecycle.
  • Ensure proper instrumentation of services and push for high-quality observability data collection.
  • Stay up-to-date with the latest observability tooling and technologies, such as Prometheus, OpenTelemetry, Grafana, Tempo, and Victoria Metrics.
  • Collaborate with UI/UX designers to create intuitive and informative observability dashboards and visualizations.
  • Develop and maintain OLAP databases, Kafka, ETL pipelines, and streaming platforms to support observability data processing and analysis.
  • Explore and implement graph-databases, ontologies, and semantic layers to enhance observability and knowledge-base capabilities.
  • Provide expertise and guidance to engineering teams on best practices for distributed AI infrastructure observability.
  • Actively participate in code reviews and provide feedback to ensure proper instrumentation and observability practices are followed.

Requirements

  • Expert knowledge of observability tooling, including Prometheus, OpenTelemetry, Grafana, Tempo, and Victoria Metrics.
  • Experience with OLAP databases, Kafka, ETL pipelines, and streaming platforms (Spark, Flink) is essential.
  • Must be hands-on in designing and implementing code; Experience with Kubernetes; Experience with cloud technologies
  • Familiarity with graph-databases, ontologies, semantic layers, and knowledgebases is highly advantageous.
  • Understanding of distributed systems and AI infrastructure at scale.
  • Ability to work closely with engineering teams and drive adoption of observability practices.
  • Excellent communication and collaboration skills to work effectively with cross-functional teams.
  • Experience in leading and mentoring a team of observability experts is desirable.
  • A proven track record of implementing and optimizing observability stacks in large-scale projects.
  • Strong problem-solving and analytical skills, with the ability to identify and resolve complex issues.
  • A passion for staying updated with the latest advancements in observability and monitoring technologies.
  • Master’s degree in Computer Science, Engineering, or a related field; PhD preferred.

Advancing connectivity to secure a brighter world.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on diversityjobs.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

1:36 min

Visualizing memory limits and isolating suspicious endpoints

Dina Matveev Dina Matveev · Europe 2026 Virtual

3:34 min

Augmenting codebases with semantic architectures

Zaak Chalal Zaak Chalal · WWC Europe 2026

2:48 min

Daily responsibilities and alignment practices for technical engineering leadership

Edoardo Dusi · LIVE

12:33 min

Exploring advanced observability stacks and distributed infrastructure challenges

Pawel Piwosz · LIVE

1:16 min

Combining saas and open-source observability for extreme scale

Mikael Robert Mikael Robert · WWC 2025

Videos

See all

Related articles

See all