AI Infra Advisory Researcher

Lenovo
Morrisville, NC, United States
1 day ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
7 years minimum
Working hours
Regular working hours

Tech stack

Artificial Intelligence Algorithm Design Application Frameworks Cloud Computing Program Optimization Nvidia CUDA Computer Engineering Data Infrastructure Extract Transform Load (ETL) Data Transformation Decision-Making Software Microprocessors
+31 more
Distributed Computing Environment Distributed Systems Fault Tolerance Graph Database Machine Learning Systems Development Life Cycle Reliability Engineering Tensorflow Signal Processing Software Deployment Software Engineering Product Software Implementation Methods Data Streaming Systems Architecture System Software AI Infrastructure Graphics Processing Unit (GPU) Data Ingestion Pytorch Apache Spark Deep Learning AI Platforms Kubernetes Information Technology Apache Flink Apache Kafka Free and Open-Source Software Data Management Machine Learning Operations Data Lakehouse Data Pipelines

Job description

The Advisory Researcher in AI Compute and Data Infrastructure will provide senior technical leadership for the research, architecture, and development of intelligent, high-performance, and resilient Hybrid AI systems. This position combines research depth, hands-on software development, system architecture expertise, and the ability to translate emerging technologies into production infrastructure and differentiated product capabilities., * Define technical directions and lead major research and development initiatives in AI compute and data infrastructure, distributed AI systems, and intelligent infrastructure management.

  • Identify high-impact technical opportunities based on infrastructure challenges, emerging technologies, product requirements, and business value.
  • Architect end-to-end AI infrastructure solutions spanning hardware, system software, data platforms, distributed training and inference, and application workloads.
  • Lead hardware/software co-analysis and co-optimization across GPUs, accelerators, CPUs, memory hierarchy, storage, networking, runtimes, frameworks, and AI applications.
  • Define optimization strategies for GPU utilization, workload placement, resource orchestration, memory and cache efficiency, communication, data movement, storage access, and model execution.
  • Lead the architecture and optimization of large-scale data pipelines for data ingestion, preprocessing, transformation, storage, retrieval, and delivery to AI workloads.
  • Define intelligent observability and diagnostic technologies for anomaly detection, root-cause analysis, performance regression, capacity planning, system health assessment, and predictive maintenance.
  • Develop resilient and fault-tolerant infrastructure architectures, including failure isolation, checkpointing and recovery, redundancy, retry, failover, graceful degradation, and automated remediation.
  • Apply machine learning and deep learning to system modeling, infrastructure control, workload forecasting, resource optimization, failure prediction, and operational intelligence.
  • Apply time-series modeling and signal processing to telemetry analytics, event detection, change-point detection, capacity forecasting, and system health monitoring.
  • Apply causal inference and causal discovery to root-cause analysis, performance attribution, intervention evaluation, and automated decision-making.
  • Define knowledge graph architectures for modeling infrastructure topology, hardware/software dependencies, workloads, operational events, and failure relationships.
  • Make architecture-level trade-offs involving performance, scalability, reliability, availability, energy consumption, cost, security, and maintainability.
  • Lead technical design reviews, architecture reviews, performance investigations, and resolution of complex cross-layer system issues.
  • Provide hands-on technical guidance in algorithm design, software implementation, system optimization, experimental validation, and production deployment.
  • Establish reusable frameworks, engineering practices, evaluation methodologies, and technical standards for AI infrastructure development.
  • Partner with research, engineering, architecture, product, and business teams to transition research into Enterprise AI and Personal AI platforms and products.
  • Mentor Staff Researchers and engineers and contribute to patents, publications, technical standards, and differentiated intellectual property.

Requirements

This candidate MUST be a US citizen or US national; US permanent residents or candidates requiring sponsorship cannot be considered., * Master’s or PhD degree in computer science, computer engineering, artificial intelligence, electrical engineering, applied mathematics, or a related field, or equivalent practical experience.

  • Seven or more years of relevant experience in AI compute and data infrastructure, machine learning systems, distributed systems, data platforms, performance engineering, reliability engineering, or advanced software development.
  • Demonstrated leadership of technically complex research, architecture, or system development initiatives.
  • Strong hands-on programming and system architecture capabilities.
  • Deep expertise in at least three of the following areas:
  • Machine learning or deep learning
  • GPU or accelerator optimization
  • Distributed AI training and inference
  • Large-scale and streaming data processing
  • Hardware/software co-optimization
  • Time-series modeling or signal processing
  • Infrastructure diagnostics, reliability, and fault tolerance
  • Causal inference
  • Knowledge graphs or graph-based reasoning
  • Proven record of delivering advanced software, infrastructure, or research technologies into products or production environments.
  • Ability to resolve ambiguous, cross-layer technical problems and make informed architecture-level trade-offs.
  • Strong technical communication, mentorship, and cross-organizational influence., * Experience architecting enterprise-scale cloud, edge, on-premises, or hybrid AI infrastructure.
  • Experience with PyTorch, TensorFlow, JAX, CUDA, ROCm, Spark, Flink, Ray, Kafka, Kubernetes, or related technologies.
  • Experience with MLOps, observability platforms, distributed computing, data lakehouse architectures, or cloud-native systems.
  • Strong publication record, patent portfolio, open-source contributions, or demonstrated commercial product impact.

About the company

Why Work at Lenovo We are Lenovo. We do what we say. We own what we do. We WOW our customers. Lenovo is a US$83 billion revenue global technology powerhouse, ranked #153 in the Fortune Global 500, and serving millions of customers every day in 180 markets. Focused on a bold vision to deliver Smarter Technology for All, Lenovo has built on its success as the world’s largest PC company with a full-stack portfolio of AI-enabled, AI-ready, and AI-optimized devices (PCs, workstations, smartphones, tablets), infrastructure (server, storage, edge, high performance computing and software defined infrastructure), software, solutions, and services. Lenovo’s continued investment in world-changing innovation is building a more equitable, trustworthy, and smarter future for everyone, everywhere. Lenovo is listed on the Hong Kong stock exchange under Lenovo Group Limited (HKSE: 992) (ADR: LNVGY). This transformation together with Lenovo’s world-changing innovation is building a more inclusive, trustworthy, and smarter future for everyone, everywhere. To find out more visit www.lenovo.com, and read about the latest news via our StoryHub.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on diversityjobs.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · WWC 2023

1:39 min

Fundamentals of tensors and the TensorFlow library

Håkan Silfvernagel · LIVE

2:28 min

Understanding Kubernetes architecture and core cluster components

Marc Nimmerrichter · WWC 2022

2:37 min

Optimizing technical profiles for AI sourcing and recruitment

Mina Golesorkhi Mina Golesorkhi · WWC Europe 2026

1:06 min

Compiling PyTorch environments for advanced time forecasting

Christoph Lohrmann Christoph Lohrmann +1 · WWC Europe 2026

2:14 min

Exploring internal AI product initiatives and global engineering roles

Maria Apazoglou · Coffee With Developers

Videos

See all

Related articles

See all