> Markdown version of [/jobs/ext/1340389-sr-software-engineer-ai-infrastructure](https://www.wearedevelopers.com/jobs/ext/1340389-sr-software-engineer-ai-infrastructure). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Sr. Software Engineer, AI Infrastructure - **Company:** LinkedIn Corporation - **Location:** Sunnyvale, CA, United States - **Experience:** Expert - **Salary:** $139,000.0 - $229,000.0 - **Contract:** Permanent contract - **Skills:** Adobe Flash, Java (Programming Language), Artificial Intelligence, Algorithm Design, Big Data, C Sharp (Programming Language), C++ (Programming Language), Software Debugging, Distributed Computing Environment, Distributed Systems, Information Retrieval, Python (Programming Language), Machine Learning, Open Source Technology, Recommender Systems, Tensorflow, AI Infrastructure, Rust (Programming Language), Feature Engineering, Pytorch, Large Language Models, Apache Spark, Deep Learning, AI Platforms, Kubernetes, Information Technology, Low Latency, Apache Flink, HuggingFace, Machine Learning Operations, Functional Programming, Golang, Programming Languages - **Published:** July 18, 2026 - **Apply:** https://www.juju.com/job/00000000ghgm7a ## About the Role + Bachelor's Degree in Computer Science or related technical discipline, or equivalent practical experience + 2+ years of experience in the industry with leading/ building deep learning systems. + 2+ years of experience with Java, C++, Python, Go, Rust, C# and/or Functional languages such as Scala or other relevant coding languages + Hands-on experience developing distributed systems or other large-scale systems. Preferred Qualifications + BS and 5+ years of relevant work experience, MS and 4+ years of relevant work experience, or PhD and 2+ years of relevant work experience + Previous experience working with geographically distributed co-workers. + Outstanding interpersonal communication skills (including listening, speaking, and writing) and ability to work well in a diverse, team-focused environment with other SRE/SWE Engineers, ---Project Managers, etc. + Experience building ML applications, LLM serving, GPU serving. + Experience with distributed data processing engines like Flink, Beam, Spark etc., feature engineering, + Experience with search systems or similar large-scale distributed systems + Expertise in machine learning infrastructure, including technologies like MLFlow, Kubeflow and large scale distributed systems + Co-author or maintainer of any open-source projects + Familiarity with containers and container orchestration systems + Expertise in deep learning frameworks and tensor libraries like PyTorch, Tensorflow, JAX/FLAX Suggested Skills + ML Algorithm Development + Experience in Machine Learning and Deep Learning + Experience in Information retrieval / recommendation systems / distributed serving / Big Data is a plus. ## Description Model Training Infrastructure: As an engineer on the AI Training Infra team, you will play a crucial role in building the next-gen training infrastructure to power AI use cases. You will design and implement high performance data I/O, work with open source teams to identify and resolve issues in popular libraries like Huggingface, Horovod and PyTorch, enable distributed training over 100s of billions of parameter models, debug and optimize deep learning training, and provide advanced support for internal AI teams in areas like model parallelism, tensor parallelism, Zero++ etc. Finally, you will assist in and guide the development of containerized pipeline orchestration infrastructure, including developing and distributing stable base container images, providing advanced profiling and observability, and updating internally maintained versions of deep learning frameworks and their companion libraries like Tensorflow, PyTorch, DeepSpeed, GNNs, Flash Attention. PyTorch Lightning and more and more. Model Serving Infrastructure: this team builds low latency high performance applications serving very large & complex models across LLM and Personalization models. As an engineer, you will build compute efficient infra on top of native cloud, enable GPU based inference for a large variety of use cases, cuda level optimizations for high performance, enable on-device and online training. Challenges include scale (10s of thousands of QPS, multiple terabytes of data, billions of model parameters), agility (experiment with hundreds of new ML models per quarter using thousands of features), and enabling GPU inference at scale. As a Sr. Software Engineer, you will have first-hand opportunities to advance one of the most scalable AI platforms in the world. At the same time, you will work together with our talented teams of researchers and engineers to build your career and your personal brand in the AI industry. Responsibilities: + Owning the technical strategy for broad or complex requirements with insightful and forward-looking approaches that go beyond the direct team and solve large open-ended problems. + Designing, implementing, and optimizing the performance of large-scale distributed serving or training for personalized recommendation as well as large language models. + Improving the observability and understandability of various systems with a focus on improving developer productivity and system sustenance. + Mentoring other engineers, defining our challenging technical culture, and helping to build a fast-growing team. + Working closely with the open-source community to participate and influence cutting edge open-source projects (e.g., vLLMs, PyTorch, GNNs, DeepSpeed, Huggingface, etc.). + Functioning as the tech-lead for several concurrent key initiatives AI Infrastructure and defining the future of AI Platforms., A request for an accommodation will be responded to within three business days. However, non-disability related requests, such as following up on an application, will not receive a response. LinkedIn will not discharge or in any other manner discriminate against employees or applicants because they have inquired about, discussed, or disclosed their own pay or the pay of another employee or applicant. However, employees who have access to the compensation information of other employees or applicants as a part of their essential job functions cannot disclose the pay of other employees or applicants to individuals who do not otherwise have access to compensation information, unless the disclosure is (a) in response to a formal complaint or charge, (b) in furtherance of an investigation, proceeding, hearing, or action, including an investigation conducted by LinkedIn, or (c) consistent with LinkedIn's legal duty to furnish information. San Francisco Fair Chance Ordinance Pursuant to the San Francisco Fair Chance Ordinance, LinkedIn will consider for employment qualified applicants with arrest and conviction records. Pay Transparency Policy Statement As a federal contractor, LinkedIn follows the Pay Transparency and non-discrimination provisions described at this link: https://lnkd.in/paytransparency. ## Related Videos - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [Developing an AI.SDK](https://www.wearedevelopers.com/videos/198-developing-an-ai-sdk) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [What Industries Outside of AI Are Hiring The Most AI Experts?](https://www.wearedevelopers.com/magazine/98-what-industries-outside-of-ai-are-hiring-the-most-ai-experts) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift)