> Markdown version of [/jobs/ext/2383206-senior-manager-technology-consulting](https://www.wearedevelopers.com/jobs/ext/2383206-senior-manager-technology-consulting). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Manager, Technology Consulting - **Company:** EPAM Systems, Inc. - **Location:** Philadelphia, PA, United States (Remote available) - **Experience:** Expert - **Salary:** $180,000.0 - $220,000.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, C++ (Programming Language), Code Generation, Nvidia CUDA, Information Technology Consulting, Extract Transform Load (ETL), Python (Programming Language), Machine Learning, Performance Tuning, Software Architecture, Regression Testing, Tensorflow, Software Engineering, Toolchain, Graphics Processing Unit (GPU), Pytorch, Free and Open-Source Software - **Published:** August 16, 2026 - **Apply:** https://www.phillyjobs.com/job.asp?id=3355970492&tx=YT1212THT&pt=1&aff=0B19D771-A501-4A5E-8338-2A822B784D54&utm_source=Job%20Feed&utm_medium=textkernel&utm_campaign=DE&utm_term=0B19D771-A501-4A5E-8338-2A822B784D54 ## About the Role * Bachelor's degree or equivalent practical experience * Overall 12+ years of industry experience; 5 years of experience with software development in C++ or Python * 3 years of experience testing, maintaining, or launching software products, and 1 year of experience with software design and architecture * Experience with performance optimization at the kernel level * Experience optimizing TPU/GPU code, using low-level kernel languages like Pallas, Compute Unified Device Architecture (CUDA), or Triton * Knowledge of ML Frameworks (JAX/PyTorch), common operations like attention and Mixture of Experts (MoEs), including model optimization and low-precision formats * Understanding of modern accelerators (e.g., data movement, pipelining, heterogeneous compute, and scale-out) * Understanding of compiler principles (optimization, code generation) and toolchains such as MLIR, OpenXLA * Demonstrate a record of building developer infrastructure, including Open-Source Software (OSS) libraries, flexible high-performance APIs, and easy-to-consume documentation to empower the community * Excellent investigative and problem-solving capabilities with communication skills across cross-functional teams ## Description * Design and optimize high-performance kernels (using languages like Pallas, Mosaic, and Triton) targeting Tensor Processing Unit (TPU) and Graphics Processing Unit (GPU) architectures for critical Machine Learning (ML) operations, redefining what's possible from massive training runs to high-speed inference * Architect infrastructure such as benchmarking suites, autotuning frameworks, performance analysis tools, regression testing, and documentation, transforming how the developer community interacts with increasingly critical custom kernels in key Open-Source Software (OSS) libraries * Track the latest advancements in hardware architectures, compiler technologies, and AI models to identify new opportunities for performance optimization through custom kernels * Engage with ML researchers, framework developers (Just After eXecution (JAX), PyTorch), and compiler engineers (Accelerated Linear Algebra (XLA)) to enhance adoption, identify new requirements, and address bottlenecks by providing appropriate solutions ## Related Videos - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Speeding up Web Apps performance with WebAssembly and Emscripten](https://www.wearedevelopers.com/videos/1985-speeding-up-web-apps-performance-with-webassembly-and-emscripten) - [Machine learning in the browser with TensorFlowjs](https://www.wearedevelopers.com/videos/155-machine-learning-in-the-browser-with-tensorflowjs) - [The weekly developer show: Boosting Python with CUDA, CSS Updates & Navigating New Tech Stacks](https://www.wearedevelopers.com/videos/1293-the-weekly-developer-show-boosting-python-with-cuda-css-updates-navigating-new-tech-stacks) - [From Model to Metal: An Open Source Stack for Accelerating Intelligence](https://www.wearedevelopers.com/videos/1636-from-model-to-metal-an-open-source-stack-for-accelerating-intelligence) - [Serverless deployment of (large) NLP models ](https://www.wearedevelopers.com/videos/158-serverless-deployment-of-large-nlp-models) ## Related Articles - [The Best X (Twitter) Accounts for Developers](https://www.wearedevelopers.com/magazine/294-the-best-x-twitter-accounts-for-developers) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers)