> Markdown version of [/jobs/ext/2976600-telecommute-principal-machine-learning-systems-engineer](https://www.wearedevelopers.com/jobs/ext/2976600-telecommute-principal-machine-learning-systems-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # TELECOMMUTE Principal Machine Learning Systems Engineer - **Company:** Atlassian - **Location:** Washington, DC, United States (Remote available) - **Experience:** Expert - **Salary:** $236,700.0 - $309,025.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Automated Storage and Retrieval Systems, Microsoft Azure, Program Optimization, Distributed Systems, Python (Programming Language), Machine Learning, Open Source Technology, Rapid Prototyping Process, Tensorflow, Google Cloud, Cloud Platform System, High Performance Computing, Pytorch, Delivery Pipeline, Large Language Models, Backend, Kubernetes, Information Technology, Low Latency, HuggingFace, Atlassian Tools, Machine Learning Operations, Software Coding, Docker - **Published:** September 18, 2026 - **Apply:** https://www.dice.com/job-detail/42128755-5e2a-4a81-9ddb-c3859e8102e6 ## About the Role * 6+ years in ML systems engineering, backend engineering, or infrastructure roles. * Strong track record of building and scaling ML-powered services in production. * Experience with large-scale model training, inference pipelines, or search/retrieval systems. Skills * Proficiency in backend systems and ML frameworks (Python, PyTorch, TensorFlow, Hugging Face). * Experience with vector databases (Weaviate, Pinecone, FAISS), orchestration frameworks (LangChain, LlamaIndex). * Strong coding skills and ability to optimize systems for performance and reliability. * Familiarity with cloud environments (AWS, Google Cloud Platform, Azure) and container/orchestration tools (Kubernetes, Docker). Education * Bachelor's or Master's in Computer Science, Machine Learning, or related field-or equivalent industry experience. Nice to Have * Background in distributed systems, high-performance computing, or GPU optimization. * Familiarity with search/GenAI evaluation metrics (e.g., NDCG, groundedness, latency benchmarks). * Experience with monitoring, observability, and reliability practices for ML systems. * Contributions to open-source infra or ML systems frameworks. ## Description * Architect and implement scalable systems for training, fine-tuning, and serving large language models and embeddings. * Build efficient retrieval, hybrid search, and RAG pipelines integrated with knowledge-grounded data. * Develop tools and infra to support rapid experimentation, evaluation, and deployment of prototypes. Enable Rapid Prototyping & Applied Research * Partner with applied scientists to bring new ideas to life in robust, production-ready pipelines. * Build proof-of-concept (POC) systems and evolve them into reliable, scalable services. * Optimize latency, throughput, and resource efficiency for GenAI workloads. Collaborate Across Disciplines * Work closely with ML engineers, backend developers, and product teams to ship end-to-end innovations. * Contribute to best practices in model deployment, monitoring, and evaluation. * Help establish the team as a world-class hub for GenAI systems innovation., In line with local law, identity verification (which may include use of biometric data) is a condition of employment with Atlassian for employment fraud purposes. ## Related Videos - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Developing the Backend with Stefan Lingler, CTO at Shpock](https://www.wearedevelopers.com/videos/100360-developing-the-backend-with-stefan-lingler-cto-at-shpock) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Introduction to Azure Machine Learning](https://www.wearedevelopers.com/videos/368-introduction-to-azure-machine-learning) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production)