> Markdown version of [/jobs/ext/2251053-machine-learning-engineer-golang](https://www.wearedevelopers.com/jobs/ext/2251053-machine-learning-engineer-golang). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Machine Learning Engineer (GoLang) - **Company:** Xfinity - **Location:** Washington, DC, United States (Remote available) - **Experience:** Expert - **Salary:** $142,651.0 - $213,977.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Amazon Web Services, Automation of Tests, Cloud Computing, Continuous Delivery, Data Security, Distributed Computing Environment, Identity and Access Management, Key Management, Machine Learning, Named Entity Recognition, Object Detection, Cloud Services, Azure Machine Learning, Search Technologies, Amazon Simple Notification Service (SNS), Software Engineering, Video Editing, Data Logging, Cloud Platform System, Large Language Models, Multi-Agent Systems, Indexer, Backend, Event Driven Architecture, AI Platforms, Kubernetes, Infrastructure Automation Frameworks, Low Latency, Apache Kafka, Machine Learning Operations, Amazon Simple Queue Service (SQS), Terraform, Stream Processing, Network Server, Data Pipelines, Golang, Programming Languages - **Published:** August 26, 2026 - **Apply:** https://www.careerjet.com/job/us3483e4449430f069fdeed1090841550c/eaa ## About the Role * **3-6 years** of professional software engineering experience. * Strong backend engineering experience with **Golang**. * Experience building and operating **APIs** (REST and/or gRPC) in production. * Hands-on experience with **Kubernetes** in production environments. * Experience using **Terraform** for infrastructure provisioning and deployment. * Solid working knowledge of **AWS** cloud services and core architectural concepts. * Experience building or supporting **ML processing pipelines** (video, image, or document). * Practical experience using **LLMs** in production systems. * Experience developing **agents** and/or **MCP servers**, or equivalent tool-integration platforms. Preferred / Nice-to-Have Qualifications * Experience with **Milvus** or other vector databases in production. * Familiarity with GPU-backed workloads and ML inference optimization. * Experience with messaging/streaming systems (Kafka, SQS, SNS, etc.). * Knowledge of secure system design for AI platforms (IAM, secrets management, least-privilege access). * Experience working on internal developer platforms or ML infrastructure teams., Kubernetes; Cloud Platform; Collaboration; Large Language Models (LLMs); Model Context Protocol; Go Programming Language, Bachelor's Degree While possessing the stated degree is preferred, Comcast also may consider applicants who hold some combination of coursework and experience, or who have extensive related professional experience. Relevant Work Experience 5-7 Years ## Description Job Summary Multimodal Analysis Framework (MAF)** is an end-to-end platform designed to process diverse content sources-including **video, images, audio, and documents**-to generate rich, structured metadata. The platform unifies multiple ML/AI models to extract curated insights at scale, tailored to specific business needs. MAF supports both **on-demand** workloads (batch uploads, ad-hoc analysis) and **real-time streaming** workflows, enabling continuous metadata generation for live content streams. Customers can define their metadata requirements-such as entity extraction, scene segmentation, object detection, transcription, summarization, or multimodal correlation-and the framework orchestrates the appropriate models and toolchains to deliver high-quality outputs. Through flexible APIs and UI-based workflows, customers and internal teams can visualize metadata, trigger enrichment, monitor processing, and integrate results into downstream applications. The platform emphasizes modularity, scalability, and extensibility to support new ML models, LLM-based agents, and cross-modal inference as use cases evolve. We are looking for a **mid-level Backend Engineer** to join our **Machine Learning Platform team**. This role focuses on building **scalable backend systems** that power ML workloads, including **video, image, and document processing**, and enable **LLM-driven applications** through **agents and MCP servers**. You will work primarily in **Golang**, deploy and operate services on **Kubernetes**, manage infrastructure with **Terraform**, and build on **AWS**. A core part of the role is designing platform capabilities that allow **LLMs to safely and reliably interact with tools, data, and services** via **agent frameworks and MCP servers**., Backend Engineering (Golang) * Design, build, and maintain **high-performance backend services** in **Golang** for ML and AI platform use cases. * Develop **REST and gRPC APIs** for inference, processing pipelines, orchestration, and platform services. * Implement asynchronous and distributed processing patterns (workers, queues, event-driven systems). * Ensure backend services meet production standards for **scalability, reliability, and security**. ML Platform & Processing Pipelines * Build and operate backend systems supporting: * Video processing** (frame extraction, metadata generation, embeddings, indexing). * Image processing** (OCR, classification, detection, embedding generation). * Document processing** (parsing, layout analysis, chunking, OCR, retrieval pipelines). * Integrate ML inference services into backend workflows with attention to **latency, throughput, and cost**. * Work closely with ML engineers and data scientists to productionize models and pipelines. LLMs, Agents, and MCP Servers * Build **LLM-enabled backend services** using structured prompting, tool/function calling, and retrieval-augmented generation (RAG). * Design and implement **agentic workflows** (multi-step reasoning, tool orchestration, retries, guardrails). * Develop and operate **MCP servers** that expose internal platform capabilities (search, retrieval, processing, data access) to LLM-based applications. * Enforce **security, access control, and observability** for agent and MCP interactions. Vector Search & Retrieval * Design and maintain vector-based retrieval systems using **Milvus**. * Implement embedding ingestion, indexing, and query pipelines at scale. * Optimize retrieval quality, latency, and relevance for downstream LLM applications. Cloud, Kubernetes & Infrastructure * Deploy and operate backend and ML services on **Kubernetes** (scaling, rollouts, resource management). * Use **Terraform** for infrastructure provisioning and continuous delivery of cloud resources. * Build and operate primarily on **AWS**, leveraging services such as: * Compute, networking, and IAM * Object storage * Managed Kubernetes * Logging and monitoring services Reliability, Quality & Operations * Implement observability using logs, metrics, and traces; define SLOs and alerts. * Write automated tests (unit, integration) and contribute to CI/CD pipelines. * Participate in on-call rotations and incident response; drive post-incident improvements., Distinguished AI Engineer (Remote Eligible) At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an indus… + 2 days ago, Lead Machine Learning Engineer (Manager IC) At Capital One, we are changing banking for good by creating responsible and reliable AI-powered systems. Our investments in technology … + 7 days ago + ## Related Videos - [Machine Learning for Software Developers (and Knitters)](https://www.wearedevelopers.com/videos/154-machine-learning-for-software-developers-and-knitters) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Developing the Backend with Stefan Lingler, CTO at Shpock](https://www.wearedevelopers.com/videos/100360-developing-the-backend-with-stefan-lingler-cto-at-shpock) - [Optimizing Discovery: PostgreSQL's Role in Transforming GetYourGuide's Search](https://www.wearedevelopers.com/videos/1647-optimizing-discovery-postgresql-s-role-in-transforming-getyourguide-s-search) - [The Road to MLOps: How Verivox Transitioned to AWS](https://www.wearedevelopers.com/videos/1050-the-road-to-mlops-how-verivox-transitioned-to-aws) - [Scoring 2000 Products per Request: Performance Pitfalls in Golang](https://www.wearedevelopers.com/videos/2073-scoring-2000-products-per-request-performance-pitfalls-in-golang) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development)