> Markdown version of [/jobs/ext/2822203-devops-platform-engineer](https://www.wearedevelopers.com/jobs/ext/2822203-devops-platform-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Devops Platform Engineer - **Company:** Advanced Micro Devices, Inc. - **Location:** San Jose, CA, United States - **Experience:** Expert - **Salary:** $136,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Bash Shell, Cloud Computing, Computer Programming, Computer Engineering, Continuous Integration, Linux, DevOps, Distributed Systems, Monitoring of Systems, Identity and Access Management, Python (Programming Language), Key Management, NetApp Applications, Octopus Deploy, Reliability Engineering, Ansible, Prometheus, Software Engineering, TypeScript, Datadog, Data Logging, Pulumi, Scripting, Graphics Processing Unit (GPU), Cloud Platform System, Istio, System Availability, Grafana, Cloudformation, Containerization, Kubernetes, Information Technology, Data Analytics, Machine Learning Operations, Hardware Infrastructure, Api Gateway, Terraform, Splunk, Dynatrace, Docker, Golang - **Published:** September 10, 2026 - **Apply:** https://jobs.localjobnetwork.com/apply/add/88290553/1 ## About the Role * Experience designing internal developer platforms or large-scale shared infrastructure. * Experience with Prometheus, Python, Bash, Grafana, OpenTelemetry, Loki, Tempo, Elastic, Datadog, Splunk, or comparable observability systems. * Experience with IBM LSF and Netapp Storage * Experience with GitOps operating models. * Experience supporting AI/ML infrastructure, model serving, GPU scheduling, Kubernetes operators, or AI workload orchestration. * Knowledge of platform security practices, including identity and access management, secrets management, policy enforcement, image security, and supply-chain security. * Experience with service mesh, API gateways, ingress, networking, or multi-cluster Kubernetes architectures. * Demonstrated ownership of production systems and a bias toward automation, operational excellence, and continuous improvement. * Semiconductor experience very helpful, EDA experience a plus., * Bachelor's degree in Computer Science, Computer Engineering, or a related technical field, or equivalent practical experience. * 6+ years of experience in software engineering, DevOps, SRE, cloud infrastructure, or platform engineering. * Strong hands-on experience operating and automating Kubernetes environments. * Experience with container technologies and deployment tooling, such as Docker, Helm, Kustomize, Argo CD, Flux, or similar. * Experience implementing infrastructure as code using Terraform, Pulumi, CloudFormation, Ansible, or comparable tools. * Practical experience with observability platforms and concepts, including metrics, logging, distributed tracing, alerting, dashboards, and SLOs. * Proficiency in at least one programming or scripting language such as Python, Go, Bash, or TypeScript. * Experience with CI/CD systems and automated software delivery practices. * Strong troubleshooting skills across Linux, networking, containers, cloud or on-premise infrastructure, and distributed systems. * Clear written and verbal communication skills, with an ability to work effectively across engineering disciplines. ## Description We are seeking a hands-on Platform Engineer to build and operate the infrastructure foundations that enable engineering teams to deliver reliable, scalable, and secure products. This role will shape the internal developer platform across Kubernetes, observability, AI-enabled workloads, CI/CD, and infrastructure as code., You will work closely with software, DevOps, security, and AI/ML teams to turn platform capabilities into a simple and dependable developer experience. The ideal candidate is comfortable moving between architecture decisions, production operations, automation, and practical enablement of partner teams., * Design, build, and operate scalable Kubernetes-based platform capabilities for development, test, and production workloads. * Create reusable infrastructure-as-code modules, patterns, and automated workflows using tools such as Terraform, Ansible, Helm, Kustomize, or equivalent technologies. * Develop and evolve platform observability: metrics, logs, traces, dashboards, alerting, service-level objectives, and operational runbooks. * Improve platform reliability, capacity management, security posture, and cost efficiency through automation and data-driven operational practices. * Build the infrastructure required to deploy, operate, observe, and govern AI/ML workloads, including model-serving, GPU-enabled compute, data-access, and workload-isolation patterns where applicable. * Enable self-service for product and engineering teams through well-designed platform APIs, templates, documentation, golden paths, and CI/CD integrations. * Partner with application, security, infrastructure, and AI teams to define platform standards and remove delivery bottlenecks. * Investigate complex production issues, lead root-cause analysis, and implement durable preventative improvements. * Contribute to technical roadmaps, architecture reviews, and engineering standards for the platform., AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD's "Responsible AI Policy" is available here. ## Related Videos - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Rate-limiting using eBPF and Istio: How to protect your SaaS customers from themselves](https://www.wearedevelopers.com/videos/100220-rate-limiting-using-ebpf-and-istio-how-to-protect-your-saas-customers-from-themselves) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Scoring 2000 Products per Request: Performance Pitfalls in Golang](https://www.wearedevelopers.com/videos/2073-scoring-2000-products-per-request-performance-pitfalls-in-golang) - [Get started with securing your cloud-native Java microservices applications](https://www.wearedevelopers.com/videos/123-get-started-with-securing-your-cloud-native-java-microservices-applications) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [DevOps Engineer Salary [2023]](https://www.wearedevelopers.com/magazine/203-devops-engineer-salary-2023) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline)