> Markdown version of [/jobs/ext/2960154-senior-cloud-infrastructure-engineer-senior-virtualization-engineer](https://www.wearedevelopers.com/jobs/ext/2960154-senior-cloud-infrastructure-engineer-senior-virtualization-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Cloud Infrastructure Engineer / Senior Virtualization Engineer - **Company:** Megan Soft, Inc. - **Location:** Dearborn, MI, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Cloud Computing, Cloud Engineering, Cyber Security, DevOps, Disaster Recovery, VMware ESX Servers, Expert Systems, Firmware, Github, Monitoring of Systems, Issue Tracking Systems, Python (Programming Language), Knowledge Management, Kernel-Based Virtual Machine, Linux System Administration, Networking Basics, OpenShift, Windows PowerShell, Role-Based Access Control, Reliability Engineering, Ansible, Prometheus, Virtual Machines, Virtualization Technology, Data Logging, Scripting, Google Cloud, Delivery Pipeline, Grafana, Event Driven Architecture, Kubernetes, Information Technology, Dynatrace, Vmware - **Published:** September 17, 2026 - **Apply:** https://www.careerjet.com/jobad/us298369d74a1afb0afc10cf1e5a26ca0d ## About the Role * Scripting, Automation, Kubernetes, Root Cause Analysis, Troubleshooting (Problem Solving), Cloud Architecture, IT Solutions, GitHub, Cloud Infrastructure, Change Management, Technical Analysis, Developer, Tekton, Utilization Management, VMware, VMware ESX Servers, Platform Support, Infrastructure Architectures, * Senior Engineer Exp: Prac. In 2 coding lang. or adv. Prac. in 1 lang.; guides. 10+ years in IT; 8+ years in development Understanding of VMware and Kubernetes concepts Experience with Linux administration and networking fundamentals Proficiency in scripting languages for automation Experience with monitoring tools and logging solutions Understanding of virtualization concepts and technologies (e.g., KVM, VMware) Excellent problem-solving skills and the ability to troubleshoot complex issues across multiple layers of the stack Knowledge of CI/CD pipelines and DevOps methodologies Strong communication and collaboration skills Self-starter. Be on a mission to go where the work is. Look for opportunities to evolve services ## Description * Ansible, GCP, Dynatrace, Powershell, Access Controls, Python, Information Security, Automation, Artificial Intelligence & Expert Systems, * Responsible for engineering, deployment, operational administration, and protection of Global enterprise solutions to meet the Virtualization Server Hosting requirements for a variety of infrastructure systems, line of business and third-party application needs * Architect/design and support the installation, and administration of virtualization (VMware and OpenShift Virtualization (OSV)) suite of products * Responsible for the entire lifecycle of technologies globally including infrastructure security vulnerability patching, planning, designing, implementation, maintenance, upgrades and decommissioning of hardware and software * Engineer, test and document procedures, monitoring, logging, disaster recovery process and security policies and guidelines * Demonstrates knowledge of hardware and software products * Research industry best practices and trends * Global large-scale deployments of virtualization technologies * Support HPE Synergy/ProLiant ILO and firmware field testing Capacity Management * Conduct capacity planning and forecasting for the platforms, including Compute/Virtual Machine (VM), memory, storage, and network resources, to ensure scalability and prevent resource exhaustion * Analyze resource utilization trends and make recommendations for infrastructure scaling, consolidation, or optimization * Collaborate with application teams and stakeholders to understand future demand and project capacity needs * Develop and maintain capacity models and reports to support strategic planning Automation & Efficiency * Develop automation solutions (scripts, playbooks) for repetitive VMware/OSV tasks, including configuration changes, VM management (like snapshot removal), auditing, remediation and integration with ticketing systems * Leverage automation to enable delivering operator updates and changes efficiently at scale * Implement Site Reliability Engineering (SRE) principles and practices to improve overall platform stability, performance, and operational efficiency * Role Based Access Control deployment and auditing * Namespace and Resource Quota management (CPU, Disk and Storage) Observability, Monitoring, logging and Troubleshooting * Implement and maintain comprehensive end to end observability solutions (monitoring, logging, tracing) for the VMware/OSV environment, including integration with tools like Dynatrace, RHACM and Prometheus/Grafana * Explore and implement Event Driven Architecture (EDA) for enhanced real time monitoring and response * Develop capabilities to flag and report abnormalities and identify ""blind spots"" in observability * Perform deep dive Root Cause Analysis (RCA), potentially utilizing available tooling, to quickly identify and resolve issues across the global compute environment * Find the needle in a haystack/unhealthy bits in the compute universe (Globally) for faster time to resolution * Monitor VM health, resource usage, and performance metrics proactively * Monitor for unusual activity that might indicate a compromise or misconfiguration Solution Design & Consulting * Provide technical consulting and expertise to application teams requiring VMware/OSV solutions * Design, implement, and validate custom or dedicated OSV clusters and VM solutions for critical applications with unique or complex requirements (e.g., specialized appliances) Knowledge Management * Create, maintain, and update comprehensive internal documentation and customer facing content to facilitate self service and clearly articulate platform capabilities Support * Participate in L1 L3 level support to Operations teams environmental related issues. Monthly after hours and weekend work will be required, The Senior Digital Infrastructure Engineer is responsible for ensuring the reliability, security, scalability, and continuous improvement of NGNRI's digital infrastructure. This ro… + 28 days ago ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [WebAssembly: The Next Frontier of Cloud Computing](https://www.wearedevelopers.com/videos/972-webassembly-the-next-frontier-of-cloud-computing) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [My journey into DevOps world - How it all started!](https://www.wearedevelopers.com/videos/545-my-journey-into-devops-world-how-it-all-started) - [Bringing AI Model Testing and Prompt Management to Your Codebase with GitHub Models](https://www.wearedevelopers.com/videos/1536-bringing-ai-model-testing-and-prompt-management-to-your-codebase-with-github-models) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers)