> Markdown version of [/jobs/ext/626711-software-engineer-ii-ai-infrastructure](https://www.wearedevelopers.com/jobs/ext/626711-software-engineer-ii-ai-infrastructure). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer II - AI Infrastructure - **Company:** Staffed4U LLC - **Location:** Annapolis Junction, MD, United States - **Experience:** Expert - **Salary:** $193,000.0 - $306,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Application Performance Management, Encodings, Information Systems, Computer Engineering, DevOps, Distributed Systems, Monitoring of Systems, Python (Programming Language), Machine Learning, Performance Tuning, Software Tools, Prometheus, Search Technologies, Software Deployment, Software Engineering, Systems Integration, Web Applications, AI Infrastructure, Data Logging, Cloud Platform System, High Performance Computing, Large Language Models, Grafana, Generative AI, Infrastructure as Code (IaC), AI Platforms, Kubernetes, Information Technology, Deployment Automation, Virtual Agents - **Published:** June 24, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=bcaed31fa173c4a3 ## About the Role Do you have a Bachelor's degree?, * Bachelor's degree in Computer Science, Software Engineering, Computer Engineering, Information Systems, or a related technical discipline., * Four (4) additional years of directly related experience may be substituted for a bachelor's degree., * Eight (8) or more years of software engineering experience. * Proven experience building and supporting production systems at scale. * Experience designing and supporting high-volume web applications. * Experience integrating complex systems across multiple technologies and platforms. * Experience supporting cloud-native infrastructure in AWS. * Experience administering and deploying applications within Kubernetes environments. Technical Skills * Strong Python development skills. * AWS Cloud Engineering * Kubernetes * Infrastructure as Code (IaC) * CI/CD Pipelines * DevOps Methodologies * Monitoring and Observability Platforms * Distributed Systems Architecture * Performance Optimization * Systems Integration Observability Technologies Experience with one or more of the following: * OpenTelemetry * Grafana * Prometheus * Application Performance Monitoring (APM) Solutions, * Strong problem-solving and analytical abilities. * Ability to thrive in ambiguous and rapidly evolving environments. * Strong organizational influence and change management skills. * Excellent written and verbal communication skills. * Ability to work independently and collaboratively within highly technical teams., * Experience with AI inference serving technologies such as: + vLLM + LiteLLM + Similar inference platforms * Experience with agentic AI frameworks such as: + LangChain + LangGraph + Similar orchestration frameworks * Experience with: + Vector databases + Embedding systems + Semantic search technologies * Knowledge of: + High-Performance Computing (HPC) + Distributed Computing Systems * Experience supporting production AI/ML environments. ## Description We are seeking an experienced Software Engineer II to support an advanced AI Infrastructure Team responsible for developing and maintaining the platform that serves as the foundation for enterprise AI capabilities. This role focuses on AI inference services while supporting a broader ecosystem of AI-enabled applications, including Retrieval-Augmented Generation (RAG), autonomous agents, and emerging machine learning technologies. The ideal candidate is a highly skilled engineer who can independently design, build, deploy, and operate scalable infrastructure solutions while helping shape the future of AI adoption across mission-critical environments., * Design, implement, and optimize infrastructure supporting AI model inference at scale. * Develop, deploy, and maintain production AI services and applications. * Support emerging AI technologies, including: + Retrieval-Augmented Generation (RAG) + Agentic AI Systems + Large Language Model (LLM) Platforms + AI Inference Services * Build highly available, reliable, and scalable AI platform components. * Navigate ambiguous requirements and define practical, scalable technical solutions., * Implement monitoring, logging, and observability solutions across AI services and infrastructure. * Develop operational dashboards and alerting capabilities using: + Grafana + Prometheus + OpenTelemetry + Application Performance Monitoring (APM) tools * Support incident response, troubleshooting, and root cause analysis efforts. DevOps & Automation * Develop and maintain CI/CD pipelines. * Improve deployment automation and operational efficiency. * Promote DevOps best practices across engineering teams. * Drive adoption of modern engineering tools and methodologies. Security & Collaboration * Contribute to secure AI system design and implementation. * Support compliance with organizational security requirements. * Provide technical guidance and informal mentorship to junior engineers. * Collaborate with software engineers, data scientists, platform engineers, and mission stakeholders., Choose from three comprehensive medical plans through Aetna. The company pays 80% of monthly premiums for employees. ## Related Videos - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [A Brief History of Data Storage](https://www.wearedevelopers.com/videos/974-a-brief-history-of-data-storage) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) - [DevOps for AI: running LLMs in production with Kubernetes and KubeFlow](https://www.wearedevelopers.com/videos/1222-devops-for-ai-running-llms-in-production-with-kubernetes-and-kubeflow) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud)