> Markdown version of [/jobs/ext/2506197-software-engineer-inference](https://www.wearedevelopers.com/jobs/ext/2506197-software-engineer-inference). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer - Inference - **Company:** Bitwise - **Location:** Laurel, MD, United States - **Experience:** Experienced - **Salary:** $197,500.0 - $233,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Continuous Integration, Python (Programming Language), Octopus Deploy, Reliability Engineering, Cloud Services, Prometheus, AI Infrastructure, Large Language Models, Grafana, Containerization, Kubernetes, Docker, Programming Languages - **Published:** August 11, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=0b473b61bb2c4cdc ## About the Role * 3 years of relevant experience with a B.S. in a technical discipline, or 7 years of experience in lieu of a degree * Experience with Python and/or other modern programming languages * Familiarity with Argo CD and/or other CI/CD frameworks * Experience with Kubernetes and Helm * Familiarity with AWS or other cloud service providers * Ability to learn new technologies quickly and adapt as requirements evolve * Strong communication skills and a genuine willingness to ask questions * Active TS/SCI clearance with polygraph Preferred Skills * Experience with vLLM, LiteLLM, or similar inference-serving frameworks * Experience with other LLM hosting frameworks and practices * Experience supporting production software using Site Reliability Engineering (SRE) best practices * Experience with Elastic, Grafana/Prometheus, or other observability frameworks and practices * Experience with Docker and containerization * Experience in traffic shaping and quality-of-service engineering * Knowledge of and genuine interest in hosting AI capabilities ## Description We're looking for a Software Engineer to join our AI infrastructure team and help build the next generation of capabilities powering our customer's most critical missions. In this role, you'll be a key contributor to the foundation that makes cutting-edge AI accessible - ensuring users throughout the organization have reliable access to the highest-quality large language models (LLMs) available. This isn't a role where you'll be handed a tightly scoped checklist and told to execute. Mission needs shift fast, new technologies emerge constantly, and you'll be expected to turn loosely defined problems into working solutions. You'll sharpen your skills continuously and grow alongside a team that takes both the mission and each other seriously. Day to day, you will: * Procure, configure, and test new inference models, preparing them for release to users * Develop in-house services and techniques to ensure continual, high-quality inference for our customer * Work with model vendor teams and representatives to build reliable pipelines for closed-source model usage * Collaborate with teammates on surge efforts to support short-term, high-priority inference needs * Partner with other teams across the organization to establish solid infrastructure and integrate LLM-powered tools that serve real user needs ## Related Videos - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Data binning and understanding histograms](https://www.wearedevelopers.com/videos/2086-data-binning-and-understanding-histograms) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) - [Keycloak case study: Making users happy with service level indicators and observability](https://www.wearedevelopers.com/videos/1599-keycloak-case-study-making-users-happy-with-service-level-indicators-and-observability) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Dev Digest 196: AI Killed DevOps, LLM Political Bias & AI Security](https://www.wearedevelopers.com/magazine/659-dev-digest-196-ai-killed-devops-llm-political-bias-ai-security) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this)