> Markdown version of [/jobs/ext/2427594-remote-llm-devops-inference-engineer](https://www.wearedevelopers.com/jobs/ext/2427594-remote-llm-devops-inference-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Remote- LLM DevOps/Inference Engineer - **Company:** Apetan Consulting - **Location:** New York, United States (Remote available) - **Contract:** Temporary contract - **Skills:** Abstraction Layers, Artificial Intelligence, Amazon Web Services, Amazon Elastic Compute Cloud, Audit Trail, Nvidia CUDA, Continuous Integration, DevOps, Identity and Access Management, Python (Programming Language), Data Logging, Delivery Pipeline, Large Language Models, Amazon Virtual Private Cloud (VPC), Rate Limiting, Kubernetes, Low Latency, Health Level Seven International, Machine Learning Operations, Cloudwatch, Terraform - **Published:** August 2, 2026 - **Apply:** https://www.dice.com/job-detail/25f497d2-8188-4507-b421-fd93ace0683e ## About the Role Build and maintain the AWS infrastructure for the evaluation platform as code, including networking, compute, orchestration, secrets, observability, and CI/CD. Strong AI & LLM Recent healthcare industry exp (HIPAA, hl7, etc), * AWS infrastructure at production scale: EKS or ECS, EC2 GPU instance families (G5/G6, P4d/P5) and their capacity realities, VPC design, IAM, KMS, Secrets Manager, ECR, CloudWatch, and Service Quotas. * Infrastructure as code: Terraform. No console-clicked production resources. * Model serving and inference optimization: Hands-on experience working with LLMs. Practical command of batching strategy, KV cache behavior, quantization tradeoffs, and multi-GPU sharding. * Container orchestration and GPU scheduling: ECS/EKS with GPU workloads, node autoscaling, and image build pipelines for CUDA-dependent stacks. * Reliability and cost engineering. SLOs, alerting, and a demonstrated track record of optimizing cloud spend without cutting capability. Desirable Skills * AWS SageMaker endpoints and Bedrock. * Hands-on experience with Python to contribute directly to the harness and the provider adapter layer. * Confidential computing fundamentals: Enclaves, remote attestation, sealed key release, and the security boundaries of TEEs. * Healthcare compliance posture: HIPAA-eligible service selection, BAA scope, audit logging, and the access-control mechanisms for PHI data. * Security hardening, including image scanning and network egress control for a closed-loop environment. Nice to Have * Prior experience hosting medical imaging or multimodal models. * Nitro Enclaves in production, or comparable TEE work. * Experience supporting self-service environments, clean tenancy boundaries, and fast credential provisioning. ## Description We are looking for a DevOps/Inference Engineer for one of our clients building a healthcare-focused AI benchmark and evaluation suite. The initial target is for clinical prediction tasks including sepsis onset, days-to-death, and lab value trend forecasting, evaluated across multiple frontier and vertical-specific models. Role Summary You will be responsible for the infrastructure the benchmark harness runs on, the selfhosted model serving stack, and the reliability of the platform. This is a role for someone who can provision a GPU cluster in the morning and tune the inferencing model engine in the afternoon. What You will Own * Build and maintain the AWS infrastructure for the evaluation platform as code, including networking, compute, orchestration, secrets, observability, and CI/CD. * Stand up the self-hosted inference track for the long tail of vertical healthcare models. This involves provisioning infrastructure for models serving on GPU compute with sensible batching, quantization where appropriate, autoscaling, and a standard onboarding path so adding new models takes hours, not weeks. * Build the provider abstraction layer alongside the AI engineers so that APIbased models (OpenAI, Anthropic, Gemini, and the growing list beyond) and self-hosted models present a uniform interface to the harness. Rate limiting, retry and backoff, quota management, request/response logging, and cost attribution per run are your responsibility. * Make benchmark runs reproducible and cost-optimized with pinned model and container versions, captured configuration, spot and reserved capacity strategy, and idle GPU elimination. * Build the observability story with throughput, latency, token and GPU-hour cost, failure taxonomy, and per-model dashboards for monitoring. * Support the surge model that the platform must let a burst of AI engineers land, run experiments, and leave without breaking anything or leaving orphaned resources behind. * Contribute to Trusted Execution Environment (TEE) architecture. Evaluate AWS Nitro Enclaves and comparable confidential computing approaches for the bring-your-own-data / bring-your-own-model scenario, including attestation- gated key release and the practical constraints of running model inference inside an enclave. ## Related Videos - [From Black Box to Glass Box : Bedrock AgentCore Observability](https://www.wearedevelopers.com/videos/2126-from-black-box-to-glass-box-bedrock-agentcore-observability) - [Infrastructure as Code: The Developer's Secret Weapon](https://www.wearedevelopers.com/videos/1221-infrastructure-as-code-the-developer-s-secret-weapon) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [30 powerful AWS hacks in just 30 minutes: Boost your developer productivity](https://www.wearedevelopers.com/videos/1624-30-powerful-aws-hacks-in-just-30-minutes-boost-your-developer-productivity) - [Implementing Feature Environments with AWS and Terraform](https://www.wearedevelopers.com/videos/531-implementing-feature-environments-with-aws-and-terraform) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers)