Junior Software Engineer - Inference

NEURAL SOLUTIONS LLC
Columbia, MD, United States
10 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Starter
Experience required
3 years minimum
Compensation
$138,000.0 - $163,000.0
Working hours
Regular working hours

Tech stack

Artificial Intelligence Amazon Web Services Continuous Integration Python (Programming Language) Octopus Deploy Reliability Engineering Cloud Services Prometheus AI Infrastructure Large Language Models Grafana Containerization
+3 more
Kubernetes Docker Programming Languages

Job description

We’re seeking a software engineer to support our AI infrastructure team at Columbia, MD. In this role, you’ll help build and maintain the foundation for customer AI capabilities while supporting a broader ecosystem of AI-enabled applications. Your focus will be ensuring access to the highest available quality LLMs to users throughout the inference software stack., * Procure, configure, and test new inference models, preparing them for release to our user base.

  • Develop in-house services and techniques to guarantee continual high-quality inference service for our customer.
  • Work with model vendor teams and representatives to create reliable pipelines for closed-source model usage.
  • Collaborate with teammates on surge efforts to support short-term, high-priority inference needs from our customer.
  • Engage with other teams in our organization to establish solid infrastructure for our services and integrate LLM-powered tools for user needs.

Requirements

  • Experience with Python and/or other modern programming languages.
  • Familiarity with Argo CD and/or other CI/CD frameworks.
  • Experience with Kubernetes/Helm.
  • Familiarity with AWS or other cloud service providers.
  • Ability to learn new technologies quickly.
  • Strong communication skills and willingness to ask questions.

Nice to Have:

  • Experience with vLLM, LiteLLM, or similar inference-serving frameworks.
  • Experience with other LLM hosting frameworks and practices.
  • Experience supporting production software using Site Reliability Engineering (SRE) best practices.
  • Experience with Elastic, Grafana/Prometheus, or other observability frameworks and practices.
  • Experience with Docker and containerization.
  • Experience in traffic shaping and quality-of-service engineering.
  • Knowledge of and interest in hosting AI capabilities.

Experience Required: 3 years with Bachelor’s degree in a technical discipline or 7 years without degree

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.clearancejobs.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

1:05 min

Measuring system availability utilizing Prometheus and straightforward PromQL

Alexander Schwartz Alexander Schwartz · WWC 2025

1:24 min

Comprehensive AI infrastructure stacks at the Linux Foundation

Matt White Matt White · WWC 2025

13:07 min

Configuring application observability with Micrometer and Prometheus

Aleksandr Kalikov · LIVE

4:47 min

Automating frontend performance metrics with Google Lighthouse

Miki Lombardi · JS Congress

Videos

See all

Related articles

See all