Software Engineer (AI Infrastructure)

BigBear.ai, Inc.
Columbia, MD, United States
30 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
4 years minimum
Working hours
Regular working hours

Tech stack

Artificial Intelligence Amazon Web Services Cloud Engineering Encodings DevOps Distributed Systems Python (Programming Language) Performance Tuning Prometheus Systems Integration Web Applications AI Infrastructure
+8 more
Data Logging High Performance Computing System Availability Grafana AI Platforms Kubernetes Infrastructure Automation Frameworks BIG-IP Access Policy Manager (APM)

Job description

  • Design, implement, and optimize infrastructure for AI model inference at scale
  • Support the development and maintenance of production AI services and applications, including retrieval augmented generation (RAG), autonomous agents, and emerging technologies
  • Navigate ambiguity and define solutions for underspecified systems and requirements
  • Drive adoption of new technologies and practices across engineering teams
  • Implement monitoring, logging, and observability solutions for AI services
  • Automate infrastructure provisioning and configuration using Infrastructure-as-Code (IaC) principles
  • Ensure high availability, reliability, and performance of AI platform components
  • Contribute to security best practices for AI systems and data
  • Provide technical guidance and informal mentorship to junior engineers

Requirements

  • 8+ years of relevant experience, or Bachelor’s degree in a technical discipline + 4+ years of experience
  • Clearance: TS/SCI w/Poly
  • Proven experience building and maintaining production systems at scale
  • Experience with high-volume web application architecture and performance optimization
  • Strong background in systems integration across diverse technologies and platforms
  • Hands-on experience with cloud engineering in AWS
  • Proficiency with Kubernetes administration and deployment patterns
  • Strong Python programming skills
  • Experience implementing observability solutions (APM, OpenTelemetry, Grafana, Prometheus)
  • Familiarity with CI/CD pipelines and DevOps practices
  • Strong change management and organizational influence skills
  • Ability to thrive in ambiguous environments and create structure where needed
  • Excellent communication and collaboration skills

What we’d like you to have

  • Experience with AI inference serving technologies (vLLM, LiteLLM, etc.)
  • Previous experience with agentic frameworks (LangChain)
  • Knowledge of vector databases and embedding systems
  • Experience with high-performance computing or distributed systems

About the company

BigBear.ai is a leading provider of AI-powered decision intelligence solutions for national security, supply chain management, and digital identity. Customers and partners rely on Bigbear.ai’s predictive analytics capabilities in highly complex, distributed, mission-based operating environments. Headquartered in McLean, Virginia, BigBear.ai is a public company traded on the NYSE under the symbol BBAI. For more information, visit https://bigbear.ai/ and follow BigBear.ai on LinkedIn: @BigBear.ai and X: @BigBearai.

BigBear.ai is an Equal opportunity employer all protected groups, including protected veterans and individuals with disabilities.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.clearancejobs.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:14 min

Exploring internal AI product initiatives and global engineering roles

Maria Apazoglou · Coffee With Developers

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

3:17 min

Optimizing character encoding with Kim variable byte encoding

Douglas Crockford Douglas Crockford · WWC 2024

2:12 min

Navigating technical clarity as a global black belt

Chris Heilmann +2 · LIVE

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

Videos

See all

Related articles

See all