Staff AI Infrastructure Engineer

SEEKR TECHNOLOGIES INC.
Reston, VA, United States
3 days ago
Apply on www.careerbuilder.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Working hours
Regular working hours

Tech stack

Artificial Intelligence Amazon Web Services Microsoft Azure C++ (Programming Language) Cloud Computing Cloud Engineering Cloud Storage Computer Networks Continuous Delivery Continuous Integration Distributed Systems Python (Programming Language)
+33 more
Machine Learning Enterprise Messaging Systems Octopus Deploy Oracle (Applications) Performance Tuning Software Architecture Systems Development Life Cycle Prometheus Software Engineering Systems Architecture AI Infrastructure Scripting Graphics Processing Unit (GPU) Google Cloud Cloud Platform System Spring Cloud Large Language Models Grafana Multi-Agent Systems Caching Event Driven Architecture AI Platforms Kubernetes Infrastructure Automation Frameworks Information Technology Low Latency Hardware Acceleration Machine Learning Operations TensorRT Hardware Infrastructure Oracle Cloud Infrastructure Docker Programming Languages

Job description

Seekr is building the infrastructure that powers the next generation of enterprise AI. As a Staff AI Infrastructure Engineer, you will design, build, and operate the platforms that enable large-scale training, serving, evaluation, and deployment of foundation models and autonomous AI agents.

You will work across distributed systems, Kubernetes, GPU infrastructure, high-performance inference, and enterprise AI platforms to build secure, scalable, and highly reliable systems capable of serving workloads ranging from edge AI deployments to trillion-parameter foundation models.

This role requires deep expertise in distributed systems, cloud-native infrastructure, AI platform engineering, and production software development. You will collaborate with research scientists, software engineers, product teams, and infrastructure engineers to define the architecture and technical direction of Seekr’s AI platform., * Design, develop, deploy, and maintain production AI infrastructure supporting model training, fine-tuning, inference, evaluation, and agentic AI workloads.

  • Design and operate scalable Kubernetes-based infrastructure supporting GPU-accelerated workloads across cloud, on-premises, hybrid, and edge environments.
  • Architect and optimize high-performance inference platforms capable of serving models ranging from resource-constrained edge deployments to trillion-parameter foundation models, with a focus on latency, throughput, scalability, reliability, and cost efficiency.
  • Build and maintain distributed systems that enable reliable scheduling, orchestration, deployment, monitoring, and lifecycle management of AI workloads.
  • Develop enterprise platforms supporting autonomous and multi-agent AI systems, including secure tool execution, orchestration, memory, evaluation, governance, and observability.
  • Design, implement, and automate AI infrastructure using Infrastructure-as-Code, GitOps, CI/CD pipelines, and modern software engineering practices.
  • Evaluate and integrate emerging AI infrastructure technologies, model serving frameworks, hardware accelerators, and cloud-native platforms to improve platform performance, scalability, and reliability.
  • Collaborate with engineering, research, product, and cross-functional teams to deliver secure, scalable, and production-ready AI platforms.
  • Lead technical design discussions, perform architecture reviews, mentor engineers, and establish engineering standards and best practices across the AI Infrastructure organization.
  • Participate in production support activities, including troubleshooting complex distributed systems, performance tuning, incident response, and continuous operational improvement.

Requirements

  • 8-12+ years of professional software engineering experience building distributed systems, cloud infrastructure, or large-scale platform services
  • Architects systems, drives technical direction cross-functionally
  • 4 year or higher degree or additional relevant experience, in addition to years of work experience
  • Demonstrated success designing and operating production Kubernetes environments supporting cloud-native applications and distributed services.
  • Strong software engineering skills using Python and one or more modern programming languages such as Go, Rust, or C++.
  • Proven ability to design, build, and operate production AI or machine learning infrastructure.
  • Expertise developing and optimizing large-scale AI inference platforms, including GPU utilization, distributed inference, batching, caching, quantization, and accelerator performance.
  • Familiarity with modern AI serving technologies such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, Ray Serve, or similar platforms.
  • Knowledge of distributed computing, networking, storage systems, cloud-native architectures, and infrastructure automation using technologies such as Kubernetes, Helm, Argo CD, Docker, Prometheus, Grafana, OpenTelemetry, and Infrastructure-as-Code tools.
  • Experience developing enterprise AI platforms, autonomous agents, or multi-agent systems, including orchestration, tool execution, governance, observability, and evaluation.
  • Familiarity with event-driven architectures, distributed messaging systems, and public cloud platforms including AWS, Azure, Oracle Cloud Infrastructure, or Google Cloud Platform.
  • Demonstrated technical leadership, including driving architectural decisions, mentoring engineers, and leading complex technical initiatives across cross-functional teams.
  • Demonstrated ability to analyze, profile, and optimize AI systems for performance, scalability, reliability, and cost across distributed compute environments., Amazon Web Services (AWS), Analysis Skills, Architectural Services, Artificial Intelligence (AI), Artificial Intelligence (AI) Agents, Automation, Best Practices, C++ Programming Language, Caching, Cloud Applications, Cloud Architecture, Cloud Computing, Cloud Storage, Computer Networks, Computer Systems, Continuous Deployment/Delivery, Continuous Integration, Cross-Functional, Data Quality, Distributed Applications, Distributed Computing, Docker, Emerging Technology, Engineering, Finance, GPU (Graphics Processing Unit), Government, Identify Issues, Incident Response, Machine Learning, Memory Hardware, Mentoring, Messaging Technology, Microsoft Windows Azure, Operational Improvement, Oracle, Performance Management, Performance Tuning/Optimization, Problem Solving Skills, Production Machining, Production Support, Production Systems, Programming Languages, Public Cloud, Python Programming/Scripting Language, Rust Programming Language, Scalable System Development, Scientific Research, Software Development, Software Engineering, System Architecture, Systems Maintenance, Team Player, Technical Leadership, Technical/Engineering Design

Benefits & conditions

Company Benefits:

  • Meaningful Mission & Impact - Work with a deeply talented, collaborative team solving some of the toughest AI challenges that matter.
  • Equity Ownership - RSUs that let you share directly in Seekr’s long-term success and growth.
  • Time Off That Respects Real Life - Unlimited PTO plus 14 paid company holidays to truly recharge.
  • Work Your Way - A flexible hybrid work environment with offices in Reston, VA and Austin, TX.
  • Competitive Total Rewards - A role-appropriate compensation structure that supports long-term growth, including base salary, bonuses, or commission plans depending on role.
  • 401(k) with Company Match - Build your future with a retirement plan that includes employer matching.
  • Comprehensive Health & Wellness - Medical, dental, vision, and life insurance coverage starting day one-for you and your family.
  • Parental Leave - Paid parental leave to support employees as they welcome a new child through birth, adoption, or foster placement.

About the company

Seekr is a leader in explainable and trustworthy artificial intelligence designed to power mission-critical decisions in enterprises, government, and regulated industries. SeekrFlow, our end-to-end AI platform, provides secure, auditable AI solutions tailored to sectors where transparency, accuracy, and compliance are paramount. Available across cloud, on-premises, and edge environments, SeekrFlow reduces bias, strengthens data integrity, and simplifies model oversight so organizations can rely on trusted AI decisions in high-stakes settings that impact society’s most sensitive and vital systems. Trusted by leading enterprises and government agencies, we partner with defense, finance, telecom, and critical infrastructure leaders to enable AI solutions that drive real-world results with unmatched transparency and control. We are a team of strategic thinkers and problem-solvers tackling the toughest challenges facing critical infrastructure and global enterprises through best-in-class AI models and customer deployment. Our team operates with unwavering commitment to our core values and mission:

  • We are driven by outcomes-our customers’’ success is what we strive for every day.
  • We believe trust is earned, which is why we build explainability and transparency into the entire AI lifecycle.
  • We take our responsibility to deliver secure AI seriously.
  • We believe innovation drives progress-we are building the technologies that power the systems our society depends on.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerbuilder.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:14 min

Exploring internal AI product initiatives and global engineering roles

Maria Apazoglou · Coffee With Developers

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

3:15 min

Reversing the caching model for artifact delivery

Thijs Feryn Thijs Feryn · World Congress 2026 Europe

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · World Congress 2026 Europe

Videos

See all

Related articles

See all