AI Platform Engineer

Propertyvalue Quantum Technologies Llc
Virginia Beach, VA, United States
1 day ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
6 years minimum
Compensation
$145,600.0
Working hours
Regular working hours
Job source

Tech stack

Java (Programming Language) Application Programming Interfaces (APIs) Artificial Intelligence Amazon Web Services Amazon Elastic Compute Cloud Amazon S3 Software Applications Cloud Computing Cloud Computing Security Databases Cursor (Graphical User Interface Elements) Software Design Patterns
+32 more
Programming Tools Disaster Recovery Distributed Systems Identity and Access Management Python (Programming Language) Machine Learning Routing Platform as a Service (PAAS) Software Architecture Reliability Engineering Tensorflow Prometheus Software Engineering Datadog Data Logging Computer Networking Systems Cloud Platform System GitHub Copilot Pytorch Autoscaling Large Language Models Grafana Caching Amazon Virtual Private Cloud (VPC) Backend Rate Limiting AI Platforms Scikit Learn Kubernetes Information Technology Api Gateway Terraform

Job description

As a Senior Software Engineer, you will focus on the reliability and resilience of mission-critical AI platforms.

You will design systems that withstand provider and dependency failures, establish observability and service-level objectives, and keep the platform current as the AI ecosystem evolves.

The team s scope includes Kubernetes-based platform-as-a-service frameworks, model-agnostic AI/LLM gateways, hybrid networking, observability, and self-service developer tooling.

We are looking for a collaborative, self-motivated engineer who is comfortable with ambiguity, takes ownership, and enjoys solving complex infrastructure challenges.

Legal AI is one of the most exciting and fast-moving areas in technology today.

If you are interested in building the foundational platforms that power the next generation of AI products, we’d love to hear from you.

What You ll Do

  • Improve platform reliability and resilience by designing for failure, defining and meeting SLOs, leading incident response, and reducing operational toil.
  • Design, build, and operate Kubernetes-based PaaS frameworks, AI/LLM gateways, APIs, and self-service tools for AI applications across Bloomberg.
  • Develop model-agnostic gateway capabilities for providers such as OpenAI, Anthropic, Gemini, and AWS Bedrock, including routing, fallback, retries, rate limiting, and cost controls.
  • Build observability systems covering metrics, logs, traces, dashboards, and alerting to detect and resolve issues before they affect clients.
  • Develop networking solutions that connect applications across public-cloud and on-premises environments.
  • Provision and manage cloud infrastructure using Terraform and modern software engineering practices.
  • Keep platforms secure and current through dependency patching, runtime upgrades, migrations, and provider-integration updates.
  • Create frameworks, templates, and workflows that improve developer productivity and reduce operational overhead.
  • Evaluate emerging AI technologies and adapt the platform to support new development patterns and use cases.

Requirements

  • 6+ years of professional software engineering experience.
  • Strong Python skills and experience developing production-grade backend services and APIs; Java experience is a plus.
  • Experience designing and operating distributed systems in public-cloud environments, with a strong understanding of failure modes and resilient design patterns.
  • Hands-on AWS experience, including services such as EC2, S3, IAM, and container-based workloads.
  • Experience with Infrastructure as Code, preferably Terraform.
  • Experience with production operations, including metrics, logging, tracing, alerting, SLOs, and incident response.
  • Strong knowledge of software architecture, databases, networking, cloud infrastructure, and modern application development.
  • A degree in computer science, engineering, or a related field, or equivalent practical experience., * Experience building or operating API gateways, LLM gateways, or similar proxy layers with routing, fallback, rate limiting, caching, and cost tracking.
  • Experience with Open Telemetry, Prometheus, Grafana, Datadog, or similar observability tools.
  • Experience with Kubernetes and autoscaling technologies, preferably Amazon EKS and Karpenter.
  • Knowledge of AWS networking and security, including VPC, Direct Connect, IAM, and cloud security controls.
  • Experience with chaos engineering, load and failure testing, capacity planning, or disaster recovery.
  • Experience developing AI-powered applications, agent-based systems, model-inference services, or AI serving platforms.
  • Familiarity with AI development tools such as Claude Code, Cursor, or GitHub Copilot.
  • Working knowledge of machine learning concepts and the ML development lifecycle. Experience with SageMaker, Bedrock, PyTorch, TensorFlow, or scikit-learn is a plus.
  • The ability to learn quickly and independently lead large technical initiatives from concept through production.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

1:52 min

Structuring and scaling the backend engineering team

Stefan Lingler Stefan Lingler +1 · Coffee With Developers

3:15 min

Reversing the caching model for artifact delivery

Thijs Feryn Thijs Feryn · World Congress 2026 Europe

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

1:12 min

Choosing TypeScript for complex backend applications

Maximilian Otto Maximilian Otto · World Congress 2024

Videos

See all

Related articles

See all