AI Platform Engineer

Insight Global
Dallas, TX, United States
19 days ago
Apply on www.dallasjobsite.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours

Tech stack

Artificial Intelligence Amazon Web Services Microsoft Azure Cloud Computing Data Infrastructure Linux Image Management Python (Programming Language) PostgreSQL Query Optimization Reliability Engineering Site Reliability Engineering Practices
+12 more
Prometheus Software Engineering Autoscaling Large Language Models Grafana AI Platforms Kubernetes Machine Learning Operations Api Gateway Decoding Serverless Computing Docker

Job description

One of our largest Telecom customers is seeking a skilled software engineer to join their Data Platform team within the Chief Data Office. This Software Engineer will play a key role in building, operating, and scaling a production-grade AI inferencing platform that supports high-throughput, low-latency large language model (LLM) workloads. The position combines software engineering, platform engineering, and reliability engineering responsibilities, with a focus on developing Python-based services, automating model deployments, and optimizing model serving performance using vLLM, Kubernetes, KServe, and Knative. The engineer will be responsible for containerization with Docker and Podman, maintaining and tuning PostgreSQL data models, troubleshooting production issues, and driving platform reliability through observability and automation. Working closely with ML engineers and platform architects, this individual will help scale mission-critical AI services, support new model rollouts, and continuously improve the infrastructure that powers enterprise AI applications.

We are a company committed to creating diverse and inclusive environments where people can bring their full, authentic selves to work every day. We are an equal opportunity/affirmative action employer that believes everyone matters. Qualified candidates will receive consideration for employment regardless of their race, color, ethnicity, religion, sex (including pregnancy), sexual orientation, gender identity and expression, marital status, national origin, ancestry, genetic factors, age, disability, protected veteran status, military or uniformed service member status, or any other status or characteristic protected by applicable laws, regulations, and ordinances. If you need assistance and/or a reasonable accommodation due to a disability during the application or recruiting process, please send a request to HR@insightglobal.com.To learn more about how we collect, keep, and process your private information, please review Insight Global’s Workforce Privacy Policy: https://insightglobal.com/workforce-privacy-policy/.

Requirements

Strong proficiency in Python, with experience building and maintaining production services.

-Hands-on experience with Docker and Podman for containerization and image management.

-Practical experience with Postgres, including query optimization and schema design.

–Solid working knowledge of Kubernetes, including deployments, services, autoscaling, and

troubleshooting workloads in production.

-Direct experience with vLLM for LLM inference serving.

-Experience deploying models with KServe and experience with Knative for serverless workloads or event-driven scaling.

-Comfortable working in a Linux environment and using standard CLI tooling.

-Cloud platform experience, such as Azure/AKS, AWS/EKS, or GCP/GKE.

-Strong debugging skills across the stack - from application code to infrastructure -Experience with GPU-aware scheduling and resource management in Kubernetes.

-Familiarity with LLM-specific optimizations: speculative decoding, quantization (FP8/INT8), KV cache

management, and continuous batching.

-Experience with API gateways or LLM routing layers, such as LiteLLM or similar.

-Background in SRE practices: SLOs/SLAs, incident response, and on-call tooling.

-Experience with Helm, ArgoCD, or other GitOps deployment tooling.

-Familiarity with observability stacks, including Prometheus, Grafana, and OpenTelemetry.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dallasjobsite.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

1:24 min

Comprehensive AI infrastructure stacks at the Linux Foundation

Matt White Matt White · World Congress 2025

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

Videos

See all

Related articles

See all