AI ML Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+15 more
Job description
This role focuses on the design, implementation, and operation of high-throughput, real-time AI and computer vision systems in a manufacturing environment. The AI ML Engineer will work on distributed Python services, real-time data/vision streaming, performance engineering, observability, and automated remediation. This position collaborates closely with cross-functional engineering teams to help deliver reliable, secure, and scalable production systems., * Design and implement backend services using Python for high-concurrency, low-latency workloads.
- Configure and maintain containerized services deployed on Kubernetes using Helm.
- Tune resource limits/requests, autoscaling, and networking for stateful and streaming workloads.
- Implement secure deployment patterns using enterprise registries and CI/CD pipelines.
Real-Time Data & Vision Streaming
- Build and maintain ingestion services for real-time data streams, including industrial video protocols.
- Integrate multi-camera and multi-stream inputs into Python backends for monitoring, diagnostics, and analytics.
- Optimize streaming, buffering, and processing behavior to meet strict latency targets.
Numerical & Statistical Computing
- Implement and optimize matrix-heavy and numerical workloads using NumPy, SciPy, and OpenCV.
- Translate statistical methods into robust, production-grade code.
- Design calculations to distinguish normal process variation from true degradation or drift in data and signals.
Automated Remediation & Diagnostics
- Implement backend decision logic to separate software-correctable issues from physical or environmental issues.
- Develop automation scripts and workflows for programmatic fixes.
- Generate precise diagnostic payloads for maintenance and operations teams., * Design and run load, stress, and soak tests that include high-FPS streams, bursty workloads, and multi-session interactions.
- Implement and tune rate limiting, request throttling, queuing, and back-pressure strategies.
- Apply resiliency patterns for failed or partial executions.
Observability & Monitoring
- Configure and extend observability stacks for metrics, logs, and distributed tracing.
- Enable visualization and inspection of complex execution graphs and data flows.
- Contribute to definition and tracking of SLOs/SLIs around system health.
APIs & Integration
- Expose clean, stable APIs and schemas for downstream consumers.
- Provide diagnostic outputs to support operator-facing tools.
- Collaborate with frontend and integration teams to ensure APIs are well-documented, versioned, and testable.
Requirements
- 6+ years in Infrastructure/SRE, Core Platform Engineering, MLOps, Vision/Edge Platform Engineering, or high-throughput data/backend systems.
- Experience delivering and supporting production systems with strict performance, reliability, and availability requirements.
Python & Backend
- Expert-level production Python skills.
- Strong understanding of memory management, concurrency, and profiling in Python.
Containers, Kubernetes & CI/CD
- Deep experience with Docker and Kubernetes.
- Practical experience with Helm for templating and deploying applications.
- Hands-on experience with CI/CD tools for automated builds, tests, deployments, and rollbacks.
Streaming & Vision
- Hands-on experience ingesting and processing industrial camera/video streams.
- Strong proficiency with OpenCV and image/matrix operations.
Numerical & Statistical Skills
- Advanced proficiency with NumPy and SciPy.
- Experience implementing statistical techniques in production pipelines.
Performance, Testing & Observability
- Proven use of load and stress testing tools for data-heavy or streaming APIs.
- Experience with observability tools and concepts.
Collaboration & Ways of Working
- Experience working as part of a cross-functional engineering team.
- Strong communication skills for documenting designs, APIs, diagnostics, and test results.
- Ability to work in an iterative environment with clear deliverables, reviews, and handoffs., * Experience with industrial IoT, manufacturing environments, or other edge/plant-floor systems.
- Background in automated rollback, closed-loop control, or other remediation/alerting frameworks.
- Familiarity with modern AI/ML or LLM/agent systems from an infrastructure or performance perspective.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Navigating the AI Shift
MLOps – What’s the deal behind it?
MLOps And AI Driven Development
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again