> Markdown version of [/jobs/ext/577306-staff-machine-learning-engineer-system-integration](https://www.wearedevelopers.com/jobs/ext/577306-staff-machine-learning-engineer-system-integration). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Staff Machine Learning Engineer System Integration - **Company:** ELLKAY, LLC. - **Location:** United States (Remote available) - **Experience:** Expert - **Salary:** $210,000.0 - $230,000.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), A/B Testing, Application Programming Interfaces (APIs), Artificial Intelligence, Airflow, Amazon S3, Computing Platforms, Automated Storage and Retrieval Systems, Audit Trail, Cloud Engineering, Software Quality, Computer Programming, Databases, Continuous Integration, Extract Transform Load (ETL), Data Transformation, Decision Support Systems, Distributed Systems, Amazon DynamoDB, Middleware, Interoperability, Python (Programming Language), Key Management, PostgreSQL, Machine Learning, Amazon Simple Notification Service (SNS), Software Engineering, Systems Integration, TypeScript, Management of Software Versions, Web Application Frameworks, AWS Cdk, Data Logging, Enterprise Software Applications, Real Time Systems, Apache Camel, Fast Healthcare Interoperability Resources, Large Language Models, Software Security, State Machines, AWS Lambda, Fastapi, Servicebus, Event Driven Architecture, AI Platforms, Kubernetes, Information Technology, Health Level Seven International, AWS Fargate, Integration Frameworks, Apache Kafka, Free and Open-Source Software, Build Tools, Api Design, Api Gateway, Amazon Simple Queue Service (SQS), Stream Processing, Data Pipelines, Dynatrace, Automation Anywhere, Mulesoft - **Published:** June 20, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=5b946f3bdb540bba ## About the Role Do you have experience in Systems integration?, * 10+ years of experience in software engineering, machine learning engineering, or closely related fields, including meaningful ownership of production systems. * Strong grounding in machine learning fundamentals, including model development, evaluation, deployment, and performance trade-offs. * Proven experience building distributed systems, integration platforms, or complex workflow-driven applications. * Expertise designing workflows across heterogeneous systems, including APIs, event streams, databases, and enterprise applications. * Strong programming ability in Python and working knowledge of Java or Go. * Experience with orchestration frameworks such as Airflow, Temporal, or Prefect. * Experience with messaging and streaming systems such as Kafka or Pub/Sub. * Experience designing and operating REST or gRPC APIs and data transformation pipelines. * Familiarity with modern AI application patterns, including LLMs, retrieval systems, vector databases, and model-integrated application workflows. * Strong communication skills and the ability to collaborate effectively across product, platform, security, and data teams., * Experience building or extending integration engines or middleware platforms similar to Iguana, MuleSoft, or Apache Camel. * Experience in healthcare or another regulated industry with complex data movement, auditability, and compliance needs. * Experience with event-driven architectures and real-time processing systems at scale. * Familiarity with Amazon Bedrock, Claude family models, or comparable enterprise AI platforms. * Contributions to open-source software, interoperability initiatives, or internal developer platforms. * Degree in computer science, machine learning, or a related technical field. ## Description ELLKAY is seeking a staff-level machine learning engineer who can do more than build models. This role is for an engineer who can design, integrate, and operate AI-powered systems across products, data pipelines, APIs, enterprise applications, and workflow engines in production settings. The ideal candidate combines strong machine learning fundamentals with deep systems thinking. This person understands how to connect models, rules, messaging, data transformations, and operational software into reliable end-to-end solutions, especially in environments where interoperability, security, and uptime matter. This role sits at the intersection of machine learning engineering, backend engineering, systems integration, and platform architecture. The focus is not on isolated model experimentation alone, but on building production-grade AI workflows that move data across systems, trigger downstream actions, and create measurable business value. The scope includes workflow orchestration, API-driven integration, event-based processing, evaluation and observability, secure deployment, and reusable platform components that allow AI capabilities to operate inside larger enterprise ecosystems., * Design and implement AI-enabled workflows that span internal platforms, third-party systems, APIs, databases, and operational tools. * Build integration pipelines for ingestion, transformation, model inference, decisioning, and downstream actions in batch and real-time environments. * Architect reusable system components such as connectors, transformation layers, orchestration patterns, and event-driven services that make AI capabilities easier to deploy across the organization. * Develop production services and APIs that expose AI functionality safely and reliably to enterprise applications and users. * Partner with product, infrastructure, and data teams to operationalize ML and LLM capabilities in business-critical workflows. * Establish engineering standards for reliability, observability, versioning, testing, evaluation, and governance of AI systems in production. * Mentor engineers and help shape the technical roadmap for AI systems integration and platform architecture. Healthcare and Data Interoperability * Build systems that work with healthcare interoperability standards such as FHIR R4, HL7 v2, and USCDI data elements. * Integrate clinical terminologies and ontology services including LOINC, SNOMED, and RxNorm to support normalization, retrieval, and decision support workflows. * Design solutions that protect PHI and align with HIPAA and broader security requirements for regulated environments. Platform and Infrastructure * Build high-performance backend services using modern Python frameworks such as FastAPI and Pydantic, with strong error handling, retry logic, and asynchronous execution patterns. * Design cloud-native architectures using services such as AWS Lambda, API Gateway, Step Functions, EventBridge, SQS, SNS, ECS Fargate, Aurora PostgreSQL, DynamoDB, and S3. * Write production-grade infrastructure as code using AWS CDK in Python or TypeScript, including support for multi-tenant deployments and cross-account environments. * Implement secure and observable systems using structured logging, distributed tracing, metrics, alarms, encryption, secrets management, and least-privilege access controls. Evaluation and Quality * Design rigorous evaluation frameworks for ML and AI systems, including benchmark creation, held-out test sets, label-quality controls, leakage prevention, and model comparison methods. * Implement prompt versioning, controlled experiments, and A/B testing approaches to continuously improve model and workflow performance. * Drive strong engineering discipline in code quality, testing, deployment hygiene, and production support. ## Related Videos - [Enterprise Integration Is Dead! Long Live AI-Driven Integration with Apache Camel](https://www.wearedevelopers.com/videos/1606-enterprise-integration-is-dead-long-live-ai-driven-integration-with-apache-camel) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Intro to FastAPI](https://www.wearedevelopers.com/videos/462-intro-to-fastapi) - [Building Reliable Serverless Applications with AWS CDK and Testing](https://www.wearedevelopers.com/videos/812-building-reliable-serverless-applications-with-aws-cdk-and-testing) - [Why make use of an integration platform in today's software developments and infrastructure?](https://www.wearedevelopers.com/videos/758-why-make-use-of-an-integration-platform-in-today-s-software-developments-and-infrastructure) - [AI Model Management Life Circles: ML Ops For Generative AI Models From Research to Deployment](https://www.wearedevelopers.com/videos/1152-ai-model-management-life-circles-ml-ops-for-generative-ai-models-from-research-to-deployment) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production)