> Markdown version of [/jobs/ext/2866276-senior-ml-engineer](https://www.wearedevelopers.com/jobs/ext/2866276-senior-ml-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior ML Engineer - **Company:** CloudBolt Software - **Location:** Rockville, MD, United States - **Experience:** Expert - **Salary:** $170,000.0 - $220,000.0 - **Contract:** Permanent contract - **Skills:** Algorithm Design, Amazon Web Services, Amazon S3, Data Transformation, Java Virtual Machine (JVM), Python (Programming Language), Machine Learning, Regression Testing, Prometheus, Software Engineering, Autoscaling, Prophet, Deep Learning, Numerical Computing, Kubernetes, Information Technology, Production Code, Machine Learning Operations, Cloud Optimization - **Published:** September 12, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=4eaf55966311cd12 ## About the Role Must have strong experience…? * Master's degree or higher in a quantitative field (Computer Science, Machine Learning, Statistics, Applied Mathematics). * 5+ years of software engineering experience, with at least 3 years building and operating machine learning or statistical systems in production. * Expert-level Python: you write typed, tested, production-grade code, and you're fluent in numpy or similar array-based numerical computing. * Hands-on experience with time-series analysis and forecasting: seasonality, trend decomposition, anomaly detection, and classical statistical methods (percentiles, distributions, smoothing), not just deep learning. * Experience testing ML systems rigorously: regression testing against known-good baselines, behavioral validation, and reasoning about numerical reproducibility. * Working knowledge of Kubernetes: resource requests and limits, autoscaling behavior, and what happens to a workload when it's under-provisioned (OOM kills, CPU throttling). * Comfort owning a production service, not just a model: queues, caches, retries, observability, and debugging issues in customer environments from logs and metrics. * Clear written and verbal communication: as an ML engineer on the team, so you must be able to explain model behavior and tradeoffs to platform engineers, product managers, and customers. Experience in the following is beneficial * Experience with Prophet or similar forecasting libraries. * Prometheus/PromQL and experience working with metrics at scale. * Cloud cost optimization, capacity planning, or infrastructure efficiency background * AWS (S3, Managed Prometheus). * Experience being the ML domain expert on a team of generalists. ## Description * Own the recommendation engine end to end: model selection, algorithm design, preprocessing, and the guardrails that keep recommendations safe to apply to live production workloads. * Design, evaluate, and productionize time-series forecasting and statistical models (e.g., Prophet, percentile-based estimation) that right-size Kubernetes workloads across CPU, memory, GPU, and JVM heap. * Build and maintain the data-quality layer: detecting and filtering anomalies, load-test windows, startup spikes, and autoscaling artifacts from production telemetry before it reaches a model. * Define and continuously improve how we measure recommendation quality: regression testing against golden datasets, behavioral validation, and accuracy/safety metrics in production. * Investigate and resolve recommendation quality issues reported from customer environments, tracing them through data, preprocessing, and model behavior. * Serve as the team's machine learning authority: guide technical direction on ML questions, make model-vs-heuristic tradeoff calls, and clearly communicate them to platform engineers, product, and leadership. * Write production-grade Python for models and pipelines alike and share ownership of the surrounding service (message consumption, metrics ingestion, caching) with the rest of the team. * Prototype and validate new optimization capabilities (new resource types, new algorithms, new workload classes) from research through gradual, feature-flagged rollout. * Stay current on time-series forecasting and resource optimization techniques, and pragmatically evaluate which are worth adopting. ## Related Videos - [Effective Machine Learning - Managing Complexity with MLOps](https://www.wearedevelopers.com/videos/185-effective-machine-learning-managing-complexity-with-mlops) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [WeAreDevelopers LIVE - CSS is DOOMed](https://www.wearedevelopers.com/videos/1838-wearedevelopers-live-css-is-doomed) - [Optimizing your AI/ML workloads for sustainability](https://www.wearedevelopers.com/videos/570-optimizing-your-ai-ml-workloads-for-sustainability) - [Instant KAI Sandboxes with vCluster: Multi-Tenant, Multi-Scheduler GPU Sharing](https://www.wearedevelopers.com/videos/100333-instant-kai-sandboxes-with-vcluster-multi-tenant-multi-scheduler-gpu-sharing) - [Hate organising your photos? Try it with 5 Terabytes](https://www.wearedevelopers.com/videos/79-hate-organising-your-photos-try-it-with-5-terabytes) ## Related Articles - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [7 Most Popular Web Developer Jobs in Europe](https://www.wearedevelopers.com/magazine/163-7-most-popular-web-developer-jobs-in-europe) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models)