Sr Big Data/Machine Learning Engineer

Randstad
United States
6 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$157,310.0 - $174,096.0
Working hours
Regular working hours
Job source

Tech stack

Training Data Java (Programming Language) Artificial Intelligence Amazon Web Services Amazon S3 Microsoft Azure Big Data BigQuery Cloud Engineering Databases Information Engineering Data Governance
+46 more
DevOps Distributed Data Store Data Flow Control Apache Hadoop Python (Programming Language) Machine Learning MongoDB NoSQL Tensorflow Software Construction Software Engineering SQL Databases Pinecone Feature Store Qdrant Google Cloud Feature Engineering Pytorch DeepSpeed Great Expectations (Foster Youth College-readiness and Support Program) Retrieval-Augmented Generation Large Language Models Snowflake Apache Spark Deep Learning Generative AI Pyspark Scikit Learn Kubernetes Information Technology Apache Flink Deployment Automation Cassandra Production Code Xgboost Apache Kafka Milvus Ollama Machine Learning Operations Drift Detection Megatron Data Pipelines Docker Databricks Programming Languages Microservices

Job description

Data Engineering & Big Data Architecture:

  • Design, build, and maintain high-throughput streaming and batch data processing pipelines to support training data preparation and feature generation at scale.
  • Manage big data storage, transformation, and querying frameworks across multi-terabyte dataset systems.

Machine Learning & Feature Engineering:

  • Implement feature engineering processes and feature store architectures ensuring training-serving consistency.
  • Convert research-level ML prototypes into optimized, scalable production code.
  • Fine-tune, benchmark, and scale traditional ML, Deep Learning, or Generative AI models for real-time and batch inference.

MLOps, Deployment & Infrastructure:

  • Build and maintain robust MLOps framework tools, continuous training (CT) pipelines, and CI/CD pipelines for automated deployment.
  • Expose ML models via REST/gRPC microservices, ensuring low-latency inference and target availability SLA compliance.
  • Implement comprehensive model monitoring tools for data drift, concept drift, system performance, and output quality checks.

Technical Leadership & Governance:

  • Mentor junior and mid-level engineers, enforcing software engineering best practices across data and ML codebases.
  • Ensure system compliance with enterprise security, data privacy, and AI governance standards.

Requirements

Experience: 5+ years of software engineering, big data engineering, and applied machine learning engineering experience in production environments.Programming Languages: Advanced proficiency in Python and Scala or Java. Big Data Frameworks: Hands-on experience with Apache Spark (PySpark/Scala), Kafka, Flink, Hadoop, Databricks, or Snowflake.Machine Learning Frameworks: Proficiency with frameworks like PyTorch, TensorFlow, scikit-learn, or XGBoost. MLOps & DevOps: Experience with MLflow, Kubeflow, Feature Stores (e.g., Feast, Hopsworks), Docker, Kubernetes, and CI/CD automation.Cloud Architecture: Deep experience with AWS (S3, SageMaker, EMR), Google Cloud Platform (BigQuery, Vertex AI, Dataflow), or Azure (Databricks, Synapse).Databases & Storage: Mastery of SQL, NoSQL databases (Cassandra, MongoDB), and Vector Databases (Pinecone, Milvus, Qdrant). Education: Bachelor’s or Master’s degree in Computer Science, Data Science, Software Engineering, or a related quantitative field.

skills:

S3,experience with AWS,Flink,Hadoop,Kafka,Apache Spark,AI,Big Data,BigQuery,Cassandra,Cloud Architecture,Dataflow,data quality frameworks,data processing pipelines,Databases,Databricks,Deep Learning,DeepSpeed,automated deployment,DevOps,distributed data,Docker,concept drift,feature generation,Feature Engineering,Feature Stores,Generative AI models,Generative AI,Great Expectations,Data Engineering,Computer Science,Java,Kubeflow,Kubernetes,Large Language Models,LLMs,Machine Learning,ML models,applied machine learning,ML infrastructure,MLOps,MLflow,Megatron,microservices,Azure,Milvus,MongoDB,NoSQL databases,Ollama,Pinecone,production code,Programming Languages,PySpark,proficiency in Python,PyTorch,Qdrant,RAG,SQL,scikit-learn,Snowflake,software engineering best practices,software engineering,Machine Learning Frameworks,TensorFlow,training data,XGBoost,resilient,data drift,Architecture,automation,continuous training,enterprise security,prototypes,data privacy,Deequ,Data Science,Vertex AI,Governance,AI governance,Infrastructure,system compliance,Mentor,model training,production systems,quality checks,system implementation,Technical Leadership,Vector Databases

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Loading talks and stories from around this role…