> Markdown version of [/jobs/ext/3533279-big-data-machine-learning-engineer](https://www.wearedevelopers.com/jobs/ext/3533279-big-data-machine-learning-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Big Data/Machine Learning Engineer - **Company:** Randstad - **Location:** Richmond, VA, United States - **Experience:** Expert - **Salary:** $139,755.0 - $154,669.0 - **Contract:** Permanent contract - **Skills:** Training Data, Java (Programming Language), Artificial Intelligence, Amazon Web Services, Amazon S3, Microsoft Azure, Big Data, BigQuery, Cloud Engineering, Databases, Information Engineering, Data Governance, DevOps, Distributed Data Store, Data Flow Control, Apache Hadoop, Python (Programming Language), Machine Learning, MongoDB, NoSQL, Tensorflow, Software Construction, Software Engineering, SQL Databases, Pinecone, Feature Store, Qdrant, Google Cloud, Feature Engineering, Pytorch, DeepSpeed, Great Expectations (Foster Youth College-readiness and Support Program), Retrieval-Augmented Generation, Large Language Models, Snowflake, Apache Spark, Deep Learning, Generative AI, Pyspark, Scikit Learn, Kubernetes, Information Technology, Apache Flink, Deployment Automation, Cassandra, Production Code, Xgboost, Apache Kafka, Milvus, Ollama, Machine Learning Operations, Drift Detection, Megatron, Data Pipelines, Docker, Databricks, Programming Languages, Microservices - **Published:** October 2, 2026 - **Apply:** https://www.dice.com/job-detail/ca018168-7673-442d-b3e5-3a997ad715c0 ## About the Role Experience: 5+ years of software engineering, big data engineering, and applied machine learning engineering experience in production environments.Programming Languages: Advanced proficiency in Python and Scala or Java. Big Data Frameworks: Hands-on experience with Apache Spark (PySpark/Scala), Kafka, Flink, Hadoop, Databricks, or Snowflake.Machine Learning Frameworks: Proficiency with frameworks like PyTorch, TensorFlow, scikit-learn, or XGBoost. MLOps & DevOps: Experience with MLflow, Kubeflow, Feature Stores (e.g., Feast, Hopsworks), Docker, Kubernetes, and CI/CD automation.Cloud Architecture: Deep experience with AWS (S3, SageMaker, EMR), Google Cloud Platform (BigQuery, Vertex AI, Dataflow), or Azure (Databricks, Synapse).Databases & Storage: Mastery of SQL, NoSQL databases (Cassandra, MongoDB), and Vector Databases (Pinecone, Milvus, Qdrant). Education: Bachelor's or Master's degree in Computer Science, Data Science, Software Engineering, or a related quantitative field. skills: S3,experience with AWS,Flink,Hadoop,Kafka,Apache Spark,AI,Big Data,BigQuery,Cassandra,Cloud Architecture,Dataflow,data quality frameworks,data processing pipelines,Databases,Databricks,Deep Learning,DeepSpeed,automated deployment,DevOps,distributed data,Docker,concept drift,feature generation,Feature Engineering,Feature Stores,Generative AI models,Generative AI,Great Expectations,Data Engineering,Computer Science,Java,Kubeflow,Kubernetes,Large Language Models,LLMs,Machine Learning,ML models,applied machine learning,ML infrastructure,MLOps,MLflow,Megatron,microservices,Azure,Milvus,MongoDB,NoSQL databases,Ollama,Pinecone,production code,Programming Languages,PySpark,proficiency in Python,PyTorch,Qdrant,RAG,SQL,scikit-learn,Snowflake,software engineering best practices,software engineering,Machine Learning Frameworks,TensorFlow,training data,XGBoost,resilient,data drift,Architecture,automation,continuous training,enterprise security,prototypes,data privacy,Deequ,Data Science,Vertex AI,Governance,AI governance,Infrastructure,system compliance,Mentor,model training,production systems,quality checks,system implementation,Technical Leadership,Vector Databases ## Description Data Engineering & Big Data Architecture: * Design, build, and maintain high-throughput streaming and batch data processing pipelines to support training data preparation and feature generation at scale. * Manage big data storage, transformation, and querying frameworks across multi-terabyte dataset systems. Machine Learning & Feature Engineering: * Implement feature engineering processes and feature store architectures ensuring training-serving consistency. * Convert research-level ML prototypes into optimized, scalable production code. * Fine-tune, benchmark, and scale traditional ML, Deep Learning, or Generative AI models for real-time and batch inference. MLOps, Deployment & Infrastructure: * Build and maintain robust MLOps framework tools, continuous training (CT) pipelines, and CI/CD pipelines for automated deployment. * Expose ML models via REST/gRPC microservices, ensuring low-latency inference and target availability SLA compliance. * Implement comprehensive model monitoring tools for data drift, concept drift, system performance, and output quality checks. Technical Leadership & Governance: * Mentor junior and mid-level engineers, enforcing software engineering best practices across data and ML codebases. * Ensure system compliance with enterprise security, data privacy, and AI governance standards.