Code Data Validation Consultant (Machine Learning & Data Processing)
US Tech Solutions, Inc.
San Jose, United States of America
yesterday
Role details
Contract type
Permanent contract Employment type
Full-time (> 32 hours) Working hours
Regular working hours Languages
English Experience level
IntermediateJob location
San Jose, United States of America
Tech stack
API
Artificial Intelligence
Data analysis
Bash
Big Data
Data Validation
Information Engineering
Software Debugging
Github
JSON
Python
Machine Learning
NumPy
Open Source Technology
Cloud Services
Jupyter Notebook
Data Processing
Scripting (Bash/Python/Go/Ruby)
Data Ingestion
Generative AI
Pandas
Information Technology
HuggingFace
Build Tools
Data Management
GPT
Data Pipelines
Data Selection
Job description
- Join our team to enable cutting-edge AI/ML innovation by building robust data pipelines and automation tools.
- You'll work closely with human data operators and generative AI teams to process, analyze, and optimize high-quality datasets for training machine learning models.
- Your work will directly impact the efficiency and performance of AI systems, from automating data quality checks to designing infrastructure that scales with evolving model requirements.
- This role is ideal for a problem-solver who thrives in fast-paced environments and enjoys bridging data engineering with machine learning.
Responsibilities:
- Data Pipeline Development:
- Design and implement Python-based automation tools to process, clean, and transform raw data for ML training.
- Build custom scripts to streamline data ingestion and preprocessing workflows.
- Quality Analysis & Reporting:
- Conduct manual and automated quality assessments to identify high/low-impact data for model training.
- Generate reports detailing experimental results, data effectiveness, and recommendations for improvement.
- ML Model Integration:
- Train and evaluate open-source ML models (e.g., Gemma) to assess data impact on model performance.
- Collaborate with AI teams to refine data selection strategies based on model feedback.
- Infrastructure Optimization:
- Develop scalable solutions in Colab/Jupyter Notebooks to automate data validation and filtering.
- Troubleshoot and debug data formatting issues (e.g., code-comment relevance, dataset consistency).
Requirements
- Preferred: 2-3+ years in data analysis/validation/engineering, ML engineering, or automation-focused roles.
- Bonus: PhD graduates with hands-on ML/data processing projects.
Required (Desired):
- Exposure to Generative AI models (e.g., GPT, Llama) or large-scale datasets.
- Bash/Shell Scripting: Ability to automate repetitive tasks.
- Familiarity with APIs for data ingestion/processing.
- Experience contributing to open-source projects or public GitHub repositories.
- Knowledge of cloud services., * Technical Expertise:
- Python: Medium to Advanced proficiency (scripting, automation, data processing libraries like Pandas/NumPy).
- Hands-on experience writing, executing and reviewing code. (Preferably using Colab/Jupyter Notebooks)
- Data & ML Skills:
- Experience training/fine-tuning ML models and analyzing their performance.
- Familiarity with public data platforms (Hugging Face, GitHub) and data formats (JSON, CSV).
- Analytical Skills.
- Proven ability to assess data quality and build tools to automate quality checks., * Bachelor's degree in Computer Science, Data Science, Engineering, or related STEM field.