Data Analyst
Dminds Solutions
Oaks, United States of America
7 days ago
Role details
Contract type
Permanent contract Employment type
Full-time (> 32 hours) Working hours
Regular working hours Languages
English Experience level
SeniorJob location
Oaks, United States of America
Tech stack
Artificial Intelligence
Azure
Big Data
ETL
Data Warehousing
R
Hadoop
Hadoop Distributed File System
MapReduce
Hive
Python
Matlab
NoSQL
Cloudera
SQL Databases
Apache Yarn
Spark
Generative AI
Build Management
Apache Flume
Information Technology
Data Lineage
HuggingFace
Mahout
Data Pipelines
Job description
- Must have 5+ years of Strong Proficiency in SQL and Python.
- Big Data, Data Science & Analytics (Apache Hadoop HDFS, Base/Hive/Pig/Mahout/Flume/Scoop/MapReduce/Yarn, Cloudera HD, NoSQL, R, Spark/Shark/Milb, MATLAB)
Day to Day job Duties: (what this person will do on a daily/weekly basis)
- Design and build scalable ingestion, transformation, and processing pipelines for AI/ML.
- Develop and maintain data storage, such as vector databases, that support Retrieval-Augmented Generation (RAG).
- Implement data lineage, validation, and security controls to ensure high-quality data for AI models.
- Clean and structure complex datasets for operational and analytical systems
- Support AI Driven analytics and mapping
- Work on Azure platform
Requirements
Basic Qualifications: (what are the skills required to this job with minimum years of experience on each)
- 5+ years of Strong Proficiency in SQL and Python.
- 5+ years of experience in working on Azure cloud environment
- 5+ years of experience in working with AI Driven analytics
- AI Tools: Familiarity with vector databases (e.g., Pinecone, Milvus) and frameworks like Hugging Face or LangChain.
- Expertise in ETL/ELT pipeline design and data warehousing
- Bachelors in Computer Science or equivalent work experience