Data Scientist

Tekshapers Inc
Raritan, United States of America
4 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English
Experience level
Senior

Job location

Raritan, United States of America

Tech stack

Artificial Intelligence
Amazon Web Services (AWS)
Azure
Big Data
Google BigQuery
Cloud Computing
Cloud Computing Security
Cloud Storage
Data Governance
ETL
Data Visualization
Distributed Computing Environment
Data Flow Control
Identity and Access Management
Python
Machine Learning
Systems Development Life Cycle
TensorFlow
Software Engineering
SQL Databases
Data Processing
Google Cloud Platform
PyTorch
Large Language Models
Scikit Learn
Kubernetes
Information Technology
Data Lineage
Google Cloud Functions
GraphQL
Machine Learning Operations
Api Design
Software Version Control
Data Pipelines
GXP
Docker
Microservices

Requirements

We are seeking a seasoned Senior Data Scientist with overall 10-12 years and at least 5-7 years of hands-on experience in developing GenAI/machine learning models and deploying them in a cloud environment, preferably on Google Cloud Platform (Google Cloud Platform).

The ideal candidate will design microservice-based solutions, containerize deployments (e.g., GKE), and drive end-to-end SDLC practices. Experience in the pharma domain is a strong advantage.

Required Qualifications

Overall 10-12yrs and Minimum 5-7 years of hands-on experience developing GenAI/ML models and deploying them in a cloud environment.

Proficiency with Google Cloud Platform (Google Cloud Platform) and its AI/ML offerings (e.g., Vertex AI, BigQuery, Dataflow, Cloud Storage, Pub/Sub, Cloud Run, GKE).

Must have experience working with any agentic framework Knowledge of Retrieval-Augmented Generation (RAG) concepts and processes

Strong software engineering skills: Python (primary), experience with ML frameworks (TensorFlow, PyTorch, scikit-learn), and API development (REST/GraphQL).

Experience designing and deploying microservices architectures and containerized solutions (Docker, Kubernetes; preference for GKE).

Solid experience in MLOps: model versioning, experiments, automated training, feature stores, model registries, monitoring, and governance.

Data processing and analytics expertise: SQL, data pipelines, ETL/ELT concepts, data quality, and data visualization support.

Excellent problem-solving, communication, and collaboration skills; ability to work with cross-disciplinary teams.

Understanding of cloud security concepts, IAM, and basic principles of data privacy and compliance.

Demonstrated ability to translate business problems into scalable ML solutions and to communicate technical concepts to non-technical stakeholders.

Preferred Qualifications

Experience in the pharmaceutical/pharma domain or regulated industries; familiarity with GxP, or similar data governance requirements.

Exposure to other cloud providers (AWS/Azure) is a plus, but a strong preference for Google Cloud Platform.

Experience with distributed training, large-scale data processing, and fine-tuning of large language models.

Knowledge of privacy-preserving ML methods (differential privacy, synthetic data) and data lineage tools.

Education:

  • Minimum qualification: Graduate degree in Information Technology.
  • Preferred: Higher education (e.g., Master's degree in Computer Science, Information Technology, Data Science, or a related field) or relevant professional degrees/certifications.

Apply for this position