Data Scientist - INDIA

VYTWO TECHNOLOGIES INC.
Prosper, TX, United States
6 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence Data Analysis Microsoft Azure Big Data Continuous Integration Extract Transform Load (ETL) Distributed Computing Environment Memory Management Github Monitoring of Systems Machine Learning Tensorflow
+31 more
Azure Machine Learning Search Technologies Unstructured Data Real Time Systems Feature Engineering GitHub Copilot Pytorch Large Language Models Prompt Engineering Deep Learning Containerization Data Lakes Pyspark Git Flow Core Data Kubernetes HuggingFace Production Code Apache Kafka Azure AKS Spark Streaming Data Management Machine Learning Operations Video Streaming Virtual Agents Restful APIs Document Classification Software Version Control Data Pipelines Docker Databricks

Job description

Structured Data - Machine Learning & Analytics Build, deploy, and optimize ML models for predictive analytics, forecasting, classification, and regression. Perform large-scale feature engineering using PySpark and Big Data tools. Work on batch pipelines, model versioning, and experiment tracking. Develop cost estimation and risk/likelihood models using statistical and ML techniques. Text Data / NLP / GenAI Build NLP pipelines using deep learning frameworks such as PyTorch, TensorFlow, or similar. Develop real-time, low-latency inference systems for text classification, embeddings, semantic search, summarization, and retrieval. Create prompts, context graphs, and agentic workflows for LLM-based systems. Apply knowledge of prompt engineering, context engineering, and autonomous agent frameworks to production systems. Core Data Science Engineering & MLOps Work in Databricks for ETL, feature engineering, ML training, and orchestration. Use Azure services for model deployment, data pipelines, and infrastructure. Collaborate using Git-based workflows; leverage tools like GitHub Copilot, Claude Code, etc. Implement model monitoring, observability, drift detection, and performance tracking.

Requirements

We are seeking a highly skilled Data Scientist (3-7 years of experience) to join our team and work across two major data science domains: Structured Data (80-90%) - Predictive analytics, forecasting, cost estimation, likelihood modeling, and batch-oriented machine learning pipelines. Text / Unstructured Data (NLP & GenAI) - Building low-latency real-time systems using deep learning, LLMs, prompt engineering, and agentic AI frameworks. This role requires strong expertise in Big Data processing, modern ML tools, and the ability to build scalable, production-ready data science solutions., Core Skills Strong hands-on experience with Databricks (Delta Lake, MLflow, Job Orchestration). Excellent PySpark skills for large-scale distributed data processing. Proficiency in Azure cloud services (ADF, Azure ML, AKS, Databricks on Azure). Strong understanding of ML algorithms, statistical methods, and data analysis. Experience with deep learning frameworks: PyTorch TensorFlow Transformers (HuggingFace) Experience with model monitoring and ML observability. Ability to write clean, optimized code and leverage AI code assistants. NLP / GenAI Specific Skills Prompt engineering (task prompts, chain of thought, tool calling, retrieval prompts). Context engineering (retrieval pipelines, RAG, memory management, context structuring). Knowledge of LLM-based agentic frameworks (LangChain, Semantic Kernel, CrewAI, AutoGen, etc.). Experience with vector databases and embedding models is a plus. Good to Have Skills Experience with containerization (Docker, Kubernetes, AKS). Experience deploying models to production (REST APIs, real-time endpoints). Knowledge of streaming technologies (Kafka, EventHub, Spark Streaming). Understanding of CI/CD for ML (Azure DevOps / GitHub Actions). Who You Are A problem solver who is comfortable working with both structured and unstructured data. Someone who enjoys using modern AI tools to accelerate development. A data scientist who writes clean, production-grade code. A collaborator who thrives in cross-functional teams and fast-paced environments. Flexible work from home options available.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.wayup.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · WWC 2023

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

3:14 min

Structuring career paths and localized data architectures

Ulrich Wurstbauer +1 · LIVE

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · WWC Europe 2026

1:59 min

Evolving roles in AI driven software teams

Ignacio Riesgo Ignacio Riesgo +1 · WWC 2024

Videos

See all

Related articles

See all