Data Scientist / Machine Learning Engineer

aKube Inc
Las Vegas, NV, United States
about 2 months ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
4 years minimum
Compensation
$176,800.0
Working hours
Regular working hours
Job source

Tech stack

Data Deduplication Python (Programming Language) Machine Learning SQL Databases Management of Software Versions Large Language Models Model Validation Pandas Build Management Pyspark Machine Learning Operations Text Analysis
+2 more
Data Pipelines Databricks

Job description

  • Build and deploy NLP classification models for customer communications.
  • Develop intent, topic, sentiment, and multi-label taxonomies.
  • Clean and prepare transcript and message data for modeling.
  • Handle short-text cases, duplicate records, system messages, and speaker identification.
  • Build trend and anomaly detection methods using baselines, seasonality, and channel mix.
  • Design maintainable Python and PySpark data pipelines.
  • Define sampling strategies and annotation guidelines for labeled datasets.
  • Support reviewer adjudication and dataset quality validation.
  • Track model precision, recall, confusion patterns, confidence scores, and drift.
  • Implement secure processing for customer communications containing sensitive data.

Requirements

  • 4-6+ years of data science or machine learning experience
  • NLP classification for customer messages or call transcripts
  • Intent, topic, sentiment, and multi-label classification
  • Confidence scoring and model evaluation
  • Text cleaning, deduplication, speaker handling, and PII-safe processing
  • Trend and anomaly detection
  • Python, PySpark, SQL, and pandas
  • Labeled dataset design and annotation workflows
  • Precision, recall, confusion matrix, and drift monitoring, * 4-6+ years of relevant machine learning, NLP, or data science experience.
  • Proven experience deploying NLP models into production.
  • Strong experience with classification systems and text analytics.
  • Advanced Python development and testing skills.
  • Hands-on experience with PySpark, SQL, pandas, and scalable data pipelines.
  • Experience creating and validating labeled datasets.
  • Strong understanding of model evaluation, monitoring, and false-alert reduction.
  • Experience working with governed or PII-bearing data.

Nice to Have:

  • Databricks
  • Unity Catalog
  • Databricks Workflows
  • MLflow
  • Model and data versioning
  • Retrieval and embedding models
  • LLM-assisted classification with evaluation and guardrails
  • Contact-center or customer-support analytics
  • Property-management or real-estate data experience

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

2:27 min

Managing traffic and tracking costs with Databricks Unity Catalog

Viktoria Semaan Viktoria Semaan · World Congress 2026 Europe

4:32 min

Harnessing Spark with Python using PySpark and Py4J

Ayon Roy · LIVE

2:03 min

Accelerating pandas dataframes using cudf module plugins

Ankit Patel Ankit Patel · World Congress 2024

2:50 min

Executing LoRA fine-tuning using serverless Databricks AI runtimes

Viktoria Semaan Viktoria Semaan · World Congress 2026 Europe

3:37 min

Scaling machine learning pipelines from prototypes to petabytes

Julian Joseph · LIVE

Videos

See all

Related articles

See all