> Markdown version of [/jobs/ext/1507355-data-scientist](https://www.wearedevelopers.com/jobs/ext/1507355-data-scientist). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Scientist - **Company:** ScienceLogic, Inc. - **Location:** United States (Remote available) - **Experience:** Experienced - **Salary:** $140,000.0 - $165,000.0 - **Contract:** Permanent contract - **Skills:** Data Analysis, Big Data, Cloud Computing, Information Technology Operations, Python (Programming Language), Standard Sql, Large Language Models, Mttr, Information Technology, ArcSight Event Correlation - **Published:** July 30, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=6a82e9f517c20ef3 ## About the Role * Bachelor's or Master's in Data Science, Computer Science, Statistics, Mathematics, or a related field or equivalent experience. * 3+ years in data science, ML, or applied quantitative analysis. * Strong applied statistics, with the judgment to design sound experiments and significance tests on noisy, non-deterministic outputs (not just clean A/B conversion). * Experience building, deploying, and monitoring predictive or time-series models in production: forecasting, anomaly detection, or trend analysis, including recalibration as data shifts. * Demonstrated work evaluating, analyzing, or improving LLM or NLP systems: eval design, quality measurement, retrieval evaluation, or agent analysis. * Proficiency in Python. * Strong SQL and comfort querying large analytical datasets. * Fluency with foundation models and hands-on experience with the modern LLM evaluation and tooling layer - eval/harness frameworks, judge pipelines, and the libraries used to serve, prompt, and test models. * Ability to build analysis and visualization in code., * Experience getting strong results out of small or self-hosted/local models under compute, memory, or latency constraints - quantization-aware evaluation, prompt and context optimization, or model routing. * Experience with retrieval-augmented systems and retrieval evaluation at scale. * Experience with agentic frameworks and tool-use/orchestration analysis, including human-in-the-loop and replayable-state patterns. * Familiarity with red-teaming or adversarial robustness for LLMs. * Domain background in IT operations - AIOps, NOC, ITSM, observability, or anomaly detection on logs and telemetry. * Experience with large-scale analytical and big-data stores. * Cloud experience for data science and ML workloads. * Exposure to enterprise security and compliance constraints in a delivery context. ## Description Evaluation & Response Quality * Design and own evaluation harnesses for LLM and agentic outputs - golden sets, regression suites, and rubric-based scoring. * Build and calibrate LLM-as-judge pipelines; validate judges against human labels and control for their bias and variance. * Define and track response-quality metrics: faithfulness/groundedness, hallucination rate, answer relevance and completeness, instruction-following, and persona adherence. * Curate, version, and grow evaluation datasets as the product and its surfaces evolve. * Benchmark the models in the suite against each other to decide which model handles which task, and quantify the quality cost of running smaller, local models versus larger alternatives. Adversarial & Robustness Testing * Red-team the system: prompt injection, jailbreaks, tool-misuse, and edge-case discovery. * Design chaos and stress tests that probe model and agent reliability under degraded or hostile conditions. * Characterize failure modes and feed them back into guardrails and regression coverage. Retrieval & Agentic Trajectory Analysis * Evaluate retrieval quality over the document corpus - recall@k, MRR/nDCG, context precision and recall - and run experiments on chunking, indexing, and hybrid retrieval strategies. * Analyze multi-step agent trajectories: tool-call correctness, trajectory efficiency, replayable-state inspection, and guardrail-breach behavior. * Assess intent classification and routing quality as measurable components, not black boxes. Behavioral Regression & Drift * Build standing evaluation that catches quality and behavioral regressions when a model in the suite is swapped, upgraded, or re-quantized, or when prompts and pipelines change. * Monitor output-distribution and quality drift in production; distinguish genuine regressions from noise on stochastic outputs. * Recommend and validate fixes through the levers available with local models - prompt changes, retrieval and grounding adjustments, routing changes, or model selection. Predictive & Trend Modeling * Build, ship, and own production models that forecast and surface trends from operational telemetry - capacity and resource forecasting, anomaly prediction, and early-warning signals on metrics and logs. * Take these from prototype to production and keep them healthy: deployment, monitoring, recalibration, and retraining as data and behavior shift. * Define accuracy and lead-time metrics that matter operationally - precision/recall on predicted incidents, forecast error, how far ahead a signal fires - not just offline scores. * Wire predictive signals into the LLM and agentic layer so forecasts and trends feed reasoning, advisories, and operator-facing recommendations. Domain & Value Analytics * Apply AIOps/NOC analysis where it's the product: log anomaly detection, event correlation, and root-cause and problem analysis. * Quantify the economics of the system - cost and token consumption per interaction, interaction-type taxonomies - and connect them to customer-facing value metrics like MTTR and operator-hours. * Communicate findings to engineering and product stakeholders through clear, in-context analysis. Method & Innovation * Use LLM-assisted workflows to scale the work itself - drafting analyses, generating synthetic evaluation cases, and bootstrapping labeled data for human refinement. * Track and adopt state-of-the-art evaluation, retrieval, and agentic-analysis techniques; bring the useful ones into the team's workflow. ## Related Videos - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [How building an industry DBMS differs from building a research one](https://www.wearedevelopers.com/videos/768-how-building-an-industry-dbms-differs-from-building-a-research-one) - [What Developers Get Wrong About Application Quality](https://www.wearedevelopers.com/videos/233-what-developers-get-wrong-about-application-quality) - [Why and when should we consider Stream Processing frameworks in our solutions](https://www.wearedevelopers.com/videos/1085-why-and-when-should-we-consider-stream-processing-frameworks-in-our-solutions) - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering)