> Markdown version of [/jobs/ext/3278957-senior-data-analyst](https://www.wearedevelopers.com/jobs/ext/3278957-senior-data-analyst). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Data Analyst - **Company:** RIGHTSLINE, INC. - **Location:** El Segundo, CA, United States - **Experience:** Expert - **Salary:** $110,000.0 - $120,000.0 - **Contract:** Permanent contract - **Skills:** A/B Testing, Artificial Intelligence, Airflow, Amazon Web Services, Data Analysis, Applications Architecture, Software as a Service, Cursor, Programming Tools, Python (Programming Language), Microsoft SQL Server, NoSQL, Redis, Power BI, Standard Sql, Software Engineering, SQL Databases, Tableau (Software), Management of Software Versions, Data Logging, GitHub Copilot, Retrieval-Augmented Generation, Large Language Models, Git, Pandas, Information Technology, Low Latency, Data Analytics, Machine Learning Operations, Api Gateway, Amazon Simple Queue Service (SQS), Looker Analytics - **Published:** September 30, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=355bd600a0ce4716 ## About the Role * 5+ years in data analytics, analytics engineering, or data science, including hands-on work with LLM-based systems in production * Strong SQL and Python (including pandas and notebook-based analysis), with the discipline to build repeatable pipelines rather than one-off queries * Hands-on experience with LLM observability and evaluation tooling such as LangSmith, Langfuse, LangChain/LangGraph, or equivalent * Practical understanding of LLM application architecture - prompting, retrieval-augmented generation, tool calling, and agent workflows - and how each of them fails * Experience designing evaluation methodology: golden datasets, scoring rubrics, LLM-as-judge grading, human annotation, and A/B testing on production traffic * Working knowledge of experiment design and statistics, enough to say confidently whether a difference is real * Ability to work independently, meet deadlines, and be accountable for your work, while operating within established standards * Excellent written and verbal communication skills, and a strong sense of ownership, urgency, and initiative - you can explain a quality regression to an engineer and to an executive in the same afternoon * BS in Computer Science, Statistics, Data Science, or equivalent knowledge, * Experience with AWS services such as Lambda, SQS/SNS, API Gateway, Step Functions, and Amazon Bedrock * Familiarity with dbt, Airflow or a similar orchestrator, Git, and BI tools such as Looker, Power BI, or Tableau * Working knowledge of MS SQL, Redis or another NoSQL store, and vector stores used for retrieval * Exposure to model providers and platforms (Anthropic, OpenAI, Amazon Bedrock), guardrail and PII-redaction tooling, and prompt versioning and release workflows * Experience with coding agents or AI developer tools (Claude Code, GitHub Copilot, Cursor, or similar) in day-to-day analysis work ## Description * Instrument Rightsline's LLM and agent workflows end to end with tracing and structured logging, using tools such as LangSmith, Langfuse, or equivalent * Define and maintain the metrics that describe AI system health - latency, token and cost per request, error and fallback rates, tool-call success, and retrieval hit rates * Build and maintain dashboards and alerting so quality and cost regressions surface in hours rather than in customer tickets Evaluation and Quality * Design and maintain offline evaluation suites - golden datasets, scoring rubrics, and LLM-as-judge grading - for prompts, retrieval pipelines, and agent flows * Run online evaluations and A/B experiments on production traffic, and report results with enough rigor to support a ship / no-ship decision * Track quality across model, prompt, and retrieval changes, and maintain the benchmark history that shows whether the product is actually improving * Partner with engineers on human review and annotation workflows, including labeling guidelines and inter-rater agreement Data Analysis and Reporting * Build and maintain the pipelines that move trace, evaluation, and usage telemetry into the analytics stack using SQL and Python * Analyze how customers actually use AI features - adoption, drop-off, and failure patterns - and translate findings into prioritized recommendations * Produce recurring reporting on AI quality, cost, and usage for engineering, product, and leadership audiences Cost, Reliability, and Troubleshooting * Monitor token spend and model usage, and identify where prompt, model, or caching changes cut cost without hurting quality * Investigate quality incidents - bad or unsafe outputs, hallucinations, tool failures, latency spikes - from trace to root cause * Follow and help extend Rightsline's development best practices and standards in the pipelines, queries, and evaluation code you ship Team Contribution * Attend and contribute to team meetings, and present telemetry and evaluation findings in a way non-specialists can act on * Raise the team's LLMOps literacy - help engineers instrument their own features, write their own evals, and read the dashboards, At Rightsline, we look for these qualities in every hire, and we hold ourselves to them every day. * Be Curious and Smart Being smart isn't about having all the answers - it's about asking the right questions, seeking out new ideas, and never assuming you already know enough., Ideas are easy. Execution is the job. We hold ourselves to what we commit to, push through the hard parts, and measure success by our actual accomplishments. * Be Humble Nobody here has a monopoly on good ideas. We listen more than we talk, share credit freely, and check our egos at the door. The best outcome matters more than whose idea it was. ## Related Videos - [Leveraging Real time data in FSIs](https://www.wearedevelopers.com/videos/806-leveraging-real-time-data-in-fsis) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Reducing LLM Calls with Vector Search Patterns - Raphael De Lio (Redis)](https://www.wearedevelopers.com/videos/1714-reducing-llm-calls-with-vector-search-patterns-raphael-de-lio-redis) - [Why LLMs Need Observability and How to Do It](https://www.wearedevelopers.com/videos/2117-why-llms-need-observability-and-how-to-do-it) - [Git for Code Reviews](https://www.wearedevelopers.com/videos/429-git-for-code-reviews) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this)