AI/ML Ops Engineer - Data & Reporting

Landmarkit Llc
Austin, TX, United States
21 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Temporary contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
2 years minimum
Working hours
Regular working hours
Job source

Tech stack

A/B Testing Data Analysis Continuous Integration Software Debugging Distributed Systems Github Graph Database Statistical Hypothesis Testing Python (Programming Language) Machine Learning Prometheus Software Engineering
+12 more
TypeScript Datadog Delivery Pipeline Large Language Models Grafana Multi-Agent Systems Build Management Gitlab-ci Machine Learning Operations Virtual Agents Artificial Intelligence Markup Language (AIML) Jenkins

Requirements

5+ years of experience in ML engineering, MLOps, platform engineering, or SRE, including 2+ years working hands-on with LLMs or LLM-powered applications in production.

Demonstrated experience building evaluation systems for ML or LLM applications: test harnesses, benchmark datasets, automated scoring (including LLM-as-judge approaches), and regression detection.

Strong software engineering skills in Python (and ideally TypeScript), with a track record of building reliable, well-tested internal platforms and tooling.

Deep familiarity with CI/CD systems (e.g., GitHub Actions, GitLab CI, Jenkins, Buildkite) and experience embedding automated quality gates into deployment pipelines.

Experience with observability and monitoring stacks (e.g., OpenTelemetry, Datadog, Grafana/Prometheus) and, ideally, LLM-specific observability tools (e.g., LangSmith, Langfuse, Arize Phoenix, Braintrust, W&B Weave).

Proven ability to debug complex distributed systems under pressure, including production incident response, root-cause analysis, and blameless postmortems.

Excellent cross-functional communication: able to translate evaluation results into clear findings and recommendations for both engineers and non-technical stakeholders.

Comfort with ambiguity and a builder s mindset: this role starts with a blank page and ends with the evaluation platform the whole organization relies on.

Experience with agentic frameworks and orchestration patterns (e.g., multi-agent systems, tool use, RAG pipelines) and their distinct failure modes.

Experience with prompt management, model routing, or fine-tuning workflows and evaluating changes across model versions and providers.

Background in statistics or experimentation (A/B testing, significance testing, sampling strategies for human review).

Design and build reusable AI agent skills, plugins and maintain internal marketplace infrastructure to extend and scale Data, AIML capabilities across the organization.

Expertise in causal inference and measurement strategy including causal graphs, ontologies, and knowledge graphs to drive rigorous, decision grade data analysis.

Experience operating in regulated or high-stakes domains where agent errors carry real business or customer impact.

A production-grade evaluation framework is in place, with automated evaluation suites covering every critical agent.

CI/CD quality gates catch agent regressions before release, with clear pass/fail criteria trusted by agent teams.

Production agents have real-time quality monitoring and alerting, with documented runbooks and severity-based escalation paths.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

6:36 min

Funding open source through GitHub Accelerator and Sponsors

Stormy Peters · World Congress 2023

1:02 min

Applying an ETL methodology to infrastructure configuration management

Axel Barbier · World Congress 2023

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

2:08 min

Essential engineering roles in the generative AI space

Mary Grygleski Mary Grygleski · LIVE

2:40 min

Using GitHub primitives for internal documentation and corporate operations

Kyle Daigle · Coffee With Developers

3:22 min

Evaluating advanced artificial intelligence platforms for daily recruitment

Rudi Bauer Rudi Bauer +1 · Cappuccino with HR

Videos

See all

Related articles

See all