> Markdown version of [/jobs/ext/2305631-applied-ai-engineer-developer-experience](https://www.wearedevelopers.com/jobs/ext/2305631-applied-ai-engineer-developer-experience). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Applied AI Engineer - Developer Experience - **Company:** Opplane, Inc. - **Location:** Santa Clara, CA, United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Airflow, Amazon Web Services, Data Analysis, Confluence, JIRA, Audit Trail, Code Review, Continuous Integration, Data-Flow Analysis, Python (Programming Language), Program Analysis, Systems Development Life Cycle, Software Engineering, SQL Databases, Strategies of Testing, Large Language Models, Multi-Agent Systems, Model Validation, Gitlab, Git, Pandas, Gitlab-ci, Scikit Learn, Statistics Packages, Operational Systems, Software Coding, Restful APIs - **Published:** August 30, 2026 - **Apply:** https://arc.dev/remote-jobs/j/redirect/pg2lrmjxac ## About the Role * 7+ years spanning software engineering and quantitative analysis. This role needs both; a strong background in one and a passing acquaintance with the other will not carry it. * Production experience with LLM applications: prompting, tool and function calling, context management, evaluation, and knowing where models fail in practice. * Experimental design and causal inference - randomized and quasi-experimental designs, difference-in-differences, instrumental variables, hierarchical models - and the judgment to say when a design does not support the claim being asked of it. * Strong Python and SQL, with a statistical stack (pandas, statsmodels, scikit-learn, or R). Data collection: instrumenting and extracting from operational systems and APIs, designing sampling that survives scrutiny, and knowing when a source cannot answer the question being asked of it. * Aggregation: resolving identity across systems, joining sources never designed to be joined, and modeling the summary tables reporting reads from. You need not own the pipeline, but you must be able to build one when the answer depends on it. * Analytics: exploratory analysis, distributions rather than averages, cohort and time-series work, and reports that state their own coverage and limits. * Real familiarity with the software delivery lifecycle - code review, CI/CD, test strategy, release and change management - sufficient to hold a credible conversation with the teams you are measuring. * Git and GitLab at instrumentation depth: merge request and pipeline data models, diffs and SHAs, what merge, squash, rebase, and cherry-pick do to line-level analysis, and the API and hook surfaces available for capturing it. * Jira and Confluence integration experience - REST APIs, changelog and page version history, the GitLab-Jira development panel, and the field and label conventions that determine whether the resulting data means anything. * Care with personnel-adjacent data: aggregate reporting by default, and a clear sense of what should not be built even when it is technically easy. * Communication that works in both directions - an executive audience that wants a number, and engineers who will dispute it. Preferred * MCP servers and clients, or comparable connector frameworks. * Agent frameworks - LangGraph, LangChain, Bedrock Agents, Strands, or equivalent. * Enterprise deployment of coding assistants, and their telemetry. * Server-side Git hooks, GitLab CI, and system or webhook-driven capture on a self-managed instance. * Confluence and Jira as MCP-connected systems - permission propagation, scoped credentials, and audit logging. Evaluation tooling - Ragas, DeepEval, Bedrock model evaluation - and LLM observability such as LangFuse, Arize, or OpenTelemetry-based tracing. * Amazon Bedrock, and AWS cost and usage data. * Engineering productivity frameworks - DORA, DX Core 4, SPACE - and a working view of their limits. * Program analysis, test generation, or developer tooling research. * dbt, Airflow, Dagster, or equivalent transformation and orchestration; warehouse or lakehouse modeling. * BI and visualization tooling, and the discipline of building on summary tables rather than raw events. * Queueing and flow analysis: utilization, batch economics, and constraint identification. ## Description Opplane is applying AI across the software delivery lifecycle - not only to writing code, but to testing, review, documentation, migration, incident response, and the validation stages where delivery is usually constrained. Measuring what that produces is the starting project, not the whole of it. You will own the analytical and applied half of that work: designing the experiments that establish what actually helps, building the classification and evaluation pipelines the measurement program depends on, and then taking the findings back into the delivery lifecycle as capabilities teams can use. This role pairs with a data engineer who owns extraction, identity, and the metric pipeline. You own what the numbers mean and what to do about them. Key Responsibilities * Design and run experiments. Model tier routing, MCP coverage, permission configuration, repository context quality, budget headroom - randomized across teams and reported with their limits stated. These are the cleanly causal questions available once a tool is deployed, and where the returns are. * Own the analytical layer of the measurement program: work classification over model traffic, evaluation design, longitudinal within-unit analysis, and the staggered-adoption estimates that connect delivery outcomes to adoption timing. * Build and validate LLM-as-judge and classification pipelines - sampling strategy, hand-labeled ground truth, precision and recall measured and published, and revalidation whenever the taxonomy or the model changes. * Extend AI beyond code authoring into the stages that constrain delivery: test authoring and maintenance, environment and data setup, migration and modernization, code review assistance, security remediation, and evidence assembly for certification. * Work directly with the constrained teams. Where validation or certification is the bottleneck, coding assistance produces little regardless of how well it works - find where the constraint actually sits and aim the capability at it. * Build evaluations for internal AI capabilities: golden sets, regression suites, groundedness and answer-quality scoring, and the cost and latency telemetry beside them. * Turn findings into practice. Identify what the most effective practitioners do differently, document it, and teach it - publishing practices rather than rankings. * Partner with the platform team on Claude Code configuration, MCP servers, gateway telemetry, and the model registry, so what you learn becomes the default rather than folklore. * Report to engineering leadership and finance: what is working, what is not, where delivery is actually constrained, and what each claim does and does not establish. ## Related Videos - [Building a Multi-Agent Orchestration Engine That Actually Follows the Rules](https://www.wearedevelopers.com/videos/100159-building-a-multi-agent-orchestration-engine-that-actually-follows-the-rules) - [How a Small Team Shrank a Microsoft Monorepo by 94%](https://www.wearedevelopers.com/videos/1236-how-a-small-team-shrank-a-microsoft-monorepo-by-94) - [Improving quality with Agentic AI with Rovo Dev and Xray](https://www.wearedevelopers.com/videos/2005-improving-quality-with-agentic-ai-with-rovo-dev-and-xray) - [WeAreDevelopers LIVE - Modern DevOps for IoT Devices and More](https://www.wearedevelopers.com/videos/1805-wearedevelopers-live-modern-devops-for-iot-devices-and-more) - [Are Code Reviews Worth It? Insights from 16 Years of Review Data](https://www.wearedevelopers.com/videos/1135-are-code-reviews-worth-it-insights-from-16-years-of-review-data) - [Enabling automated 1-click customer deployments with built-in quality and security](https://www.wearedevelopers.com/videos/83-enabling-automated-1-click-customer-deployments-with-built-in-quality-and-security) ## Related Articles - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [A 5-Step Open-Source Setup for Agentic Engineering](https://www.wearedevelopers.com/magazine/738-a-5-step-open-source-setup-for-agentic-engineering)