> Markdown version of [/jobs/ext/3608665-agentic-ai-analyst](https://www.wearedevelopers.com/jobs/ext/3608665-agentic-ai-analyst). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Agentic AI Analyst - **Company:** Invoca, Inc. - **Location:** New York, NY, United States - **Experience:** Experienced - **Salary:** $94,000.0 - $134,000.0 - **Contract:** Permanent contract - **Skills:** LangGraph Framework, Application Programming Interfaces (APIs), Google AdWords, Artificial Intelligence, Computer Vision, Software as a Service, Customer Data Management, Data Files, JSON, Python (Programming Language), Automation of Marketing, Regression Analysis, Operational Databases, Performance Tuning, Systems Integration, Design of Telemetry/logging Agents, Enterprise Software Applications, Retrieval-Augmented Generation, Large Language Models, Multi-Agent Systems, Agentic-AI, Data Analytics, Data Management, Webhooks, Call Tracing, Agent2Agent Protocol - **Published:** October 7, 2026 - **Apply:** https://diversityjobs.com/career/18552998/Agentic-Ai-Analyst-New-York-New-York ## About the Role * 3+ years of experience in product, analytics, or technical customer-facing roles within SaaS or enterprise software. * Hands-on experience with agentic AI systems in production environments, including multi-step reasoning, tool use, conversational agents and human-in-the-loop agents. * Experience designing or integrating tools for AI agents, such as building or configuring MCP servers, defining function/tool schemas, or connecting agents to APIs and enterprise systems. * Experience with agent design and architecture, authoring prompts, skills, or agent instructions and iterating on them based on measured results. * Experience designing and running evals for Agentic systems, including building golden data sets, defining success criteria, and building scorers. * Experience with simulation or synthetic testing approaches for agents (e.g., synthetic users, mocked tools, scenario generation). * Experience with AI orchestration frameworks (e.g., LangChain, LangGraph). * Demonstrated ability to work independently on medium-scope initiatives and balance tradeoffs across requirements, user value, and implementation effort. * Skilled at breaking down ambiguous problems, gathering data, and proposing thoughtful, data-driven solutions. * Strong communicator who can lead and coordinate messaging across technical and non-technical stakeholders. * Highly collaborative and comfortable working across Engineering, GTM, and Design partners. * Familiarity with agent-to-agent (A2A) protocols or multi-agent orchestration patterns. * Experience designing for long-running or stateful agent sessions (context persistence, session resumption, drift over time). Strong Plus * Working comfort with Python, JSON/JSON Schema, and reading API documentation, enough to prototype and inspect tool calls and traces. * Prior experience building or managing integrations between SaaS platforms (e.g., CRM, marketing automation, contact center, or data platforms) via APIs, webhooks, or iPaaS tools. * Experience in marketing or advertising technology, such as paid media, call tracking, attribution, marketing analytics, or ad platforms (Google Ads, Meta, etc.). ## Description As an Agentic AI Analyst at Invoca, you will play a pivotal role in shaping and refining the emerging category of agentic AI applications. Within our Product organization, you'll collaborate closely with a cross-functional team, including Product Managers, Engineers, Data Scientists, and Customer Success, to identify, evaluate, and operationalize agentic use cases that directly improve customer outcomes and advance Invoca's strategic AI vision. A core focus of this role is the integration layer that makes agents useful in the real world: designing and validating the tools agents call (via MCP and similar protocols), authoring the skills that shape agent behavior, building the evaluation harnesses that measure whether agents actually work, and running simulations that stress-test agents before customers ever touch them. This is an opportunity to grow your product skillset at the intersection of customer needs, agentic AI system design, and real-world AI performance. You will contribute to building new features, evaluating agent performance in production, and distilling insights from field engagements and system data to inform product priorities. You should bring strong problem-solving ability, curiosity around emerging AI capabilities, and confidence working across stakeholders to translate ambiguity into clarity and execution. What You'll Do * Design and Validate Agentic Integrations (MCP Tools): Define the tools our agents use to read and act on customer data. Author tool specifications (names, descriptions, input/output schemas, error behavior), prototype MCP servers and tool integrations, and validate that agents select and invoke tools correctly across realistic scenarios. Partner with Engineering on tool ergonomics that reduce hallucinated calls and improve task completion. * Author and Maintain Agent Skills: Design, write, and maintain reusable skills: structured instructions, workflows, and domain knowledge that govern how agents approach specific jobs. Test skill triggering accuracy, iterate on instructions based on eval results, and manage a versioned library of skills for reuse across products and customers. * Build and Run Agent Evals: Design evaluation frameworks for agentic behaviors, including golden datasets, task-level success criteria, grading rubrics, and LLM-as-judge pipelines. Run evals across model, prompt, tool, and skill changes to catch regressions, quantify improvements, and inform release decisions. Report results in a way that both Engineering and GTM can act on. * Simulate Agent Behavior Before Production: Build and operate simulation environments (synthetic callers, mock tool backends, adversarial and edge-case scenarios) to exercise multi-step agent workflows at scale. Use simulation results to uncover failure modes, calibrate human-in-the-loop handoff points, and establish confidence thresholds before beta and GA. * Design End-to-End Agent Behavior: Own how an agent reasons through a task from start to finish, including planning logic, state management across steps, escalation and human-in-the-loop decision points, and fallback behavior when tools or data are unavailable. * Validate Use Cases with Customers and GTM: Work with Sales Engineering, Customer Success, and customers to surface high-value problems, and mine call transcripts, usage data, and tool-call traces to identify where agents can deliver measurable outcomes. Translate those findings into scoped, experiment-ready proposals. * Monitor Agent Performance in Production: Track and interpret key metrics, such as task completion rates, tool-call success rates, fallback frequency, and customer feedback, to assess agent efficacy and prioritize improvements. * Optimize Retrieval and Model Behavior: Partner with Engineering and ML teams to tune the components behind agent quality, including retrieval and RAG strategies, context management, and model selection, using eval and production data to guide changes. * Triage Production Issues and Regressions: Own the process for capturing, reproducing, and documenting unintended agent behaviors reported from production. Work cross-functionally to root-cause issues, land fixes, and add coverage to evals and simulations so they don't recur. * Document Agent Designs and Best Practices: Own internal documentation for agent designs, tool specifications, skill libraries, eval methodology, and troubleshooting guides so that patterns can be reused and scaled across teams. * Enable Cross-Functional Teams and Customers: Support Product Marketing and Enablement by developing clear collateral and training on agentic AI features. Represent Product in customer meetings, agile ceremonies, and internal demos. * Design Multi-Agent and Long-Running Workflows: Architect how agents hand off tasks to other agents (via A2A or similar protocols) and how they maintain context and state across long-running or multi-session buyer interactions - not just single-turn exchanges.