AI Engineer, Search & Knowledge Systems

STONE, RANDY
San Francisco, CA, United States
about 1 month ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Automated Storage and Retrieval Systems Automation of Tests Big Data Cloud Computing Security Static Program Analysis Databases Software Debugging Programming Tools Graph Database Information Retrieval
+26 more
Python (Programming Language) Knowledge Management Knowledge-Based Systems PostgreSQL Metadata Natural Language Processing Named Entity Recognition Search Technologies Software Engineering TypeScript Unstructured Data Workflow Management Systems AI Infrastructure Enterprise Search Datadog Data Ingestion Retrieval-Augmented Generation Large Language Models Topic Modeling Indexer Backend Knowledge Representation Search Engines Front End Software Development Data Pipelines Docker

Job description

We are looking for an AI Engineer specializing in search, retrieval, knowledge systems, and relationship discovery.

You will design and build the systems that help Pi understand and connect security-relevant context across code, pull requests, tickets, documents, incidents, findings, cloud resources, and customer environments. Your work will power the retrieval, grounding, provenance, and relationship modeling behind Pi’s agentic security workflows.

This role is ideal for someone who combines strong software engineering with deep interest in information retrieval, applied AI, knowledge representation, ranking, evaluation, and production systems.

What You’ll Do

  • Build AI-powered search and discovery systems across structured and unstructured security and engineering data.
  • Develop retrieval-augmented generation pipelines using embeddings, hybrid search, reranking, chunking, metadata filtering, grounding, and citation-aware generation.
  • Build knowledge systems that represent entities, relationships, events, decisions, vulnerabilities, controls, code ownership, services, and provenance.
  • Improve relevance, recall, precision, ranking quality, and answer accuracy across search, investigation, and agentic workflows.
  • Design systems for entity extraction, entity resolution, ontology design, relationship inference, and semantic enrichment.
  • Evaluate and combine lexical search, semantic search, hybrid search, graph-based retrieval, and agentic retrieval patterns.
  • Build evaluation frameworks for retrieval quality, hallucination reduction, grounding, freshness, citation accuracy, and user satisfaction.
  • Build ingestion and indexing pipelines that normalize, enrich, connect, and refresh data from multiple customer and product sources.
  • Monitor production AI systems, debug retrieval failures, improve latency, and optimize cost/performance tradeoffs.
  • Partner with product, backend, frontend, platform, and security teams to turn ambiguous customer needs into reliable knowledge systems.
  • Help create the foundation that lets Pi preserve institutional security memory and prevent recurring vulnerability classes., * Build a hybrid search system that combines keyword search, semantic search, metadata filters, and graph traversal.
  • Design a knowledge system that connects repositories, services, pull requests, tickets, findings, vulnerabilities, cloud resources, owners, and decisions.
  • Build a RAG system that produces grounded answers with citations, confidence signals, and traceable source context.
  • Create pipelines for extracting entities and relationships from code, tickets, documents, security findings, logs, and cloud metadata.
  • Develop relevance evaluation datasets and automated tests for retrieval quality, grounding, and answer accuracy.
  • Improve agentic workflows by giving AI systems better context, better retrieval, and better understanding of customer-specific security history.

Success In This Role Looks Like

  • Users can find the right security and engineering context faster and with higher confidence.
  • AI-generated answers are grounded, cited, and reliable.
  • Relationships that were previously hidden across code, tickets, documents, findings, and infrastructure become discoverable and useful.
  • Search relevance, retrieval accuracy, grounding quality, and system latency measurably improve over time.
  • Knowledge systems are maintainable, observable, and extensible as new data sources are added.
  • The product helps customers understand risk, act faster, and prevent the same security issues from recurring.

Requirements

  • Strong software engineering experience in Python, TypeScript, or similar languages.
  • Experience building production search, recommendation, knowledge management, or AI retrieval systems.
  • Hands-on experience with RAG architectures, embedding models, vector search, rerankers, and LLM-backed workflows.
  • Strong understanding of information retrieval concepts such as indexing, ranking, query expansion, relevance scoring, recall/precision, BM25, dense retrieval, and hybrid search.
  • Experience working with structured and unstructured data, including code, documents, tickets, logs, metadata, databases, APIs, and event streams.
  • Experience designing evaluation methods for search relevance, retrieval quality, and AI-generated answers.
  • Ability to build reliable, observable, production-grade systems.
  • Strong product judgment: you can translate ambiguous user needs into practical search, knowledge, and retrieval systems.
  • Strong security instincts around authorization, tenant isolation, data exposure, provenance, and safe handling of customer context.
  • Ability to work in a fast-moving startup environment with ownership, autonomy, and good judgment.

Technologies We Use

  • Python
  • TypeScript
  • Embedding models
  • Rerankers
  • Lexical, semantic, and hybrid search
  • Vector search
  • PostgreSQL
  • Graph-based data modeling
  • Workflow orchestration systems
  • Data ingestion and indexing pipelines
  • Evaluation and observability tooling
  • Docker

Nice To Have

  • Experience with knowledge graphs, graph databases, graph embeddings, ontology design, or taxonomy management.
  • Experience with entity linking, entity resolution, relationship extraction, or semantic enrichment.
  • Experience with LLM orchestration, agentic search, tool use, or multi-step reasoning systems.
  • Experience with NLP techniques such as named entity recognition, classification, summarization, clustering, topic modeling, or semantic similarity.
  • Experience with data pipelines for ingesting, transforming, indexing, and refreshing large datasets.
  • Experience with cloud platforms and production AI infrastructure.
  • Experience with security products, developer tools, code analysis, cloud security, enterprise search, legal tech, finance, healthcare, or research platforms.

Example Projects

About the company

Pi is building an agentic product security platform for teams that need to secure software at the speed they build it.

Modern development is accelerating, but security knowledge is still scattered across code, tickets, documents, incidents, reviews, and the people who remember why decisions were made. Pi turns that context into institutional security memory, helping teams triage faster, remediate in context, prevent repeat vulnerability classes, and embed security guardrails where engineering work already happens.

We are building for a future where security is not a blocker at the end of the development process. It is part of how software gets designed, reviewed, shipped, and improved.

Read about Pi Security on Forbes!

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:37 min

Optimizing technical profiles for AI sourcing and recruitment

Mina Golesorkhi Mina Golesorkhi · WWC Europe 2026

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

1:52 min

Structuring and scaling the backend engineering team

Stefan Lingler Stefan Lingler +1 · Coffee With Developers

2:36 min

Analyzing limitations with PostgreSQL bitmap heap scans

Dharin Shah Dharin Shah · WWC 2025

3:22 min

Creating dedicated AI assistants for advanced candidate sourcing

José Kadlec José Kadlec · WWC 2025

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · WWC Europe 2026

Videos

See all

Related articles

See all