> Markdown version of [/jobs/ext/2702803-software-engineer-artificial-intelligence-llm](https://www.wearedevelopers.com/jobs/ext/2702803-software-engineer-artificial-intelligence-llm). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer, Artificial Intelligence/LLM - **Company:** Beacons AI Inc. - **Location:** San Carlos, CA, United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** A/B Testing, Amazon Web Services, Amazon S3, Continuous Integration, Software Debugging, DevOps, Amazon DynamoDB, Search Technologies, Management of Software Versions, Retrieval-Augmented Generation, Large Language Models, TensorRT, Serverless Computing - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/software-engineer-artificial-intelligence-llm-multiple-seniority-levels-beacon-ai-8838969 ## About the Role * Experience with Bedrock, OpenSearch Serverless, pgvector, Pinecone, or Weaviate. * Prompt versioning, guardrails, and provider routing in production. * Multimodal work with time series or video. * Familiarity with GPU inference, Triton, or TensorRT-LLM. * Aviation or other safety-critical domain exposure. * DevOps basics for CI/CD, IaC, and secure secrets handling. ## Description We're hiring across levels. Senior engineers own features and services. Staff engineers own systems, standards, and cross-team technical direction., * Ship APIs and workers in Python or TypeScript with clear contracts, streaming, and backoff. * Add caching, request shaping, prompt templates, and context packing to control latency and cost. * Integrate with AWS Bedrock, OpenAI, Anthropic, or self-hosted endpoints as needed. Retrieval and data prep * Collaborate with infrastructure teammates to develop chunking, embeddings, and indexing capabilities for documents, time series, and multimedia. * Choose and tune vector backends such as OpenSearch, pgvector, or Pinecone. * Keep knowledge bases fresh with data syncs from S3, Aurora, DynamoDB, and external sources. Evaluation and quality * Create offline evals and golden sets for prompts, retrievers, and tools. * Stand up online metrics for task success, hallucination rate, retrieval precision/recall, p95 latency, and cost per request. * Run A/B tests and prompt/version rollouts with guardrails and canaries. Safety, privacy, and compliance * Implement content and policy checks, PII detection and redaction, access controls, and auditing. * Design human-in-the-loop paths for sensitive actions. * Handle aviation data with care and follow internal security standards. Operate what you build * Add tracing, logs, and dashboards for model calls, token usage, errors, and saturation. * Debug tricky failures across retrieval, prompts, tools, and providers., * Strong builder: Comfortable writing production code, tests, and docs. You keep things simple and observable. * RAG and tools depth: You understand embeddings, chunking, vector search tradeoffs, and function calling. * Quality mindset: You design evals, define success metrics, and iterate based on evidence. * Cost and latency aware: You track p95, hit SLAs, and reduce cost without hurting quality. * Clear communicator: You explain tradeoffs and align partners across product, infra, and security., * Transform an internal knowledge base into a low-latency RAG service, complete with explicit schemas and evaluations. * Add tool-calling to automate a repetitive cockpit or ops workflow with guardrails and audit trails. * Reduce the cost per request through improved chunking, caching, and prompt refactoring, while maintaining task success rates. ## Related Videos - [LLMs in the wild: Building an AI agent that survives production](https://www.wearedevelopers.com/videos/100319-llms-in-the-wild-building-an-ai-agent-that-survives-production) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Tour de Force: Open-Source LLM Inference Optimization from Simple to Sophisticated](https://www.wearedevelopers.com/videos/100099-tour-de-force-open-source-llm-inference-optimization-from-simple-to-sophisticated) - [WeAreDevelopers LIVE - CSS is DOOMed](https://www.wearedevelopers.com/videos/1838-wearedevelopers-live-css-is-doomed) - [Efficient deployment and inference of GPU-accelerated LLMs​](https://www.wearedevelopers.com/videos/929-efficient-deployment-and-inference-of-gpu-accelerated-llms) - [You are not an AI developer](https://www.wearedevelopers.com/videos/1148-you-are-not-an-ai-developer) ## Related Articles - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 132 - Binging WADFlix?](https://www.wearedevelopers.com/magazine/473-dev-digest-132-binging-wadflix) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Dev Digest 129 - Now that's what I call private data!](https://www.wearedevelopers.com/magazine/468-dev-digest-129-now-that-s-what-i-call-private-data)