Software Engineer, Artificial Intelligence/LLM

Beacons AI Inc.
San Carlos, CA, United States
10 days ago
Apply on startup.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Working hours
Regular working hours
Job source

Tech stack

A/B Testing Amazon Web Services Amazon S3 Continuous Integration Software Debugging DevOps Amazon DynamoDB Search Technologies Management of Software Versions Retrieval-Augmented Generation Large Language Models TensorRT
+1 more
Serverless Computing

Job description

We’re hiring across levels. Senior engineers own features and services. Staff engineers own systems, standards, and cross-team technical direction., * Ship APIs and workers in Python or TypeScript with clear contracts, streaming, and backoff.

  • Add caching, request shaping, prompt templates, and context packing to control latency and cost.
  • Integrate with AWS Bedrock, OpenAI, Anthropic, or self-hosted endpoints as needed.

Retrieval and data prep

  • Collaborate with infrastructure teammates to develop chunking, embeddings, and indexing capabilities for documents, time series, and multimedia.
  • Choose and tune vector backends such as OpenSearch, pgvector, or Pinecone.
  • Keep knowledge bases fresh with data syncs from S3, Aurora, DynamoDB, and external sources.

Evaluation and quality

  • Create offline evals and golden sets for prompts, retrievers, and tools.
  • Stand up online metrics for task success, hallucination rate, retrieval precision/recall, p95 latency, and cost per request.
  • Run A/B tests and prompt/version rollouts with guardrails and canaries.

Safety, privacy, and compliance

  • Implement content and policy checks, PII detection and redaction, access controls, and auditing.
  • Design human-in-the-loop paths for sensitive actions.
  • Handle aviation data with care and follow internal security standards.

Operate what you build

  • Add tracing, logs, and dashboards for model calls, token usage, errors, and saturation.
  • Debug tricky failures across retrieval, prompts, tools, and providers., * Strong builder: Comfortable writing production code, tests, and docs. You keep things simple and observable.
  • RAG and tools depth: You understand embeddings, chunking, vector search tradeoffs, and function calling.
  • Quality mindset: You design evals, define success metrics, and iterate based on evidence.
  • Cost and latency aware: You track p95, hit SLAs, and reduce cost without hurting quality.
  • Clear communicator: You explain tradeoffs and align partners across product, infra, and security., * Transform an internal knowledge base into a low-latency RAG service, complete with explicit schemas and evaluations.
  • Add tool-calling to automate a repetitive cockpit or ops workflow with guardrails and audit trails.
  • Reduce the cost per request through improved chunking, caching, and prompt refactoring, while maintaining task success rates.

Requirements

  • Experience with Bedrock, OpenSearch Serverless, pgvector, Pinecone, or Weaviate.
  • Prompt versioning, guardrails, and provider routing in production.
  • Multimodal work with time series or video.
  • Familiarity with GPU inference, Triton, or TensorRT-LLM.
  • Aviation or other safety-critical domain exposure.
  • DevOps basics for CI/CD, IaC, and secure secrets handling.

Benefits & conditions

Perks & Benefits (Full-Time Employees)

  • Healthcare: 100%* of employee medical premiums covered; 25% for dependents
  • Time Off: 3 weeks PTO plus 13+ paid company holidays
  • Stipends: Monthly phone and wellness benefits
  • 401(k): Offered (no current employer match, but we are committed to enhancing this benefit in the future).

Due to U.S. export control regulations, we can only hire U.S. Persons (U.S. citizens, Green Card holders, lawful permanent residents, or individuals granted asylum or refugee status). We are unable to provide visa sponsorship or support visa transfers. All work must be performed in the United States.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on startup.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:43 min

The enduring legacy of the amazon S3 storage API

Chris Heilmann +3 · LIVE

4:52 min

Essential phases in building and refining language models

Anshul Jindal Anshul Jindal +1 · World Congress 2025

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

2:14 min

Exploring internal AI product initiatives and global engineering roles

Maria Apazoglou · Coffee With Developers

3:44 min

Automating storage savings with S3 intelligent tiering

Sébastien Stormacq · World Congress 2021

2:32 min

Core libraries driving inference engines and multi-GPU networking

Adolf Hohl Adolf Hohl · World Congress 2024

Videos

See all

Related articles

See all