> Markdown version of [/jobs/ext/2710207-ai-engineer](https://www.wearedevelopers.com/jobs/ext/2710207-ai-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # AI Engineer - **Company:** Postman, Inc. - **Location:** United States - **Contract:** Internship / Graduate position - **Skills:** Application Programming Interfaces (APIs), Cloud Computing, Continuous Integration, Data Structures, Cursor (Graphical User Interface Elements), Github, Python (Programming Language), Machine Learning, Open Source Technology, Software Safety, Smoke Testing, TypeScript, Web Services, Graphics Processing Unit (GPU), Postman, Pytorch, Large Language Models, Deep Learning, Information Technology, Virtual Agents, Docker - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/ai-engineer-internship-summer-2026-applications-open-now-postman-8848434 ## About the Role * Currently pursuing a BS, MS, or PhD in Computer Science, Data Science, or a related quantitative field. * Hands-on experience training or evaluating ML models: course projects, research, hackathons, or a prior internship all count. * Solid Python fundamentals: data structures, functions, basic testing; comfortable writing and reviewing code outside of notebooks. * Working knowledge of at least one deep-learning framework (PyTorch preferred). * Clear written and verbal communication, and a habit of documenting what you build., * Experience fine-tuning open-weight LLMs (SFT, LoRA, RL, or distillation), with the improvement measured on a benchmark. * Experience building LLM agents (tool calling, multi-step loops) or LLM evaluation harnesses/benchmarks, and reporting results with statistical rigor. * A track record of shipping real software end-to-end: APIs and services, CLIs, Docker, CI/CD, cloud; public code on GitHub is a big plus. * Interest or experience in AI safety and robustness: red-teaming, prompt injection, agent security, fairness, or interpretability. * Exposure to model-efficiency work: quantization, low-bit inference, or serving optimization. * Evidence of rigor and initiative: publications, technical blog posts, ablation studies, or self-driven side projects with quantified results. * Fluency with AI coding tools (Claude Code, Cursor, Codex) to ship fast while still deeply understanding the systems you build. ## Description You'll work directly with the AI team, taking responsibility for well-scoped pieces of real systems, with mentorship from senior engineers. Benchmarks & Evaluation * Contribute to APIFlow-Bench, our open-source benchmark for real API-development work: design and review benchmark tasks and their mock API environments, extend the evaluation harness and task-generation pipeline in Python, and help maintain the public multi-model leaderboard with statistical confidence intervals. * Help build a new action-level AI safety benchmark: instead of grading what a model says, it scores what an agent actually does inside a simulated enterprise API environment. You'll work on scenario design, threat modeling (prompt injection, data exfiltration, permission overreach), and auditable evaluation design. Model Training & Efficiency * Fine-tune open-weight models for tool calling and agentic tasks (SFT, distillation, and RL) using PyTorch and the open-source training ecosystem, on both managed training platforms and self-managed cloud GPUs. * Design and run experiments with rigor: evaluate every training run on our benchmarks, support ablation studies and error analysis, track experiments, and report results honestly, including cost. * Evaluate ultra-low-bit quantized models for on-device use: extend our quantized vs. full-precision benchmark comparisons and analyze where and why they diverge. Agent Systems & Engineering Practice * Help build the next generation of Postman's in-product AI agent (Agent Mode): a deliberately minimal agent architecture that calls LLM APIs directly (tool loops, multi-step execution, checkpointing), primarily in TypeScript. No prior TypeScript is required; strong Python fundamentals transfer quickly. * Read the source code of open-source agent harnesses and turn what you learn into design specs and prototypes. * Document experiments, design decisions, and runbooks so your work is legible to the next person; flag safety, fairness, or privacy concerns you observe in model or agent behavior. ## Related Videos - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [How AI Models Get Smarter](https://www.wearedevelopers.com/videos/1374-how-ai-models-get-smarter) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) - [One AI API to Power Them All](https://www.wearedevelopers.com/videos/1601-one-ai-api-to-power-them-all) ## Related Articles - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [What is Agentic Programming and Why Should Developers Care?](https://www.wearedevelopers.com/magazine/625-what-is-agentic-programming-and-why-should-developers-care) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this)