Jobs · Training · California

AI Engineer, Internship - Summer 2026 - Applications Open Now

Find Data Science Jobs · San Francisco Bay Area · 1 wk ago
TrainingVolunteer

About Us

Postman is the world’s leading API platform, used by more than 45 million developers and 500,000 organizations, including 98% of the Fortune 500. We help developers and professionals across the globe build the API-first world by simplifying each step of the API lifecycle and streamlining collaboration—enabling users to create better APIs, faster. Headquartered in San Francisco, Postman has offices in Boston, New York, Austin, Tokyo, London, and Bangalore, where the company was founded. Postman is privately held, with funding from Battery Ventures, BOND, Coatue, CRV, Insight Partners, and Nexus Venture Partners.

About the Role

We're seeking an AI Engineer Intern to work alongside our AI team on large-scale AI and Agentic systems from data pipeline to production deployment. This role is scoped for someone with foundational experience who wants to deepen it: you'll own discrete pieces of real systems under the mentorship of senior engineers, not shadow work or isolated coursework-style projects.

Responsibilities

  • Work directly with the AI team, taking responsibility for well-scoped pieces of real systems, with mentorship from senior engineers.
  • Benchmarks & Evaluation
    • Contribute to APIFlow-Bench, our open-source benchmark for real API-development work: design and review benchmark tasks and their mock API environments, extend the evaluation harness and task-generation pipeline in Python, and help maintain the public multi-model leaderboard with statistical confidence intervals.
    • Help build a new action-level AI safety benchmark: instead of grading what a model says, it scores what an agent actually does inside a simulated enterprise API environment. Work on scenario design, threat modeling (prompt injection, data exfiltration, permission overreach), and auditable evaluation design.
  • Model Training & Efficiency
    • Fine-tune open-weight models for tool calling and agentic tasks (SFT, distillation, and RL) using PyTorch and the open-source training ecosystem, on both managed training platforms and self-managed cloud GPUs.
    • Design and run experiments with rigor: evaluate every training run on our benchmarks, support ablation studies and error analysis, track experiments, and report results honestly, including cost.
    • Evaluate ultra-low-bit quantized models for on-device use: extend our quantized vs. full-precision benchmark comparisons and analyze where and why they diverge.
  • Agent Systems & Engineering Practice
    • Help build the next generation of Postman’s in-product AI agent (Agent Mode): a deliberately minimal agent architecture that calls LLM APIs directly (tool loops, multi-step execution, checkpointing), primarily in TypeScript. No prior TypeScript is required; strong Python fundamentals transfer quickly.
    • Read the source code of open-source agent harnesses and turn what you learn into design specs and prototypes.
    • Document experiments, design decisions, and runbooks so your work is legible to the next person; flag safety, fairness, or privacy concerns you observe in model or agent behavior.

Requirements

  • Currently pursuing a BS, MS, or PhD in Computer Science, Data Science, or a related quantitative field.
  • Hands-on experience training or evaluating ML models: course projects, research, hackathons, or a prior internship all count.
  • Solid Python fundamentals: data structures, functions, basic testing; comfortable writing and reviewing code outside of notebooks.
  • Working knowledge of at least one deep-learning framework (PyTorch preferred).
  • Clear written and verbal communication, and a habit of documenting what you build.

Preferred Qualifications

  • Experience fine-tuning open-weight LLMs (SFT, LoRA, RL, or distillation), with the improvement measured on a benchmark.
  • Experience building LLM agents (tool calling, multi-step loops) or LLM evaluation harnesses/benchmarks, and reporting results with statistical rigor.
  • A track record of shipping real software end-to-end: APIs and services, CLIs, Docker, CI/CD, cloud; public code on GitHub is a big plus.
  • Interest or experience in AI safety and robustness: red-teaming, prompt injection, agent security, fairness, or interpretability.
  • Exposure to model-efficiency work: quantization, low-bit inference, or serving optimization.
  • Evidence of rigor and initiative: publications, technical blog posts, ablation studies, or self-driven side projects with quantified results.
  • Fluency with AI coding tools (Claude Code, Cursor, Codex) to ship fast while still deeply understanding the systems you build.

Benefits

  • Full medical coverage.
  • Flexible PTO.
  • Wellness reimbursement.
  • Monthly lunch stipend.
  • Wellness programs to support physical and mental health.
  • Frequent team-building events.
  • Donation-matching program for causes you care about.

Schedule

This is an in-office role, 5 days a week for all hub locations (San Francisco Bay Area, Boston, Austin, New York City, Tokyo, and London). For Bangalore, employees currently work in the office three days a week and will transition to five days per week by the end of the year.

Values

  • Curiosity: We create with the same curiosity that we see in our users.
  • Transparency: We value transparency and honest communication about not only successes, but also failures.
  • Focus: In our work, we focus on specific goals that add up to a larger vision.
  • Inclusivity: Our inclusive work culture ensures that everyone is valued equally as important pieces of our final product.
  • Dedication: We are dedicated to delivering the best products we can.

Similar jobs