Member of Technical Staff, Performance & Capacity
Physical Superintelligence is a startup with roots at Google, NVIDIA, Harvard, Meta, MIT, Oxford, Johns Hopkins, Cambridge, and the Perimeter Institute building AI systems to discover new physics at scale. We are seeking engineers to build platform infrastructure at the intersection of computational science, AI systems, and software engineering. Our mission is to discover and commercialize transformative physics breakthroughs at scale with artificial superintelligence, safely, verifiably, and for broad public benefit.
Role and Responsibilities
- Measure the fleet instead of estimating it—utilization, goodput, and training efficiency taken on real workloads so capacity decisions rest on observed numbers.
- Identify and eliminate platform capacity waste.
- Own the capacity model: convert research workloads into defensible node counts with stated assumptions and error bars, and defend them when real money is committed.
- Own the workload mix over time: document fleet run patterns at given hours, preemptible fraction as a measured quantity, and checkpointing economics.
- Maintain a fast data path: parallel file systems, data locality feeding GPUs during training, and interconnect behavior at multi‑node scale.
- Find where time actually goes, not just what a profiler summary reports.
- Own the model and the measurements feeding it.
What We're Looking For
- Five or more years with GPU and large‑scale compute workloads, including real multi‑node experience in distributed training performance, interconnect and collective‑communication behavior, and the patience to find where scale breaks down.
- Deep knowledge of GPU performance characteristics at the system level: which workloads justify an H100 versus a B200, how to optimize across clusters rather than within one, and where the bottleneck actually sits.
- Profound understanding of AI training and inference performance: step time, MFU, communication patterns, bandwidth‑bound decoding, batching benefits, and latency floors.
- Ability to locate bottlenecks in either regime and quantify their cost.
- Experience owning a capacity decision with money attached and being held to the resulting numbers; ability to explain assumptions that proved wrong and the changes made.
- Knowledge of parallel file systems and their impact on training throughput.
- Experience with InfiniBand or RDMA‑class networking behavior.
- Skill in explaining performance results to researchers without requiring them to learn internals.
- Deliverable numbers that are usable by people who never open a profiler.
Nice to Have
- Time on a data center floor: thermal and power envelopes, physical failure domains, and capacity planning against real hardware rather than a console.
- Energy‑ or preemption‑aware scheduling, treating power, time of day, or interruptible capacity as real variables in workload placement.
- Capacity sourcing across cloud and specialist GPU providers, and the economics of moving workloads between them.
- Background in scientific computing or HPC environments where efficiency was measured because someone paid for the machine.
How We Work
We hold a high technical bar and give people full ownership of their work, from spec to ship to on‑call. We write contracts before logic, test against real systems instead of mocks, and favor simple designs that ship over clever ones that do not. Our development process is AI‑native: we work with agentic coding tools daily, write specs that are legible to humans and agents alike, and lead with leverage.
Location and Compensation
This role is based in Boston. We will consider remote candidates on a case‑by‑case basis. We offer competitive compensation including salary, benefits, and meaningful early‑stage equity. Evaluation is based on technical breadth, systems thinking, scientific curiosity, and shipping velocity. We are an equal‑opportunity employer and value diverse perspectives in building platforms for AI‑driven discovery.