Role details
- Perplexity is looking for a technical program manager to be the connective tissue between our model providers, engineering, and product teams, driving our core inference platform forward.
- Perplexity runs one of the highest-throughput inference stacks in the industry, serving Ask, Computer, and API traffic across a large and constantly shifting portfolio of first-party and third-party models.
- This role sits at the intersection of product, engineering, and finance: you'll orchestrate across model providers and internal teams to keep new models and capacity moving smoothly into production, while executing the roadmap for the inference platform itself.
- The ideal candidate has strong technical judgment, thrives coordinating across teams and external partners with competing timelines, and is energized by building the operating model for a function that doesn't have much precedent yet.
OUR MISSION
Perplexity's mission is to power curiosity. Curious people are the people who drive change in the world. Driving change is a continuous cycle of learning, building, and integrating.
Repeat. For curious people this is a cycle that never ends.
WHAT YOU'LL DO
- Execute the roadmap for the inference platform — request handling, rate limits and quotas, usage controls, and the reliability and observability surface engineering and product teams depend on
- Be the connective tissue between model providers and Perplexity's engineering and product teams — coordinating onboarding, launch readiness, and rollout for new models and capacity
- Drive latency, throughput, uptime, and cost-efficiency as core execution metrics, surfacing tradeoffs between them rather than letting them become side effects
- Run the operating model for model-release and optimization programs, including day-zero launches, across performance engineering, infrastructure, and product teams
- Lead cross-functional delivery for inference-stack changes, from planning through launch and post-launch validation
- Build the mechanisms that make releases predictable — rituals, dashboards, launch checklists — so inference releases stay low-risk at Perplexity's scale
- Partner with GPU capacity and compute teams to reconcile execution decisions against cost, capacity, and vendor constraints
QUALIFICATIONS
- Strong experience with technical program management or product management in infrastructure, distributed systems, or ML/model-serving products
- Direct experience with production LLM or ML inference — understanding what makes serving fast, reliable, and cheap rather than just what a roadmap slide says about it
- Comfort orchestrating across external partners and internal engineering teams with competing priorities and timelines
- Experience with data and metrics, and the judgment to surface difficult tradeoffs between latency, throughput, uptime, and cost
- Thrives in a small, agile team; has initiative and desire for ownership without much precedent to lean on
- 6+ years of combined technical program management or product management experience