The posting, in OpenAI's own words
archived Sep 17, 2026About the Team
The Product & Platform teams at OpenAI are responsible for delivering the company’s most impactful offerings—such as ChatGPT, our API platform, and new enterprise capabilities—to a global and diverse customer base. These systems must perform at scale and deliver exceptional experiences to developers, consumers, and businesses alike. The ChatGPT Multimodal team works across voice, image generation, and other multimodal experiences to turn frontier research capabilities into reliable products. The team connects product usage and failure patterns with research, evaluation, data, inference, capacity, and external partnerships so that model and product improvements translate into better experiences for users.
About the Role
We are seeking a Technical Program Manager to build the flywheel that helps ChatGPT multimodal products learn from real-world usage and improve quickly. You will lead programs spanning production-signal mining, evaluation and data pipelines, research-to-production parity, multimodal capacity planning, and complex cross-functional dependencies for voice and image-generation launches. You will work closely with product engineering, research, Human Data, inference and capacity teams, safety partners, and external vendors or product partners. Success requires technical depth, strong systems thinking, comfort with ambiguity, and the ability to turn fragmented or manual work into durable mechanisms that teams adopt. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees.
Read the full posting ↓
In this role, you will:
Build a system for mining production conversations and product signals to identify representative multimodal workflows, user needs, and failure modes. Establish and maintain evaluations for the highest-priority multimodal behaviors and use cases, with clear coverage, quality standards, and ownership. Package production signals into decision-ready data and evaluations that research teams can use to improve model behavior. Measure whether model, prompt, configuration, and product changes produce meaningful improvements in multimodal evaluations and user outcomes. Close gaps between research and production environments, including system prompts, sampling behavior, multimodal configurations, inference differences, and other sources of parity drift. Create a repeatable process for reproducing product failures with research partners and validating fixes in the shipped experience. Lead multimodal capacity planning by forecasting demand, translating it into GPU and serving needs, and managing headroom and reallocation tradeoffs for voice and image-generation workloads. Improve the tooling and operating processes used to plan, launch, and operate multimodal capabilities as demand and model behavior evolve. Coordinate targeted multilingual data collection across research, Human Data, and external vendors. Drive cross-functional programs that multimodal launches depend on, including multimodal actor recruitment and selection and voice-related product partnerships across vehicles, smart speakers, and headphone ecosystems. Create clear operating cadences, decision rights, metrics, risk management, and executive-ready communication across complex, time-sensitive programs.