Principal Technical Program Manager (TPM) - AI Infrastructure Operations

External

Nscaleoperationsukltd · US

Full-timeOn-site1w ago

AgileCapacity PlanningCI/CDLeadershipLinuxRisk Management

Cover Letter Connect

Prepare for this interview

Elite

AI-generated questions, company research, and talking points tailored to this role

Responsibilities

Program Leadership: Own the planning, execution, and delivery of strategic operational programs, including new data center AI infrastructure build-outs, large-scale fleet software/firmware rollouts, and the implementation of new operational tooling (in partnership with SRE).
Process Engineering: Analyze and optimize operational workflows across Fleet Operations, Network Operations, and SRE teams. Drive the standardization of incident management, change management, and postmortem processes to reduce toil and improve Mean Time to Recovery (MTTR).
Cross-Functional Coordination: Serve as the primary liaison between engineering teams (Hardware, Compute Platform, Network), Data Center Operations, and external vendors (GPU, Network hardware). Proactively identify and resolve dependencies, risks, and roadblocks.
Risk Management: Proactively identify technical, schedule, and resource risks related to AI infrastructure scaling and stability. Develop mitigation strategies and communicate impacts clearly to stakeholders.
Required Qualifications
Experience: 5+ years of experience in a Technical Program Management role, successfully driving large-scale, complex infrastructure or software engineering programs.
Technical Domain Knowledge: Strong foundational understanding of data center infrastructure, distributed systems, Linux, and networking concepts.
Program Management Rigor: Proven expertise in modern program management methodologies (Agile, Scrum, PMP certification preferred). Exceptional organizational, communication, and presentation skills.
Metrics-Driven Approach: Demonstrable experience in defining, tracking, and improving system performance based on operational metrics (e.g., Uptime, Availability, MTTR, SLOs/SLIs).
Execution in Ambiguity: Ability to thrive in a fast-paced, high-growth environment, managing multiple priorities and adapting to evolving technical requirements.

Requirements

Direct experience managing programs related to data center infrastructure build-outs and hardware commissioning processes.
Specific domain knowledge of AI/HPC infrastructure , including NVIDIA GPUs, InfiniBand/RDMA networks, and the challenges of tightly-coupled systems.
Experience in a hyperscale or public cloud environment supporting 24/7 mission-critical services.
Familiarity with SRE principles, automation tooling, and continuous integration/continuous deployment (CI/CD) pipelines for infrastructure.
A Bachelor's or Master's degree in a technical field (Computer Science, Engineering, etc.) or equivalent practical experience.
What We Can Offer You
At Nscale, you'll find a collaborative, supportive, and innovative environment where your contributions spark real impact. We're building something

Benefits

Health insurance

Additional Information

. About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior results by reducing the complexity of AI development. Our GPU cloud bolsters technical capabilities and directly supports strategic business outcomes, including cost management, rapid innovation, and environmental responsibility. We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you'll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you'll be contributing to building the technology that powers the future. Role Overview As a Technical Program Manager (TPM) for AI Infrastructure Operations, you will be the operational backbone of our high-scale, high-performance AI and High-Performance Computing (HPC) environment. You will be responsible for driving complex, cross-functional programs that ensure the stability, availability, and growth of our cutting-edge GPU fleet and Infiniband network fabrics. This role requires a blend of deep technical understanding, rigorous program management, and a relentless focus on delivering against key operational metrics (SLAs, Uptime, Availability). You will bridge the gap between engineering execution and strategic business goals, directly impacting our ability to serve customer workloads at scale.

Your Match

How well this role fits your profile.

Company Intel

What employees say

Worked at nscaleoperationsukltd? Share your experience

Interested in this role?

Apply on the company's website.

Cover Letter Connect