Industrial AI Cloud - Infrastructure Engineer (REF5500F)
ExternalPrepare for this interview
EliteAI-generated questions, company research, and talking points tailored to this role
Responsibilities
- Coordinate Operations with Data Center Teams: Coordinate and support hardware lifecycle activities (installs, GPU upgrades, storage expansion, firmware updates) and manage server/network interconnections and related documentation (NetBox).
- Server & Node Management: Provision and maintain bare-metal servers and GPU nodes (PXE boot, OS installs, firmware updates).
- Design & Operate NVIDIA AI related infrastructure stack
- Automation & IaC: Develop and maintain Ansible and Terraform playbooks to automate provisioning, configuration, and deployments.
- OS & Firmware Management: Maintain Debian-based environments, apply patches, and manage firmware upgrades at scale.
- Identity & Access Management: Integrate and maintain Keycloak, Entra ID / CAIMAN, and AD for user authentication and authorization.
- Run AI & HPC Workloads: Support and operate distributed AI workloads within bare metal hosts and Kubernetes environments.
- Monitoring & Observability: Operate Prometheus and Grafana stacks for proactive infrastructure monitoring and alerting.
- Storage Administration: Manage high-performance storage environments (WEKA by Hitachi).
- ITIL Processes: Follow and improve incident, problem, and change management workflows; document runbooks and standard operating procedures. Adhere to ZERO Outage guidelines.
- Consult and provide project deliverables to fulfil the project scope with focus on Nvidia technology stack.
Benefits
Additional Information
General description/ Purpose NVIDIA and Deutsche Telekom are jointly developing the world's first industrial AI cloud for European manufacturers. This AI factory in Germany will host 10,000 GPUs across NVIDIA DGX B200 systems and RTX Pro Servers. Deutsche Telekom provides secure, sovereign and fast infrastructure, including data centers, operations, security, and AI solutions. Role Overview We are seeking an Infrastructure Engineer to build, automate, and operate compute, network, and storage environment of the Industrial AI Cloud. In this role you will provision and maintain servers, manage networking (on server OS level) and storage, automate deployments, implement monitoring, and ensure reliable day-to-day operations of large-scale GPU clusters. You'll be working and coordinating between multiple teams to deliver and continuously improve infrastructure services following ITIL processes.
Your Match
How well this role fits your profile.
Company Intel
What employees say
Worked at Deutschetelekomitsolutions? Share your experience