Principal Networking AI Systems Architect
We are seeking a highly skilled Principal Networking AI System Architect to join as a key contributor to the Applied Networking AI group. In this role you will scope and lead AI based solutions for networking technologies and drive their integration across teams.You’ll lead a portfolio and roadmap of projects that encompass agentic-AI for data-center management, predictive-resiliency, optimization and more. By collaborating closely with subject-matter-experts (SMEs), applied-researchers, product-managers, architects, data-engineers and other stakeholders you will push the envelope forward in using cutting-edge technologies and data-driven insights to improve NVIDIA's products.
What you'll be doing
Build a shared roadmap and vision for AI based data-center management solutions spanning LLM intelligence for troubleshooting, predictive-resiliency and AIOPS, black-box optimization and performance tuning. Work closely with engineering and reliability teams to scope and define workflows utilizing and benefitting from AI/ML. Drive the integration of AI capabilities into system architecture and engineering workflows. Identify system-level opportunities for failure management, automated troubleshooting, performance improvement, and resource optimization. Translate system behavior, dependencies, data, and operational constraints into formulated research problems.
What we need to see
Ph.D in electrical engineering, machine-learning, computer-science or another relevant field. 10+ years of deep technical experience in high-performance network architecture, data center networking, or distributed systems design. Mastery of high-speed interconnect protocols including InfiniBand and/or advanced Ethernet architectures. Thorough experience driving high-impact projects centered on modern AI/ML such as LLMs/agents, deep-learning, black-box optimization or another relevant field. Deep knowledge of AI/ML and networking-hardware/system-architecture. Excellent ability to convey and communicate data-based insights to stakeholders and management. Experience demonstrating an excellent track of collaboration with hands-on teams.
Ways to stand out from the crowd
Demonstrated track record of architecting and deploying multi-thousand-node GPU clusters for hyperscale cloud environments. Deep knowledge of NVIDIA networking technologies, including BlueField DPUs, Quantum InfiniBand switches, and Spectrum Ethernet platforms. Expertise in in-network computing, telemetry, adaptive routing, and telemetry-driven network optimization. High energy and a positive, proactive and curious approach.
Every role in the library, ranked against your CV.
Rewritten from your real, matching experience for the job you pick.
Your whole pipeline in one place, with a fresh move every morning.