About the role
Our Host Management Team focuses on AI factory management infrastructure. We define the next-generation management architecture of the data center — top-down, clusters, racks, platforms and multi-SoC trays. The architecture covers end-to-end flows for power-on, reset and boot sequences, power saving, management protocols to keep the system operating securely and reliably.
We are seeking a Senior System Management Architect to join our growing team. You will be the resident expert for system management, ensuring that AI factory management components (BMCs, MCUs, CPUs, network switches and CPLDs) operate in perfect unison. You will own the lifecycle management specifications for AI factory platforms — from factory provisioning to day-one deployment. If you enjoy studying a high-level system diagram and then immediately diving into the system-level buses and interface signals topology and the associated host management protocols modeling required to make it work, this role is for you.
What you'll be doing:
System Architecture & Flow Definition: Define and document comprehensive production, provisioning, update, recovery, and reset flows for AI factory sub-systems.
Ensure system-level architectural alignment and orchestrate seamless coordination between all the AI-Factory building blocks: GPU, DPU, CPU, Switch ASIC, NIC ASIC, BMC, and peripheral components.Protocol Implementation & Guidance: Specify the use of standard management protocols (Red-Fish, NSM, PLDM, SPDM, MCTP, NC-SI) across sub-system interfaces to implement the management tasks.
Hardware/Software Intersect: Define system-level hardware and software behavior, including CPLD logic requirements, power-sequencing dependencies, and reset orchestration (e.g., handling PCIe PERST#, SBR, and forced resets).
Interface Management: Architect intra-board communication pathways using PCIe, I2C/I3C, USB, and SPI, ensuring robust out-of-band (OOB) and in-band management connectivity.
Cross-Functional Leadership: Collaborate with Production, Hardware and Firmware Engineering, ASIC FW/Driver teams, SW/NOS teams and QA to translate architectural specifications into implementable, testable solutions.
Technical Documentation: Author rigorous technical specifications, test plans, and architectural design documents to guide development and factory production.
What we need to see:
Deep System-Level Perspective: Proven ability to understand complex computing or networking systems top-to-bottom, with a background in networking architectures and efficiency and power saving methods.
Protocol Fluency: Strong working knowledge of modern platform management and security protocols.
Hardware Interface Knowledge: Familiarity with board-level communication buses (I2C, I3C, SPI, USB) and system interconnects (PCIe architecture, including link states and reset mechanisms).
Analytical Documentation Skills: Exceptional ability to write clear, unambiguous technical specifications, state-machine descriptions, and sequence diagrams.
Experience: BS/MS in Electrical Engineering, Computer Engineering, Computer Science, or a related field, with 5+ years of relevant industry experience.
Ways to stand out from the crowd:
Experience with power-saving methods.
Background with Redfish RESTful API definitions and OOB network infrastructure.
Familiarity with hardware RoT concepts and secure boot architectures.
Experience with production line tooling, factory provisioning, and manufacturing test flows.
Background in networking principles, Switch, NIC/SmartNIC data-plane and offload operations, or high-speed data center topologies.
Aplyr's read
NVIDIA is a pioneering force in GPUs and AI, attracting top talent in engineering and innovation-driven roles across various tech domains.
What's promising
- •NVIDIA leads the GPU market, crucial for gaming and AI applications.
- •The company invests heavily in AI and deep learning, driving technological advancements.
- •NVIDIA's strong market position offers stability and growth opportunities for employees.
What to watch
- •High competition in the semiconductor industry can impact market share.
- •Rapid technological changes require constant adaptation and learning.
- •Intense workload and high expectations may affect work-life balance.
Why NVIDIA
- •NVIDIA's GPUs are industry benchmarks in gaming and professional graphics.
- •The company's AI research is at the forefront of deep learning innovation.
- •NVIDIA's culture emphasizes cutting-edge technology and engineering excellence.
Aplyr’s read is generated by AI from public sources. Was it useful?
About NVIDIA
NVIDIA is a leading technology company known for its graphics processing units (GPUs) for gaming and professional markets, as well as its advancements in artificial intelligence and deep learning.
Similar roles
Senior System Management Architect
NVIDIA
Staff Engineer Software (Flight Management System Software Architect)
Northrop Grumman
JBOSS Business Rules Management System (BRMS) Architect Consultant
Prosidian Consulting
Senior System Architect - Client Lifecycle Management (CLM) 100% (f/m/d) - (Contract through our external payroll partner for 12 months with possible extension)
Julius Baer
IT Application Architect, Laboratory Information Management System, LIMS
Edwards Lifesciences
Strategic Training Partner, Senior Manager (Workforce Capability Architecture – Global Quality Management System)
Vertex Pharmaceuticals