HPC Data Center Production Engineer
Confirmed live in the last 24 hours
Jump Trading
Job Description
HPC Data Center Production Engineer
Location: Chicago or New York (On-site 5 days/week)
Jump Trading Group is committed to world class research. We empower exceptional talents in Mathematics, Physics, and Computer Science to seek scientific boundaries, push through them, and apply cutting edge research to global financial markets. Our culture is unique. Constant innovation requires fearlessness, creativity, intellectual honesty, and a relentless competitive streak. We believe in winning together and unlocking unique individual talent by incenting collaboration and mutual respect. At Jump, research outcomes drive more than superior risk adjusted returns. We design, develop, and deploy technologies that change our world, fund start-ups across industries, and partner with leading global research organizations and universities to solve problems.
Trading Infrastructure is a global organization of Engineers who architect, build and maintain our world-class infrastructure. From colo design/implementation, to optimizing our exchange connectivity, to building world class low latent Wide Area Networks, we leverage research and automation to consistently adapt and innovate our infrastructure to scale and drive our trading and evolving business.
We are looking for an HPC Data Center Production Engineer to build and own the automation and tooling that powers Jump's HPC data center operations. This is a development-heavy role focused on automating the onboarding and lifecycle management of data center hardware—servers, switches, rack PDUs, CDUs, and environmental sensors—and building tools for capacity planning, outage simulation, monitoring, and metrics integration. You will work hand-in-hand with the HPC Planning, Engineering, and Operations leads to turn tooling and monitoring vision into production-ready systems. Heavy, daily use of AI tools is expected to accelerate development and raise the quality bar on everything you build.
What You'll Do:
Hardware Onboarding Automation
-
Design, develop, and maintain automation to onboard new hardware devices into Jump's HPC data centers, including servers, network switches, rack PDUs, CDUs, and environmental sensors.
-
Build end-to-end provisioning workflows that take hardware from racked-and-cabled through discovery, configuration, validation, and production-ready state with minimal manual intervention.
-
Extend and adapt onboarding automation as new hardware platforms and device types are introduced.
Data Center Tooling Development -
Develop tools for power and cooling capacity planning—enabling the operations and planning teams to model current utilization, forecast growth, and identify constraints before they become problems.
-
Build outage simulation tooling to model the impact of power, cooling, or network failures across HPC facilities and validate redundancy/failover configurations.
-
Develop and maintain operational tooling that supports day-to-day data center workflows such as hardware lifecycle tracking, data center inventory/spares, change management, and diagnostics.
Monitoring & Metrics Integration -
Build and maintain monitoring integrations for HPC data center infrastructure—pulling telemetry from servers, switches, PDUs, CDUs, environmental sensors, and facility systems into centralized observability platforms.
-
Integrate metrics feeds from colocation and data center providers into Jump's monitoring stack, normalizing data for alerting and capacity reporting.
-
Work with the Operations Lead to implement the monitoring and alerting strategy, translating requirements into deployed, production-grade instrumentation.
Cross-T
Similar Jobs
Linxon
HR Systems Analyst (SAP SuccessFactors)
Synchrony Financial
VP-Privileged Access Management (PAM) Engineer (L12)Engineer (L12)
Nasdaq
Information Security Senior Analyst
Nasdaq
Workday Integration Specialist
Sun Life
AVP, Enterprise Architecture Strategy & Execution
Sun Life