Software Engineer, ML Networking
Confirmed live in the last 24 hours
Anthropic
Compensation
$280,000 - $850,000/year
Job Description
About Anthropic
Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.
Role Summary
A systems-level engineer specializing in network infrastructure and network optimization, with expertise in building and maintaining software that interacts with networks. You will be responsible for writing and maintaining software that interfaces between our accelerators and our high-speed networks. This role requires deep technical knowledge of network protocols, kernel-space and/or user-space networks, interfacing with hardware, and the ability to debug and optimize distributed software at the network level.
You may be a good fit if you have:
Networking Systems Engineering:
- Expert-level proficiency with network protocols and networking concepts
- Deep kernel networking: TCP/IP stack internals, XDP, eBPF, io_uring, and epoll
- User-space networking: DPDK, RDMA, kernel bypass techniques
- Understanding of how to build higher-level abstractions like collectives and RPC
- Skilled at diagnosing and resolving networking issues in distributed systems, especially at OSI model layers 2-4
Low-Level Systems and OS Programming:
- Strong programming skills in a systems programming language, including memory management, lock-free data structures, and NUMA-aware programming
- Software, driver, and OS performance optimization tools and techniques
- Comfort with or desire to learn Rust
Strong candidates may have:
- Understanding of ML accelerators and accelerator drivers
- Demonstrated ability to design new network protocols
- Experience with PCIe and drivers for PCIe devices
- Expertise in algorithms used in networking, including compression and graph algorithms
- Experience programming on SmartNICs
- 5+ years of experience in systems programming or network programming
- Often comes from backgrounds in: HPC, telecommunications, host networking software, OS/kernel engineering, or embedded systems
- Strong debugging mindset with patience for complex, multi-layered issues
Representative Projects:
- Build a system for accelerator-initiated tensor movement over the network
- Benchmark software for a new networking environment
- Implement a new collective algorithm to improve latency
- Optimize congestion control algorithms for large-scale synchronous workloads
- Debug kernel-level network latency spikes
The annual compensation range for this role is listed below.
For sales roles, the range provided is the role’s On Target Earnings ("OTE") range, meaning that the range includes both the sales commissions/sales bonuses target and annual base salary for the role.
Logistics
Minimum education: Bachelor’s degree or an equivalent combination of education, training, and/or experience
Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience
Minimum years of experience: Years of experience required will correlate with the internal job level requirements for the position
Location-based hybrid policy: Currently, we expect all staff to be in one of our offices at
Similar Jobs
Roku
Sr Manager, SW Engineering - ML
Anthropic
ML Infrastructure Engineer, Safeguards
Aurora Innovation
Software Engineer, Behavior Planning ML Platform
Databricks
Specialist Solutions Architect - AI & ML (Communications, Media, Entertainment & Games)
Databricks
Specialist Solutions Architect - AI & ML (Financial Services)
Annapurna Labs (U.S.) Inc.