About the role
fal is the generative media ecosystem powering the next generation of AI products. We build the infrastructure, tools, and model access that teams need to move from idea to production, and do it at scale without compromise. For developers and enterprises, fal is the foundation that makes generative media not just possible, but practical: a unified platform where high-performance inference, orchestration, and observability come together to unlock new categories of AI-native products.
As generative media reshapes industries across a market projected to grow by hundreds of billions over the next decade, fal is becoming the ecosystem that ambitious teams build on.
About this role:
You build the custom compute environments we deliver to customers — bare metal or virtual machines with GPU passthrough, dedicated Kubernetes clusters, and the networking that ties them together. You work across the full stack from Linux image building to overlay network design to cluster bootstrapping.
Key responsibilities
- Build and deliver custom environments with excellent GPU performance for customer workloads
- Leverage AI to an extreme level to automate provisioning, alerting and recovery
- Provision and configure dedicated Kubernetes clusters tailored to customer requirements
- Design and implement overlay networking (VLAN, VXLAN) and routing configurations (ECMP, BGP) and tunnels (strongSwan, IPSEC) for tenant isolation and performance
- Build and maintain Linux images
- Set up network monitoring and diagnostics for customer environments
- Automate the end-to-end lifecycle of customer compute environments: creation, configuration, validation, and teardown
Requirements
- 5+ years experience with Linux virtualization: KVM/QEMU, libvirt, VFIO device passthrough, hugepages, NUMA, CPU pinning
- Strong networking fundamentals: VXLAN, VLAN, ECMP, BGP, ARP, and the ability to debug packet-level issues (tcpdump, Wireshark)
- Production experience building and operating Kubernetes clusters on bare metal (MetalLB)
- Proficiency with Linux image building and OS provisioning (kickstart, cloud-init, PXE/iPXE)
- Proficiency in Python, Bash, Ansible and Terraform
- Deep experience with NVIDIA GPUs: drivers, MIG, container runtimes (nvidia-container-toolkit), InfiniBand, RDMA/RoCEv2 and GPUDirect for high-performance AI networking
- Excellent communication and ability to drive technical decisions across teams
- Self-starter who executes quickly, takes ownership, and constantly seeks improvement
Nice to have
- Experience with SR-IOV, DPDK, or other high-performance networking technologies
- Experience with shared network storage (Ceph, Lustre, Weka)
- Experience with network automation tools (Netbox, Nautobot, Nornir)
Compensation
- $180,000-250,000 plus equity + benefits (This range encompasses 2 levels Senior and Staff)
Location
-
San Francisco, CA
What we offer at fal
- Interesting and challenging work
- A lot of learning and growth opportunities
- We are currently hiring in downtown San Francisco.
- We offer relocation assistance to San Francisco.
- Health, dental, and vision insurance (US)
- Regular team events and offsites
Aplyr's read
Fal.ai is an AI-driven platform revolutionizing productivity and collaboration through intelligent automation, attracting tech-savvy professionals focused on cutting-edge innovation.
What's promising
- •Fal.ai leverages AI to significantly enhance workplace productivity and collaboration.
- •The company is expanding rapidly, with diverse roles in engineering and sales.
- •Fal.ai's focus on intelligent automation attracts top talent in AI and tech.
What to watch
- •Limited public information about Fal.ai's financial stability and long-term viability.
- •The competitive AI landscape poses significant challenges for sustained differentiation.
- •Rapid expansion may strain company resources and impact employee work-life balance.
Why fal.ai
- •Fal.ai specializes in intelligent automation for productivity and collaboration.
- •The company offers a wide range of engineering roles, indicating a strong tech focus.
- •Fal.ai's emphasis on AI-driven solutions sets it apart in the productivity sector.
Aplyr’s read is generated by AI from public sources. Was it useful?
About fal.ai
Fal is an AI-driven platform that focuses on enhancing productivity and collaboration through intelligent automation.
Similar roles
Senior Software Engineer, Sandboxes & Virtualization
CoreWeave
Senior Software Engineer – Virtualization & SIL Integration
General Motors
Vice President, Virtualization Engineering & SRE
Mitsubishi UFG
Principal Software Engineer - Virtualization
Red Hat
Senior Embedded System Software Engineer - Hypervisor and Virtualization
NVIDIA
Hybrid Virtualization & Private Cloud Engineer
Mitsubishi UFG