The posting, in xAI's own words
archived Sep 2, 2026SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.
ABOUT THE ROLE:
SpaceXAI is looking for an exceptional network engineer with experience in mission-critical, large-scale production environments to support the design, build-out, and operation of networks that power our AI supercomputer campuses. As a member of the Supercomputer Infrastructure / Network Engineering team, you will provide design and operational support for the fabrics used by GPU training and inference clusters, site operations, automation and controls, and facilities teams. The ideal candidate thrives in intense, high-flux environments, brings a strong sense of urgency balanced with operational excellence, communicates clearly, and demonstrates high technical acumen.
Read the full posting ↓
RESPONSIBILITIES:
Design and implement highly available, low-latency, high-bandwidth networks, carefully balancing routing, congestion control, and redundancy technologies for AI training fabrics, inference front-ends, storage, and site/OT networks. Design and maintain supercomputer data center and campus networks in accordance with company network standards. Collaborate with adjacent infrastructure, compute, storage, SiteOps, and enterprise teams. Evaluate, procure, and deploy network hardware including data-center class switches, NICs, firewalls, optical multiplexers, and related appliances supporting 400G/800G and beyond. Contribute to maturing network automation tooling; implement configuration analysis, linting, validation, and scalable deployment frameworks (GitOps / IaC). Plan and coordinate network change windows with stakeholders to perform software updates, hardware refreshes, cluster expansions, and general maintenance (including evenings and weekends when required by compute schedules). Troubleshoot and resolve network-related issues affecting cluster health and job performance; publish root cause analysis (RCA) documentation and host retrospectives. Provide direct networking support during cluster bring-up, expansion, and production training/inference campaigns; serve as on-call or networking responsible engineer during operations. Proactively tailor network monitoring and telemetry (fabric health, congestion, packet loss, NCCL/collective performance) so issues are detected before they impact training or inference. Continuously create and update network documentation, including architecture overviews, design drawings, fiber/cable plant records, and operational procedures. Collaborate with cross-functional teams to identify and resolve potential design issues, especially systemic or cascading failure modes and false redundancy in AI fabrics and site networks. Perform job walks with customers, vendors, and contractors to gather requirements and produce implementation plans for new halls, rows, and campus interconnects. Ensure networks are configured and maintained in compliance with industry and cybersecurity standards (e.g., ITAR, ISO, NIST), with particular attention to segmentation between compute fabrics, storage, OT/controls, and corporate networks.