The posting, in xAI's own words
archived Sep 2, 2026SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. SpaceXAI is looking for an exceptional network engineer with experience in mission-critical and large-scale production environments to support high-density AI training and inference clusters as well as high-reliability data center networks. As a member of the Data Center Network Engineering team, you will provide operational and design services for networks used by compute clusters, automation & controls engineering, and facilities teams. The ideal candidate thrives in intense, high-flux environments, brings a strong sense of urgency balanced with operational excellence, communicates effectively, and demonstrates high levels of technical acumen.
RESPONSIBILITIES:
Design and implement highly available, high-bandwidth networks for AI data centers, carefully balancing routing, overlay, and redundancy technologies (including spine-leaf / Clos fabrics) to meet unique requirements of GPU training clusters, inference, storage, and management planes. Design and maintain data center and campus networks in accordance with company network standards. Collaborate with adjacent compute, infrastructure, facilities, and enterprise teams. Evaluate, procure, and deploy network hardware including high-speed switches, optics, firewalls, multiplexing, and related appliances. Contribute to ever-maturing network automation tooling; implementing configuration analysis, linting, and scalability into the deployment framework. Plan and coordinate network maintenance windows with stakeholders to perform software updates, hardware refreshes, and general network work (sometimes on weekends and evenings) in live production environments. Troubleshoot and resolve network-related issues, publishing root cause analysis (RCA) documentation and hosting retrospective reviews. Provide direct networking support during cluster bring-up, capacity expansions, and high-load operations; participate in on-call or serve as networking responsible engineer during critical events. Proactively tailor network monitoring and telemetry to detect congestion, packet loss, and other issues before they impact training or inference workloads. Continuously create and update network documentation, including architecture overviews, design drawings, and operational procedures. Collaborate with cross-functional teams to proactively identify and resolve potential technical issues with network designs, especially systemic and cascading failure modes and insufficient redundancy. Perform job walks with customers, vendors, and contractors to gather network and connectivity requirements and create implementation plans for new halls and expansions. Ensure networks are configured and maintained in compliance with industry and cybersecurity standards, with particular attention to segmentation between compute fabrics, management, and facilities/OT networks.