SpaceXAI logo
SpaceXAI

Hardware Failure Analysis Engineer - Memphis

USAPosted 4 weeks ago

Apply opens SpaceXAI's site. When you're back, we'll ask whether you applied.

Job type
Not listed
Work mode
Not listed
Level
Not listed
Department
Engineering
Experience
2+ years experience
Posted
Sep 3, 2026

About the role

ABOUT THE ROLE:

Determine true root cause of fleet hardware failures — component and system level — and drive fixes through vendors to the manufacturer. Turn "swap it again" into "vendor redesign / firmware fix / manufacturing escape found."

RESPONSIBILITIES:

  • Own named failure classes across GPU trays/baseboards, NIC/DPU, motherboard/PCIe switch, memory, power, thermal, cables/connectors, and rack-scale patterns.
  • Run recurrence analysis and fleet-wide defect clustering; detect systemic patterns before they become fleet-scale loss.
  • Build vendor technical escalation packages with evidence quality that forces action; partner with OEM/ODM/component manufacturers on firmware, bring-up, and manufacturing escapes through committed CAPA.
  • Feed findings into RMA policy, spare strategy, and "do not reseat forever" stop-rules.
  • Work the FA intake queue on rotation; keep queue age within SLA

BASIC QUALIFICATIONS:

  • Bachelor's degree in Systems Engineering, Electrical Engineering, Computer Science, or a related field (or equivalent experience).
  • 2+ years of experience in hardware reliability engineering, preferably in high-performance computing or data center environments.
  • Proven expertise in firmware analysis, hardware specifications review, and release validation.
  • Strong experience with RMA processes, including filing claims, vendor negotiations, and pushing for resolutions outside standard protocols.
  • Demonstrated ability to diagnose and prove complex hardware failures, including grey or intermittent issues, using tools, logic analyzers, or diagnostic software.
  • Familiarity with data center hardware components (e.g., servers, GPUs, networking equipment) and emerging technologies.
  • Proficiency in scripting (Python, Bash) for automation and analysis, plus general experience in at least one systems language (C, C++, Java, Rust, or similar). Not required to be expert in all of them.
  • Excellent problem-solving skills with a data-driven approach to reliability engineering.
  • Ability to work collaboratively with cross-functional teams, including operations technicians.

PREFERRED SKILLS AND EXPERIENCE:

  • Experience in AI/ML infrastructure or supercomputing environments.
  • Knowledge of vendor ecosystems (e.g., NVIDIA, Dell, HP, Supermicro) and supply chain management.
  • Certifications in hardware engineering or reliability (e.g., CRE, CompTIA Server+).
  • Prior work in a fast-paced startup or tech company like SpaceXAI.