- Job type
- Not listed
- Work mode
- Not listed
- Level
- Not listed
- Department
- Engineering
- Experience
- 2+ years experience
- Posted
- Sep 3, 2026
About the role
ABOUT THE ROLE:
Determine true root cause of fleet hardware failures — component and system level — and drive fixes through vendors to the manufacturer. Turn "swap it again" into "vendor redesign / firmware fix / manufacturing escape found."
RESPONSIBILITIES:
- Own named failure classes across GPU trays/baseboards, NIC/DPU, motherboard/PCIe switch, memory, power, thermal, cables/connectors, and rack-scale patterns.
- Run recurrence analysis and fleet-wide defect clustering; detect systemic patterns before they become fleet-scale loss.
- Build vendor technical escalation packages with evidence quality that forces action; partner with OEM/ODM/component manufacturers on firmware, bring-up, and manufacturing escapes through committed CAPA.
- Feed findings into RMA policy, spare strategy, and "do not reseat forever" stop-rules.
- Work the FA intake queue on rotation; keep queue age within SLA
BASIC QUALIFICATIONS:
- Bachelor's degree in Systems Engineering, Electrical Engineering, Computer Science, or a related field (or equivalent experience).
- 2+ years of experience in hardware reliability engineering, preferably in high-performance computing or data center environments.
- Proven expertise in firmware analysis, hardware specifications review, and release validation.
- Strong experience with RMA processes, including filing claims, vendor negotiations, and pushing for resolutions outside standard protocols.
- Demonstrated ability to diagnose and prove complex hardware failures, including grey or intermittent issues, using tools, logic analyzers, or diagnostic software.
- Familiarity with data center hardware components (e.g., servers, GPUs, networking equipment) and emerging technologies.
- Proficiency in scripting (Python, Bash) for automation and analysis, plus general experience in at least one systems language (C, C++, Java, Rust, or similar). Not required to be expert in all of them.
- Excellent problem-solving skills with a data-driven approach to reliability engineering.
- Ability to work collaboratively with cross-functional teams, including operations technicians.
PREFERRED SKILLS AND EXPERIENCE:
- Experience in AI/ML infrastructure or supercomputing environments.
- Knowledge of vendor ecosystems (e.g., NVIDIA, Dell, HP, Supermicro) and supply chain management.
- Certifications in hardware engineering or reliability (e.g., CRE, CompTIA Server+).
- Prior work in a fast-paced startup or tech company like SpaceXAI.