Search for More Jobs
Get alerts for jobs like this Get jobs like this tweeted to you
Company: AMD
Location: Secaucus, NJ
Career Level: Hourly
Industries: Technology, Software, IT, Electronics

Description



ADVANCE YOUR CAREER. ADVANCE THE WORLD. 

At AMD, we believe technology has the power to solve the world's most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMD is shaping the future. 

 

Whether you're designing next-gen processors, enabling AI breakthroughs, or bringing leading edge products to market, every role at AMD contributes to something bigger — technology that moves the world forward. Join us and, together, we'll advance your career.



THE ROLE:

Own end-to-end FA across hardware, firmware, silicon, and integration for Helios server platform bring-up and system-level failures. This role drives firmware-aware hardware debug across BIOS, BMC, CPLD, PMBus, POST, PCIe, high-speed interconnect initialization, power sequencing, register state, and platform-level interactions. The engineer will determine whether failures are driven by firmware behavior, hardware design, silicon behavior, component quality, power delivery, configuration, or cross-domain interaction issues. Success in this role requires serving as the firmware SME for FA, enabling independent isolation of systemic system issues while owning interfaces with BIOS, BMC, silicon, validation, design, manufacturing, supplier quality, and customer-facing teams to drive corrective actions and improve platform quality.

THE PERSON:

The ideal candidate is a hands-on systems integrator and failure analysis technical leader with deep platform bring-up experience on complex server or hyperscale systems. They can navigate ambiguous failures across BIOS, BMC, CPLD, firmware, silicon, motherboard, PDB, power delivery, PCIe, high-speed interconnects, diagnostics, and system configuration boundaries. This person is not expected to be a firmware coder; instead, they must understand firmware-controlled hardware behavior well enough to isolate whether the failure is firmware, hardware, or an interaction between domains. They should be comfortable leading structured debug, reviewing register dumps and logs, validating behavior with scopes and logic analyzers, defining diagnostic strategy, communicating clear RCA conclusions, and driving corrective actions across cross-functional engineering teams.

KEY RESPONSIBILITIES:

  • Lead end-to-end failure analysis and root cause ownership for Helios server platform bring-up, factory, customer, and system-level failures.
  • Own platform bring-up debug across BIOS, BMC, CPLD, PMBus, POST, PCIe enumeration, high-speed interconnect initialization, resets, clocks, power sequencing, and system configuration domains.
  • Determine whether failures are caused by firmware behavior, hardware design, component quality, silicon behavior, power delivery, manufacturing process, configuration, or cross-domain interaction issues.
  • Develop and execute structured debug plans using register dumps, firmware and BIOS logs, BMC event logs, telemetry, POST codes, diagnostic results, schematics, board layouts, oscilloscope captures, and logic analyzer traces.
  • Perform power sequencing and platform readiness debug, including rail enable timing, reset behavior, clock availability, PMBus communication, voltage/current telemetry, and fault propagation analysis.
  • Validate firmware-controlled hardware behavior using scopes, logic analyzers, protocol tools, register reads, and data-driven correlation across boot, initialization, and failure states.
  • Own technical interfaces with silicon, BIOS, BMC, firmware, validation, diagnostics, hardware design, manufacturing, and supplier teams to lead cross-functional RCA and drive corrective action closure.
  • Automate debug data collection, log parsing, register analysis, and failure correlation using Python, Linux tools, scripting, and data analysis workflows.
  • Create clear technical reports, executive summaries, debug timelines, and 8D-style documentation that communicate failure mode, evidence, root cause, impact, and recommended actions.
  • Define diagnostic strategy and instrumentation needs by identifying gaps in bring-up procedures, telemetry, register visibility, platform logs, factory screens, and customer debug processes to accelerate systemic issue isolation.

PREFERRED SKILLS AND EXPERIENCE:

  • Strong experience debugging server, hyperscale, GPU/accelerator, rack-scale, or comparable complex compute platforms through bring-up, validation, factory, or customer failure analysis phases.
  • Deep platform bring-up competency across BIOS, BMC, CPLD, PMBus, POST, PCIe enumeration, resets, clocks, power sequencing, register state, and high-speed interconnect initialization flows.
  • Ability to isolate failures across firmware-controlled hardware behavior, silicon, motherboard/PDB design, power delivery, component quality, diagnostics, manufacturing process, and system configuration boundaries.
  • Hands-on proficiency with oscilloscopes, logic analyzers, protocol analyzers, digital multimeters, power supplies, telemetry tools, register access utilities, and lab validation equipment.
  • Experience analyzing register dumps, BIOS/BMC/FW logs, POST codes, event logs, telemetry streams, diagnostic outputs, schematic evidence, and electrical captures to build evidence-based RCA conclusions.
  • Working knowledge of PCIe, high-speed interfaces, interconnect initialization, retimers, link training, signal path dependencies, and failure modes common to dense server platforms.
  • Experience using Python, shell scripting, Linux tools, SQL or data analysis methods to automate log parsing, register review, debug triage, and failure correlation.
  • Ability to own technical interfaces and lead cross-functional RCA reviews with silicon, BIOS, BMC, firmware, validation, diagnostics, design, manufacturing, and supplier teams, driving corrective actions to closure.
  • Experience with structured problem solving, 8D, platform bring-up debug, validation escapes, factory issue resolution, supplier quality, or customer failure analysis processes.
  • Ability to understand firmware behavior and hardware control flows without being a pure firmware developer; this role requires firmware-aware systems integration and failure analysis expertise.

ACADEMIC CREDENTIALS:

  • Bachelor's or Master's degree in Electrical Engineering, Computer Engineering, Computer Science, or related discipline.
  • Equivalent hands-on experience in server platform bring-up, systems integration debug, firmware-aware hardware failure analysis, validation, or complex system-level RCA will be considered.

 

#LI-LB1



Benefits offered are described:  AMD benefits at a glance.

 

AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law.   We encourage applications from all qualified candidates and will accommodate applicants' needs under the respective laws throughout all stages of the recruitment and selection process.

 

AMD may use Artificial Intelligence to help screen, assess or select applicants for this position.  AMD's “Responsible AI Policy” is available here.

 

This posting is for an existing vacancy.


 Apply on company website