Description
ADVANCE YOUR CAREER. ADVANCE THE WORLD.
At AMD, we believe technology has the power to solve the world's most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMD is shaping the future.
Whether you're designing next-gen processors, enabling AI breakthroughs, or bringing leading edge products to market, every role at AMD contributes to something bigger — technology that moves the world forward. Join us and, together, we'll advance your career.
THE ROLE:
As a Principal System Debug Engineer, you will lead the system-level debug of our advanced microprocessors within complex customer applications and system design environments. You will act as the critical bridge between our internal silicon/software engineering teams and our top-tier OEM, ODM, and Hyperscaler customers.
This is a highly visible, cross-functional leadership role that requires deep technical expertise in CPU architecture, high-speed I/O interfaces, firmware/BIOS, system software, drivers, OS/kernels and software stacks, and platform hardware design, signal integrity, power integrity. You will drive the investigation, root-cause analysis, and resolution of the most complex, system-level CPU issues in the context of new system bring-up and customer applications and platform level debug.
THE PERSON:
You are a subject matter expert and strong technical contributor with strong processor architecture, hardware and software expertise, and extensive system level debug experience. You excel as part of a team where technical leadership, customer engagement, communication and team skills are highly valued.
KEY RESPONSIBILITIES:
- Technical Leadership & Debug: Lead the triage, investigation, and root-cause analysis of complex system-level issues involving CPU silicon, platform hardware, firmware, drivers and OS/hypervisor interactions in customer environments.
- System Bring-Up: Drive early platform bring-up activities for next-generation CPUs, both in internal labs and on-site at customer facilities, ensuring rapid time-to-market.
- Customer Engagement: Serve as the primary technical liaison for strategic customers. Guide customer engineering teams through platform design, debug methodologies, and issue resolution.
- Cross-Functional Collaboration: Partner closely with internal Silicon Design, Architecture, Firmware (BIOS/UEFI/BMC), OS/Kernel, and Validation teams to drive systemic fixes and influence future CPU architectures based on customer feedback.
- Debug Methodology & Tooling: Architect and develop advanced debug methodologies, scripts, and tools to accelerate issue isolation across the hardware/software boundary.
- Escalation Management: Act as the technical task force leader during critical customer escalations, providing clear executive updates and driving technical action plans under tight deadlines.
PREFERRED EXPERIENCE:
- CPU Architecture: Deep, foundational knowledge of modern CPU architectures (x86 mandatory), including pipelines, cache coherency, memory controllers, and power management (C-states, P-states). RAS features, MCA (Machine Check Architecture), SoC and platform security features, RoT (root-of-trust).
- I/O & Interconnects: Extensive expertise in the architecture and protocol-level debug of high-speed I/O interfaces, including PCIe (Gen 4/5/6), CXL, DDR4/DDR5, Ethernet, USB, and low-speed buses (I2C, SPI, I3C, eSPI).
- Firmware & Software Stacks: Strong understanding of the full system software stack, including BIOS/UEFI, BMC/IPMI, Linux/Windows kernel internals, device drivers, and hypervisors (KVM, VMware, Hyper-V).
- Platform Design: Solid grasp of system-level hardware design, including board schematics, PCB layout, power delivery networks (VRMs), clocking, and signal integrity fundamentals.
- Hands-On Debug: Proven track record of performing complex system bring-up and debug using hardware tools (JTAG/ITP debuggers, oscilloscopes, logic analyzers, protocol analyzers) and software debuggers (GDB, WinDbg, kernel panics/crash dump analysis).
- Experience working directly with Cloud Service Providers (Hyperscalers) and Tier-1 Server/Client OEMs.
- Proficiency in scripting and programming languages (Python, C, C++) for test automation and debug tool development.
- Hands-on expertise in debugging hardware (system bring-up, signal integrity, power integrity issues) and debugging firmware/software running on actual hardware platforms.
- Experience with RAS (Reliability, Availability, and Serviceability) architecture and machine check exception (MCE) analysis.
ACADEMIC CREDENTIALS:
Bachelor's, Master's, or PhD in Electrical Engineering, Computer Engineering, Computer Science, or a related field.
LOCATION:
Austin, TX or San Jose, CA preferred, hybrid option available.
This role is not eligible for visa sponsorship.
#LI-MV1
#HYBRID
Benefits offered are described: AMD benefits at a glance.
AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants' needs under the respective laws throughout all stages of the recruitment and selection process.
AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD's “Responsible AI Policy” is available here.
This posting is for an existing vacancy.
Apply on company website