Ziwei Li
GPU Design Engineer @ Apple
About
Hardware Accelerator for Machine Learning and AI workload GPU RTL design and Micro-Architecture design. (TileBuffer; Clipper; Cull; Viewport Transfer; Setup) Experience with Verilog and System Verilog coding. Experience with low power design Experience with area reduction Experience with timing fix Experience with ECO Experience with GPU architecture Experience with the use of Verdi, Lint, Synthesis. Understanding of semiconductor physics.
United States
Newark
Semiconductors
Git, GPU, clock gating, perforce, Verdi, DVE, 机器学习, Emulation, SystemVerilog, VLSI, ASIC, Verilog, system verilog, Modelsim, Perl, ASIC design, Cadence Virtuoso 工具, Cadence Encounter, PrimeTime, NCSim
Experience

ASIC Design Engineer
Mountain View, California, United States
• Experienced in designing Hardware Accelerator for Machine Learning/AI workload. • Responsible for Micro-architecture spec and RTL design of fully deterministic Instruction Control Units (ICUs) that handle instruction parsing, decoding and dispatching in the Tensor Streaming Processor pipeline. • Developed ICUs to support chip level die-to-die communications and multi-chip synchronizations. • Proposed and implemented new ISA to support chip level clock masking, clock period synthesis, chip level synchronization and preemptive voltage compensation. • Supported RAS by parity check implementation and fault detection. • Worked closely with Compiler team and innovated a new timing structure enable compiler program, which tremendously simplify the software deterministic instruction scheduling effort. • Analyzed the complexity of supporting several different numerical formats. • Improved dynamic power by 30% with optimized clock control and clock-domain splitting. • Helped top, cluster and block level DV debugging the compiler program issue using DVE and Verdi. • Collaborated with DV and Compiler team to analyze Emulation result via Palladium. • Worked on Logical Synthesis flow with PD to fix the Congestion issues and Timing violations. • Implemented ICUs Functional Verification Coverage.

Senior GPU RTL Design Engineer
SAMSUNG AUSTIN SEMICONDUCTIOR
San Jose, California, United States
• Owned U-arch spec and RTL design of synthesize-friendly and power-efficient Primitive Assembler that handles Clipping, Culling, Set-up process for interpolation, and Viewport Transfer in the GPU pipeline. • Owned U-arch spec and RTL design of Medium Grain Clock Gating via ICG cell, Fine Grain Clock Gating and clock-domain splitting for power reduction to meeting the goal. • Owned U-arch spec and RTL design of throughput doubling . • Reduced the area consumption with the methods of splitting memory, retiming logic removal and simplifying redundant logics. • Refined RTL code and re-organized the pipeline logic to meeting the timing requirements. • Worked with EMULATION team and PD team to help in emulation debug and ECO to minimize the impacted logics after RTL Freeze and Post-Silicon debug and ECO after tape-out. • Worked with FDV, AM, Formal and CSIM team to debug using Verdi, hit the target of functional coverage, resolve timing violation and performance loss. • Worked with other RTL unit design team for interface handshake and inter-unit’s communication.

GPU RTL Designer
Samsung Austin Semiconductor
San Jose, California
Tile Buffer Micro-Arch design/Low power design. • Owned U-arch spec and RTL design of synthesize-friendly and power-efficient smart memory in storing Depth, Stencil and Color information from Rasterization and Color blender units using System Verilog and modifying Yaml file. • Owned U-arch spec and RTL design of Packed Format/YUV format/MSAA/Adaptive De-sampling/Block Linear/Pitch Linear Flush, Normal/BINZ/Decompression Preload, Fast clear Data-path Control Logic and FSM. • Owned U-arch spec and RTL design of pipeline mode Sync Token Control Logic and timing synchronous mode NNOC Packet Control Logic. • Implemented D16 to D24 and D16 to FP32 Conservative Format Conversion Logic. • Fixed the timing violations by re-organizing pipeline logic and simplifying existed logic to meet timing requirements. • Worked with EMULATION team and PD team to help emulation debug and ECO to provide cheapest choice for the 2 tape-outs. • Worked with FDV, Formal and DVM team to fix SEQ mismatch and worked with Physical and Integration team for congestion issues. Primitive Assembler (clipper, cull, view port transfer, setup) Micro-Arch design/Low power design. • Owned U-arch spec and RTL design of synthesize-friendly and power-efficient Primitive Assembler that handles Clipping, Culling, Set-up process for interpolation, and Viewport Transfer in the GPU pipeline.

Trainee of Bidding Department(Intership)
Sieyuan Qingneng Electronics Corp
Shanghai, China
Assisted in the editing and review of a project application form Managed the bidding documents Visited workshops to acquire information
Ziwei Li's Contact Information
Phone
Find the Right Leads
Find Verified Contact Data
What LeadContact does well
Find verified emails, phone numbers, and decision-makers with 98% accuracy.
Find Leads
Find the right people by company, role, industry, location, and more.
925M+ professional profiles

Find Emails
Access verified email addresses for your target contacts.
657M+ emails

Find Phone Numbers
Get cross-validated phone data from multiple top sources.
239M+ phone numbers

More Accurate. Lower Cost.
Find contact data in 1 tool with 98% accuracy
LeadContact integrates leading enrichment tools to deliver more accurate contact data—without paying for each one.
Great conversations start with the right contact.
It’s time to find yours.




