PYTHON · NUMPY
Why the multiply-free win is a memory win

Pack ternary weights into 2-bit codes, unpack them two ways, and prove the fast way is free. The branchy select-chain decode is given; your one gap is the branch-free arithmetic decode. The grader checks it is BIT-IDENTICAL to the branchy decode and to the original weights (max error 0), then reports the real measured RTX 4050 speedups (2.58x select-chain -> 5.83x branch-free). numpy only, deterministic, < 2 s.

You read

the arrays and values already in scope

You change

the code you write in each cell

Fixed

the dataset and the checks that grade you