Counts are per token, for 1 feed-forward block · 346-token prompt · M1 Max, native wgpu · output scores were bitwise identical across all 3