Engineering log

The NPU didn’t make it faster

We moved our on-device language model from the CPU to the RK3588’s NPU expecting a speed-up. Token generation came out a dead heat. We kept the NPU anyway, and the reason is more useful than the speed-up would have been.