sarvam-translate on three Windows laptops: 17.5 to 73.6 tok/s
I tested the public sarvam-translate model on three Windows 11 laptops with NVIDIA GPUs from 2014 and 2023. It ran fully GPU-offloaded on all three, with llama.cpp generation ranging from about 17.5 to 73.6 tok/s.
Jul 22, 2026