The author revised their 3× Tesla P100 budget build results after discovering that the GPU core clocks were stuck at 405 MHz during the three‑card test, invalidating the earlier performance figures. Consequently, the reported token generation (~21 t/s) and prompt processing (~67 t/s) rates, power draw (~160 W), and temperatures (38‑51 °C) should be disregarded. The claim that three cards performed about four times slower than two is also withdrawn.
Read original
reddit/r/LocalLLM