Users have successfully deployed DeepSeek V4 Flash with speculative decoding on a single AMD Ryzen AI MAX+ 395 featuring 128 GB of unified memory. The implementation achieved a performance rate of up to 32 tokens per second. The accompanying code for this setup has been released under the Apache-2.0 license.
Read original
reddit/r/LocalLLaMA