The post reports benchmark results for the AMD Ryzen AI Max+ 395 with 96 GB VRAM, using Lemonade Server to compare inference performance of Gemma‑4 and Qwen‑3.6 in chat and coding tasks, and evaluate Multi‑Token Prediction throughput. The system featured a Ryzen AI Max+ 395 CPU, Radeon 8060S iGPU, and 128 GB system memory.
Read original
reddit/r/LocalLLM