The author is deciding between adding an M5 Ultra 256GB or dual DGX Spark units to their existing RTX 5090 for local LLM workloads, leaning toward the dual Spark configuration for its higher throughput in agentic coding and parallel request processing. They note the challenge of obtaining comparable benchmarks due to the fast‑changing local LLM landscape and reference a Mindstudio article that compares the two options.
Read original
reddit/r/LocalLLaMA