The team has open-sourced the world's fastest WebGPU kernels for running local AI models directly in browsers via WebGPU. The collection covers over 200 common machine learning operations that execute entirely client-side, with ongoing efforts to integrate these optimizations into frameworks like Transformers.js and ONNX Runtime Web. Detailed implementations and performance benchmarks are available at the Hugging Face kernel repository.
Read original
reddit/r/LocalLLaMA