A Java‑based framework uses TornadoVM to compile Java code to CUDA and cuTile, achieving up to 90% of the performance of llama.cpp for local LLM inference on NVIDIA GPUs. The system includes the jitLLM inference engine and has been presented in a deep‑dive talk. Source code for both TornadoVM and jitLLM is hosted on GitHub.

Read original