The nano-llm-posttraining repository provides a framework for conducting minimal LLM post-training experiments on hardware with only 8GB of VRAM. It supports various alignment techniques, including Supervised Fine-Tuning (SFT), Direct Preference Optimization (DPO), and Group Relative Policy Optimization (GRPO).
Read original
hackernews