Wnuan introduces a three-stage pipeline for enterprise question answering that preserves general language capabilities while incorporating proprietary knowledge. The process builds task-oriented supervision from documents, applies supervised fine‑tuning with general-data replay, and then uses reinforcement learning to correct residual errors. On the 707-question WnuanBench, the 32B model’s acceptable‑answer rate improves from 52.76% pre‑adaptation to 80.06% after fine‑tuning and 91.51% following reinforcement learning.
Read original
huggingface/daily-papers