A user reports having a dense ~9.4‑billion‑parameter language model ready for training, featuring a 1/2/3 Engram table, Moonshot's AttnRes attention, and RoPE/NoPE layering at a 3:1 ratio, initialized with the Llama 3 tokenizer and LM head. The model is being trained on logit‑level extracts from a Llama 3 teaching model, and initial training steps have confirmed stability. They inquire whether there remains strong community interest in such a dense 9B model.
Read original
reddit/r/LocalLLaMA