Model "distillation" accusations are getting way overblown at this point
Article automatically generated from technical news.
Every time a strong open model drops, the same cycle plays out: ai bro's claims it's "just distilled from GPT4/Claude/whatever," case closed, move on. I think this take doesn't hold up as well as people assume. A few points worth separating out: Training on outputs isn't the same as real distillation. Proper token level distillation needs access to logits, the full probability distribution over the vocabulary, not just the final text response. Nobody gets that from a public A
Fonte originale