WeVisDoc is a two-stage data‑centric framework designed for robust end‑to‑end document parsing. Its first stage broadens semantic, structural, and appearance coverage to counteract biases in training corpora toward common document types and clean pages. The approach aims to improve parser capability beyond mere coverage expansion.
Read original
huggingface/daily-papers