Datalab's Marker 2 achieves 76.0 on the olmOCR-bench, outperforming MinerU, Docling, and LiteParse while maintaining 5× MinerU's throughput. It is designed as a high-performance PDF parser pipeline for converting documents into structured formats like markdown or JSON. The system prioritizes both accuracy and sustained processing speed
reddit/r/machinelearningnews