This paper introduces Document Retrieval-Aware Chunking (D-RAC), a universal ingestion framework for enterprise RAG systems that handles heterogeneous document formats including PDFs, Word documents, presentations, and scans. D-RAC addresses the limitations of rule-based extraction and OCR—which destroy reading order, flatten tables, and lose heading hierarchy—by leveraging PDF normalization and multimodal Markdown conversion to preserve document structure. The approach reduces the token costs and hallucination risks associated with fully agentic chunking over extracted text, enabling more faithful retrieval-aware processing of complex enterprise documents.
Read original
huggingface/daily-papers