We benchmarked 18 RAG pipeline variants on Google's FRAMES benchmark using identical model, embeddings, and documents across 824 multi‑hop questions. The highest‑scoring pipeline achieved 78.9% accuracy, while an agent loop equipped with retrieval tools reached 92.7%—comparable to providing the model with the correct articles upfront. The study evaluated hybrid search, reranking, query decomposition, and query expansion to assess their individual contributions.

Read original