The authors introduce RGBD20K, a large‑scale benchmark designed to advance RGB‑D semantic segmentation by providing richer semantic coverage, more training data, and cleaner annotations. The dataset contains 20,000 synchronized RGB‑D image pairs, each labeled with one of 160 fine‑grained semantic categories—substantially exceeding the 40 classes of NYUv2 and the 37 classes of SUN RGB‑D. To ensure annotation quality, the team performed a rigorous re‑evaluation and correction of existing labels, eliminating long‑standing noise and establishing a reliable ground‑truth foundation. Leveraging this high‑fidelity data, they propose a novel score‑purified fusion (SPF) method that effectively combines RGB and depth cues; experiments show SPF attains state‑of‑the‑art performance across all evaluated segmentation benchmarks, confirming the benefit of high‑quality multimodal information. The dataset and associated code are publicly available via the provided URL.
Read original
huggingface/daily-papers