LingBot‑Map introduces a feed‑forward 3D foundation model for streaming reconstruction that centers on a Geometric Context Transformer (GCT) which jointly grounds coordinates, integrates dense geometric cues, and corrects long‑range drift via anchor context, a pose‑reference window, and trajectory memory. The model employs paged KV‑cache attention to achieve stable inference at roughly 20 FPS on 518×378 resolution for sequences longer than 10 000 frames, outperforming both prior streaming methods and iterative optimization baselines on benchmarks including KITTI, Oxford Spires, VBR, Droid‑W, TUM‑D, 7‑scenes, ETH3D, Tanks and Temples, and NRGBD. Training uses video RoPE on 320 views; for longer inputs a windowed mode (--mode windowed --window_size 128) with configurable keyframe intervals (e.g., --keyframe_interval 10) and overlap keyframes (--overlap_keyframes 8) resets the KV cache to maintain pose stability. The system relies on PyTorch 2.8.0 + CUDA 12.8, FlashInfer for efficient attention, NVIDIA Kaolin for batch rendering, and an ONNX sky‑segmentation model (skyseg.onnx) for outdoor masking. Installation proceeds via a conda environment, pip‑installable packages, and optional CUDA extensions; demo scripts (demo.py, demo_render/batch_demo.py) enable interactive or offline point‑cloud flythrough generation with configurable rendering flags. The balanced checkpoint (lingbot‑map) and its stage‑1 variant are hosted on Hugging Face under Apache 2.0, with the accompanying ECCV 2026 paper detailing the Geometric Context Transformer.
Read original
github-trending/python