The Lotus‑Diffusion‑Dense‑Prediction model adapts a pretrained text‑to‑image diffusion backbone to dense prediction by directly estimating scene annotations (e.g., depth) in a single diffusion step rather than predicting noise across multiple steps, a design the authors claim reduces variance, simplifies optimization and speeds inference. A “detail preserver” tuning strategy is added to improve fine‑grained predictions. The Replicate endpoint exposes only zero‑shot depth estimation: it accepts an image URI and optional parameters—output_type (color or grayscale depth map, default color), resample_method (default bicubic), processing_resolution (default 0, interpreted as the maximum resolution used), and disparity (boolean default true) which merely reverses the color mapping for disparity visualization without altering the underlying depth output. The model returns a URI pointing to the depth map; no metric scale, units, or calibration is documented, and the schema does not specify image dimensions, file type, or MIME type. No latency, cost, hardware, or image‑size limits are provided for this deployment, though the repository’s tested environment used Ubuntu 20.04 LTS, Python 3.10, CUDA 12.3 and an NVIDIA A800‑SXM4‑80GB GPU. The model is released under Apache‑2.0, with the latest version (ID 26033a14ffa3b967b54daaf425adeba698e9fa67260422b9e41c4907ae0e45f5) created on 2025‑01‑09.

Read original