# LingBot-Depth 2.0 发布：深度误差减半，12/16 基准第一

- 来源：Rohan Paul (@rohanpaul_ai)
- 发布时间：2026-07-07 03:37
- AIHOT 分数：58
- AIHOT 链接：https://aihot.virxact.com/items/cmr9mwo2w008eiha3676hr3qm
- 原文链接：https://x.com/rohanpaul_ai/status/2074216188964634804

## AI 摘要

深度补全模型 LingBot-Depth 2.0 发布，专攻玻璃、镜面、透明物体等传统深度相机失效的场景。训练数据从 3M 扩展到 150M（50 倍），在 12/16 个深度补全基准中排名第一，最难室内场景 RMSE 从 0.132 降至 0.062（误差减半）。模型基于视觉基础模型 LingBot-Vision 构建，后者已完全开源，训练时利用物体边缘几何信息且无需人工边界标签。

## 正文

The robot’s “eyes” just received a big upgrade.

LingBot-Depth 2.0, a depth-completion model with half the depth error just dropped. 12/16 benchmarks topped.

Glass, mirrors, and transparent objects are so easy for us humans, but so hard for robots, because they do not behave like ordinary surfaces in a camera pipeline.

A robot that misunderstands a balcony window or a table edge, will have a completely false planning inside a false world. Huge implecation.

LingBot-Depth 2.0 takes an RGB image plus a broken depth map from a sensor and then outputs a cleaner depth map and a usable 3D point cloud.

Numbers on LingBot-Depth 2.0
• Excels on glass, mirrors & transparent objects — where traditional depth cameras fail
• Training data: 3M → 150M (50x scale-up)
• 12 out of 16 first-place rankings on depth completion benchmarks
• RMSE cut in half: 0.132 → 0.062 on the hardest indoor scenes

LingBot-Vision trained on boundaries, because object edges carry the geometry robots need. No human boundary labels are used, which makes this approach easier to scale.

The open-sourced LingBot-Vision is the general vision backbone, and LingBot-Depth 2.0 is the depth model built on it.

### 引用推文

> Robbyant：🪞 Glass. Mirrors. Transparent objects. — The nightmare of every depth camera. We just solved it! Introducing LingBot-Depth 2.0: 150M-scale training, half the d...
