实验 32:A26 Depth-confidence Regional Lock

把 regional visible taxonomy 与 VGGT depth/confidence 聚合成风险分数,诊断 case-level lock 是否足够

实验定位 诊断负结果 / 指向局部 depth map

A26 接着 A25 做 depth-confidence regional lock:把 VGGT depth delta、visible loss、under-coverage、IoU worse 和 normal delta 聚合成 case-level risk score,并用小网格搜索诊断无 GT 规则的折中。GT 只用于回顾评估,不能把该阈值当作泛化结论。

结果:A26 消除了 A25 regional 的 false accept(2→0),但 false reject 仍为 4,没有优于 A24。它还错过了 gso_002 这种 depth 改善但 coverage taxonomy 看起来有风险的 case,说明 case-level risk score 仍太粗。

实验设计(Image-2 风格)

实验设计图:A26 depth-confidence regional lock
实验设计图:A26 depth-confidence regional lock
模块设计图:case-level risk is not enough
模块设计图:case-level risk is not enough

模块设计(Image-2 风格)

Hypothesis

问题:A25 regional visible gate reduces false rejects but accepts silhouette-benign GT-harmful repairs; depth-coupled variants overcorrect.

Depth-confidence lock:Accept repair only when depth is not hard-failing, regional visible support is not deleted, and there is positive depth or silhouette evidence.

限制:The threshold grid is retrospective on 10 cases, so it diagnoses signal structure rather than proving generalization.

Selected Diagnostic Thresholds

risk threshold:0.025

evidence threshold:0.000

阈值只用于 10-case retrospective 诊断,不作为最终方法超参。

实验结果(表格)

策略选择 repair 数false rejectfalse acceptselected Chamfer Δselected F@5 Δ选择的 repair case
A24 calibrated340-0.000540.0067gso_000_input4, gso_002_input4, gso_009_input4
A25 regional visible722-0.000290.0051gso_000_input4, gso_001_input4, gso_002_input4, gso_003_input4, gso_004_input4, gso_005_input4, gso_007_input4
A26 depth-confidence risk340-0.000210.0019gso_000_input4, gso_003_input4, gso_004_input4

A26 逐 case 决策

caseA26 gateriskevidenceVGGT depth Δmean IoU Δmax visible lossGT Chamfer ΔGT F@5 Δ原因
gso_000_input4pass-0.06940.1450-0.023300.14000.0538-0.001600.0202[]
gso_001_input4reject0.03940.0002-0.00020-0.00130.0080-0.001050.0082['risk score above threshold']
gso_002_input4reject0.03520.0304-0.006560.02280.0386-0.003770.0502['risk score above threshold']
gso_003_input4pass0.00710.00000.00193-0.00010.0022-0.00005-0.0041[]
gso_004_input4pass0.01750.0001-0.00009-0.00060.0020-0.000410.0032[]
gso_005_input4reject0.03220.00000.000080.00000.00000.00119-0.0079['risk score above threshold']
gso_006_input4reject0.21140.00230.03191-0.08310.04770.001440.0204['depth lock: candidate moves visible geometry away from VGGT', 'risk score above threshold']
gso_007_input4reject0.02740.00150.001940.00130.00860.00281-0.0188['risk score above threshold']
gso_008_input4reject0.20240.00050.01519-0.02370.08930.00474-0.0168['depth lock: candidate moves visible geometry away from VGGT', 'regional lock: severe visible support loss', 'risk score above threshold']
gso_009_input4reject0.09900.01020.00017-0.00270.0722-0.00006-0.0029['risk score above threshold']

可视化结果

A26 与 A24/A25 的错误类型对比
A26 与 A24/A25 的错误类型对比
A26 与 A24/A25 的回顾质量对比
A26 与 A24/A25 的回顾质量对比

实验结论

下一步想法