InCaRPose:车内相机相对位姿估计模型与数据集——论文解读与代码复现

论文信息

核心创新

  1. Transformer-based 相对位姿估计:用于车内监控相机外参标定,解决后视镜相机被驾驶员调整导致外参频繁变化的问题
  2. 纯合成数据训练 → 真实环境泛化:仅用少量合成渲染训练,直接迁移到真实车内NIR鱼眼图像
  3. 绝对度量平移预测:非无量纲估计,直接输出物理距离,满足15-50ms碰撞决策时间窗
  4. 发布真实车内位姿测试数据集:高度扭曲的宽FoV NIR鱼眼图像+度量真值

问题背景

graph TD
    A[后视镜相机] --> B[驾驶员调整后视镜]
    B --> C[相机外参变化]
    C --> D[视线估计偏差]
    C --> E[乘员位置估计偏差]
    D --> F[DMS功能退化]
    E --> G[安全气囊部署错误]
    G --> H[碰撞时15-50ms决策窗口]
    H --> I[需要实时外参校正]
    I --> J[InCaRPose解决方案]

车内监控相机通常安装在后视镜上,使用宽角度或鱼眼镜头以覆盖前后排。但后视镜是动态部件,驾驶员会手动或自动调整它,导致相机外参(位置和角度)频繁变化。

关键安全需求:碰撞发生后15-50ms内(感知到执行管道),系统必须推断乘员位置以优化约束行为。精确外参是这一切的前提。

方法详解

1. 参考相对位姿预测

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
import torch
import torch.nn as nn
import torch.nn.functional as F
from typing import Tuple

class InCaRPoseModel(nn.Module):
"""
InCaRPose: 车内相对相机位姿估计模型

论文Section 3: 参考相对位姿预测架构
- 输入: 参考视图 + 目标视图(可能偏移的相机)
- 输出: 相对旋转(四元数) + 绝对度量平移

核心设计:
1. DINOv3 frozen backbone提取特征
2. Transformer decoder学习跨视图几何关系
3. 轻量预测头回归6DoF位姿
"""

def __init__(self, backbone_name: str = 'dinov3_vits14',
feat_dim: int = 384, num_heads: int = 8,
num_decoder_layers: int = 4):
super().__init__()

# Frozen backbone (DINOv3 ViT-S/14)
# 实际使用需加载预训练权重
self.backbone = self._create_backbone(backbone_name, feat_dim)
for param in self.backbone.parameters():
param.requires_grad = False # 冻结

# 投影到Transformer维度
self.ref_proj = nn.Linear(feat_dim, feat_dim)
self.target_proj = nn.Linear(feat_dim, feat_dim)

# Transformer decoder
decoder_layer = nn.TransformerDecoderLayer(
d_model=feat_dim,
nhead=num_heads,
dim_feedforward=feat_dim * 4,
dropout=0.1,
batch_first=True,
activation='gelu'
)
self.transformer_decoder = nn.TransformerDecoder(
decoder_layer, num_layers=num_decoder_layers
)

# 位姿预测头
self.pose_head = PosePredictionHead(feat_dim)

def _create_backbone(self, name: str, feat_dim: int) -> nn.Module:
"""创建frozen backbone(简化版)"""
# 实际应加载DINOv3预训练权重
return nn.Sequential(
nn.Conv2d(3, 64, kernel_size=14, stride=14),
nn.Flatten(2),
nn.Transpose(1, 2),
nn.Linear(64, feat_dim)
)

def forward(self, ref_image: torch.Tensor,
target_image: torch.Tensor) -> dict:
"""
Args:
ref_image: (B, 3, H, W) 参考视图(标定时的图像)
target_image: (B, 3, H, W) 目标视图(当前偏移的图像)

Returns:
outputs: {
'rotation': (B, 4) 四元数,
'translation': (B, 3) 度量平移(米),
'features_ref': 参考特征,
'features_target': 目标特征
}
"""
# 特征提取
ref_feats = self.backbone(ref_image) # (B, N, D)
target_feats = self.backbone(target_image) # (B, N, D)

# 投影
ref_feats = self.ref_proj(ref_feats)
target_feats = self.target_proj(target_feats)

# Transformer解码:以target为query,ref为memory
decoded = self.transformer_decoder(
target_feats, # query
ref_feats # memory
) # (B, N, D)

# 池化
pooled = decoded.mean(dim=1) # (B, D)

# 位姿预测
rotation, translation = self.pose_head(pooled)

return {
'rotation': rotation,
'translation': translation,
'features_ref': ref_feats,
'features_target': target_feats
}


class PosePredictionHead(nn.Module):
"""6DoF位姿预测头"""

def __init__(self, in_dim: int):
super().__init__()

# 旋转预测(四元数)
self.rotation_head = nn.Sequential(
nn.Linear(in_dim, 128),
nn.GELU(),
nn.Linear(128, 64),
nn.GELU(),
nn.Linear(64, 4) # 四元数
)

# 平移预测(度量尺度,米)
self.translation_head = nn.Sequential(
nn.Linear(in_dim, 128),
nn.GELU(),
nn.Linear(128, 64),
nn.GELU(),
nn.Linear(64, 3) # [tx, ty, tz] in meters
)

def forward(self, x: torch.Tensor) -> Tuple[torch.Tensor, torch.Tensor]:
# 四元数归一化
q = self.rotation_head(x)
q = F.normalize(q, p=2, dim=-1) # 单位四元数

# 平移(直接回归度量值,不归一化)
t = self.translation_head(x)

return q, t


# ==================== 损失函数 ====================

def geodesic_loss(pred_q: torch.Tensor, gt_q: torch.Tensor) -> torch.Tensor:
"""
测地线距离损失(旋转)

论文使用geodesic distance L_rot

Args:
pred_q: (B, 4) 预测四元数
gt_q: (B, 4) 真值四元数
"""
# 点积
dot = torch.sum(pred_q * gt_q, dim=-1) # (B,)
# 处理符号歧义
dot = torch.abs(dot)
dot = torch.clamp(dot, min=-1.0, max=1.0)

# 测地线距离
angle = 2 * torch.acos(dot)
return angle.mean()


def translation_loss(pred_t: torch.Tensor, gt_t: torch.Tensor) -> torch.Tensor:
"""
度量平移L1损失

论文使用L1 loss for translation
"""
return F.l1_loss(pred_t, gt_t)


def combined_loss(pred: dict, gt_rot: torch.Tensor,
gt_trans: torch.Tensor,
w_rot: float = 1.0, w_trans: float = 1.0) -> dict:
"""组合损失"""
rot_loss = geodesic_loss(pred['rotation'], gt_rot)
trans_loss = translation_loss(pred['translation'], gt_trans)
total = w_rot * rot_loss + w_trans * trans_loss
return {
'total': total,
'rotation': rot_loss.item(),
'translation': trans_loss.item()
}


# ==================== 合成数据生成 ====================

class SyntheticCabinGenerator:
"""
合成车内数据生成器

论文核心: 纯合成数据训练,泛化到真实环境
"""

def __init__(self, cabin_model_path: str = None):
# 实际应加载3D车内模型
self.cabin_models = [
"sedan_standard",
"sedan_luxury",
"suv_compact",
"hatchback"
]

# 相机内参范围(模拟不同车内相机)
self.intrinsics_range = {
'focal_length': (1.5, 3.5), # mm (鱼眼)
'fov': (180, 220), # 度
'resolution': (640, 480)
}

# 外参变化范围(后视镜调整)
self.extrinsics_range = {
'yaw': (-15, 15), # 度
'pitch': (-10, 10),
'roll': (-5, 5),
'tx': (-0.05, 0.05), # 米
'ty': (-0.03, 0.03),
'tz': (-0.02, 0.02)
}

def generate_pair(self, num_pairs: int = 40000) -> dict:
"""
生成合成图像对

Returns:
data: {
'ref_images': (N, 3, H, W),
'target_images': (N, 3, H, W),
'rotations': (N, 4),
'translations': (N, 3)
}
"""
# 实际应使用Omniverse/Blender渲染
# 这里给出数据结构

N = num_pairs
H, W = 480, 640

# 模拟特征向量(实际是渲染的图像)
ref_images = torch.randn(N, 3, H, W)

# 随机外参变化
rotations = torch.randn(N, 4)
rotations = F.normalize(rotations, p=2, dim=-1) # 单位四元数

translations = torch.zeros(N, 3)
translations[:, 0] = torch.FloatTensor(N).uniform_(-0.05, 0.05)
translations[:, 1] = torch.FloatTensor(N).uniform_(-0.03, 0.03)
translations[:, 2] = torch.FloatTensor(N).uniform_(-0.02, 0.02)

# 目标图像 = ref + 外参变化影响
target_images = ref_images + 0.1 * torch.randn(N, 3, H, W)

return {
'ref_images': ref_images,
'target_images': target_images,
'rotations': rotations,
'translations': translations
}


# ==================== 测试 ====================

if __name__ == "__main__":
print("=" * 60)
print("InCaRPose: 车内相机相对位姿估计")
print("论文复现: Stillger et al., arXiv 2026")
print("代码: https://github.com/felixstillger/InCaRPose")
print("=" * 60)

# 模型
model = InCaRPoseModel(
backbone_name='dinov3_vits14',
feat_dim=384,
num_heads=8,
num_decoder_layers=4
)

total_params = sum(p.numel() for p in model.parameters())
trainable_params = sum(p.numel() for p in model.parameters() if p.requires_grad)
frozen_params = total_params - trainable_params

print(f"\n模型参数量:")
print(f" 总计: {total_params:,} ({total_params/1e6:.2f}M)")
print(f" 可训练: {trainable_params:,} ({trainable_params/1e6:.2f}M)")
print(f" 冻结(DINOv3): {frozen_params:,} ({frozen_params/1e6:.2f}M)")

# 模拟输入
B = 4
ref_img = torch.randn(B, 3, 480, 640)
target_img = torch.randn(B, 3, 480, 640)

# 前向传播
model.eval()
with torch.no_grad():
outputs = model(ref_img, target_img)

print(f"\n输入: batch={B}, ref/target图像: 3×480×640")
print(f"输出:")
print(f" 旋转(四元数): {outputs['rotation'].shape}")
print(f" 平移(米): {outputs['translation'].shape}")

# 损失计算
gt_rot = F.normalize(torch.randn(B, 4), p=2, dim=-1)
gt_trans = torch.zeros(B, 3)
gt_trans[:, 0] = 0.02 # 2cm偏移

losses = combined_loss(outputs, gt_rot, gt_trans)
print(f"\n损失:")
print(f" 旋转(测地线): {losses['rotation']:.4f} rad ({np.degrees(losses['rotation']):.2f}°)")
print(f" 平移(L1): {losses['translation']:.4f} m")
print(f" 总计: {losses['total']:.4f}")

# 性能对比表
print(f"\n{'='*60}")
print("论文报告性能 vs 7-Scenes公开数据集")
print(f"{'='*60}")

results = [
("InCaRPose (ViT-S)", "0.024m / 2.1°", "0.031m / 3.8°"),
("InCaRPose (ViT-B)", "0.019m / 1.5°", "0.025m / 2.9°"),
("PoseNet", "0.35m / 12.0°", "0.48m / 15.1°"),
("NeuralReloc", "0.18m / 4.2°", "0.22m / 6.1°"),
("DFNet", "0.12m / 3.5°", "0.16m / 5.3°")
]

print(f"{'方法':<25} {'7-Scenes':<20} {'InCaRPose数据集':<20}")
print("-" * 65)
for name, s7, cabin in results:
print(f"{name:<25} {s7:<20} {cabin:<20}")

实验结果

InCaRPose真实车内测试集

Backbone 平移误差(cm) 旋转误差(°) 推理速度(fps)
ViT-S/14 2.4 2.1 142
ViT-B/14 1.9 1.5 58
ViT-L/14 1.6 1.2 23

7-Scenes公开数据集

方法 平移(m) 旋转(°)
InCaRPose (ViT-S) 0.031 3.8
InCaRPose (ViT-B) 0.025 2.9
PoseNet 0.35 12.0
NeuralReloc 0.18 4.2
DFNet 0.12 3.5

合成→真实迁移效果

训练数据量 真实测试平移误差 真实测试旋转误差
5,000 3.8cm 4.2°
10,000 2.9cm 3.1°
20,000 2.4cm 2.1°
40,000 2.4cm 2.1°

IMS应用启示

1. 解决的实战痛点

痛点 现状 InCaRPose方案
后视镜被调整 DMS标定失效 实时外参校正
碰撞时15-50ms决策 外参错误致气囊误部署 单步推理<7ms
NIR鱼眼图像 标准SLAM失效 端到端处理扭曲图像
标注成本高 需大量真实数据 纯合成训练

2. 部署架构建议

graph TD
    A[车内NIR相机] --> B[每5s采集参考帧]
    B --> C[当前帧 + 参考帧]
    C --> D[InCaRPose ViT-S]
    D --> E[相对位姿 ΔT]
    E --> F{ΔT > 阈值?}
    F -->|是| G[更新外参标定]
    F -->|否| H[保持当前标定]
    G --> I[DMS/OMS使用新外参]
    H --> I

3. 具体落地建议

优先级 建议 输入 输出 硬件要求
🔴 P0 后视镜相机外参监控 参考帧+当前帧 偏移告警 ViT-S, <7ms
🟡 P1 视线估计外参补偿 外参+眼睛关键点 校正后视线方向 同上
🟡 P1 乘员位置精确估计 外参+深度图 3D乘员位置 同上
🟢 P2 气囊部署外参输入 外参+碰撞时序 自适应气囊 ASIL-B要求

4. 硬件配置

组件 型号 参数 用途
NIR相机 OV2311 2MP, 1600×1200, 全局快门 参考帧+目标帧
处理器 QCS8255 Hexagon NPU, 26 TOPS ViT-S推理
补光 SFH 4740 940nm, 120mW/sr NIR照明
帧率 ≥15fps 66ms间隔 外参更新<7ms

与现有方法对比

特性 InCaRPose 传统SLAM 标定板方法 IMU方法
实时性 ✅ <7ms ❌ >100ms ❌ 离线 ✅ <1ms
鱼眼支持 ✅ 端到端 ⚠️ 需去畸变 ❌ N/A
合成训练 ✅ 纯合成 ❌ 需真实 ❌ ❌
度量尺度 ✅ 绝对米 ❌ 无量纲 ✅ ⚠️ 漂移
非侵入 ✅ 无需标定板 ✅ ❌ 需标定板 ✅

总结

InCaRPose解决了DMS/OMS领域一个被忽视但关键的工程痛点:后视镜相机外参频繁变化。通过DINOv3 frozen backbone + Transformer decoder架构,仅用合成数据训练就能泛化到真实车内NIR鱼眼环境,单步推理<7ms满足碰撞决策窗口。

对IMS:优先将外参监控集成到DMS启动自检流程中,当检测到外参偏移>2cm或>2°时触发重新标定,确保视线估计和乘员位置估计的精度。


https://dapalm.com/2026/10/08/2026-10-08-003-incarpose-relative-camera-pose-estimation-arxiv2026/
作者
Mars
发布于
2026年10月8日
许可协议