FIFA: 基于细粒度帧间注意力的驾驶员视频视线估计 | CVPR 2025 论文解读

论文信息

核心创新

FIFA 提出了一种细粒度帧间注意力机制,专门针对驾驶员视频视线估计任务设计。与传统的单帧视线估计不同,FIFA 利用视频帧之间的细粒度时序信息,通过注意力机制捕捉驾驶员视线在连续帧中的微妙变化,从而提高动态驾驶场景下的视线估计精度。

方法详解

1. 问题定义

驾驶员视线估计目标是从车载摄像头捕获的视频序列中估计驾驶员的注视方向或注视区域。与传统单帧方法不同,FIFA 将其建模为时序问题:

1
2
3
输入: 视频帧序列 X = {x_1, x_2, ..., x_T}
输出: 每帧视线方向 G = {g_1, g_2, ..., g_T}
其中 g_t ∈ R^3 (3D视线方向向量)

2. 细粒度帧间注意力机制

FIFA 的核心是帧间注意力模块,设计动机为:

  • 相邻帧的视线变化通常是微妙的:驾驶员从看路面到看手机,中间有过渡帧
  • 传统时序方法(LSTM/3D-CNN)容易丢失细粒度信息
  • 注意力机制可以自适应地关注关键过渡帧
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
import torch
import torch.nn as nn
import torch.nn.functional as F

class FineGrainedInterFrameAttention(nn.Module):
"""
FIFA: 细粒度帧间注意力模块

论文: FIFA: Fine-grained Inter-frame Attention for
Driver's Video Gaze Estimation (CVPR 2025)

核心思想: 在相邻视频帧之间建立细粒度注意力,
捕捉驾驶员视线的微妙时序变化
"""

def __init__(self, feat_dim=512, num_heads=8, window_size=5):
super().__init__()
self.feat_dim = feat_dim
self.num_heads = num_heads
self.window_size = window_size # 时序窗口大小

# 帧内特征提取 (per-frame feature)
self.frame_encoder = nn.Sequential(
nn.Conv2d(3, 64, kernel_size=7, stride=2, padding=3),
nn.BatchNorm2d(64),
nn.ReLU(inplace=True),
nn.MaxPool2d(kernel_size=3, stride=2, padding=1),
nn.Conv2d(64, 128, kernel_size=3, stride=2, padding=1),
nn.BatchNorm2d(128),
nn.ReLU(inplace=True),
nn.AdaptiveAvgPool2d((1, 1)),
)

# 细粒度帧间注意力
self.inter_frame_attn = nn.MultiheadAttention(
embed_dim=feat_dim,
num_heads=num_heads,
batch_first=True
)

# 差异编码器: 编码相邻帧间的差异
self.diff_encoder = nn.Sequential(
nn.Linear(feat_dim, feat_dim // 2),
nn.ReLU(inplace=True),
nn.Linear(feat_dim // 2, feat_dim),
)

# 视线回归头
self.gaze_regressor = nn.Sequential(
nn.Linear(feat_dim * 2, feat_dim),
nn.ReLU(inplace=True),
nn.Dropout(0.3),
nn.Linear(feat_dim, 3) # 3D gaze direction
)

self.norm = nn.LayerNorm(feat_dim)

def forward(self, video_clip):
"""
前向传播

Args:
video_clip: 视频片段, shape=(B, T, C, H, W)
B=batch, T=帧数, C=3, H/W=分辨率

Returns:
gaze_pred: 视线方向预测, shape=(B, T, 3)
attn_weights: 注意力权重, shape=(B, T, T)
"""
B, T, C, H, W = video_clip.shape

# 1. 逐帧特征提取
frames_flat = video_clip.view(B * T, C, H, W)
frame_feats = self.frame_encoder(frames_flat) # (B*T, 128, 1, 1)
frame_feats = frame_feats.view(B, T, -1) # (B, T, 128)

# 维度对齐
if frame_feats.size(-1) != self.feat_dim:
frame_feats = F.pad(
frame_feats,
(0, self.feat_dim - frame_feats.size(-1))
)

# 2. 计算相邻帧差异
frame_diffs = torch.zeros_like(frame_feats)
frame_diffs[:, 1:] = frame_feats[:, 1:] - frame_feats[:, :-1]
diff_embeds = self.diff_encoder(frame_diffs)

# 3. 细粒度帧间注意力
# 将原始特征与差异特征融合
fused = self.norm(frame_feats + diff_embeds)

# 多头注意力
attn_out, attn_weights = self.inter_frame_attn(
fused, fused, fused
)

# 残差连接
attn_out = self.norm(attn_out + frame_feats)

# 4. 视线回归
gaze_input = torch.cat([attn_out, diff_embeds], dim=-1)
gaze_pred = self.gaze_regressor(gaze_input) # (B, T, 3)

# 归一化为单位向量
gaze_pred = F.normalize(gaze_pred, p=2, dim=-1)

return gaze_pred, attn_weights


class FIFALoss(nn.Module):
"""FIFA 训练损失函数"""

def __init__(self, alpha=1.0, beta=0.3):
super().__init__()
self.alpha = alpha # 视线方向损失权重
self.beta = beta # 时序一致性损失权重

def forward(self, gaze_pred, gaze_gt, attn_weights):
"""
Args:
gaze_pred: 预测视线, (B, T, 3)
gaze_gt: 真实视线, (B, T, 3)
attn_weights: 注意力权重, (B, T, T)
"""
# 角度损失 (cosine distance)
cos_sim = F.cosine_similarity(gaze_pred, gaze_gt, dim=-1)
angle_loss = (1 - cos_sim).mean()

# 时序一致性损失
temporal_diff = gaze_pred[:, 1:] - gaze_pred[:, :-1]
temporal_loss = torch.norm(temporal_diff, p=2, dim=-1).mean()

# 注意力熵正则化
attn_entropy = -(attn_weights *
torch.log(attn_weights + 1e-8)).sum(-1).mean()

total_loss = (
self.alpha * angle_loss +
self.beta * temporal_loss -
0.01 * attn_entropy
)

return total_loss, {
'angle_loss': angle_loss.item(),
'temporal_loss': temporal_loss.item(),
'attn_entropy': attn_entropy.item()
}


# 实际测试
if __name__ == "__main__":
# 模拟驾驶视频片段
batch_size = 4
num_frames = 5 # 5帧窗口
model = FineGrainedInterFrameAttention(
feat_dim=512, num_heads=8, window_size=5
)

# 模拟输入: 4个视频片段,每个5帧
video = torch.randn(batch_size, num_frames, 3, 224, 224)

# 前向传播
gaze_pred, attn_weights = model(video)

print(f"输入形状: {video.shape}")
print(f"视线预测形状: {gaze_pred.shape}")
print(f"注意力权重形状: {attn_weights.shape}")
print(f"预测视线向量范数: {torch.norm(gaze_pred, p=2, dim=-1)[0]}")

# 计算与真实标签的损失
criterion = FIFALoss()
gaze_gt = F.normalize(torch.randn_like(gaze_pred), p=2, dim=-1)
loss, metrics = criterion(gaze_pred, gaze_gt, attn_weights)

print(f"\n损失: {loss.item():.4f}")
for k, v in metrics.items():
print(f" {k}: {v:.4f}")

3. 时序窗口策略

FIFA 使用滑动窗口策略处理长视频:

窗口大小 延迟 精度 适用场景
3帧 100ms 较低 实时性要求高
5帧 167ms 中等 平衡模式
7帧 233ms 最高 离线分析

4. 与现有方法对比

方法 输入 时序建模 参数量 角度误差
GazeNet 单帧 无 5.2M 7.8°
GazeTR 单帧 无 8.1M 6.5°
LSTM-Gaze 视频 LSTM 6.3M 6.2°
FIFA 视频 帧间注意力 5.8M 5.1°

IMS 开发启示

1. 算法集成建议

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
# IMS 集成 FIFA 的接口设计
class IFAGazeEstimator:
"""IMS 视线估计模块 - 基于 FIFA"""

def __init__(self, model_path, device='cuda'):
self.model = FineGrainedInterFrameAttention()
self.model.load_state_dict(
torch.load(model_path, map_location=device)
)
self.model.to(device).eval()
self.frame_buffer = deque(maxlen=5) # 5帧缓冲

def process_frame(self, frame):
"""
处理单帧,返回视线估计结果

Args:
frame: BGR图像, (H, W, 3)

Returns:
gaze_direction: 3D视线方向 (yaw, pitch, roll)
confidence: 置信度
"""
rgb = cv2.cvtColor(frame, cv2.COLOR_BGR2RGB)
tensor = self._preprocess(rgb)
self.frame_buffer.append(tensor)

if len(self.frame_buffer) < 5:
return None, 0.0

clip = torch.stack(list(self.frame_buffer)).unsqueeze(0)
with torch.no_grad():
gaze, attn = self.model(clip.to(self.device))

return gaze[0, -1].cpu().numpy(), 0.95

2. 部署优化

  • 模型量化:FP32 → INT8,精度损失 <2%
  • 窗口优化:推理时窗口=5帧,内存占用约 50MB
  • 帧率适配:支持 15/25/30fps 输入

3. 与 IMS 疲劳/分心检测的协同

IMS 模块 FIFA 输出 用途
分心检测 视线方向 视线偏离道路 >3s 触发警告
疲劳检测 注意力权重 注意力分散模式分析
危险行为 帧间差异 突然视线转移检测

实验结果

数据集

  • DDP (Driver Drive Pose): 50名驾驶员,100小时视频
  • DAISA: 分心/警觉标注数据集
  • 自建数据集: 30名驾驶员,包含白天/夜晚/隧道场景

性能对比

指标 GazeNet GazeTR FIFA
角度误差(°) 7.8 6.5 5.1
注视区域准确率 82.3% 85.7% 89.2%
帧率(RTX3090) 120fps 95fps 85fps
帧率(Jetson) 28fps 22fps 20fps

总结

FIFA 的核心贡献在于将细粒度帧间注意力引入驾驶员视线估计,相比传统单帧方法:

  1. 精度提升:角度误差降低 1.4°-2.7°
  2. 时序鲁棒:利用视频时序信息,减少抖动
  3. 部署友好:参数量 5.8M,支持边缘部署

对 IMS 的直接价值:

  • 可替代现有单帧视线估计模块,提升分心检测精度
  • 注意力权重可作为认知分心的辅助特征
  • 帧间差异编码器可检测突然视线转移(危险行为)

https://dapalm.com/2026/10/03/2026-10-03-231-fifa-fine-grained-inter-frame-attention-driver-gaze-cvpr2025/
作者
Mars
发布于
2026年10月3日
许可协议