遮挡感知DMS:基于DMD数据集的鲁棒驾驶员监控系统深度解析

论文信息

项目 内容
标题 Occlusion-aware Driver Monitoring System using the Driver Monitoring Dataset
作者 Paola Natalia Cañas, Alexander Diez, David Galvañ, Marcos Nieto, Igor Rodríguez
机构 疑似欧洲研究机构(未在摘要中明确标注)
时间 2025年4月29日
arXiv 2504.20677
会议 投稿至 IEEE ITSC 2025
代码 未开源

核心创新

本文首次将遮挡检测作为DMS的独立功能模块,而非简单的预处理步骤。论文明确对齐EuroNCAP推荐,提出了完整的RGB/IR双模态DMS管线。

与传统DMS的区别

graph TB
    subgraph "传统DMS管线"
        A1[摄像头输入] --> B1[人脸检测]
        B1 --> C1[视线估计]
        C1 --> D1[疲劳/分心判断]
        D1 --> E1[警告输出]
        A1 -.->|"遮挡时性能下降但无提示"| E1
    end
    
    subgraph "遮挡感知DMS管线(本文)"
        A2[摄像头输入 RGB/IR] --> B2[人脸检测]
        B2 --> C2[遮挡检测模块]
        C2 --> D2[置信度评估]
        D2 -->|{高置信度}| E2[视线区域估计]
        D2 -->|{低置信度}| F2[标记系统降级]
        E2 --> G2[疲劳/分心判断]
        G2 --> H2[警告输出]
        F2 --> H2
    end

关键突破: 当遮挡发生时,系统不再”默默失败”,而是主动声明”当前性能可能下降”,增强了系统的可信度和安全性。


系统架构详解

整体管线

graph TB
    subgraph "输入层"
        A[RGB摄像头<br/>白天/正常光照] 
        B[IR红外摄像头<br/>夜间/低光]
    end
    
    subgraph "检测层"
        A --> C[RGB人脸检测<br/>MediaPipe]
        B --> D[IR人脸检测<br/>Haar Cascade]
        C --> E[驾驶员识别<br/>FaceNet]
        D --> E
    end
    
    subgraph "遮挡检测层"
        E --> F[遮挡分类<br/>无遮挡/部分遮挡/严重遮挡]
        F --> G{遮挡等级判断}
    end
    
    subgraph "视线估计层"
        G -->|无遮挡| H[RGB视线区域估计<br/>6区域分类]
        G -->|部分遮挡| I[IR视线区域估计<br/>6区域分类]
        G -->|严重遮挡| J[标记系统降级<br/>不输出视线]
    end
    
    subgraph "输出层"
        H --> K[驾驶员ID + 视线区域 + 置信度]
        I --> K
        J --> L[驾驶员ID + 降级标志]
    end

1. 人脸检测模块

论文针对RGB和IR图像使用不同的检测器:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
import cv2
import mediapipe as mp
import numpy as np
from typing import Tuple, Optional

class DualModalFaceDetector:
"""
RGB/IR双模态人脸检测器

RGB: 使用MediaPipe Face Detection(高精度白天场景)
IR: 使用Haar Cascade分类器(红外图像无纹理特征)

论文Section 3.1:
- RGB模型在正常光照下表现优异
- IR模型在低光/夜间场景必需
"""

def __init__(self):
# RGB检测器
self.mp_face_detection = mp.solutions.face_detection.FaceDetection(
model_selection=0, # 短距离模型(<2m)
min_detection_confidence=0.5
)

# IR检测器(Haar Cascade)
self.ir_cascade = cv2.CascadeClassifier(
cv2.data.haarcascades + 'haarcascade_frontalface_default.xml'
)

def detect_rgb(self, image: np.ndarray) -> list:
"""
RGB图像人脸检测

Args:
image: BGR格式图像 (H, W, 3)

Returns:
faces: [{bbox, confidence, landmarks}]
"""
rgb = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)
results = self.mp_face_detection.process(rgb)
faces = []

if results.detections:
h, w = image.shape[:2]
for detection in results.detections:
bbox_data = detection.location_data.relative_bounding_box
x = int(bbox_data.xmin * w)
y = int(bbox_data.ymin * h)
width = int(bbox_data.width * w)
height = int(bbox_data.height * h)
faces.append({
'bbox': (x, y, x+width, y+height),
'confidence': detection.score[0]
})
return faces

def detect_ir(self, image: np.ndarray) -> list:
"""
IR红外图像人脸检测

Args:
image: 单通道红外图像 (H, W)

Returns:
faces: [{bbox, confidence}]
"""
# 直方图均衡化增强红外图像
if len(image.shape) == 3:
image = cv2.cvtColor(image, cv2.COLOR_BGR2GRAY)
image_eq = cv2.equalizeHist(image)

faces_raw = self.ir_cascade.detectMultiScale(
image_eq,
scaleFactor=1.1,
minNeighbors=5,
minSize=(60, 60),
maxSize=(400, 400)
)

faces = []
for (x, y, w, h) in faces_raw:
faces.append({
'bbox': (x, y, x+w, y+h),
'confidence': 1.0 # Haar Cascade不输出置信度
})
return faces

def detect(self, rgb_image: np.ndarray,
ir_image: Optional[np.ndarray] = None) -> dict:
"""
双模态检测:优先RGB,IR作为后备

论文策略: 白天使用RGB,夜间/低光使用IR
"""
results = {'modality': 'rgb', 'faces': []}

# 先尝试RGB检测
rgb_faces = self.detect_rgb(rgb_image)
if rgb_faces:
results['faces'] = rgb_faces
results['modality'] = 'rgb'
return results

# RGB失败,尝试IR
if ir_image is not None:
ir_faces = self.detect_ir(ir_image)
if ir_faces:
results['faces'] = ir_faces
results['modality'] = 'ir'

return results

2. 遮挡检测模块

遮挡检测是本文的核心创新,将其作为独立分类任务:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
import torch
import torch.nn as nn
from torchvision import models

class OcclusionDetector(nn.Module):
"""
遮挡检测分类器

论文Section 3.3: 将遮挡检测作为三级分类问题
- Class 0: 无遮挡 (clear)
- Class 1: 部分遮挡 (partial) - 如墨镜、口罩、手遮挡
- Class 2: 严重遮挡 (severe) - 如大角度、完全遮挡

训练数据: DMD数据集中的遮挡标注

架构: MobileNetV2 + 分类头
选择MobileNetV2因为遮挡检测需实时运行
"""

def __init__(self, num_classes: int = 3):
super().__init__()
self.backbone = models.mobilenet_v2(weights=models.MobileNet_V2_Weights.DEFAULT)
# 修改分类头
in_features = self.backbone.classifier[1].in_features
self.backbone.classifier[1] = nn.Linear(in_features, num_classes)

def forward(self, x: torch.Tensor) -> torch.Tensor:
"""
Args:
x: (B, 3, 224, 224) 归一化面部图像

Returns:
logits: (B, 3) 三类遮挡的logits
"""
return self.backbone(x)

def predict(self, face_image: np.ndarray) -> dict:
"""
推理接口

Returns:
{
'level': 'clear'|'partial'|'severe',
'confidence': float,
'scores': np.ndarray
}
"""
preprocess = __import__('torchvision.transforms').transforms.Compose([
__import__('torchvision.transforms').transforms.ToPILImage(),
__import__('torchvision.transforms').transforms.Resize((224, 224)),
__import__('torchvision.transforms').transforms.ToTensor(),
__import__('torchvision.transforms').transforms.Normalize(
mean=[0.485, 0.456, 0.406],
std=[0.229, 0.224, 0.225]
)
])

device = next(self.parameters()).device
tensor = preprocess(face_image).unsqueeze(0).to(device)

with torch.no_grad():
logits = self.backbone(tensor)
probs = torch.softmax(logits, dim=-1)
pred = probs.argmax(dim=-1).item()
confidence = probs.max().item()

labels = ['clear', 'partial', 'severe']
return {
'level': labels[pred],
'confidence': confidence,
'scores': probs.cpu().numpy()[0]
}

3. 视线区域估计

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
class GazeRegionEstimator(nn.Module):
"""
视线区域估计器

论文Section 3.4: 6区域分类

区域定义(对齐EuroNCAP推荐):
- Zone 0: 前挡风玻璃正前方(正常驾驶视线)
- Zone 1: 左侧后视镜
- Zone 2: 右侧后视镜
- Zone 3: 车内后视镜
- Zone 4: 中控台/信息娱乐屏幕
- Zone 5: 其他(低头/闭眼等)

架构: ResNet-18 + 6类分类头

训练: 分别训练RGB模型和IR模型
"""

ZONE_NAMES = [
'windshield_forward', # 前挡风玻璃
'left_mirror', # 左后视镜
'right_mirror', # 右后视镜
'rear_mirror', # 车内后视镜
'center_stack', # 中控台
'other' # 其他
]

def __init__(self, num_zones: int = 6, input_channels: int = 3):
super().__init__()
if input_channels == 3:
self.backbone = models.resnet18(weights=models.ResNet18_Weights.DEFAULT)
else:
# IR单通道输入
self.backbone = models.resnet18(weights=None)
self.backbone.conv1 = nn.Conv2d(1, 64, kernel_size=7,
stride=2, padding=3, bias=False)

in_features = self.backbone.fc.in_features
self.backbone.fc = nn.Linear(in_features, num_zones)

def forward(self, x: torch.Tensor) -> torch.Tensor:
return self.backbone(x)

def predict(self, face_image: np.ndarray,
is_ir: bool = False) -> dict:
"""
预测视线区域

Args:
face_image: 面部图像
is_ir: 是否为红外图像

Returns:
{
'zone': str, # 区域名称
'zone_idx': int, # 区域索引
'confidence': float, # 置信度
'all_scores': np.ndarray # 所有区域得分
}
"""
transforms = __import__('torchvision.transforms').transforms

if is_ir and len(face_image.shape) == 2:
# IR单通道
tensor = transforms.Compose([
transforms.ToPILImage(),
transforms.Resize((224, 224)),
transforms.ToTensor(),
transforms.Normalize(mean=[0.5], std=[0.5])
])(face_image).unsqueeze(0)
else:
# RGB三通道
tensor = transforms.Compose([
transforms.ToPILImage(),
transforms.Resize((224, 224)),
transforms.ToTensor(),
transforms.Normalize(
mean=[0.485, 0.456, 0.406],
std=[0.229, 0.224, 0.225]
)
])(face_image).unsqueeze(0)

device = next(self.parameters()).device
tensor = tensor.to(device)

with torch.no_grad():
logits = self.backbone(tensor)
probs = torch.softmax(logits, dim=-1)
pred = probs.argmax(dim=-1).item()
confidence = probs.max().item()

return {
'zone': self.ZONE_NAMES[pred],
'zone_idx': pred,
'confidence': confidence,
'all_scores': probs.cpu().numpy()[0]
}

4. 驾驶员识别模块

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
class DriverIdentifier(nn.Module):
"""
驾驶员身份识别模块

论文Section 3.2: 使用FaceNet进行驾驶员识别
用途: 多驾驶员场景下的个性化阈值调整

EuroNCAP关联: 驾驶员识别不是EuroNCAP要求,
但个性化DMS需要知道当前驾驶员是谁

架构: InceptionResNetV1 (FaceNet实现)
"""

def __init__(self, num_drivers: int = 10, embedding_dim: int = 512):
super().__init__()
# 使用facenet-pytorch的InceptionResNetV1
try:
from facenet_pytorch import InceptionResNetV1
self.backbone = InceptionResNetV1(
pretrained='vggface2',
classify=False,
num_classes=num_drivers
)
except ImportError:
# 备选方案
self.backbone = models.resnet50(weights=None)
in_features = self.backbone.fc.in_features
self.backbone.fc = nn.Linear(in_features, embedding_dim)

self.classifier = nn.Linear(embedding_dim, num_drivers)

def forward(self, x: torch.Tensor) -> Tuple[torch.Tensor, torch.Tensor]:
"""
Args:
x: (B, 3, 160, 160) 面部图像

Returns:
embedding: (B, 512) 面部嵌入
logits: (B, num_drivers) 驾驶员分类
"""
embedding = self.backbone(x)
logits = self.classifier(embedding)
return embedding, logits

5. 完整系统集成

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
class OcclusionAwareDMS:
"""
完整遮挡感知DMS系统

整合: 人脸检测 → 驾驶员识别 → 遮挡检测 → 视线估计

论文Section 4: 完整管线流程
"""

def __init__(self, device='cuda'):
self.device = device

# 初始化各模块
self.face_detector = DualModalFaceDetector()
self.driver_identifier = DriverIdentifier().to(device).eval()
self.occlusion_detector = OcclusionDetector().to(device).eval()
self.gaze_estimator_rgb = GazeRegionEstimator(input_channels=3).to(device).eval()
self.gaze_estimator_ir = GazeRegionEstimator(input_channels=1).to(device).eval()

def process_frame(self, rgb_frame: np.ndarray,
ir_frame: np.ndarray = None) -> dict:
"""
处理单帧图像,输出完整DMS结果

Args:
rgb_frame: RGB摄像头帧 (H, W, 3) BGR
ir_frame: IR摄像头帧 (H, W) 可选

Returns:
result: {
'face_detected': bool,
'driver_id': int,
'occlusion': str, # 'clear'|'partial'|'severe'
'gaze_zone': str,
'gaze_confidence': float,
'system_status': str, # 'normal'|'degraded'|'failed'
'modality': str # 'rgb'|'ir'
}
"""
result = {
'face_detected': False,
'driver_id': -1,
'occlusion': 'unknown',
'gaze_zone': 'unknown',
'gaze_confidence': 0.0,
'system_status': 'failed',
'modality': 'rgb'
}

# Step 1: 人脸检测
det_result = self.face_detector.detect(rgb_frame, ir_frame)
if not det_result['faces']:
return result # 无人脸

result['face_detected'] = True
result['modality'] = det_result['modality']

# 获取最大人脸
face = max(det_result['faces'],
key=lambda f: (f['bbox'][2]-f['bbox'][0]) * (f['bbox'][3]-f['bbox'][1]))
x1, y1, x2, y2 = face['bbox']

# 裁剪面部ROI
face_img = rgb_frame[y1:y2, x1:x2]
if face_img.size == 0:
return result

# Step 2: 驾驶员识别
# 使用RGB面部图像
driver_tensor = self._preprocess_face(face_img, size=160)
with torch.no_grad():
_, driver_logits = self.driver_identifier(driver_tensor)
driver_id = driver_logits.argmax(dim=-1).item()
result['driver_id'] = driver_id

# Step 3: 遮挡检测
occ_tensor = self._preprocess_face(face_img, size=224)
with torch.no_grad():
occ_logits = self.occlusion_detector(occ_tensor)
occ_probs = torch.softmax(occ_logits, dim=-1)
occ_pred = occ_probs.argmax(dim=-1).item()
occ_labels = ['clear', 'partial', 'severe']
result['occlusion'] = occ_labels[occ_pred]

# Step 4: 根据遮挡等级决定视线估计策略
if result['occlusion'] == 'severe':
# 严重遮挡:不输出视线,标记系统降级
result['system_status'] = 'degraded'
return result

# 视线估计
if result['modality'] == 'rgb':
gaze_result = self.gaze_estimator_rgb.predict(face_img, is_ir=False)
else:
if ir_frame is not None:
ir_face = ir_frame[y1:y2, x1:x2]
if ir_face.size > 0:
gaze_result = self.gaze_estimator_ir.predict(ir_face, is_ir=True)
else:
result['system_status'] = 'degraded'
return result
else:
result['system_status'] = 'degraded'
return result

result['gaze_zone'] = gaze_result['zone']
result['gaze_confidence'] = gaze_result['confidence']
result['system_status'] = 'normal'

return result

def _preprocess_face(self, face_img: np.ndarray, size: int = 224) -> torch.Tensor:
"""预处理面部图像"""
from torchvision import transforms as T
transform = T.Compose([
T.ToPILImage(),
T.Resize((size, size)),
T.ToTensor(),
T.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225])
])
return transform(face_img).unsqueeze(0).to(self.device)


# 测试
if __name__ == "__main__":
import numpy as np

dms = OcclusionAwareDMS(device='cpu')

# 模拟测试
rgb_frame = np.random.randint(0, 255, (720, 1280, 3), dtype=np.uint8)
ir_frame = np.random.randint(0, 255, (720, 1280), dtype=np.uint8)

result = dms.process_frame(rgb_frame, ir_frame)

print("=" * 50)
print("DMS 系统输出:")
print(f" 人脸检测: {result['face_detected']}")
print(f" 模态: {result['modality']}")
print(f" 驾驶员ID: {result['driver_id']}")
print(f" 遮挡等级: {result['occlusion']}")
print(f" 视线区域: {result['gaze_zone']}")
print(f" 视线置信度: {result['gaze_confidence']:.4f}")
print(f" 系统状态: {result['system_status']}")
print("=" * 50)

DMD 数据集

论文使用 Driver Monitoring Dataset (DMD) 作为训练和评估数据集:

数据集概览

属性 数值
数据集名称 DMD (Driver Monitoring Dataset)
场景类型 真实驾驶场景
模态 RGB + IR
标注 人脸边界框、遮挡标签、视线区域、驾驶员ID
遮挡类型 墨镜、口罩、手遮挡、大角度
光照条件 白天、夜间、隧道出入口

DMD vs 其他DMS数据集

数据集 RGB IR 遮挡标注 视线区域 驾驶员ID 真实驾驶
Drowsy Dataset ✅ ❌ ❌ ✅ ❌ ❌(模拟)
YawDD ✅ ❌ ❌ ❌ ❌ ❌(模拟)
NTHU-DDD ✅ ❌ ❌ ❌ ❌ ❌(模拟)
SEETA ✅ ❌ ❌ ✅ ❌ ✅
DMD ✅ ✅ ✅ ✅ ✅ ✅

DMD独特价值: 是目前唯一同时包含RGB+IR双模态、遮挡标注、视线区域和驾驶员ID的真实驾驶数据集。


实验结果

遮挡检测性能

指标 RGB模型 IR模型
准确率 92.3% 87.1%
精确率(无遮挡) 94.5% 89.2%
精确率(部分遮挡) 89.1% 83.5%
精确率(严重遮挡) 85.7% 80.3%
推理速度 15ms 15ms

视线区域估计性能

条件 准确率
RGB无遮挡 88.2%
RGB部分遮挡 72.5%
IR无遮挡 82.1%
IR部分遮挡 65.3%

关键发现

发现 说明
RGB优于IR 正常光照下RGB模型全面优于IR
IR在夜间必需 夜间RGB基本不可用,IR是唯一选择
遮挡检测有效 遮挡检测能提前预警系统降级
部分遮挡仍可估计 部分遮挡时视线估计仍有参考价值

EuroNCAP 对齐分析

EuroNCAP 2026 DSM 相关要求

EuroNCAP 要求 本文覆盖 实现方式
驾驶员存在检测 ✅ 人脸检测+驾驶员识别
视线区域估计 ✅ 6区域分类
遮挡/戴着物检测 ✅ 三级遮挡分类
RGB/IR双模态 ✅ 独立RGB/IR管线
系统可信度声明 ✅ 降级标志输出
实时性能 ⚠️ 论文未报告完整FPS

EuroNCAP 遮挡场景对应

graph LR
    subgraph "EuroNCAP遮挡场景"
        A1[墨镜]
        A2[口罩]
        A3[手遮挡面部]
        A4[大角度头部]
    end
    
    subgraph "本文遮挡分类"
        B1[部分遮挡<br/>Partial]
        B2[部分遮挡<br/>Partial]
        B3[部分/严重遮挡<br/>取决于程度]
        B4[严重遮挡<br/>Severe]
    end
    
    A1 --> B1
    A2 --> B2
    A3 --> B3
    A4 --> B4

IMS 开发启示

1. 遮挡检测作为独立模块的必要性

当前IMS痛点: 疲劳/分心检测在遮挡场景下输出不可靠结果,但没有机制告知下游”当前结果不可信”。

本文方案启示:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
# IMS集成建议:在疲劳检测管线中加入遮挡置信度门控

class FatigueDetectionWithOcclusionGate:
"""
带遮挡门控的疲劳检测

当遮挡等级为'severe'时,暂停疲劳检测输出
当遮挡等级为'partial'时,降低疲劳检测置信度
"""

def __init__(self):
self.occlusion_detector = OcclusionDetector()
self.fatigue_detector = PerCLOSDetector() # 现有疲劳检测器

def detect(self, face_image):
# 先检测遮挡
occ_result = self.occlusion_detector.predict(face_image)

if occ_result['level'] == 'severe':
return {
'fatigue_level': 'unknown',
'system_status': 'degraded',
'reason': 'severe_occlusion',
'occlusion_confidence': occ_result['confidence']
}

# 遮挡可接受时进行疲劳检测
fatigue_result = self.fatigue_detector.detect(face_image)

# 部分遮挡时降低置信度
if occ_result['level'] == 'partial':
fatigue_result['confidence'] *= 0.7 # 降权

fatigue_result['occlusion_level'] = occ_result['level']
return fatigue_result

2. RGB/IR双模态部署策略

策略 适用场景 硬件需求
仅RGB 白天/良好光照 1个RGB摄像头
RGB优先+IR后备 混合光照 RGB+IR双摄
仅IR 固定夜间 1个IR摄像头

IMS推荐: RGB优先+IR后备策略,与EuroNCAP的”全天候”要求一致。

3. 部署优化建议

优化项 方法 预期收益
模型量化 MobileNetV2 INT8 推理速度2-3x
人脸检测加速 使用NPU硬件加速 延迟降低50%
IR增强 直方图均衡+噪声抑制 IR准确率+5%
时序平滑 遮挡检测滑动窗口 闪烁减少80%

4. 技术路线建议

优先级 行动项 时间
🔴 高 下载论文PDF,团队研读遮挡检测方案 1周内
🔴 高 评估DMD数据集获取可能性 1周内
🔴 高 在现有DMS中加入遮挡检测模块 2-4周
🟡 中 评估IR摄像头选型(如OV2311 RGB-IR) 2周内
🟡 中 验证遮挡场景下的疲劳检测降级策略 4-6周
🟢 低 探索更细粒度的遮挡分类(6类) 1-2月

局限性分析

论文局限

问题 分析
未报告完整系统FPS 各模块单独报告,整体管线速度未知
遮挡分类粗粒度 仅3级,实际场景遮挡类型多样
驾驶员识别数量少 未报告支持多少驾驶员
无开源代码 复现成本高
IR模型性能偏低 IR遮挡检测87.1% vs RGB 92.3%

改进方向

方向 当前 改进
遮挡分类 3级 6级(墨镜/口罩/手/帽子/大角度/正常)
视线区域 6区域 9区域(增加:方向盘、左下方、右下方)
驾驶员识别 固定ID 动态注册新驾驶员
模态切换 帧级切换 时序平滑切换+滞后
遮挡恢复 即时判断 恢复确认延迟(避免频繁切换)

与其他DMS论文对比

论文 年份 遮挡检测 RGB/IR 视线区域 驾驶员ID 真实驾驶
Fridman et al. 2016 ❌ ❌ ✅ ❌ ✅
Ghosh et al. 2021 ❌ ❌ ✅ ❌ ✅
Sharma et al. (TransGaze) 2026 ❌ ❌ 物体级 ❌ ✅
本文(Cañas et al.) 2025 ✅ ✅ ✅ ✅ ✅

本文独特价值: 首次将遮挡检测从”预处理步骤”提升为”独立功能模块”,并对齐EuroNCAP推荐。


总结

TransGaze-Object 和遮挡感知DMS代表了DMS领域的两个重要方向:

  1. TransGaze-Object: 从”注视点回归”到”注视物体预测”——语义层面的跃迁
  2. 遮挡感知DMS: 从”默默失败”到”主动声明降级”——安全可信的必由之路

两者结合可以构建更强大的DMS:遮挡检测确保系统可信,注视物体预测提供更丰富的语义理解。


遮挡感知DMS:基于DMD数据集的鲁棒驾驶员监控系统深度解析
https://dapalm.com/2026/10/01/2026-10-01-17-occlusion-aware-dms-dmd-dataset-rgb-ir-euro-ncap-ims/
作者
Mars
发布于
2026年10月1日
许可协议