边缘端实时DMS:低成本硬件上的驾驶员行为识别系统深度解析

论文信息

项目 内容
标题 Real-Time In-Cabin Driver Behavior Recognition on Low-Cost Edge Hardware
作者 Vesal Ahsani, Babak Hossein Khalaj, Hamed Shah-Mansouri
机构 Sharif University of Technology, Iran
时间 2025年12月25日(2026年1月6日更新)
arXiv 搜索标题可获取
平台 Raspberry Pi 5 (CPU) + Google Coral (Edge TPU)

核心创新

本文解决了DMS从实验室到量产的关键鸿沟——在低成本边缘硬件上实现实时驾驶员行为识别。不同于大多数论文只报告GPU性能,本文在两个真实低成本平台上部署和评估。

核心问题

graph TB
    subgraph "实验室环境"
        A[GPU服务器<br/>RTX 4090] --> B[模型训练]
        B --> C[准确率95%+]
        B --> D[帧率60fps+]
    end
    
    subgraph "量产现实"
        E[低成本边缘硬件<br/><$100] --> F[模型推理]
        F --> G[??? 准确率]
        F --> H[??? 帧率]
        F --> I[??? 功耗]
    end
    
    C -.->|"巨大鸿沟"| G
    D -.->|"巨大鸿沟"| H

本文回答的核心问题: DMS模型在<$100的硬件上能跑到多快、多准?


系统架构

整体管线

graph LR
    A[单目摄像头<br/>30fps] --> B[人脸检测<br/>YOLOv8n]
    B --> C[面部ROI裁剪]
    C --> D[行为分类模型<br/>MobileNetV3]
    D --> E[输出行为类别]
    
    subgraph "行为类别"
        E --> F1[正常驾驶]
        E --> F2[手机使用-通话]
        E --> F3[手机使用-打字]
        E --> F4[调整中控]
        E --> F5[调整后视镜]
        E --> F6[疲劳-闭眼]
        E --> F7[疲劳-打哈欠]
        E --> F8[分心-视线偏离]
    end

1. 人脸检测模块

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
import cv2
import numpy as np
import time
from typing import Tuple, List, Optional

class FaceDetector:
"""
人脸检测模块 - 针对边缘硬件优化

方案1: YOLOv8n (精度优先, 需Edge TPU)
方案2: Haar Cascade (速度优先, CPU-only)
方案3: MediaPipe (平衡方案)

论文策略: 根据硬件平台选择不同检测器
"""

def __init__(self, platform: str = 'rpi5'):
self.platform = platform

if platform == 'coral':
# Google Coral: 使用Edge TPU加速的YOLOv8n
from pycoral.adapters import detect
from pycoral.utils.edgetpu import make_interpreter
self.interpreter = make_interpreter('yolov8n_face_edgetpu.tflite')
self.input_size = (320, 320)

elif platform == 'rpi5':
# 树莓派5: 使用MobileNet SSD (CPU友好)
self.detector = cv2.dnn.readNetFromCaffe(
'deploy.prototxt',
'mobilenet_ssd.caffemodel'
)
self.input_size = (300, 300)

else:
# 默认: MediaPipe (跨平台)
import mediapipe as mp
self.mp_detection = mp.solutions.face_detection.FaceDetection(
model_selection=0,
min_detection_confidence=0.5
)

def detect(self, frame: np.ndarray) -> Optional[Tuple[int, int, int, int]]:
"""
检测人脸并返回最大人脸的边界框

Args:
frame: BGR图像 (H, W, 3)

Returns:
bbox: (x1, y1, x2, y2) 或 None
"""
h, w = frame.shape[:2]

if self.platform == 'coral':
return self._detect_coral(frame, w, h)
elif self.platform == 'rpi5':
return self._detect_rpi5(frame, w, h)
else:
return self._detect_mediapipe(frame, w, h)

def _detect_rpi5(self, frame: np.ndarray,
w: int, h: int) -> Optional[Tuple[int, int, int, int]]:
"""树莓派5 CPU检测 - 使用OpenCV DNN"""
blob = cv2.dnn.blobFromImage(
frame, 0.007843, self.input_size,
(127.5, 127.5, 127.5), swapRB=False, crop=False
)
self.detector.setInput(blob)
detections = self.detector.forward()

best_face = None
best_area = 0

for i in range(detections.shape[2]):
confidence = detections[0, 0, i, 2]
if confidence > 0.5:
class_id = int(detections[0, 0, i, 1])
if class_id == 15: # 人脸类别
x1 = int(detections[0, 0, i, 3] * w)
y1 = int(detections[0, 0, i, 4] * h)
x2 = int(detections[0, 0, i, 5] * w)
y2 = int(detections[0, 0, i, 6] * h)
area = (x2 - x1) * (y2 - y1)
if area > best_area:
best_area = area
best_face = (x1, y1, x2, y2)

return best_face

def _detect_coral(self, frame: np.ndarray,
w: int, h: int) -> Optional[Tuple[int, int, int, int]]:
"""Google Coral Edge TPU检测"""
# 预处理
input_data = cv2.resize(frame, self.input_size)
input_data = cv2.cvtColor(input_data, cv2.COLOR_BGR2RGB)
input_data = input_data.reshape(1, *self.input_size, 3)
input_data = input_data.astype(np.uint8)

# Edge TPU推理
from pycoral.adapters import common
common.set_input(self.interpreter, input_data)
self.interpreter.invoke()

from pycoral.adapters import detect
results = detect.get_objects(
self.interpreter, score_threshold=0.5
)

best_face = None
best_area = 0

for obj in results:
bbox = obj.bbox
x1 = int(bbox.xmin * w / self.input_size[0])
y1 = int(bbox.ymin * h / self.input_size[1])
x2 = int(bbox.xmax * w / self.input_size[0])
y2 = int(bbox.ymax * h / self.input_size[1])
area = (x2 - x1) * (y2 - y1)
if area > best_area:
best_area = area
best_face = (x1, y1, x2, y2)

return best_face

def _detect_mediapipe(self, frame: np.ndarray,
w: int, h: int) -> Optional[Tuple[int, int, int, int]]:
"""MediaPipe检测(跨平台后备)"""
rgb = cv2.cvtColor(frame, cv2.COLOR_BGR2RGB)
results = self.mp_detection.process(rgb)

if not results.detections:
return None

best_face = None
best_area = 0

for detection in results.detections:
bbox = detection.location_data.relative_bounding_box
x1 = int(bbox.xmin * w)
y1 = int(bbox.ymin * h)
x2 = int((bbox.xmin + bbox.width) * w)
y2 = int((bbox.ymin + bbox.height) * h)
area = (x2 - x1) * (y2 - y1)
if area > best_area:
best_area = area
best_face = (x1, y1, x2, y2)

return best_face

2. 行为分类模型

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
import torch
import torch.nn as nn
from torchvision import models

class DriverBehaviorClassifier(nn.Module):
"""
驾驶员行为分类器

论文: 8类行为分类
0: 正常驾驶 (normal_driving)
1: 手机通话 (phone_call)
2: 手机打字 (phone_text)
3: 调整中控 (adjust_center)
4: 调整后视镜 (adjust_mirror)
5: 闭眼疲劳 (eye_closed)
6: 打哈欠 (yawning)
7: 视线偏离 (looking_away)

架构: MobileNetV3-Large + 分类头
选择原因: MobileNetV3在移动端最优Pareto前沿

模型大小: ~5.4M参数
量化后: ~1.4MB (INT8)
"""

BEHAVIOR_NAMES = [
'normal_driving', 'phone_call', 'phone_text',
'adjust_center', 'adjust_mirror', 'eye_closed',
'yawning', 'looking_away'
]

def __init__(self, num_classes: int = 8):
super().__init__()
# MobileNetV3-Large作为backbone
self.backbone = models.mobilenet_v3_large(weights=models.MobileNet_V3_Large_Weights.DEFAULT)

# 修改分类头
in_features = self.backbone.classifier[3].in_features
self.backbone.classifier[3] = nn.Linear(in_features, num_classes)

# Dropout防止过拟合
self.dropout = nn.Dropout(0.2)

def forward(self, x: torch.Tensor) -> torch.Tensor:
"""
Args:
x: (B, 3, 224, 224) 归一化面部图像

Returns:
logits: (B, 8) 行为分类logits
"""
features = self.backbone.features(x) # (B, 960, 7, 7)
pooled = self.backbone.avgpool(features) # (B, 960, 1, 1)
pooled = torch.flatten(pooled, 1) # (B, 960)
pooled = self.backbone.classifier[0](pooled) # FC 1280
pooled = self.backbone.classifier[1](pooled) # Hardswish
pooled = self.dropout(pooled)
logits = self.backbone.classifier[3](pooled) # (B, 8)
return logits


class EdgeTPUBehaviorClassifier:
"""
Edge TPU推理封装

将PyTorch模型转换为TFLite格式
在Google Coral Edge TPU上加速推理
"""

def __init__(self, model_path: str):
from pycoral.utils.edgetpu import make_interpreter
from pycoral.adapters import common

self.interpreter = make_interpreter(model_path)
self.interpreter.allocate_tensors()

self.input_details = self.interpreter.get_input_details()
self.output_details = self.interpreter.get_output_details()

self.input_shape = self.input_details[0]['shape'][1:3] # (H, W)

def predict(self, face_image: np.ndarray) -> dict:
"""
Edge TPU推理

Args:
face_image: (H, W, 3) BGR面部图像

Returns:
{
'behavior': str,
'confidence': float,
'inference_time_ms': float
}
"""
# 预处理
img = cv2.resize(face_image, tuple(self.input_shape))
img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)
input_data = img.reshape(1, *self.input_shape, 3)
input_data = input_data.astype(np.uint8)

# 推理
start = time.perf_counter()
self.interpreter.set_tensor(self.input_details[0]['index'], input_data)
self.interpreter.invoke()
output = self.interpreter.get_tensor(self.output_details[0]['index'])
elapsed = (time.perf_counter() - start) * 1000

# 后处理
scores = torch.softmax(torch.from_numpy(output), dim=-1)[0]
pred = scores.argmax().item()
confidence = scores.max().item()

BEHAVIOR_NAMES = [
'normal_driving', 'phone_call', 'phone_text',
'adjust_center', 'adjust_mirror', 'eye_closed',
'yawning', 'looking_away'
]

return {
'behavior': BEHAVIOR_NAMES[pred],
'confidence': confidence,
'inference_time_ms': elapsed
}


class RPICPUClassifier:
"""
树莓派5 CPU推理封装

使用TFLite CPU推理或ONNX Runtime
优化: 多线程 + NEON指令集
"""

def __init__(self, model_path: str, num_threads: int = 4):
import tflite_runtime.interpreter as tflite
self.interpreter = tflite.Interpreter(
model_path=model_path,
num_threads=num_threads
)
self.interpreter.allocate_tensors()

self.input_details = self.interpreter.get_input_details()
self.output_details = self.interpreter.get_output_details()
self.input_shape = self.input_details[0]['shape'][1:3]

def predict(self, face_image: np.ndarray) -> dict:
"""树莓派5 CPU推理"""
img = cv2.resize(face_image, tuple(self.input_shape))
img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)
input_data = img.reshape(1, *self.input_shape, 3)
input_data = (input_data.astype(np.float32) / 127.5) - 1.0

start = time.perf_counter()
self.interpreter.set_tensor(self.input_details[0]['index'], input_data)
self.interpreter.invoke()
output = self.interpreter.get_tensor(self.output_details[0]['index'])
elapsed = (time.perf_counter() - start) * 1000

scores = np.exp(output) / np.sum(np.exp(output), axis=-1, keepdims=True)
pred = np.argmax(scores[0])
confidence = scores[0][pred]

BEHAVIOR_NAMES = [
'normal_driving', 'phone_call', 'phone_text',
'adjust_center', 'adjust_mirror', 'eye_closed',
'yawning', 'looking_away'
]

return {
'behavior': BEHAVIOR_NAMES[pred],
'confidence': float(confidence),
'inference_time_ms': elapsed
}

3. 完整系统集成

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
class EdgeDMS:
"""
完整边缘DMS系统

针对不同硬件平台的优化策略:
- Coral: 全流程Edge TPU加速
- RPi5: CPU + 多线程优化
- 自适应帧率: 根据推理速度调整
"""

def __init__(self, platform: str = 'rpi5'):
self.platform = platform

# 初始化模块
self.face_detector = FaceDetector(platform=platform)

if platform == 'coral':
self.classifier = EdgeTPUBehaviorClassifier(
'mobilenet_v3_behavior_edgetpu.tflite'
)
else:
self.classifier = RPICPUClassifier(
'mobilenet_v3_behavior_cpu.tflite',
num_threads=4
)

# 性能监控
self.frame_count = 0
self.total_time = 0

def process_frame(self, frame: np.ndarray) -> dict:
"""
处理单帧

Returns:
{
'behavior': str,
'confidence': float,
'fps': float,
'face_detected': bool
}
"""
frame_start = time.perf_counter()

# Step 1: 人脸检测
face_bbox = self.face_detector.detect(frame)

if face_bbox is None:
return {
'behavior': 'no_face',
'confidence': 0.0,
'fps': 0.0,
'face_detected': False
}

x1, y1, x2, y2 = face_bbox
face_img = frame[y1:y2, x1:x2]

if face_img.size == 0:
return {
'behavior': 'error',
'confidence': 0.0,
'fps': 0.0,
'face_detected': False
}

# Step 2: 行为分类
result = self.classifier.predict(face_img)

# 计算FPS
frame_time = time.perf_counter() - frame_start
self.frame_count += 1
self.total_time += frame_time
avg_fps = self.frame_count / self.total_time

result['fps'] = avg_fps
result['face_detected'] = True

return result


# 测试
if __name__ == "__main__":
import numpy as np

# 初始化系统
dms = EdgeDMS(platform='rpi5')

# 模拟测试
frame = np.random.randint(0, 255, (480, 640, 3), dtype=np.uint8)

for i in range(10):
result = dms.process_frame(frame)
print(f"Frame {i}: behavior={result['behavior']}, "
f"conf={result['confidence']:.3f}, "
f"fps={result['fps']:.1f}")

性能对比

各平台推理速度

模块 Raspberry Pi 5 (CPU) Google Coral (Edge TPU) GPU (RTX 4090)
人脸检测 25ms 8ms 3ms
行为分类 45ms 12ms 5ms
总计 70ms (~14fps) 20ms (~50fps) 8ms (~125fps)
功耗 5W 2W+1W 350W
成本 ~$80 ~$75 ~$1500

准确率对比

平台 模型 准确率 模型大小
GPU (FP32) MobileNetV3-Large 94.2% 21MB
Coral (INT8) MobileNetV3-Large (量化) 92.8% 5.4MB
RPi5 (INT8) MobileNetV3-Large (量化) 92.5% 5.4MB
RPi5 (FP16) MobileNetV3-Small 88.1% 2.9MB

关键发现

发现 数据 启示
量化损失<2% 94.2% → 92.8% INT8量化对DMS足够
Coral比RPi5快3.5x 50fps vs 14fps Edge TPU显著提升
Coral成本仅$75 含开发板 低成本量产可行
MobileNetV3最优 Pareto前沿 架构选择正确
单摄像头足够 30fps输入 降低系统成本

EuroNCAP 对齐分析

EuroNCAP 要求 本文支持度 实现
手机使用检测 ✅ 通话+打字两类
疲劳检测 ✅ 闭眼+打哈欠
视线偏离 ✅ looking_away类
实时性 ✅ Coral 50fps
低成本部署 ✅ <$100硬件

IMS 开发启示

1. 边缘部署硬件选型

方案 芯片 成本 帧率 推荐场景
方案A Qualcomm QCS8255 ~$50 30-50fps 量产DMS
方案B Google Coral Edge TPU ~$75 50fps 开发验证
方案C 树莓派5 ~$80 14fps 原型开发
方案D TI TDA4VM ~$60 40fps 量产DMS

2. 量化策略

1
2
3
4
5
6
7
8
9
10
11
12
13
14
# 量化流程示例
# Step 1: 训练FP32模型
model = DriverBehaviorClassifier(num_classes=8)
# ... 训练 ...

# Step 2: 静态量化
import torch.quantization as quant
model.eval()
model_fused = quant.fuse_modules(model, [['backbone.features.0.0', 'backbone.features.0.1']])
model_quant = quant.convert(quant.prepare(model_fused), inplace=False)

# Step 3: 转换为TFLite
# PyTorch -> ONNX -> TFLite -> Edge TPU编译器
# 最终: mobilenet_v3_behavior_edgetpu.tflite (5.4MB, INT8)

3. 技术路线建议

优先级 行动项 时间
🔴 高 购买Google Coral开发板验证 1周内
🔴 高 将现有IMS模型量化为INT8 1-2周
🟡 中 在Coral上跑通完整DMS管线 2-4周
🟡 中 评估MobileNetV3在QCS8255上的性能 2周内
🟢 低 探索更轻量的MobileNetV3-Small 1-2月

局限性分析

问题 分析
仅8类行为 实际场景行为更复杂
单摄像头 无IR,夜间不可用
无时序建模 单帧分类,无时序信息
数据集偏差 可能在特定人群上过拟合
无遮挡处理 遮挡场景未评估

总结

本文为DMS从实验室到量产提供了关键工程数据:

  1. Coral Edge TPU 是低成本DMS的最优硬件选择
  2. INT8量化 损失<2%,对DMS应用足够
  3. MobileNetV3 在边缘Pareto前沿最优
  4. 单摄像头方案 足以满足EuroNCAP基本要求

边缘端实时DMS:低成本硬件上的驾驶员行为识别系统深度解析
https://dapalm.com/2026/10/02/2026-10-02-01-edge-dms-raspberry-pi-coral-edge-tpu-ims/
作者
Mars
发布于
2026年10月2日
许可协议