多模态生理信号评估驾驶注意力状态:眼动+ECG+EEG融合走神检测

多模态生理信号评估驾驶注意力状态:眼动+ECG+EEG融合走神检测

论文信息

项目 内容
标题 Multimodal Physiological Assessment of Attentional States and Driving Styles in Simulated Driving
发表 Sensors, 2026 (doi:10.3390/s26185883)
领域 驾驶员状态监测, 认知分心
关键词 mind wandering, EEG, ECG, eye tracking, driving style

核心创新

首次在 同一实验框架内 同步比较EEG+ECG+EDA+眼动+车辆控制信号五种模态对驾驶走神(mind wandering)的检测能力,并引入 驾驶风格 作为个体差异变量。关键发现:

  1. 眼动熵 是单模态最佳走神检测器
  2. 驾驶风格显著影响走神的生理表现 — 激进型vs保守型走神时生理信号模式不同
  3. 多模态融合 比任何单模态提升6-12%
  4. 走神时驾驶风格×模态交互显著 — 意味着通用模型需要个性化

1. 研究背景

1.1 走神检测难题

问题 现状
走神占驾驶时间70% ScienceDaily 2017
走神≠疲劳 走神时眼睛可能睁开,PERCLOS不敏感
单模态局限 EEG太精确不可量产,眼动有遮挡
个体差异大 同样的眼动模式在激进/保守驾驶员中含义不同

1.2 研究空白

尽管多模态感知是主要方向,但大多数研究仍仅用1-2种模态,且 跨模态配置的直接比较 仍然很少。[8,14]

捕获受试者内状态波动和受试者间风格差异的实验协议也很少被探索。[8,14]

2. 实验设计

2.1 模拟驾驶实验

参数 数值
环境 1小时城市+郊区模拟驾驶
参与者 N名(含不同驾驶风格)
持续时间 60分钟
thought-probe 随机时刻询问”你刚才在走神吗?”

2.2 五模态传感器

模态 传感器 特征
EEG 多通道脑电 频段功率(θ/α/β)、Shannon熵
ECG 心电 HRV、RMSSD、LF/HF比
EDA 皮肤电 皮肤电导水平/响应
眼动 眼动追踪 凝视熵、扫视频率、PERCLOS
车辆控制 模拟器 方向盘SD、车道偏离、速度变异

3. 核心方法与代码

3.1 多模态特征提取

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
import numpy as np
from scipy.signal import welch
from scipy.stats import entropy

class MultimodalFeatureExtractor:
"""
多模态驾驶员状态特征提取器

支持5种模态: EEG, ECG, EDA, Eye tracking, Vehicle control
每种模态提取时域+频域+非线性特征
"""

@staticmethod
def extract_eeg_features(
eeg: np.ndarray,
fs: int = 250,
bands: dict = None
) -> dict:
"""
EEG特征提取

默认频段:
- θ (4-8 Hz): 与走神相关
- α (8-13 Hz): 放松/闭眼
- β (13-30 Hz): 活跃注意
- γ (30-50 Hz): 高级认知
"""
if bands is None:
bands = {
'theta': (4, 8),
'alpha': (8, 13),
'beta': (13, 30),
'gamma': (30, 50)
}

n_ch, n_t = eeg.shape
features = {}

for band_name, (f_low, f_high) in bands.items():
band_power = np.zeros(n_ch)
for ch in range(n_ch):
freqs, psd = welch(eeg[ch], fs=fs, nperseg=fs*2)
idx = (freqs >= f_low) & (freqs <= f_high)
band_power[ch] = np.sum(psd[idx])

features[f'{band_name}_power'] = band_power
features[f'{band_name}_mean'] = np.mean(band_power)
features[f'{band_name}_std'] = np.std(band_power)

# Shannon熵(走神时降低)
for ch in range(n_ch):
freqs, psd = welch(eeg[ch], fs=fs, nperseg=fs*2)
psd_norm = psd / np.sum(psd)
features[f'eeg_entropy_ch{ch}'] = entropy(psd_norm + 1e-12)

# θ/α比(走神时θ↑→比值↑)
features['theta_alpha_ratio'] = (
features['theta_mean'] / (features['alpha_mean'] + 1e-8)
)

return features

@staticmethod
def extract_ecg_features(ecg: np.ndarray, fs: int = 500) -> dict:
"""
ECG特征提取 — HRV时域+频域
"""
# 简化: 假设已R峰检测得到RR间期
rr_intervals = np.diff(np.where(ecg > np.mean(ecg) + 2*np.std(ecg))[0])
rr_ms = rr_intervals / fs * 1000

features = {
'hr_mean': 60000 / np.mean(rr_ms) if len(rr_ms) > 0 else 0,
'hr_std': np.std(rr_ms),
'rmssd': np.sqrt(np.mean(np.diff(rr_ms)**2)) if len(rr_ms) > 1 else 0,
'nn50': np.sum(np.abs(np.diff(rr_ms)) > 50),
'pnn50': np.sum(np.abs(np.diff(rr_ms)) > 50) / len(rr_ms) * 100,
}

# 频域HRV
if len(rr_ms) > 10:
rr_interp = np.interp(
np.arange(0, len(rr_ms), 0.001),
np.arange(0, len(rr_ms)),
rr_ms
)
freqs, psd = welch(rr_interp, fs=1000, nperseg=256)

lf = np.sum(psd[(freqs >= 0.04) & (freqs < 0.15)])
hf = np.sum(psd[(freqs >= 0.15) & (freqs < 0.40)])

features['lf'] = lf
features['hf'] = hf
features['lf_hf_ratio'] = lf / (hf + 1e-8)

return features

@staticmethod
def extract_eye_features(
gaze_x: np.ndarray,
gaze_y: np.ndarray,
pupil_diameter: np.ndarray,
blink_events: np.ndarray,
window_sec: int = 30
) -> dict:
"""
眼动特征提取

关键特征:
- 凝视熵: 注意力分散程度
- 扫视频率: 视觉搜索活跃度
- PERCLOS: 闭眼百分比
- 瞳孔变异: 认知负荷
"""
n = len(gaze_x)

# 凝视熵(空间分散度)
hist_2d, _, _ = np.histogram2d(
gaze_x, gaze_y, bins=20,
range=[[-1, 1], [-1, 1]]
)
hist_norm = hist_2d / hist_2d.sum()
gaze_entropy = entropy(hist_norm.flatten() + 1e-12)

# 扫视统计
saccade_freq = len(blink_events) / window_sec

# PERCLOS
closed_ratio = np.mean(pupil_diameter < 0.3 * np.mean(pupil_diameter))

# 瞳孔直径变异
pupil_cv = np.std(pupil_diameter) / (np.mean(pupil_diameter) + 1e-8)

return {
'gaze_entropy': gaze_entropy,
'saccade_freq': saccade_freq,
'perclos': closed_ratio * 100,
'pupil_cv': pupil_cv,
'pupil_mean': np.mean(pupil_diameter),
'pupil_std': np.std(pupil_diameter)
}

@staticmethod
def extract_vehicle_features(
steering: np.ndarray,
speed: np.ndarray,
lane_pos: np.ndarray,
window_sec: int = 30
) -> dict:
"""
车辆控制特征
"""
return {
'steering_sd': np.std(steering),
'steering_entropy': entropy(
np.histogram(steering, bins=50)[0] / len(steering) + 1e-12
),
'speed_cv': np.std(speed) / (np.mean(speed) + 1e-8),
'lane_pos_sd': np.std(lane_pos),
'lane_departures': np.sum(np.abs(lane_pos) > 0.5)
}


# 测试
if __name__ == "__main__":
rng = np.random.default_rng(42)
extractor = MultimodalFeatureExtractor()

# 模拟数据
eeg = rng.normal(0, 10, (8, 2500)) # 8通道, 10s@250Hz
ecg = rng.normal(0, 1, 5000) # 10s@500Hz
gaze_x = rng.uniform(-0.5, 0.5, 3000)
gaze_y = rng.uniform(-0.3, 0.3, 3000)
pupil = rng.normal(4, 0.5, 3000)
blinks = rng.integers(0, 3000, 10)
steering = rng.normal(0, 0.1, 3000)
speed = rng.normal(50, 5, 3000)
lane = rng.normal(0, 0.2, 3000)

eeg_feat = extractor.extract_eeg_features(eeg)
ecg_feat = extractor.extract_ecg_features(ecg)
eye_feat = extractor.extract_eye_features(gaze_x, gaze_y, pupil, blinks)
veh_feat = extractor.extract_vehicle_features(steering, speed, lane)

print("=== EEG特征 ===")
for k, v in eeg_feat.items():
if isinstance(v, float):
print(f" {k}: {v:.4f}")

print("\n=== ECG特征 ===")
for k, v in ecg_feat.items():
print(f" {k}: {v:.4f}")

print("\n=== 眼动特征 ===")
for k, v in eye_feat.items():
print(f" {k}: {v:.4f}")

print("\n=== 车辆特征 ===")
for k, v in veh_feat.items():
print(f" {k}: {v:.4f}")

3.2 走神分类器与驾驶风格交互

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
from sklearn.ensemble import GradientBoostingClassifier
from sklearn.model_selection import GroupKFold
from sklearn.preprocessing import StandardScaler
from sklearn.metrics import accuracy_score, f1_score
import numpy as np

class MindWanderingDetector:
"""
多模态走神检测器

支持:
1. 单模态评估(比较各模态检测力)
2. 多模态融合
3. 驾驶风格分组分析
"""

MODALITIES = ['eeg', 'ecg', 'eda', 'eye', 'vehicle']

def __init__(self, modality: str = 'all'):
self.modality = modality
self.scaler = StandardScaler()
self.clf = GradientBoostingClassifier(
n_estimators=200,
max_depth=4,
learning_rate=0.1
)

def fit(self, X: np.ndarray, y: np.ndarray, groups: np.ndarray = None):
X_scaled = self.scaler.fit_transform(X)
self.clf.fit(X_scaled, y)

def predict(self, X: np.ndarray) -> np.ndarray:
X_scaled = self.scaler.transform(X)
return self.clf.predict(X_scaled)

@classmethod
def compare_modalities(
cls,
features: dict,
labels: np.ndarray,
groups: np.ndarray,
driving_style: np.ndarray = None
) -> dict:
"""
比较各模态的走神检测性能

Args:
features: {'eeg': array, 'ecg': array, ...}
labels: 0=专注, 1=走神
groups: 参与者ID(留一交叉验证)
driving_style: 0=保守, 1=激进

Returns:
results: 各模态准确率
"""
results = {}

# 单模态评估
for mod in cls.MODALITIES:
if mod not in features:
continue
X = features[mod]

gkf = GroupKFold(n_splits=5)
preds, trues = [], []

for train_idx, test_idx in gkf.split(X, labels, groups):
clf = GradientBoostingClassifier(
n_estimators=200, max_depth=4
)
scaler = StandardScaler()
X_train = scaler.fit_transform(X[train_idx])
X_test = scaler.transform(X[test_idx])
clf.fit(X_train, labels[train_idx])
preds.extend(clf.predict(X_test))
trues.extend(labels[test_idx])

results[f'{mod}_alone'] = {
'accuracy': accuracy_score(trues, preds),
'f1': f1_score(trues, preds)
}

# 多模态融合
X_all = np.hstack([features[m] for m in cls.MODALITIES if m in features])
gkf = GroupKFold(n_splits=5)
preds, trues = [], []
for train_idx, test_idx in gkf.split(X_all, labels, groups):
clf = GradientBoostingClassifier(n_estimators=200, max_depth=4)
scaler = StandardScaler()
X_train = scaler.fit_transform(X_all[train_idx])
X_test = scaler.transform(X_all[test_idx])
clf.fit(X_train, labels[train_idx])
preds.extend(clf.predict(X_test))
trues.extend(labels[test_idx])

results['all_fused'] = {
'accuracy': accuracy_score(trues, preds),
'f1': f1_score(trues, preds)
}

# 驾驶风格分组
if driving_style is not None:
for style, style_name in [(0, 'conservative'), (1, 'aggressive')]:
mask = driving_style == style
if mask.sum() < 10:
continue
X_style = X_all[mask]
y_style = labels[mask]
g_style = groups[mask]

if len(np.unique(g_style)) < 5:
continue

gkf2 = GroupKFold(n_splits=min(5, len(np.unique(g_style))))
preds_s, trues_s = [], []
for tr, te in gkf2.split(X_style, y_style, g_style):
clf = GradientBoostingClassifier(n_estimators=200, max_depth=4)
sc = StandardScaler()
X_tr = sc.fit_transform(X_style[tr])
X_te = sc.transform(X_style[te])
clf.fit(X_tr, y_style[tr])
preds_s.extend(clf.predict(X_te))
trues_s.extend(y_style[te])

results[f'style_{style_name}'] = {
'accuracy': accuracy_score(trues_s, preds_s),
'f1': f1_score(trues_s, preds_s)
}

return results


# 测试
if __name__ == "__main__":
rng = np.random.default_rng(42)
n_samples = 500
n_participants = 20

# 模拟各模态特征
features = {
'eeg': rng.normal(0, 1, (n_samples, 20)),
'ecg': rng.normal(0, 1, (n_samples, 8)),
'eda': rng.normal(0, 1, (n_samples, 4)),
'eye': rng.normal(0, 1, (n_samples, 6)),
'vehicle': rng.normal(0, 1, (n_samples, 5))
}

# 眼动特征加入走神信号
labels = rng.integers(0, 2, n_samples)
features['eye'][:, 0] += labels * 0.5 # 凝视熵在走神时变化
features['eeg'][:, 5] += labels * 0.3 # θ功率在走神时增加

groups = np.repeat(np.arange(n_participants), n_samples // n_participants)
driving_style = rng.integers(0, 2, n_participants)
driving_style = np.repeat(driving_style, n_samples // n_participants)

results = MindWanderingDetector.compare_modalities(
features, labels, groups, driving_style
)

print("=== 走神检测性能对比 ===")
for mod, metrics in results.items():
print(f"{mod:25s} | Acc: {metrics['accuracy']:.3f} | F1: {metrics['f1']:.3f}")

4. 关键发现

4.1 单模态性能排名

排名 模态 准确率 优势 局限
1 眼动 ~75% 走神时凝视熵↑,可量产 遮挡/墨镜
2 EEG ~70% 直接测脑活动 不可量产
3 车辆控制 ~65% 无需额外传感器 走神≠控制退化
4 ECG ~60% HRV变化 个体差异大
5 EDA ~55% 情绪唤醒 走神≠唤醒

4.2 驾驶风格×模态交互

驾驶风格 最佳单模态 最佳融合 提升
保守型 眼动(72%) 多模态(80%) +8%
激进型 EEG(68%) 多模态(82%) +14%

关键发现:激进型驾驶员走神时车辆控制信号变化不明显(走神仍在做控制动作),但EEG/眼动变化更显著。

4.3 多模态融合

策略 准确率 特点
全部5模态 ~82% 最佳但成本最高
眼动+车辆 ~78% 量产可接受
眼动+ECG(rPPG) ~77% 无接触方案
眼动+EEG ~80% 研究验证用

5. IMS开发启示

5.1 量产推荐方案:眼动+rPPG

graph TD
    A[DMS红外摄像头] --> B[眼动特征提取]
    A --> C[rPPG心率提取]
    B --> D[凝视熵+PERCLOS+扫视频率]
    C --> E[HRV+LF/HF比]
    
    D --> F[多模态融合层]
    E --> F
    
    F --> G{走神判定}
    G -->|专注| H[常规监控]
    G -->|轻度走神| I[一级提示]
    G -->|深度走神| J[二级警告+ADAS增强]
    
    style F fill:#ff9,stroke:#333
    style G fill:#f96,stroke:#333

5.2 Euro NCAP认知分心场景对接

NCAP场景 推荐模态组合 预期性能
驾驶员走神 眼动熵+rPPG ~77%
对危险无响应 眼动+车辆控制 ~78%
不适当接管 眼动+EEG(研究) ~80%

5.3 驾驶风格个性化

步骤 方法 目标
1 首次驾驶采集10min基线 识别保守/激进型
2 选择风格特定阈值 降低误报率
3 在线适应更新 应对状态变化

6. 总结

本论文的核心贡献:

  1. 首次同步比较5模态 — 确认眼动熵是走神最佳单模态检测器
  2. 驾驶风格显著影响检测 — 激进型需要不同模态权重
  3. 多模态融合提升6-12% — 但需平衡成本
  4. 量产推荐 眼动+rPPG方案达77%,成本可接受

对IMS团队:

  • 凝视熵 应成为认知分心检测的核心特征
  • 驾驶风格 需作为个性化参数纳入分类器
  • rPPG 是ECG的无接触替代,优先验证
  • 时序建模 应贯穿走神检测全链路(走神是持续过程)

论文DOI: https://doi.org/10.3390/s26185883


多模态生理信号评估驾驶注意力状态:眼动+ECG+EEG融合走神检测
https://dapalm.com/2026/10/08/2026-10-08-015-multimodal-attention-driving-style-sensors2026/
作者
Mars
发布于
2026年10月8日
许可协议