高铁驾驶员走神检测:序变感知的Transformer-LSTM EEG框架(SSRN 2026 论文解读)

高铁驾驶员走神检测:序变感知的Transformer-LSTM EEG框架

论文信息

项目 内容
标题 Ordinal Transition-Aware EEG-Based Detection of Mind Wandering in High-Speed Rail Drivers
作者 Zhenqi Chen, Qingqing Hu, Qiaofeng Guo, Xiang Gao, Zizheng Guo
来源 SSRN Working Paper
年份 2026
链接 https://papers.ssrn.com/sol3/papers.cfm?abstract_id=7428503
框架 Region-aware Ordinal Transition-aware Transformer-LSTM

1. 核心创新

1.1 问题定义

走神 (Mind Wandering):驾驶员注意力从驾驶任务转移到内部思维,外部表现不明显但反应能力显著下降。

与疲劳/分心的区别:

状态 外部表现 内部状态 检测难度
疲劳 闭眼、打哈欠 困倦 ✅ 可检测
分心 视线偏离 注意力外移 ✅ 可检测
走神 外观正常 注意力内移 ❌ 难以检测

1.2 核心方法

“Region-aware ordinal transition-aware Transformer-LSTM framework based on EEG signals”

三个关键创新:

  1. Region-aware:脑区感知——不同脑区对不同认知状态有不同敏感性
  2. Ordinal transition-aware:序变感知——走神不是突变的,而是渐进的
  3. Transformer-LSTM融合:Transformer捕获空间模式 + LSTM捕获时序动态

1.3 应用场景

高铁驾驶员在长时间单调驾驶中极易走神——70%的驾驶时间被报告有走神(ScienceDaily 2017研究)。高铁驾驶环境比汽车更单调,走神风险更高。


2. 方法详解

2.1 脑区划分

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
"""
Region-aware EEG processing
不同脑区在认知控制中的作用
"""

import numpy as np
import torch
import torch.nn as nn

# 脑区定义
EEG_REGIONS = {
'frontal': {
'channels': ['Fp1', 'Fp2', 'F3', 'F4', 'Fz', 'F7', 'F8'],
'function': '认知控制、注意力分配、决策',
'mind_wandering_marker': 'frontal theta 降低'
},
'central': {
'channels': ['C3', 'C4', 'Cz'],
'function': '运动准备、感觉运动整合',
'mind_wandering_marker': 'central beta 变化'
},
'parietal': {
'channels': ['P3', 'P4', 'Pz'],
'function': '注意门控、空间定位',
'mind_wandering_marker': 'parietal alpha 变化'
},
'occipital': {
'channels': ['O1', 'O2', 'Oz'],
'function': '视觉处理',
'mind_wandering_marker': 'occipital alpha 增强'
},
'temporal': {
'channels': ['T3', 'T4', 'T5', 'T6'],
'function': '听觉处理、记忆',
'mind_wandering_marker': 'temporal theta 变化'
}
}

class RegionAwareProcessor(nn.Module):
"""
脑区感知EEG处理器

对每个脑区分别提取特征,保留空间信息
"""

def __init__(self, num_channels: int = 32,
region_map: dict = None,
embed_dim: int = 64):
super().__init__()

if region_map is None:
region_map = EEG_REGIONS

# 每个脑区的特征提取器
self.region_encoders = nn.ModuleDict()
for region, info in region_map.items():
n_ch = len(info['channels'])
self.region_encoders[region] = nn.Sequential(
nn.Conv1d(n_ch, embed_dim, 7, padding=3),
nn.BatchNorm1d(embed_dim),
nn.GELU(),
nn.Conv1d(embed_dim, embed_dim, 3, padding=1),
nn.GELU()
)

# 区域间注意力
self.region_attention = nn.MultiheadAttention(
embed_dim, num_heads=4, batch_first=True
)

def forward(self, eeg: dict) -> torch.Tensor:
"""
Args:
eeg: dict of region -> tensor (B, n_channels_region, T)

Returns:
region_features: (B, n_regions, embed_dim)
"""
region_features = []
for region in self.region_encoders:
x = eeg[region] # (B, n_ch, T)
feat = self.region_encoders[region](x) # (B, embed_dim, T)
feat = feat.mean(dim=2) # (B, embed_dim) 时序池化
region_features.append(feat)

# (B, n_regions, embed_dim)
region_feat = torch.stack(region_features, dim=1)

# 区域间注意力
attended, _ = self.region_attention(
region_feat, region_feat, region_feat
)

return attended + region_feat # 残差


# 测试
if __name__ == "__main__":
processor = RegionAwareProcessor(embed_dim=64)

# 模拟各脑区EEG信号
batch_size = 2
duration = 500 # 1秒@500Hz
eeg_input = {}
for region, info in EEG_REGIONS.items():
n_ch = len(info['channels'])
eeg_input[region] = torch.randn(batch_size, n_ch, duration)

output = processor(eeg_input)
print(f"输入: {len(eeg_input)} 个脑区")
print(f"输出: {output.shape}") # (2, 5, 64)
print(f"参数量: {sum(p.numel() for p in processor.parameters())/1e6:.2f}M")

2.2 序变感知 Transformer-LSTM

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
"""
Ordinal Transition-Aware Transformer-LSTM
序变感知的Transformer-LSTM框架

核心思想:
- 走神是渐进的:专注 → 轻微走神 → 深度走神
- 这种渐进过程可以用"序变"(ordinal transition)建模
- Transformer捕获空间模式,LSTM捕获时序动态
"""

import torch
import torch.nn as nn
import torch.nn.functional as F

class OrdinalTransitionModule(nn.Module):
"""
序变模块

建模走神的渐进过程:
专注(0) → 轻微走神(1) → 中度走神(2) → 深度走神(3)

使用有序回归而非多分类
"""

def __init__(self, feature_dim: int = 64,
num_levels: int = 4):
super().__init__()

self.num_levels = num_levels

# 序变特征提取
self.transition_encoder = nn.LSTM(
input_size=feature_dim,
hidden_size=feature_dim * 2,
num_layers=2,
batch_first=True,
bidirectional=False, # 因果:只看过去
dropout=0.1
)

# 序变分类头(有序回归)
self.ordinal_head = nn.Sequential(
nn.Linear(feature_dim * 2, feature_dim),
nn.GELU(),
nn.Linear(feature_dim, num_levels - 1) # num_levels-1个阈值
)

def forward(self, x: torch.Tensor) -> tuple:
"""
Args:
x: (B, T, feature_dim) 时序特征

Returns:
ordinal_logits: (B, T, num_levels-1)
transition_probs: (B, T, num_levels) 走神等级概率
"""
# LSTM 时序建模
lstm_out, _ = self.transition_encoder(x) # (B, T, 2*dim)

# 序变logits
ordinal_logits = self.ordinal_head(lstm_out) # (B, T, num_levels-1)

# 有序概率
# P(level >= k) = sigmoid(logit_k)
cum_probs = torch.sigmoid(ordinal_logits) # (B, T, num_levels-1)

# P(level = k) = P(>=k) - P(>=k+1)
probs = torch.zeros(
*ordinal_logits.shape[:2], self.num_levels,
device=ordinal_logits.device
)

# P(level=0) = 1 - P(>=1)
probs[:, :, 0] = 1 - cum_probs[:, :, 0]
# P(level=k) = P(>=k) - P(>=k+1)
for k in range(1, self.num_levels - 1):
probs[:, :, k] = cum_probs[:, :, k-1] - cum_probs[:, :, k]
# P(level=max) = P(>=max-1)
probs[:, :, -1] = cum_probs[:, :, -1]

return ordinal_logits, probs


class MindWanderingDetector(nn.Module):
"""
完整的走神检测模型

架构:
1. Region-aware encoder(脑区感知)
2. Transformer(空间-时序模式)
3. Ordinal Transition LSTM(序变动态)
4. 有序分类头
"""

def __init__(self, num_channels: int = 32,
embed_dim: int = 64,
num_heads: int = 4,
num_transformer_layers: int = 2,
num_levels: int = 4):
super().__init__()

# 1. 脑区感知编码
self.region_encoder = RegionAwareProcessor(
num_channels=num_channels,
embed_dim=embed_dim
)

# 2. Transformer(时序自注意力)
transformer_layer = nn.TransformerEncoderLayer(
d_model=embed_dim,
nhead=num_heads,
dim_feedforward=embed_dim * 4,
dropout=0.1,
batch_first=True
)
self.transformer = nn.TransformerEncoder(
transformer_layer, num_layers=num_transformer_layers
)

# 3. 序变感知 LSTM
self.ordinal_transition = OrdinalTransitionModule(
feature_dim=embed_dim,
num_levels=num_levels
)

def forward(self, eeg: dict) -> dict:
"""
Args:
eeg: dict of region -> tensor (B, n_ch, T)

Returns:
{
'ordinal_logits': (B, T, num_levels-1),
'level_probs': (B, T, num_levels),
'attention_weights': optional
}
"""
# 1. 脑区编码
region_features = self.region_encoder(eeg) # (B, n_regions, dim)

# 扩展到时序维度
# 假设region_features已在时序上提取
# 需要调整为 (B, T, dim) 格式
# 这里简化处理:将区域特征作为序列
x = region_features # (B, n_regions, dim)

# 2. Transformer
x = self.transformer(x) # (B, n_regions, dim)

# 3. 序变 LSTM
ordinal_logits, level_probs = self.ordinal_transition(x)

return {
'ordinal_logits': ordinal_logits,
'level_probs': level_probs,
'features': x
}


# 测试
if __name__ == "__main__":
model = MindWanderingDetector(
num_channels=32,
embed_dim=64,
num_heads=4,
num_transformer_layers=2,
num_levels=4 # 专注/轻微/中度/深度走神
)

# 模拟输入
batch_size = 4
duration = 500
eeg_input = {}
for region, info in EEG_REGIONS.items():
n_ch = len(info['channels'])
eeg_input[region] = torch.randn(batch_size, n_ch, duration)

output = model(eeg_input)

print("=== 走神检测模型输出 ===")
print(f"序变logits: {output['ordinal_logits'].shape}")
print(f"等级概率: {output['level_probs'].shape}")
print(f"特征: {output['features'].shape}")

# 打印每个样本的走神等级
level_names = ['专注', '轻微走神', '中度走神', '深度走神']
for i in range(batch_size):
# 取最后一个时间步
probs = output['level_probs'][i, -1, :]
level = torch.argmax(probs).item()
print(f"样本{i}: {level_names[level]} (置信度={probs[level]:.2%})")

print(f"\n总参数量: {sum(p.numel() for p in model.parameters())/1e6:.2f}M")

2.3 从EEG到摄像头的代理指标

虽然论文使用EEG信号,但核心洞察可以转化为基于摄像头的代理指标:

EEG发现 摄像头代理指标 可行性
Frontal theta 降低 → 认知控制下降 眼动规律性(熵)降低 ✅ 可从眼动数据提取
走神渐进过程 眨眼间隔方差逐渐增大 ✅ 可从摄像头提取
区域间注意力变化 视线固定点分散度增大 ✅ 可从视线估计提取
序变模式 微表情频率变化 ⚠️ 需要高帧率摄像头
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
"""
基于摄像头的走神代理指标
"""

import numpy as np
from scipy.stats import entropy

class CameraBasedMindWanderingProxy:
"""
基于摄像头信号的走神代理检测器

使用眼动指标作为EEG的代理
"""

def __init__(self, window_size: int = 300): # 10秒@30fps
self.window_size = window_size

def compute_gaze_entropy(self, gaze_points: np.ndarray) -> float:
"""
计算视线分散熵

走神时视线分散度增大
"""
# 将视线落点离散化到网格
grid_size = 10
x_bins = np.linspace(0, 1, grid_size)
y_bins = np.linspace(0, 1, grid_size)

x = (gaze_points[:, 0] * grid_size).astype(int)
y = (gaze_points[:, 1] * grid_size).astype(int)
x = np.clip(x, 0, grid_size - 1)
y = np.clip(y, 0, grid_size - 1)

# 2D直方图
hist, _, _ = np.histogram2d(x, y, bins=grid_size)
hist = hist.flatten() + 1e-10
hist = hist / hist.sum()

return entropy(hist)

def compute_blink_regularity(self, blink_times: np.ndarray) -> float:
"""
计算眨眼规律性

走神时眨眼间隔方差增大
"""
if len(blink_times) < 2:
return 1.0

intervals = np.diff(blink_times)
mean_interval = np.mean(intervals)
std_interval = np.std(intervals)

# 变异系数(CV):越小越规律
cv = std_interval / (mean_interval + 1e-10)

# 转换为规律性分数(0-1,1=最规律)
regularity = 1.0 / (1.0 + cv)

return regularity

def compute_saccade_frequency(self, saccade_times: np.ndarray) -> float:
"""
扫视频率

走神时扫视频率降低(注意力内移)
"""
if len(saccade_times) < 2:
return 0.0

duration = self.window_size / 30 # 秒
freq = len(saccade_times) / duration
return freq

def assess_mind_wandering(self,
gaze_points: np.ndarray,
blink_times: np.ndarray,
saccade_times: np.ndarray) -> dict:
"""
综合评估走神状态

Returns:
{
'wandering_score': 0-1,
'level': 'focused'|'light'|'moderate'|'deep',
'indicators': dict
}
"""
# 计算各指标
gaze_entropy = self.compute_gaze_entropy(gaze_points)
blink_regularity = self.compute_blink_regularity(blink_times)
saccade_freq = self.compute_saccade_frequency(saccade_times)

# 走神评分
# 视线熵高 + 眨眼不规律 + 扫视频率低 = 走神
score = 0.0

# 视线熵(越高越走神)
if gaze_entropy > 2.0:
score += 0.4 * min((gaze_entropy - 2.0) / 1.0, 1.0)

# 眨眼不规律(CV大=不规律=走神)
if blink_regularity < 0.7:
score += 0.3 * (0.7 - blink_regularity) / 0.3

# 扫视频率低(走神时减少)
if saccade_freq < 1.0: # 每秒<1次扫视
score += 0.3 * (1.0 - saccade_freq) / 1.0

score = min(score, 1.0)

if score < 0.2:
level = 'focused'
elif score < 0.4:
level = 'light'
elif score < 0.7:
level = 'moderate'
else:
level = 'deep'

return {
'wandering_score': round(score, 3),
'level': level,
'indicators': {
'gaze_entropy': round(gaze_entropy, 3),
'blink_regularity': round(blink_regularity, 3),
'saccade_freq_hz': round(saccade_freq, 3)
}
}


# 测试
if __name__ == "__main__":
detector = CameraBasedMindWanderingProxy(window_size=300)

np.random.seed(42)

# 模拟专注驾驶:视线集中在道路前方
focused_gaze = np.random.dirichlet([5, 1, 1, 1, 1, 1, 1, 1, 1, 1], size=300)
focused_gaze = focused_gaze[:, :2] # 取x,y
focused_blinks = np.sort(np.random.choice(300, size=8, replace=False))
focused_saccades = np.sort(np.random.choice(300, size=30, replace=False))

# 模拟走神:视线分散
wandering_gaze = np.random.rand(300, 2) # 均匀分布
wandering_blinks = np.sort(np.random.choice(300, size=5, replace=False))
wandering_saccades = np.sort(np.random.choice(300, size=10, replace=False))

focused_result = detector.assess_mind_wandering(
focused_gaze, focused_blinks, focused_saccades
)
wandering_result = detector.assess_mind_wandering(
wandering_gaze, wandering_blinks, wandering_saccades
)

print("=== 专注驾驶 ===")
print(f"走神评分: {focused_result['wandering_score']}")
print(f"等级: {focused_result['level']}")
print(f"指标: {focused_result['indicators']}")

print("\n=== 走神驾驶 ===")
print(f"走神评分: {wandering_result['wandering_score']}")
print(f"等级: {wandering_result['level']}")
print(f"指标: {wandering_result['indicators']}")

3. 对 IMS 开发的启示

3.1 可落地的开发建议

优先级 建议 理由
🔴 P0 走神检测使用眼动熵 视线落点分散度是最直接的走神代理指标
🔴 P0 眨眼间隔规律性分析 走神时眨眼不规律,可从摄像头提取
🟡 P1 序变建模 走神是渐进的,使用有序回归而非二分类
🟡 P1 多脑区思路迁移 不同面部区域提供不同信息(眼/嘴/眉毛)
🟢 P2 EEG集成(长期) 未来座舱可集成非接触EEG

3.2 走神检测 vs 疲劳检测的差异

维度 疲劳检测 走神检测
核心指标 PERCLOS(闭眼比例) 眼动熵(视线分散度)
时序模式 眨眼频率增加 眨眼间隔不规律
视线特征 视线偏离道路 视线分散但可能在道路方向
检测窗口 60秒 10-30秒
误报来源 眨眼/光线 思考/路况复杂
严重程度 渐进(PERCLOS上升) 渐进(序变模型)

3.3 跨领域价值

高铁驾驶员走神检测的方法论可以直接迁移到汽车座舱:

高铁场景 汽车场景 迁移可行性
长时间单调驾驶 高速公路巡航 ✅ 高
EEG frontal theta 眼动熵 ✅ 中(代理指标)
序变 Transformer-LSTM 时序眼动模型 ✅ 高
脑区感知 面部区域感知 ✅ 中

4. 技术路线判断

4.1 核心洞察

走神是认知分心中最难检测的子类——因为外部行为表现正常。这篇论文的贡献在于:

  1. 序变建模:将走神视为渐进过程而非突变事件
  2. 脑区感知:不同脑区对不同认知状态有不同敏感性
  3. EEG→摄像头迁移路径:frontal theta → gaze entropy

4.2 对 Euro NCAP 的间接影响

Euro NCAP 当前不直接评估走神检测,但认知分心是 DSM 评估的一部分。走神检测的突破将直接影响 Euro NCAP DSM 评分。

4.3 长期趋势

  • 短期(2026-2027):基于眼动的走神代理指标
  • 中期(2028-2029):多模态融合(眼动+行为+生理)
  • 长期(2030+):非接触EEG + 摄像头融合

5. 参考


高铁驾驶员走神检测:序变感知的Transformer-LSTM EEG框架(SSRN 2026 论文解读)
https://dapalm.com/2026/10/06/2026-10-06-020-eeg-mind-wandering-hsr-driver-ordinal-transition-ssrn2026/
作者
Mars
发布于
2026年10月6日
许可协议