
一、档案馆环境监测网络的运维痛点
档案馆(与档案库房不同,通常包含阅览区、数字化加工区、库房区、办公区等多功能区域)的环境监测网络规模往往比单一库房更大、更复杂。一个省级档案馆的监测网络可能覆盖上万平米、接入上百个温湿度传感器节点。这种规模的网络在运维中面临的核心问题不是"采集不到数据",而是"节点悄悄离线了,没人知道"。
场景 | 发现时间 | 后果 |
|---|---|---|
夜间空调故障导致温湿度超限 | 次日上班发现 | 库房环境偏离标准8小时,纸质档案可能受损 |
传感器节点电源适配器松动 | 3天后巡检发现 | 3天数据缺失,无法证明期间环境合规 |
交换机端口故障导致整区节点离线 | 1周后系统巡检发现 | 1周数据空白,审计不通过 |
网络改造导致IP段变更 | 2天后用户投诉 | 大屏显示大面积"无数据",管理员紧急排查 |
共性特征:问题不是突然发生的,而是有一个"静默期"——节点离线后系统没有及时告警,导致问题被延迟发现。
┌─────────────────────────────────────────────────────────────────────┐
│ 为什么节点离线后没有及时告警? │
├─────────────────────────────────────────────────────────────────────┤
│ │
│ 原因1:系统只检测"有没有数据",不检测"节点是否在线" │
│ ├── 节点每5分钟上报一次数据 │
│ ├── 如果节点离线,系统要等5分钟才能发现"这次没收到数据" │
│ └── 但系统没有区分"数据延迟"和"节点离线" │
│ │
│ 原因2:告警阈值设置不合理 │
│ ├── 连续3次没收到数据才告警 → 15分钟延迟 │
│ ├── 但网络抖动也会导致偶尔丢包 → 频繁误报 → 管理员关闭告警 │
│ └── 结果:真出问题时不告警 │
│ │
│ 原因3:告警渠道单一 │
│ ├── 只在监控大屏上显示红色标记 │
│ ├── 夜间无人值守时无人看大屏 │
│ └── 没有短信/电话等主动通知 │
│ │
│ 原因4:缺乏分级告警 │
│ ├── 单个节点离线 vs 整区节点离线,严重程度不同 │
│ ├── 但系统用同一种方式处理 │
│ └── 管理员无法判断是否需要立即处理 │
│ │
│ 原因5:没有告警确认和升级机制 │
│ ├── 告警发出后,如果管理员没处理,不会升级 │
│ └── 告警被淹没在大量信息中 │
│ │
└─────────────────────────────────────────────────────────────────────┘维度 | 数据上报 | 心跳检测 |
|---|---|---|
触发方 | 节点主动 | 服务器主动查询 |
频率 | 低(1~5分钟) | 高(30~60秒) |
目的 | 获取环境数据 | 确认节点存活 |
网络开销 | 大(含温湿度值) | 小(仅状态字) |
故障发现速度 | 慢(等待超时) | 快(主动探测) |
对节点功耗影响 | 高(采样+发送) | 低(仅响应) |
核心思路:数据上报负责"采集数据",心跳检测负责"确认节点在线"。两者解耦,互不影响。
在档案馆场景中,传感器节点通常通过RS485总线或无线(LoRa/Wi-Fi)连接到区域采集器,采集器通过TCP/IP与服务器通信。心跳协议需要在采集器-节点和服务器-采集器两个层面分别设计。
┌──────────┬──────────┬────────────┬──────────────────────────────────┐
│ 偏移 │ 长度 │ 字段 │ 说明 │
├──────────┼──────────┼────────────┼──────────────────────────────────┤
│ 0 │ 2 │ 帧头 │ 0xAA55(帧同步) │
│ 2 │ 1 │ 命令 │ 0x10=心跳请求, 0x11=心跳响应 │
│ 3 │ 1 │ 节点ID │ 从站地址(1~247) │
│ 4 │ 1 │ 节点状态 │ Bit0=在线, Bit1=低电量, │
│ │ │ │ Bit2=传感器故障, Bit3=存储满 │
│ 5 │ 2 │ 运行时间 │ 节点上电后秒数(高位) │
│ 7 │ 2 │ 软件版本 │ 主版本.次版本 │
│ 9 │ 1 │ RSSI │ 无线信号强度(仅无线节点) │
│ 10 │ 1 │ 保留 │ 预留 │
│ 11 │ 2 │ CRC16 │ 校验 │
│ 总计 │ 13字节 │ │ │
└──────────┴──────────┴────────────┴──────────────────────────────────┘{
"type": "heartbeat",
"collector_id": "C-001",
"timestamp": 1705728000,
"nodes": [
{
"node_id": 1,
"status": "online",
"last_seen": 1705727995,
"rssi": -45,
"uptime": 86400,
"fw_version": "2.1.3"
},
{
"node_id": 2,
"status": "offline",
"last_seen": 1705727400,
"rssi": null,
"uptime": null,
"fw_version": null
}
],
"collector_status": {
"cpu_usage": 12,
"mem_usage": 34,
"disk_free": 2048,
"network": "ok"
}
}┌──────────────────────────────────────┐
│ │
▼ │
┌───────────┐ ┌───────────┐ ┌───────────┐ ┌───────────┐
│ UNKNOWN │───▶│ ONLINE │───▶│ DEGRADED │───▶│ OFFLINE │
│ (未知) │ │ (在线) │ │ (降级) │ │ (离线) │
└───────────┘ └───────────┘ └───────────┘ └───────────┘
▲ │ │ │
│ │ │ │
│ ▼ ▼ ▼
│ 连续3次心跳 连续1~3次心跳 连续>3次心跳
│ 正常响应 丢失/超时 丢失/超时
│ │
│ 自动恢复 ───────────────────────────────────┘
│ (节点重新上线)
│
└── 系统启动初始化状态定义:
状态 | 含义 | 运维动作 |
|---|---|---|
UNKNOWN | 系统刚启动,尚未收到任何心跳 | 等待首次心跳 |
ONLINE | 节点正常,心跳响应及时 | 无需动作 |
DEGRADED | 节点响应变慢或偶尔丢心跳 | 记录日志,持续观察 |
OFFLINE | 节点连续多次无响应 | 触发告警,派工单 |
┌─────────────────────────────────────────────────────────────────────┐
│ 告警引擎架构 │
│ │
│ ┌─────────────┐ ┌─────────────┐ ┌─────────────┐ │
│ │ 心跳接收器 │───▶│ 状态评估器 │───▶│ 告警决策器 │ │
│ │ (Collector) │ │ (Evaluator) │ │ (Decider) │ │
│ └─────────────┘ └─────────────┘ └─────────────┘ │
│ │ │ │ │
│ ▼ ▼ ▼ │
│ ┌─────────────┐ ┌─────────────┐ ┌─────────────┐ │
│ │ 节点状态表 │ │ 告警规则库 │ │ 通知分发器 │ │
│ │ (NodeTable) │ │ (RuleBase) │ │ (Dispatcher)│ │
│ └─────────────┘ └─────────────┘ └─────────────┘ │
│ │ │
│ ▼ │
│ ┌─────────────────┐ │
│ │ 通知渠道 │ │
│ │ - 短信 │ │
│ │ - 邮件 │ │
│ │ - 微信/钉钉 │ │
│ │ - 语音电话 │ │
│ │ - 监控大屏 │ │
│ └─────────────────┘ │
└─────────────────────────────────────────────────────────────────────┘告警规则采用条件+抑制+升级三层结构:
# alert_rules.yaml
rules:
- name: "单节点离线"
condition:
type: "node_offline"
duration: "5m" # 连续离线5分钟
count: 1 # 1个节点
severity: "warning"
suppress:
- "同一节点5分钟内不重复告警"
- "网络维护窗口期内抑制"
notify:
channels: ["email", "dashboard"]
escalate:
after: "30m"
if: "still_offline"
to: "supervisor"
- name: "区域节点批量离线"
condition:
type: "area_offline"
area: "库房区A"
count: ">=3" # 同一区域≥3个节点
duration: "3m"
severity: "critical"
suppress:
- "同一区域10分钟内不重复"
notify:
channels: ["sms", "email", "voice_call", "dashboard"]
escalate:
after: "15m"
if: "still_offline"
to: "manager"
- name: "采集器离线"
condition:
type: "collector_offline"
duration: "2m"
severity: "critical"
notify:
channels: ["sms", "email", "voice_call"]
escalate:
after: "10m"
to: "manager"
- name: "心跳延迟"
condition:
type: "heartbeat_delay"
latency: ">3s" # 心跳响应超过3秒
count: ">=5" # 连续5次
severity: "info"
notify:
channels: ["dashboard"]# alert_engine.py
import time
import smtplib
import requests
from datetime import datetime, timedelta
from email.mime.text import MIMEText
from email.mime.multipart import MIMEMultipart
from typing import List, Dict, Optional
from enum import Enum
import yaml
class Severity(Enum):
INFO = 0
WARNING = 1
CRITICAL = 2
EMERGENCY = 3
class AlertStatus(Enum):
ACTIVE = 1
ACKNOWLEDGED = 2
RESOLVED = 3
SUPPRESSED = 4
@dataclass
class AlertRule:
name: str
condition: Dict
severity: str
suppress: List[str]
notify_channels: List[str]
escalate_after: Optional[str] = None
escalate_to: Optional[str] = None
@dataclass
class Alert:
alert_id: str
rule_name: str
node_id: Optional[int]
area: Optional[str]
severity: Severity
status: AlertStatus
message: str
created_at: float
acknowledged_at: Optional[float] = None
resolved_at: Optional[float] = None
notified_channels: List[str] = field(default_factory=list)
escalation_level: int = 0
class AlertEngine:
def __init__(self, config_path: str):
self.rules = self._load_rules(config_path)
self.active_alerts: Dict[str, Alert] = {}
self.alert_history: List[Alert] = []
self.suppression_cache: Dict[str, float] = {}
def _load_rules(self, config_path: str) -> List[AlertRule]:
with open(config_path, 'r') as f:
data = yaml.safe_load(f)
rules = []
for r in data['rules']:
rules.append(AlertRule(
name=r['name'],
condition=r['condition'],
severity=r['severity'],
suppress=r.get('suppress', []),
notify_channels=r['notify']['channels'],
escalate_after=r.get('escalate', {}).get('after'),
escalate_to=r.get('escalate', {}).get('to')
))
return rules
def evaluate(self, node_status: Dict):
"""评估节点状态,决定是否触发告警"""
for rule in self.rules:
if self._match_condition(rule.condition, node_status):
alert_key = self._get_alert_key(rule, node_status)
if self._is_suppressed(rule, alert_key):
continue
self._trigger_alert(rule, node_status, alert_key)
def _match_condition(self, condition: Dict, node_status: Dict) -> bool:
"""匹配告警条件"""
cond_type = condition.get('type')
if cond_type == 'node_offline':
duration = condition.get('duration', '5m')
duration_sec = self._parse_duration(duration)
if node_status.get('state') == 'OFFLINE':
offline_time = time.time() - node_status.get('last_seen', 0)
return offline_time >= duration_sec
return False
elif cond_type == 'area_offline':
area = condition.get('area')
min_count = condition.get('count', '>=3')
# 实现略:统计区域内离线节点数
return False
elif cond_type == 'collector_offline':
duration = condition.get('duration', '2m')
duration_sec = self._parse_duration(duration)
if node_status.get('type') == 'collector':
return node_status.get('state') == 'OFFLINE'
return False
elif cond_type == 'heartbeat_delay':
latency = condition.get('latency', '>3s')
count = condition.get('count', '>=5')
# 实现略:检查心跳延迟
return False
return False
def _is_suppressed(self, rule: AlertRule, alert_key: str) -> bool:
"""检查是否被抑制"""
# 检查抑制规则
for suppress_rule in rule.suppress:
if '不重复' in suppress_rule:
# 提取时间窗口
if '5分钟' in suppress_rule:
window = 300
elif '10分钟' in suppress_rule:
window = 600
else:
window = 300
if alert_key in self.suppression_cache:
if time.time() - self.suppression_cache[alert_key] < window:
return True
return False
def _trigger_alert(self, rule: AlertRule, node_status: Dict, alert_key: str):
"""触发告警"""
alert_id = f"{rule.name}_{node_status.get('node_id', 'unknown')}_{int(time.time())}"
severity = Severity[rule.severity.upper()]
alert = Alert(
alert_id=alert_id,
rule_name=rule.name,
node_id=node_status.get('node_id'),
area=node_status.get('area'),
severity=severity,
status=AlertStatus.ACTIVE,
message=self._generate_message(rule, node_status),
created_at=time.time()
)
self.active_alerts[alert_id] = alert
self.suppression_cache[alert_key] = time.time()
# 发送通知
self._dispatch_notifications(alert, rule.notify_channels)
print(f"Alert triggered: {alert.message}")
def _dispatch_notifications(self, alert: Alert, channels: List[str]):
"""分发通知到各渠道"""
for channel in channels:
try:
if channel == 'email':
self._send_email(alert)
elif channel == 'sms':
self._send_sms(alert)
elif channel == 'voice_call':
self._make_voice_call(alert)
elif channel == 'dashboard':
self._update_dashboard(alert)
elif channel in ['wechat', 'dingtalk']:
self._send_webhook(alert, channel)
alert.notified_channels.append(channel)
except Exception as e:
print(f"Failed to send {channel} notification: {e}")
def _send_email(self, alert: Alert):
"""发送邮件通知"""
# 实现略
pass
def _send_sms(self, alert: Alert):
"""发送短信通知"""
# 实现略
pass
def _make_voice_call(self, alert: Alert):
"""语音电话通知"""
# 实现略
pass
def _update_dashboard(self, alert: Alert):
"""更新监控大屏"""
# 实现略
pass
def _send_webhook(self, alert: Alert, channel: str):
"""发送Webhook通知(微信/钉钉)"""
# 实现略
pass
def _generate_message(self, rule: AlertRule, node_status: Dict) -> str:
"""生成告警消息"""
node_id = node_status.get('node_id', 'unknown')
area = node_status.get('area', 'unknown')
return f"节点 {node_id} (区域: {area}) 离线超过 {rule.condition.get('duration', 'N/A')}"
def _get_alert_key(self, rule: AlertRule, node_status: Dict) -> str:
"""生成告警键,用于去重"""
return f"{rule.name}_{node_status.get('node_id', 'unknown')}"
def _parse_duration(self, duration: str) -> int:
"""解析持续时间字符串为秒"""
if duration.endswith('s'):
return int(duration[:-1])
elif duration.endswith('m'):
return int(duration[:-1]) * 60
elif duration.endswith('h'):
return int(duration[:-1]) * 3600
else:
return 300 # 默认5分钟
def acknowledge_alert(self, alert_id: str, user: str):
"""确认告警"""
if alert_id in self.active_alerts:
alert = self.active_alerts[alert_id]
alert.status = AlertStatus.ACKNOWLEDGED
alert.acknowledged_at = time.time()
print(f"Alert {alert_id} acknowledged by {user}")
def resolve_alert(self, alert_id: str):
"""解决告警"""
if alert_id in self.active_alerts:
alert = self.active_alerts.pop(alert_id)
alert.status = AlertStatus.RESOLVED
alert.resolved_at = time.time()
self.alert_history.append(alert)
print(f"Alert {alert_id} resolved")
def get_active_alerts(self, severity: Optional[Severity] = None) -> List[Alert]:
"""获取活跃告警"""
if severity is None:
return list(self.active_alerts.values())
return [a for a in self.active_alerts.values() if a.severity == severity]
def check_escalation(self):
"""检查告警升级"""
now = time.time()
for alert in list(self.active_alerts.values()):
if alert.escalation_level > 0:
continue
# 查找对应的规则
for rule in self.rules:
if rule.name == alert.rule_name and rule.escalate_after:
escalate_after = self._parse_duration(rule.escalate_after)
if now - alert.created_at >= escalate_after:
# 升级告警
self._escalate_alert(alert, rule)
break
def _escalate_alert(self, alert: Alert, rule: AlertRule):
"""升级告警"""
alert.escalation_level += 1
alert.severity = Severity.CRITICAL
# 重新发送通知
self._dispatch_notifications(alert, rule.notify_channels)
print(f"Alert {alert.alert_id} escalated to {alert.severity}")┌─────────────────────────────────────────────────────────────────────┐
│ 档案馆环境监测网络 - 实时拓扑 │
│ │
│ ┌───────────────────────────────────────────────────────────────┐ │
│ │ 监控中心 │ │
│ │ ┌─────────────┐ ┌─────────────┐ ┌─────────────┐ │ │
│ │ │ 主服务器 │ │ 备服务器 │ │ 数据库 │ │ │
│ │ │ [ONLINE] │ │ [ONLINE] │ │ [ONLINE] │ │ │
│ │ └──────┬──────┘ └─────────────┘ └─────────────┘ │ │
│ └─────────┼──────────────────────────────────────────────────────┘ │
│ │ │
│ ┌─────────┼──────────────────────────────────────────────────────┐ │
│ │ 核心交换机 [OK] │ │
│ └─────────┼──────────────────────────────────────────────────────┘ │
│ │ │
│ ┌─────────┼──────────────────────────────────────────────────────┐ │
│ │ 区域A: 库房区 (采集器 C-001) [WARN] │ │
│ │ ┌──────┴──────┐ │ │
│ │ │ 节点1 [OK] │ 节点2 [OK] │ 节点3 [WARN] │ 节点4 [OK] │ │
│ │ └─────────────┘ │ │
│ └─────────────────────────────────────────────────────────────────┘ │
│ │
│ ┌─────────────────────────────────────────────────────────────────┐ │
│ │ 区域B: 阅览区 (采集器 C-002) [OK] │ │
│ │ ┌──────┬──────┐ │ │
│ │ │节点5│节点6│ 节点7 [OK] │ 节点8 [OK] │ │
│ │ └──────┴──────┘ │ │
│ └─────────────────────────────────────────────────────────────────┘ │
│ │
│ ┌─────────────────────────────────────────────────────────────────┐ │
│ │ 区域C: 数字化加工区 (采集器 C-003) [ERR] │ │
│ │ ┌──────┬──────┐ │ │
│ │ │节点9│节点10│ 节点11 [OFFLINE] │ 节点12 [OFFLINE] │ │
│ │ └──────┴──────┘ │ │
│ └─────────────────────────────────────────────────────────────────┘ │
│ │
│ ┌─────────────────────────────────────────────────────────────────┐ │
│ │ 告警面板 │ │
│ │ [CRITICAL] 区域C: 节点11、节点12 离线超过10分钟 │ │
│ │ [WARNING] 区域A: 节点3 心跳延迟 │ │
│ │ [INFO] 系统运行正常,共监控 48 个节点 │ │
│ └─────────────────────────────────────────────────────────────────┘ │
└─────────────────────────────────────────────────────────────────────┘{
"node_id": 3,
"area": "库房区A",
"collector": "C-001",
"status": "DEGRADED",
"last_seen": "2025-01-20 14:25:30",
"uptime": "15天 3小时 22分",
"firmware": "v2.1.3",
"signal": {
"rssi": -62,
"snr": 8.5,
"tx_power": 14
},
"health": {
"cpu_temp": 42.3,
"battery": 85,
"memory_usage": 45
},
"heartbeat": {
"interval": 30,
"last_response_time": 2.8,
"avg_response_time": 1.2,
"packet_loss": 0.5
},
"alerts": [
{
"level": "WARNING",
"message": "心跳响应时间超过阈值 (2.8s > 2.0s)",
"timestamp": "2025-01-20 14:25:35"
}
]
}┌────────────┐ ┌────────────┐ ┌────────────┐ ┌────────────┐
│ 检测到节点 │ │ 尝试远程 │ │ 远程重启 │ │ 等待节点 │
│ 离线 │────▶│ 重启 │────▶│ 成功? │────▶│ 重新上线 │
└────────────┘ └────────────┘ └─────┬──────┘ └─────┬──────┘
│ │
│ 否 │ 是
▼ ▼
┌────────────┐ ┌────────────┐
│ 标记需要 │ │ 恢复成功 │
│ 人工干预 │ │ 关闭告警 │
└────────────┘ └────────────┘# self_healing.py
import time
import threading
from typing import Dict, List
class SelfHealingManager:
"""节点自愈管理器"""
def __init__(self, alert_engine, node_controller):
self.alert_engine = alert_engine
self.node_controller = node_controller # 节点控制器,用于发送重启命令
self.healing_attempts: Dict[int, int] = {} # 节点ID -> 已尝试次数
self.max_attempts = 3
self.healing_interval = 300 # 两次自愈尝试之间的最小间隔(秒)
self.last_healing_time: Dict[int, float] = {}
def handle_node_offline(self, node_id: int, area: str):
"""处理节点离线事件"""
# 检查是否已经超过最大尝试次数
if self.healing_attempts.get(node_id, 0) >= self.max_attempts:
print(f"Node {node_id} has exceeded max healing attempts")
return
# 检查是否在冷却时间内
now = time.time()
last_time = self.last_healing_time.get(node_id, 0)
if now - last_time < self.healing_interval:
print(f"Node {node_id} is in healing cooldown")
return
# 启动自愈线程
healing_thread = threading.Thread(
target=self._healing_worker,
args=(node_id, area),
daemon=True
)
healing_thread.start()
def _healing_worker(self, node_id: int, area: str):
"""自愈工作线程"""
self.healing_attempts[node_id] = self.healing_attempts.get(node_id, 0) + 1
self.last_healing_time[node_id] = time.time()
print(f"Attempting to heal node {node_id} (attempt {self.healing_attempts[node_id]})")
# 步骤1: 尝试远程重启
if self.node_controller.reboot_node(node_id):
print(f"Node {node_id} reboot command sent")
# 等待节点重新上线
for i in range(12): # 等待最多60秒
time.sleep(5)
if self.node_controller.is_node_online(node_id):
print(f"Node {node_id} is back online after reboot")
# 关闭相关告警
self.alert_engine.resolve_alert_for_node(node_id)
return
print(f"Node {node_id} did not come back after reboot")
else:
print(f"Failed to send reboot command to node {node_id}")
# 步骤2: 如果远程重启失败,尝试重启采集器端口
if self.healing_attempts[node_id] >= 2:
collector_id = self.node_controller.get_collector_for_node(node_id)
if collector_id:
print(f"Attempting to reset collector {collector_id} port for node {node_id}")
self.node_controller.reset_collector_port(collector_id, node_id)
# 等待节点重新上线
for i in range(12):
time.sleep(5)
if self.node_controller.is_node_online(node_id):
print(f"Node {node_id} is back online after port reset")
self.alert_engine.resolve_alert_for_node(node_id)
return
# 步骤3: 如果所有尝试都失败,标记为需要人工干预
if self.healing_attempts[node_id] >= self.max_attempts:
print(f"Node {node_id} requires manual intervention")
self.alert_engine.escalate_to_manual(node_id, area)# report_generator.py
from datetime import datetime, timedelta
import csv
import json
class ReportGenerator:
"""运维报表生成器"""
def __init__(self, alert_engine, node_status_db):
self.alert_engine = alert_engine
self.node_status_db = node_status_db
def generate_daily_report(self, date: datetime) -> Dict:
"""生成日报"""
start_time = date.replace(hour=0, minute=0, second=0)
end_time = start_time + timedelta(days=1)
# 查询当天的告警
alerts = self.alert_engine.get_alerts_by_time_range(start_time, end_time)
# 查询节点状态统计
node_stats = self.node_status_db.get_node_stats_by_time_range(
start_time, end_time
)
report = {
"date": date.strftime("%Y-%m-%d"),
"summary": {
"total_nodes": len(node_stats),
"online_nodes": sum(1 for n in node_stats.values() if n['state'] == 'ONLINE'),
"offline_nodes": sum(1 for n in node_stats.values() if n['state'] == 'OFFLINE'),
"total_alerts": len(alerts),
"critical_alerts": sum(1 for a in alerts if a.severity == Severity.CRITICAL),
"avg_response_time": self._calc_avg_response_time(node_stats),
"data_completeness": self._calc_data_completeness(node_stats)
},
"alerts_by_area": self._group_alerts_by_area(alerts),
"node_availability": self._calc_node_availability(node_stats),
"heartbeat_stats": self._calc_heartbeat_stats(node_stats)
}
return report
def _calc_avg_response_time(self, node_stats: Dict) -> float:
"""计算平均心跳响应时间"""
response_times = [n.get('avg_response_time', 0) for n in node_stats.values()]
return sum(response_times) / len(response_times) if response_times else 0
def _calc_data_completeness(self, node_stats: Dict) -> float:
"""计算数据完整性百分比"""
total_expected = len(node_stats) * 24 * 60 / 5 # 假设每5分钟一条数据
total_received = sum(n.get('data_count', 0) for n in node_stats.values())
return (total_received / total_expected) * 100 if total_expected > 0 else 0
def _group_alerts_by_area(self, alerts: List[Alert]) -> Dict:
"""按区域分组告警"""
result = {}
for alert in alerts:
area = alert.area or 'unknown'
if area not in result:
result[area] = []
result[area].append({
'node_id': alert.node_id,
'severity': alert.severity.name,
'message': alert.message,
'created_at': alert.created_at
})
return result
def _calc_node_availability(self, node_stats: Dict) -> Dict:
"""计算节点可用性"""
result = {}
for node_id, stats in node_stats.items():
total_time = 24 * 3600 # 一天的总秒数
offline_time = stats.get('offline_duration', 0)
availability = ((total_time - offline_time) / total_time) * 100
result[node_id] = {
'availability': availability,
'offline_count': stats.get('offline_count', 0),
'total_offline_time': offline_time
}
return result
def _calc_heartbeat_stats(self, node_stats: Dict) -> Dict:
"""计算心跳统计"""
total_nodes = len(node_stats)
if total_nodes == 0:
return {}
response_times = [n.get('avg_response_time', 0) for n in node_stats.values()]
packet_losses = [n.get('packet_loss', 0) for n in node_stats.values()]
return {
'avg_response_time': sum(response_times) / total_nodes,
'max_response_time': max(response_times),
'min_response_time': min(response_times),
'avg_packet_loss': sum(packet_losses) / total_nodes,
'nodes_with_packet_loss': sum(1 for p in packet_losses if p > 0)
}=== 档案馆环境监测网络运维日报 ===
日期: 2025-01-20
【总体概况】
监控节点总数: 48
在线节点: 46 (95.8%)
离线节点: 2 (4.2%)
当日告警总数: 5
严重: 1
警告: 3
信息: 1
平均心跳响应时间: 1.2s
数据完整性: 99.7%
【按区域统计】
库房区A: 节点15个, 在线15个, 告警1个
阅览区: 节点12个, 在线12个, 告警0个
数字化加工区: 节点12个, 在线10个, 告警4个
办公区: 节点9个, 在线9个, 告警0个
【离线节点详情】
节点11 (数字化加工区): 离线2小时15分, 原因: 网络中断
节点12 (数字化加工区): 离线2小时15分, 原因: 网络中断
【告警详情】
[CRITICAL] 14:25 数字化加工区 节点11、12 离线
[WARNING] 10:30 库房区A 节点3 心跳延迟
[WARNING] 09:15 库房区A 节点7 心跳延迟
[WARNING] 08:45 阅览区 节点5 心跳延迟
[INFO] 08:00 系统启动完成
【运维建议】
- 数字化加工区网络中断已持续2小时,建议立即排查交换机端口
- 库房区A节点3、7心跳延迟,建议检查无线信号强度关键词:档案馆,环境监测网络,运维,传感器心跳检测,离线节点,自动告警,告警引擎,自愈机制,运维报表,审计
标签:#档案馆 #环境监测 #运维 #心跳检测 #离线告警 #告警引擎 #自愈机制 #运维报表 #网络可靠性 #传感器管理
原创声明:本文系作者授权腾讯云开发者社区发表,未经许可,不得转载。
如有侵权,请联系 cloudcommunity@tencent.com 删除。