01 项目架构与模块设计
前 26 篇我们学了很多独立的技能点——SSH 远程、Python 编程、SQLite 数据库、Telegram Bot、定时任务、Web 服务……但知识只有串联起来才能发挥真正的威力。本篇我们要做一个"毕业项目":用手机搭一台 24 小时运行的自动化运维监控站。
这个监控站能同时监控多台云服务器、树莓派、Web 服务甚至手机自身状态,一旦发现异常立刻通过 Telegram 推送告警,还提供一个手机浏览器就能访问的 Web 面板看趋势图。
💡 比喻:如果前 26 篇是在学各种零件怎么用(螺丝、齿轮、马达、电路板),那这一篇就是把所有零件组装成一台会自己跑的机器——而且这台机器还能监控其他机器。
1.1 六层架构
整个系统采用经典的分层架构,从下到上分别是采集层、存储层、分析层、通知层、展示层、调度层,每层职责单一,互不干扰。
📊 Termux 监控站架构图
⏰ 调度层
termux-job-scheduler 定时采集 · Termux:Widget 一键检查
🖥️ 展示层
FastAPI Web 面板 · matplotlib 趋势图 · CLI 工具
🔔 通知层
Telegram Bot 推送 · termux-notification 本地通知
💾 存储层
SQLite 历史数据 · JSON 配置文件
📡 采集层
SSH 远程采集 · HTTP 健康检查 · termux-api 本地状态
1.2 模块拆分与文件结构
按照"一个文件只做一件事"的原则,我们把系统拆成 7 个 Python 模块,再加配置文件和数据库。这样的结构既清晰又方便调试。
monitor_station/
├── config.json # 配置:服务器列表、阈值、Token
├── monitor.db # SQLite 数据库(自动创建)
├── monitor.py # SSH 采集服务器指标
├── checker.py # HTTP 健康检查
├── alerter.py # 告警逻辑
├── notifier.py # 通知推送(Telegram + 本地)
├── webapp.py # FastAPI Web 面板
├── scheduler.py # 定时任务管理
└── cli.py # 命令行入口工具
模块间的调用关系很清晰:scheduler 定时触发 → monitor/checker 采集数据 → 数据存 SQLite → alerter 检查阈值 → 触发异常时调用 notifier 推送 → webapp 从数据库读数据展示。
1.3 配置文件设计
把所有可变参数(服务器地址、告警阈值、API Token)都放到 config.json 里,代码里只读配置,不硬编码。这是做项目的基本素养——换服务器、改阈值都不用动代码。
# config.json
{
"servers": [
{
"name": "cloud-01",
"host": "8.137.176.208",
"port": 22,
"user": "root",
"key_file": "~/.ssh/id_ed25519"
}
],
"websites": [
{"name": "blog", "url": "https://example.com"}
],
"thresholds": {
"cpu": 80,
"memory": 85,
"disk": 90,
"temp": 70
},
"telegram": {
"bot_token": "你的Bot Token",
"chat_id": "你的Chat ID"
}
}
⚠️ 安全提示:config.json 里有密钥和 Token,不要上传到公开仓库。建议加入 .gitignore,权限设为 600(只有自己能读写)。
1.4 安装依赖
项目用到的 Python 库不多,都是我们之前学过或用过的。在 Termux 中一次性装好:
# 安装系统依赖(Termux 原生)
pkg install python libffi openssl libjpeg-turbo freetype
# 安装 Python 库
pip install paramiko==5.0.0 fastapi==0.141.1 uvicorn matplotlib==3.11.1
# 验证安装
python -c "import paramiko, fastapi, matplotlib; print('OK')"
💡 小贴士:如果 Termux 原生环境安装 matplotlib 编译失败,可以切换到 proot Ubuntu 环境中运行,apt 安装 python3-matplotlib 更省心。
02 SSH 采集:服务器指标采集器
监控的第一步是"拿到数据"。对于远程服务器,最直接的方式就是 SSH 登录上去执行命令,把输出解析成结构化数据。我们用 paramiko 这个 Python SSH 库来实现。
2.1 paramiko SSH 连接
paramiko 是纯 Python 实现的 SSH2 协议库,支持密钥登录、密码登录、端口转发等功能。我们用密钥方式连接,比密码更安全也更方便自动化。
# monitor.py — SSH 采集模块
import os
import json
import paramiko
def load_config(path="config.json"):
with open(path) as f:
return json.load(f)
def ssh_exec(host, port, user, key_file, cmd, timeout=10):
"""执行远程命令,返回 stdout 文本"""
client = paramiko.SSHClient()
client.set_missing_host_key_policy(paramiko.AutoAddPolicy())
key_path = os.path.expanduser(key_file)
pkey = paramiko.Ed25519Key.from_private_key_file(key_path)
try:
client.connect(hostname=host, port=port, username=user,
pkey=pkey, timeout=timeout)
stdin, stdout, stderr = client.exec_command(cmd)
output = stdout.read().decode().strip()
return output
finally:
client.close()
💡 小贴士:用 Ed25519 密钥比 RSA 更短更安全。生成命令:ssh-keygen -t ed25519,一路回车即可。
2.2 采集核心指标
我们重点采集 5 项核心指标:CPU 使用率、内存使用率、磁盘使用率、系统负载、CPU 温度。每一项对应一条 Linux 命令,把输出解析成数字。
# monitor.py — 指标采集函数
import re
def collect_metrics(server):
"""采集一台服务器的所有指标"""
host = server["host"]
port = server.get("port", 22)
user = server["user"]
key_file = server["key_file"]
metrics = {"name": server["name"], "status": "online"}
try:
# CPU 使用率(取 1 分钟平均 idle,100 - idle)
cpu_out = ssh_exec(host, port, user, key_file,
"top -bn1 | grep 'Cpu(s)' | awk '{print $8}'")
metrics["cpu"] = round(100 - float(cpu_out.replace("id,", "")), 1)
# 内存使用率
mem_out = ssh_exec(host, port, user, key_file,
"free | grep Mem | awk '{print $3/$2*100}'")
metrics["memory"] = round(float(mem_out), 1)
# 磁盘使用率(根分区)
disk_out = ssh_exec(host, port, user, key_file,
"df -h / | tail -1 | awk '{print $5}'")
metrics["disk"] = int(disk_out.replace("%", ""))
# 1 分钟负载
load_out = ssh_exec(host, port, user, key_file, "uptime")
load_match = re.search(r"load average: ([\d.]+)", load_out)
metrics["load_1m"] = float(load_match.group(1)) if load_match else 0
# CPU 温度(树莓派等 ARM 设备可用)
temp_out = ssh_exec(host, port, user, key_file,
"cat /sys/class/thermal/thermal_zone0/temp 2>/dev/null || echo 0")
metrics["temp"] = round(int(temp_out) / 1000, 1) if temp_out.strip() else 0
except Exception as e:
metrics["status"] = "offline"
metrics["error"] = str(e)
return metrics
2.3 批量采集与数据入库
单台采集搞定了,多台就是循环调用的事。采集完的数据存进 SQLite,方便后面查历史趋势。
# monitor.py — 批量采集 + 存储
import sqlite3
from datetime import datetime
def init_db(db_path="monitor.db"):
conn = sqlite3.connect(db_path)
conn.execute("""
CREATE TABLE IF NOT EXISTS metrics (
id INTEGER PRIMARY KEY AUTOINCREMENT,
server_name TEXT NOT NULL,
status TEXT,
cpu REAL,
memory REAL,
disk INTEGER,
load_1m REAL,
temp REAL,
created_at TEXT NOT NULL
)
""")
conn.commit()
conn.close()
def save_metrics(metrics, db_path="monitor.db"):
conn = sqlite3.connect(db_path)
conn.execute("""
INSERT INTO metrics
(server_name, status, cpu, memory, disk, load_1m, temp, created_at)
VALUES (?, ?, ?, ?, ?, ?, ?, ?)
""", (
metrics["name"], metrics.get("status"),
metrics.get("cpu"), metrics.get("memory"),
metrics.get("disk"), metrics.get("load_1m"),
metrics.get("temp"),
datetime.now().isoformat()
))
conn.commit()
conn.close()
def collect_all(config):
"""采集所有服务器,返回指标列表"""
results = []
for server in config["servers"]:
m = collect_metrics(server)
save_metrics(m)
results.append(m)
return results
⚠️ 常见错误
每台服务器只执行一次 SSH 连接 — 上面的代码每条命令都重新连接,5 条命令就是 5 次握手,慢还费资源
✓ 正确:建立一次连接后连续执行多条命令,或把多条命令拼成一条用分号分隔,一次 exec_command 搞定
03 HTTP 健康检查与本地状态
除了 SSH 采集服务器内部指标,我们还要从外部视角检查网站和服务是否能正常访问——用户能打开页面才是硬道理。同时,手机本身作为监控站,也需要监控自己的状态。
3.1 HTTP 健康检查
用 requests 发 GET 请求,检查状态码和响应时间。除了 200 OK,还要看响应时间是否在可接受范围内——页面返回 200 但等了 10 秒,用户体验同样很差。
# checker.py — HTTP 健康检查
import time
import sqlite3
from datetime import datetime
import requests
def check_website(name, url, timeout=5):
"""检查网站可用性,返回状态码、响应时间、状态"""
result = {
"name": name,
"url": url,
"status_code": 0,
"response_time": 0,
"status": "down",
}
start = time.time()
try:
resp = requests.get(url, timeout=timeout,
headers={"User-Agent": "Termux-Monitor/1.0"})
result["status_code"] = resp.status_code
result["response_time"] = round(
(time.time() - start) * 1000) # 毫秒
if resp.status_code == 200:
result["status"] = "up"
elif 300 <= resp.status_code < 400:
result["status"] = "redirect"
else:
result["status"] = "error"
except Exception as e:
result["response_time"] = round(
(time.time() - start) * 1000)
result["error"] = str(e)
return result
3.2 termux-api 采集手机状态
监控站自己的健康也很重要——手机没电了、温度太高了、存储空间不够了,都得及时知道。用 Termux:API 可以读取电池、温度、存储等信息。
# checker.py — 手机本地状态采集
import subprocess
import json
def termux_battery():
"""获取电池信息"""
try:
out = subprocess.check_output(
["termux-battery-status"],
timeout=5).decode()
data = json.loads(out)
return {
"percentage": data.get("percentage", 0),
"status": data.get("status", "UNKNOWN"),
"temperature": data.get("temperature", 0),
}
except Exception:
return {"percentage": 0, "status": "ERROR"}
def termux_storage():
"""获取 Termux 存储使用情况"""
out = subprocess.check_output(
["df", "-h", "/data"]).decode()
line = out.strip().split("\n")[-1]
parts = line.split()
return {
"total": parts[1],
"used": parts[2],
"use_pct": int(parts[4].replace("%", "")),
}
⚠️ 注意:需要先安装 Termux:API App 和 termux-api 包:pkg install termux-api,并且从 F-Droid 安装 Termux:API 应用,两者缺一不可。
3.3 网站状态入库
和服务器指标一样,网站检查结果也要存进数据库,方便后续查可用性趋势和计算 SLA。
# checker.py — 批量检查 + 存储
def init_website_table(db_path="monitor.db"):
conn = sqlite3.connect(db_path)
conn.execute("""
CREATE TABLE IF NOT EXISTS website_checks (
id INTEGER PRIMARY KEY AUTOINCREMENT,
name TEXT NOT NULL,
url TEXT,
status_code INTEGER,
response_time INTEGER,
status TEXT,
created_at TEXT NOT NULL
)
""")
conn.commit()
conn.close()
def check_all(config, db_path="monitor.db"):
results = []
conn = sqlite3.connect(db_path)
for site in config.get("websites", []):
result = check_website(site["name"], site["url"])
results.append(result)
conn.execute("""
INSERT INTO website_checks
(name, url, status_code, response_time, status, created_at)
VALUES (?, ?, ?, ?, ?, ?)
""", (
result["name"], result["url"],
result["status_code"],
result["response_time"],
result["status"],
datetime.now().isoformat()
))
conn.commit()
conn.close()
return results
04 告警逻辑与通知推送
光采集数据没用,出问题了能及时通知到人才是监控的核心价值。告警不能太灵敏(误报多了就没人看了),也不能太迟钝(真出问题时已经晚了)。我们用"连续 N 次超标才告警"的策略来平衡。
4.1 阈值检测与告警去抖
告警去抖(Debounce)是监控系统的基本功——CPU 偶尔冲到 90% 一秒钟很正常,不值得半夜把人叫起来。只有连续多次超标才说明真的有问题。
# alerter.py — 告警逻辑
import json
import os
class Alerter:
def __init__(self, thresholds, state_file="alert_state.json"):
self.thresholds = thresholds
self.state_file = state_file
self.state = self._load_state()
def _load_state(self):
if os.path.exists(self.state_file):
with open(self.state_file) as f:
return json.load(f)
return {} # { server_name_metric: count }
def _save_state(self):
with open(self.state_file, "w") as f:
json.dump(self.state, f, indent=2)
def check_metrics(self, metrics):
"""检查一台服务器指标,返回告警列表"""
alerts = []
name = metrics["name"]
# 离线告警
if metrics.get("status") == "offline":
key = f"{name}_offline"
self.state[key] = self.state.get(key, 0) + 1
if self.state[key] == 2: # 连续 2 次离线才告警
alerts.append({
"level": "critical",
"msg": f"🚨 {name} 离线,连续 2 次连接失败",
})
return alerts
# 各项指标阈值检查
checks = [
("cpu", "CPU", "warning"),
("memory", "内存", "warning"),
("disk", "磁盘", "critical"),
("temp", "温度", "warning"),
]
for metric, label, level in checks:
value = metrics.get(metric, 0)
threshold = self.thresholds.get(metric, 100)
key = f"{name}_{metric}"
if value >= threshold:
self.state[key] = self.state.get(key, 0) + 1
# 连续 3 次超标才告警
if self.state[key] == 3:
alerts.append({
"level": level,
"msg": (
f"⚠️ {name} {label} {value}% "
f"(阈值 {threshold}%,连续 3 次)"
),
})
else:
# 恢复正常,重置计数
self.state.pop(key, None)
self._save_state()
return alerts
💡 小贴士:告警状态持久化到文件很重要——如果 Termux 进程被杀重启了,不会因为"第一次超标"就疯狂告警,也不会漏掉正在持续的故障。
4.2 Telegram Bot 推送告警
告警消息通过 Telegram Bot 推送到手机,这样即使不在 Termux 界面也能第一时间看到。第 25 篇我们已经学过 python-telegram-bot,这里直接用 requests 调 API 更轻量。
# notifier.py — 通知模块
import requests
import subprocess
class Notifier:
def __init__(self, bot_token, chat_id):
self.bot_token = bot_token
self.chat_id = chat_id
def send_telegram(self, text):
"""发送 Telegram 消息"""
url = f"https://api.telegram.org/bot{self.bot_token}/sendMessage"
try:
requests.post(url, json={
"chat_id": self.chat_id,
"text": text,
"parse_mode": "HTML",
}, timeout=5)
except Exception as e:
print(f"Telegram 发送失败: {e}")
def send_local(self, title, content):
"""本地通知(手机通知栏)"""
try:
subprocess.run([
"termux-notification",
"--title", title,
"--content", content,
"--priority", "high",
], timeout=5, capture_output=True)
except Exception:
pass
def alert(self, alerts):
"""批量发送告警"""
for a in alerts:
self.send_telegram(a["msg"])
# 严重告警同时发本地通知
if a["level"] == "critical":
self.send_local("监控告警", a["msg"])
4.3 串联完整告警流程
把采集、告警、通知串起来,一次完整的监控检查就是:采集数据 → 存数据库 → 检查阈值 → 触发告警 → 推送通知。
# cli.py — 命令行入口
import sys
from monitor import load_config, init_db, collect_all
from checker import init_website_table, check_all
from alerter import Alerter
from notifier import Notifier
def cmd_run():
config = load_config()
init_db()
init_website_table()
# 1. 采集服务器指标
metrics_list = collect_all(config)
# 2. 检查网站
website_results = check_all(config)
# 3. 告警检测
alerter = Alerter(config["thresholds"])
all_alerts = []
for m in metrics_list:
all_alerts.extend(alerter.check_metrics(m))
# 4. 推送告警
if all_alerts:
tg = config["telegram"]
notifier = Notifier(tg["bot_token"], tg["chat_id"])
notifier.alert(all_alerts)
print(f"告警: {len(all_alerts)} 条")
else:
print("一切正常 ✓")
if __name__ == "__main__":
if len(sys.argv) > 1 and sys.argv[1] == "run":
cmd_run()
05
FastAPI Web 监控面板
告警通知只能告诉你"出问题了",但排查问题时还需要看历史趋势和当前全貌。本节用 FastAPI 0.141.1 搭一个轻量 Web 面板,手机浏览器直接打开就能看所有服务器的实时状态和 CPU / 内存趋势图。
5.1 安装依赖与项目结构
Web 面板需要三个库:FastAPI 负责 API 路由,uvicorn 是 ASGI 服务器,matplotlib 用来画趋势图(生成 PNG 后转 base64 内嵌到页面,避免静态文件托管)。
# 安装依赖
pip install fastapi==0.141.1 uvicorn matplotlib
在 monitor 目录下新增 web.py,项目结构变成:
monitor/
├── config.json
├── collector.py # SSH 采集
├── health.py # HTTP 健康检查
├── storage.py # SQLite 存储
├── alerter.py # 告警逻辑
├── notifier.py # 通知推送
├── cli.py # 命令行入口
└── web.py # FastAPI 面板 ← 新增
5.2 API 接口设计
面板需要三个核心接口:/api/servers 返回所有服务器最新状态,/api/servers/{name} 返回单台服务器最近 N 条历史数据,/api/websites 返回网站健康状态。
web.py — API 部分
from fastapi import FastAPI
from fastapi.responses import HTMLResponse
from storage import Storage
import json
with open("config.json") as f:
config = json.load(f)
app = FastAPI(title="手机监控面板")
db = Storage(config["db_path"])
@app.get("/api/servers")
def list_servers():
"""所有服务器最新一条指标"""
rows = db.get_latest_metrics()
return {"servers": rows}
@app.get("/api/servers/{name}")
def server_detail(name: str, limit: int = 60):
"""单台服务器最近 limit 条历史"""
rows = db.get_history(name, limit)
return {"name": name, "history": rows}
@app.get("/api/websites")
def list_websites():
"""所有网站最新健康状态"""
rows = db.get_latest_websites()
return {"websites": rows}
Storage 类需要补两个查询方法:get_latest_metrics() 和 get_history()。用窗口函数 ROW_NUMBER() 取每台服务器最新一条。
5.3 趋势图生成(matplotlib)
手机上看趋势图比看数字列表直观得多。用 matplotlib 生成 PNG 后转 base64,直接内嵌到 HTML 里,不需要静态文件服务器,也不需要前端图表库。
web.py — 趋势图生成
import matplotlib
matplotlib.use("Agg") # 非交互式后端,Termux 无显示也能跑
import matplotlib.pyplot as plt
import base64
from io import BytesIO
def generate_chart(history, metric, color="#16a34a"):
"""生成指定指标的趋势图,返回 base64 PNG"""
values = [row[metric] for row in history if row[metric] is not None]
if not values:
return None
fig, ax = plt.subplots(figsize=(6, 2.5), dpi=100)
ax.plot(values, color=color, linewidth=1.5)
ax.fill_between(range(len(values)), values, alpha=0.1, color=color)
ax.set_ylim(0, 100)
ax.set_title(metric.upper(), fontsize=10)
ax.tick_params(axis="both", labelsize=8)
ax.grid(alpha=0.3)
fig.tight_layout()
buf = BytesIO()
fig.savefig(buf, format="png", bbox_inches="tight")
plt.close(fig)
buf.seek(0)
return base64.b64encode(buf.read()).decode()
💡 小贴士:matplotlib 后端
Termux 没有图形界面,必须在 import pyplot 之前设置 matplotlib.use("Agg"),否则会报 no display name 错误。这个坑在所有无显示环境(服务器、Docker)都会遇到。
5.4 HTML 面板页面
FastAPI 可以直接返回 HTML 字符串,不用模板引擎。面板设计成移动端优先的单列布局:顶部总览 → 服务器卡片列表 → 网站状态列表。
web.py — HTML 面板
@app.get("/", response_class=HTMLResponse)
def dashboard():
servers = db.get_latest_metrics()
websites = db.get_latest_websites()
# 统计在线/离线数
online = sum(1 for s in servers if s["status"] == "online")
total = len(servers)
# 生成服务器卡片 HTML
cards = ""
for s in servers:
status_color = "#16a34a" if s["status"] == "online" else "#ef4444"
status_text = "在线" if s["status"] == "online" else "离线"
cards += f'''
<div class="card">
<div class="card-header">
<span class="name">{s["name"]}</span>
<span class="status" style="background:{status_color}">{status_text}</span>
</div>
<div class="metrics">
<div>CPU: {s.get("cpu", 0)}%</div>
<div>内存: {s.get("memory", 0)}%</div>
<div>磁盘: {s.get("disk", 0)}%</div>
</div>
</div>
'''
# 网站列表
site_rows = ""
for w in websites:
dot = "🟢" if w["status"] == "up" else "🔴"
site_rows += f'<div class="site-row">{dot} {w["name"]} ' \
f'<span class="ms">{w.get("response_time", 0)}ms</span></div>'
# 完整 HTML
html = f'''<!DOCTYPE html>
<html><head><meta name="viewport" content="width=device-width,initial-scale=1">
<title>手机监控面板</title>
<style>
body {{ font-family: system-ui; background: #f5f5f5; margin: 0; padding: 12px; }}
.header {{ text-align: center; padding: 16px; }}
.header h1 {{ margin: 0; font-size: 20px; color: #333; }}
.summary {{ display: flex; justify-content: center; margin: 12px 0; }}
.summary div {{ text-align: center; margin: 0 12px; }}
.summary .num {{ font-size: 28px; font-weight: bold; color: #16a34a; }}
.summary .label {{ font-size: 12px; color: #666; }}
.card {{ background: white; border-radius: 12px; padding: 14px; margin-bottom: 10px; }}
.card-header {{ display: flex; justify-content: space-between; align-items: center; margin-bottom: 10px; }}
.name {{ font-weight: bold; font-size: 16px; }}
.status {{ color: white; padding: 2px 10px; border-radius: 12px; font-size: 12px; }}
.metrics {{ display: flex; justify-content: space-between; font-size: 13px; color: #555; }}
.site-row {{ background: white; padding: 10px 14px; border-radius: 8px; margin-bottom: 6px; font-size: 14px; display: flex; justify-content: space-between; }}
.ms {{ color: #999; font-size: 12px; }}
.section-title {{ font-size: 15px; font-weight: bold; margin: 16px 0 8px; color: #333; }}
</style>
</head><body>
<div class="header">
<h1>📱 手机监控中心</h1>
<div class="summary">
<div><div class="num">{online}/{total}</div><div class="label">服务器在线</div></div>
<div><div class="num">{len(websites)}</div><div class="label">监控站点</div></div>
</div>
</div>
<div class="section-title">服务器状态</div>
{cards}
<div class="section-title">网站健康</div>
{site_rows}
</body></html>'''
return html
启动面板,在手机浏览器访问 http://localhost:8000 就能看到效果:
# 启动 Web 面板
cd ~/monitor && uvicorn web:app --host 0.0.0.0 --port 8000
--host 0.0.0.0 允许同一 Wi-Fi 下的其他设备访问。如果想在外网访问,配合第 21 篇的 frp 内网穿透即可。
5.5 简单认证保护
监控面板暴露了服务器状态,不能谁都能看。加一层 HTTP Basic Auth,FastAPI 自带 HTTPBasic 依赖注入。
web.py — 认证保护
from fastapi import Depends, HTTPException, status
from fastapi.security import HTTPBasic, HTTPBasicCredentials
security = HTTPBasic()
def verify_auth(credentials: HTTPBasicCredentials = Depends(security)):
user = config.get("web_auth", {})
correct_user = credentials.username == user.get("username", "admin")
correct_pwd = credentials.password == user.get("password", "")
if not (correct_user and correct_pwd):
raise HTTPException(
status_code=status.HTTP_401_UNAUTHORIZED,
detail="认证失败",
headers={"WWW-Authenticate": "Basic"},
)
@app.get("/", response_class=HTMLResponse)
def dashboard(_=Depends(verify_auth)):
... # 原有逻辑不变
在 config.json 中加上认证账号:
config.json 新增字段
{
"web_auth"
: {
"username"
:
"admin"
,
"password"
:
"你的密码"
},
...
}
⚠️ 常见错误:matplotlib 中文乱码
Termux 默认没有中文字体,matplotlib 画中文标题会显示成方框。解决方案:① 标题只用英文/数字(推荐,最简单);② 安装 fonts-noto-cjk 包并指定字体路径。本篇示例标题用 CPU / MEMORY 等英文,避开这个问题。
06
部署运行与定时调度
代码写完只是第一步,真正的挑战是让它在手机上 7×24 小时稳定运行。手机不是服务器,有电池优化、内存回收、网络切换等问题。本节讲解完整的部署方案:后台常驻、定时采集、开机自启、异常重启。
6.1 tmux 后台常驻
采集脚本和 Web 面板都需要长期运行。最简单的方式是用 tmux 开两个窗口,分别跑采集循环和 Web 服务。这样即使 Termux 退到后台,进程也不会被直接杀掉。
# 启动 tmux 会话
tmux new-session -s monitor -d
# 窗口 0:运行采集循环
tmux send-keys -t monitor:0 "cd ~/monitor && python cli.py loop" C-m
# 新建窗口 1:运行 Web 面板
tmux new-window -t monitor -n web
tmux send-keys -t monitor:web "cd ~/monitor && uvicorn web:app --host 0.0.0.0 --port 8000" C-m
# 查看 / 切回会话
tmux attach -t monitor
python cli.py loop 是循环模式,需要在 cli.py 里加一个 loop 子命令,每隔 N 分钟执行一次采集:
cli.py — 新增 loop 模式
import time
def cmd_loop():
"""循环采集模式,每 interval 分钟跑一次"""
interval = config.get("interval_minutes", 5) * 60
print(f"循环模式启动,每 {interval//60} 分钟采集一次")
while True:
try:
cmd_run()
except Exception as e:
print(f"采集异常: {e}")
time.sleep(interval)
if __name__ == "__main__":
if len(sys.argv) > 1:
if sys.argv[1] == "run":
cmd_run()
elif sys.argv[1] == "loop":
cmd_loop()
6.2 保活与防杀
Android 系统会在后台杀进程释放内存,光靠 tmux 不够。需要三重保险:Wake Lock 唤醒锁 + 电池优化白名单 + 通知栏常驻。
# 1. 获取 Wake Lock(防止 CPU 休眠)
termux-wake-lock
# 2. 把 Termux 加入电池优化白名单
termux-setup-android-internal-storage # 先授权存储
# 然后手动设置:设置 → 电池 → 电池优化 → 选择 Termux → 不优化
# 3. 常驻通知栏(Termux 自带,运行时会有通知)
# 在 Termux 通知中可以看到"Acquiring wakelock"即生效
💡 小贴士:电量与性能的平衡
24 小时跑监控会加快耗电。建议:① 采集间隔设为 5 分钟而非 1 分钟,减少网络和 CPU 唤醒;② 手机插着充电器运行(旧手机当监控专用机最理想);③ 夜间可降低采集频率到 15 分钟。
6.3 termux-job-scheduler 定时调度
如果不想一直跑 while 循环,可以用 Termux 自带的定时任务调度器 termux-job-scheduler,它基于 Android JobScheduler API,系统级调度更省电,进程被杀也能按时唤醒执行。
# 安装调度器
pkg install termux-job-scheduler
# 创建采集脚本
cat
> ~/monitor/run_monitor.sh
<'EOF'
#!/data/data/com.termux/files/usr/bin/bash
cd ~/monitor
python cli.py run >> monitor.log 2>&1
EOF
chmod +x ~/monitor/run_monitor.sh
# 注册定时任务:每 5 分钟执行一次,job_id=100
termux-job-scheduler -s 100 --period-ms 300000 \
--script ~/monitor/run_monitor.sh
# 查看已注册任务
termux-job-scheduler -l
# 取消任务
termux-job-scheduler -c 100
两种方案对比:while 循环精度高但耗电且进程被杀就停;job-scheduler 由系统托管更省电,但间隔最小 5 分钟且执行时机不保证精确。监控场景推荐后者。
6.4 开机自启动
手机重启后监控不能断。用 termux-boot 包实现开机自动执行脚本。
# 安装 termux-boot
pkg install termux-boot
# 创建开机脚本目录和脚本
mkdir -p ~/.termux/boot
cat
> ~/.termux/boot/start-monitor.sh
<'EOF'
#!/data/data/com.termux/files/usr/bin/bash
# 开机自动启动监控
termux-wake-lock
# 启动 Web 面板(tmux 后台)
tmux new-session -s monitor -d
tmux send-keys -t monitor "cd ~/monitor && uvicorn web:app --host 0.0.0.0 --port 8000" C-m
# 注册定时采集任务
termux-job-scheduler -s 100 --period-ms 300000 \
--script ~/monitor/run_monitor.sh
EOF
chmod +x ~/.termux/boot/start-monitor.sh
重启手机后,Termux 会自动启动并执行 ~/.termux/boot/ 下所有脚本。注意第一次必须手动打开一次 Termux,授权"自启动"权限后才能生效。
6.5 日志与异常监控
监控系统自己也可能出问题。加一个简单的看门狗:定期检查日志文件最后更新时间,如果超过阈值说明采集卡住了,给自己发告警。
watchdog.py — 看门狗
import os
import time
from notifier import Notifier
import json
def check_alive(log_file, max_idle_min=15):
if not os.path.exists(log_file):
return False, "日志文件不存在"
mtime = os.path.getmtime(log_file)
idle_min = (time.time() - mtime) / 60
if idle_min > max_idle_min:
return False, f"日志已 {idle_min:.0f} 分钟未更新"
return True, "正常"
if __name__ == "__main__":
with open("config.json") as f:
config = json.load(f)
alive, msg = check_alive("monitor.log")
if not alive:
tg = config["telegram"]
notifier = Notifier(tg["bot_token"], tg["chat_id"])
notifier.send_telegram(f"🛑 监控系统异常: {msg}")
看门狗可以单独注册一个每小时执行的 job-scheduler 任务,形成"监控的监控"。
⚠️ 常见错误:termux-job-scheduler 不执行
常见原因:① 脚本路径用了 ~ 但调度器不认,必须写绝对路径 /data/data/com.termux/files/home/monitor/run_monitor.sh;② 脚本没有执行权限;③ Android 电池优化限制了后台任务,需加入白名单;④ 最小周期是 5 分钟(Android 限制),设更短不会生效。