模型突然會撞進來。Infra 不能等告警響了才學會關機。
上週看到 Simon Willison 整理 Anthropic Frontier Red Team 的結果,我第一個反應是 incident runbook 沒寫這一類事故。
在 100 個隨機 Binary Exploitation benchmark 任務中,GLM-5.3 有 4% 的試驗做到 control-flow hijack,Claude Mythos Preview 是 6%。較早的 Claude Opus 4.6 與 GLM-5.2 則是 0%。4% 不等於真實環境成功率,但從 0 到 4%,對 infra 來說已經是另一個風險類別。
把 capability jump 當成 production incident
模型升級不能只看 latency、token cost、錯誤率,還要加 capability gate。每次換版,先在隔離環境跑 abuse suite,測 prompt injection、secret discovery、tool misuse、sandbox escape 和 privilege escalation。
只要 tool scope 多一個 shell、cloud credential 或 production database,就算功能測試全綠也不能上。
「模型沒有 root,所以很安全」是常見自欺。如果它能建立 IAM policy、讀 CI secret、修改 Terraform state,名字不是 root 也沒差。權限拆成短效 token、單一用途、明確 resource scope,高風險動作放人工核准。agent 預設只讀測試環境,production write 經 policy engine。
kill switch 放 API gateway deny rule 或 service mesh circuit breaker。模型說「我已停止」不算關機,gateway 回 403 才算。
Canary、log 和演練
新模型先放 1% 流量,觀察一個 deploy cycle。除了 p95 latency,還要看 tool call、拒絕率、讀取 secret 和重試次數。Audit log 記模型版本、tool name、參數摘要、授權決策、結果和 trace id,不能只留「tool called」。
每季做 incident drill,5 分鐘內撤回 traffic,15 分鐘內完成 token revoke 和 tool disable。
說穿了就是,模型能力從 0% 跨到 4% 的那一刻,產品團隊看到 benchmark 新聞,infra 團隊看到的是 blast radius 變大。不要等它撞進 production 才補 runbook。
作者:CtrlC