DB 接入塊 0-5 上主幹後的收尾工作,讓 DB-on 模式可真人使用 + 補齊功能與測試。 OIDC / pairing FK 修復(接 DB 上線必要): - 新建 internal/user package(User + Store + InMemory + Postgres);OIDC callback 驗證 id_token 成功後 fail-closed upsert users(sub 直接當 users.id,MC sub 為 UUID) - pairing exchange 雲端自建 device(不動 local-tool)+ 同 tx 綁 session token; 自建 device 空 serial 寫 NULL(避免撞 partial unique) - device.SaveTx / session.CreateTx 新增 tx-aware 版本 B4 model metadata: - 轉檔 result 的 analysis_info(input_shape/classes/framework)串進 model: converter_client → flow → adapter → model.Model → PG → ModelResponse DTO - input_shape 優先用陣列、後備四維組 NCHW、缺一不亂組;全 optional 防禦性 - 前端詳細頁顯示(另 repo);轉檔端串接交接檔 b4-converter-handoff.md nginx healthz(部署層): - 新增 /healthz/deep 轉發 backend(ping PG+Redis、down 回 503)給 LB - 修掉 default_server return 444 短路 bug(docker healthcheck 長期 unhealthy 真因) storage error 統一映射(不洩漏 storage 後端細節)。 測試:補 internal/api(storage/errors handler)、cmd/api-server(seed/adapter)、 internal/db(redis)、relay/session 弱處,含 testcontainers integration。 DB 接入相關 package 真環境覆蓋達 88-94%。全程 Reviewer 審查 + 130 真 PG/Redis dbtest 綠。 Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
13 KiB
visionA 部署指南
本檔聚焦「DB 接入版 backend 部署到 stage 並接真 PG/Redis」的端到端流程與煙測結果。 stage 基礎設定(host nginx、HTTPS termination、docker daemon 連線)另見同目錄
stage-deployment-setup.md; Phase 0.6 交接見phase-0.6-handover.md。
1. 架構總覽(DB 接入後)
公司 host nginx (HTTPS termination, LE 證書)
│ stage-9527.innovedus.com:9527
▼
visiona container (image visiona:stage)
├─ 內層 nginx :80 ── 反代 ──► api-server :3721 (/api/*)
│ ── 反代 ──► frontend standalone (Next.js)
│ ── /healthz → return 200(淺層 liveness,不打 backend)
│ ── /healthz/deep → 反代 api-server /healthz(深層 readiness,ping PG+Redis)見 §5/§9
├─ api-server :3721
│ ├─ Postgres pool ──► 192.168.0.130:5432/visiona(真 PG)
│ └─ Redis client ──► visiona-redis:6379/0(真 Redis)
└─ remote-proxy :3800/:3801(tunnel)
│
docker network: visiona-stage_default
├─ visiona 172.19.0.2
└─ visiona-redis (同網段,hostname 解析靠這個 network)
DB-on 啟用條件(.env.stage 已設):
- Postgres:
VISIONA_DB_HOST+VISIONA_DB_USER+VISIONA_DB_NAME非空 → 自動建池 + auto-migrate + repository 切 Postgres。 - Redis:
VISIONA_REDIS_HOST非空(無密碼也算啟用)→ userSession 切 Redis、cookie session 持久化。 VISIONA_DB_AUTO_MIGRATE預設 true(internal/config/load.go:109)→ 啟動自動跑migrate up。
2. 前置需求
| 項目 | 內容 |
|---|---|
| docker daemon | stage host 130 開 tcp://192.168.0.130:2375,公司內網直連(VPN 大流量會卡,見 deploy-stage-v2.sh 註解) |
.env.stage |
含 DB + Redis + OIDC + storage 等 env,git-ignored,不進 repo(已驗證 git check-ignore 命中) |
| 真 PG | 192.168.0.130:5432 db=visiona user=vsausr sslmode=disable,schema 已 migrate 到 version 3 |
| 真 Redis | container visiona-redis,network visiona-stage_default,無密碼,6379 |
| 部署腳本 | scripts/deploy-stage-v2.sh(remote build 模式,build 全在 stage daemon 跑) |
3. 部署步驟(DB 接入版)
3.1 network 接 visiona-redis — 怎麼解決的(關鍵)
backend 用 hostname visiona-redis 連 Redis,必須跟 visiona-redis 同 docker network否則解析失敗。
解決方式:靠 compose project name 自動對上,不需改 compose。
deploy-stage-v2.sh用-p visiona-stage起 compose → 預設 network 名為visiona-stage_default。visiona-redis本來就在visiona-stage_default(Up,label 顯示屬於同 stack)。- 因此新
visionacontainer 起來自動落在同網段,visiona-redishostname 直接可解。
驗證(部署前):
export DOCKER_HOST=tcp://192.168.0.130:2375
docker run --rm --network visiona-stage_default redis:alpine redis-cli -h visiona-redis ping # → PONG
docker run --rm --network visiona-stage_default postgres:16-alpine pg_isready -h 192.168.0.130 -p 5432 # → accepting connections
⚠️ 注意:docker-compose.stage.yml 沒有顯式宣告 external network。目前能對上是因為「compose 自建的 default network 名稱 = visiona-redis 所在的 network 名稱」這個巧合(同 project name)。若未來 visiona-redis 改由別的 compose stack 管、或改 project name,需在 compose 顯式 networks: 接 external visiona-stage_default。建議後續硬化(見 §10)。
3.2 部署指令
# 1. 先備份當前 image 供 rollback
export DOCKER_HOST=tcp://192.168.0.130:2375
CURID=$(docker inspect visiona --format '{{.Image}}' | cut -c8-19)
docker tag "$CURID" visiona:stage-rollback-pre-db
# 2. build + deploy main HEAD(remote build,全在 stage daemon 跑)
DOCKER_HOST=tcp://192.168.0.130:2375 bash scripts/deploy-stage-v2.sh
部署會 Recreate visiona container(覆蓋舊版)。本次部署版本:main HEAD 4d0b870(DB 接入塊 0-5),image tag visiona:stage + visiona:stage-20260620-184814-4d0b870。
4. 啟動煙測(本次實測,2026-06-20)
啟動 log 關鍵行(.autoflow/07-delivery/logs/block6-container-startup-202606201848.log):
postgres pool initialized target=192.168.0.130:5432/visiona sslmode=disable max_conns=10 min_conns=2
migrate up: applied version=3
migrations applied target=192.168.0.130:5432/visiona
redis client initialized target=visiona-redis:6379/0 db=0
pairing/session token stores initialized backend=postgres
user session store initialized backend=redis
device repository initialized backend=postgres
device unpairer initialized backend=postgres-tx
model repository initialized backend=postgres
conversion service initialized converter_base_url=http://192.168.0.130:9501
file access (FAA download) initialized (Phase 0.9 功能正常)
api-server listening addr=0.0.0.0:3721
無 fatal / panic。所有 6 個 store 都顯示 backend=postgres / backend=redis(DB-on 模式生效)。
5. healthz 行為(重要)
2026-06-20 更新:§8 的 nginx /healthz 不反映 DB 健康問題已修復並重部署(見 §8、§9)。 現在 stage 有兩條對外健康路由:淺層
/healthz(存活)與深層/healthz/deep(readiness,反映 PG+Redis)。
| 路徑 | 回應(健康時) | 說明 |
|---|---|---|
https://stage-9527.innovedus.com:9527/healthz(公開,淺層) |
200 "ok" |
nginx 直接 return,不打 backend、不反映 DB。用途:liveness / 「nginx 活著」探測。 |
https://stage-9527.innovedus.com:9527/healthz/deep(公開,深層)✅ 新增 |
200 {"checks":{"postgres":"ok","redis":"ok"},"status":"ok"} |
nginx proxy_pass 到 backend /healthz,會 ping PG + Redis、down 回 503。用途:LB readiness。 |
backend :3721/healthz(容器內) |
200 {"checks":{"postgres":"ok","redis":"ok"},"status":"ok"} |
深層路由的後端來源,真正 ping PG + Redis。 |
backend healthz 實作(internal/api/health.go):每次呼叫 ping PG + Redis,任一失敗回 503。
5.1 503 fail-fast 實測(停 visiona-redis)
log:.autoflow/07-delivery/logs/block6-healthz-503-202606201848.log
docker stop visiona-redis
→ backend :3721/healthz HTTP 503 {"checks":{"postgres":"ok","redis":"down"},"status":"unavailable"}
docker start visiona-redis
→ backend :3721/healthz HTTP 200 {"checks":{"postgres":"ok","redis":"ok"},"status":"ok"}
- Redis down → 正確偵測
redis:down回 503,PG 仍ok。 - visiona container 沒有因 Redis 中斷而 crash(fail-fast 只在啟動,runtime 降級僅反映在 healthz)。
- Redis 恢復後 healthz 自動回 200。
6. 重啟資料持久化實測
log:.autoflow/07-delivery/logs/block6-persistence-202606201848.log
插入 model row (name=block6-persist-...) 進 stage PG → rows_before_restart = 1
docker restart visiona
→ 啟動 log: "migrate up: no change (already at latest version)"(migration 冪等)
→ backend healthz 200
查詢 model row → rows_after_restart = 1(同 marker,資料還在)
(測試資料已 cleanup)
證明真部署環境下 DB-on 的資料跨 container 重啟持久化。
7. 既有功能回歸
| 檢查 | 結果 |
|---|---|
GET /api/models(公開,auth-gated) |
401(api-server 活、DB-backed handler 可達,非 500/502) |
GET /(frontend) |
200 |
| conversion service | 啟動 log 顯示 initialized(converter base url 正確) |
| FAA download(Phase 0.9) | 啟動 log 顯示 initialized |
8. ✅ 已修復:nginx /healthz 不反映 DB 健康(+ docker healthcheck false 444)
原問題:backend 的 DB-aware healthz(會 ping PG/Redis、down 回 503 的 fail-fast 邏輯)實作了但 nginx 攔截不轉發。
docker/nginx.stage.conf兩個location = /healthz都return 200 "ok",不proxy_pass到 api-server。- 結果:對外
/healthz永遠 200,即使 PG/Redis 掛了。backend 的 503 邏輯形同 dead code。
併發現的第二個 bug(本次一併修):default_server block 結尾用 server-level return 444;,在 nginx rewrite phase 會「先於」location 匹配短路掉所有請求 —— 包含本該回 200 的 location = /healthz。導致 docker healthcheck(從 container 內打 localhost/healthz、來源 127.0.0.1)也被打成 444,container 長期顯示 (unhealthy)(FailingStreak 已累積到 96)。原 deploy-stage-v2.sh:196 把這當成「Host 白名單造成的 false negative」,實際成因是 return 444 的 phase 順序,不是 Host 白名單。
修復內容(2026-06-20,commit 於 docker/nginx.stage.conf):
- default_server 的 catch-all
return 444從 server-level 收進location / { return 444; }—— 讓 exact-matchlocation = /healthz依 nginx location 優先序勝出。docker healthcheck 現回 200、containerhealthy。 - 公開 server block 新增
location = /healthz/deep→proxy_pass http://visiona_api/healthz(深層 readiness,反映 PG+Redis)。 - 保留淺層
/healthz(return 200)給 docker healthcheck / liveness。
部署方式:因 nginx config 在 image build 時 COPY 進去(非 bind mount),改 source 後以 deploy-stage-v2.sh rebuild + recreate 持久化(nginx -s reload 只是 runtime 暫補、container 重建即失效)。重部署後驗證 baked config 與 source 一致。
9. healthz 接 load balancer(最終建議 — 已實作)
採 選項 A(liveness 淺 / readiness 深 分離,對齊 K8s 慣例)。已實作於 nginx:
# 公開 server block(server_name stage-9527.innovedus.com)
location = /healthz/deep {
access_log off;
proxy_pass http://visiona_api/healthz; # api-server :3721,ping PG+Redis、down 回 503
proxy_http_version 1.1;
proxy_set_header Host $host;
proxy_set_header Connection "";
proxy_connect_timeout 2s; proxy_read_timeout 2s; proxy_send_timeout 2s; # 探針不該 hang
proxy_no_cache 1; proxy_cache_bypass 1;
add_header Cache-Control "no-store" always;
}
LB 設定建議:
- health check path:
/healthz/deep← LB readiness 打這條 - 期望狀態碼:200;503 → 拉出輪替(DB/Redis 掛時自動踢除實例)
- interval 10–15s、timeout 5s、unhealthy threshold 連續 3 次(對齊 backend health.go 的 2s ping 逾時 + nginx 2s proxy 逾時)
- 淺層
/healthz保留給 docker healthcheck / liveness(DB 抖動時不會誤把實例標死)
實測(2026-06-20,log:.autoflow/07-delivery/logs/healthz-deep-verify-*.log):
PG+Redis 健康:
公開 /healthz → 200 "ok"
公開 /healthz/deep → 200 {"checks":{"postgres":"ok","redis":"ok"},"status":"ok"}
docker healthcheck → healthy(streak=0;修復前 unhealthy streak=96)
停 visiona-redis(503 路徑):
公開 /healthz/deep → 503 {"checks":{"postgres":"ok","redis":"down"},"status":"unavailable"} ← LB 會踢
公開 /healthz → 200(淺層不受影響,liveness 不誤殺)
公開 /(frontend) → 200
起 visiona-redis(恢復):
公開 /healthz/deep → 200(自動恢復)
10. 後續硬化建議(非本次範圍)
§8 healthz 接 LB:依 §9 選項 A 加✅ 已完成(2026-06-20),見 §8/§9。/healthz/deep。- §3.1 network 顯式化:在
docker-compose.stage.yml顯式宣告 external networkvisiona-stage_default,去除「靠 project name 巧合對上」的隱性依賴。 - migration 與多副本:目前單副本啟動跑 auto-migrate 沒問題;未來多副本需改
VISIONA_DB_AUTO_MIGRATE=false+ 獨立cmd/migrate步驟,避免多實例同時 migrate。
11. Rollback
export DOCKER_HOST=tcp://192.168.0.130:2375
# 回到 DB 接入前的 Phase 0.9 image
docker tag visiona:stage-rollback-pre-db visiona:stage
bash scripts/deploy-stage-v2.sh --skip-build
(注意:rollback 到 DB 接入前版本後,repository 回 in-memory;PG/Redis 資料保留但不被讀。)
12. 煙測 evidence log 清單
| log | 內容 |
|---|---|
.autoflow/07-delivery/logs/block6-deploy-202606201848.log |
build + compose up(exit 0) |
.autoflow/07-delivery/logs/block6-container-startup-202606201848.log |
啟動 log(DB/Redis/migration initialized) |
.autoflow/07-delivery/logs/block6-healthz-503-202606201848.log |
503 fail-fast 實測(停/起 redis) |
.autoflow/07-delivery/logs/block6-persistence-202606201848.log |
重啟資料持久化 + migration 冪等 |