visionA/local-tool/server/scripts/kneron_bridge.py
jim800121chen f62141a925 fix(bridge): Windows 上中文標籤與 log 亂碼
使用者截圖標籤顯示「撣�」而非「布」。實證:布 的 UTF-8 位元組
e5 b8 83 用 cp950 解碼正好得到「撣�」。

根因是 bridge 的 stdio 綁在系統 ANSI code page(繁中 Windows = cp950),
而非 UTF-8。兩個方向的表現不同:

- Go → Python(stdin):Go 的 json.Marshal 不 escape 非 ASCII,送出的是
  原始 UTF-8,Python 卻用 cp950 解碼 → 標籤壞掉。這是實際的損壞路徑。
- Python → Go(stdout):json.dumps 預設 ensure_ascii=True 會轉成 \uXXXX,
  所以碰巧沒事 —— 但那是巧合不是設計。

同一根因也造成 log 的「(source: declared) �X SDK did not report」,那個
�X 是程式碼裡的 em dash 編碼失敗。

修正兩層,各自不可省:

1. kneron_bridge.py 在 module import 時強制 stdio 為 UTF-8(早於任何 I/O,
   import 期間的 traceback 也涵蓋),並在 os.fdopen 明確指定 encoding
   —— 那是 JSON-RPC 回應通道,reconfigure() 碰不到它,目前只靠
   ensure_ascii 巧合存活。errors="replace" 是刻意的:bridge 崩潰會讓裝置
   離線,比一個壞字元嚴重得多。

2. kl720_driver.go 加 PYTHONUTF8=1。實測 PYTHONIOENCODING 只影響 sys.stdin、
   PYTHONUTF8 才管得到 os.fdopen 與 locale.getencoding(),單一層都不夠。

Go 端 stdout scanner 不需改(bufio.Scanner 是 byte-oriented,不做轉碼)。

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-23 02:00:16 +08:00

3062 lines
128 KiB
Python
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

#!/usr/bin/env python3
"""Kneron Bridge - JSON-RPC over stdin/stdout
This script acts as a bridge between the Go backend and the Kneron PLUS
Python SDK. It reads JSON commands from stdin and writes JSON responses
to stdout.
Supports:
- KL520 (USB Boot mode - firmware must be loaded each session)
- KL720 (flash-based - firmware pre-installed, models freely reloadable)
"""
import sys
import json
import base64
import time
import os
import io
import numpy as np
def _force_utf8_stdio():
"""把 stdin / stdout / stderr 一律綁成 UTF-8不理會系統預設編碼。
為什麼需要這支Windows 中文亂碼根因):
Windows 的 Python 把 stdio 綁到「系統 ANSI code page」而非 UTF-8。繁體
中文 Windows 的 ANSI code page 是 cp950。而本 bridge 的 JSON-RPC 兩個
方向並不對稱:
Go → Python (stdin)Go 的 `encoding/json` **不** escape 非 ASCII
`{"labels":[""]}` 在 wire 上就是 raw UTF-8 位元組 e5 b8 83。
以 cp950 解碼會得到 `撣` + U+FFFD使用者截圖的亂碼
嚴格模式下則直接 UnicodeDecodeError 讓整個 bridge 掛掉。
Python → Go (stdout)`json.dumps` 預設 ensure_ascii=True中文被
escape 成 \\uXXXX 純 ASCII所以這條路「目前」剛好沒壞。但那是
隱性依賴 —— 任何人加上 ensure_ascii=False 就會壞。這裡一併綁定,
把「stdout 是 UTF-8」變成顯式契約而非巧合。
stderr`_log()` 的中文訊息與 em dashU+2014等字元現在就會壞
(使用者先前看到的 `<60>X SDK did not report` 即為此)。
為什麼用 errors="replace" 而非預設的 "strict"
這是診斷用的 log 通道與協定通道。遇到極端的無效位元組時,我們寧可看到
一個 U+FFFD 也不要讓整個 bridge 因為一行 log 而崩潰 —— bridge 掛掉
會讓裝置直接失聯比一個壞字元嚴重得多。stdin 側同理Go 端送出的一定
是合法 UTF-8replace 只是最後一道防線。
為什麼在 module import 時就呼叫(而不是在 main() 裡):
必須早於任何 I/O。main() 會 os.dup() stdout、import 期間的例外也會經由
stderr 輸出,兩者都得在編碼已經正確之後才發生。
Python 3.7+ 才有 TextIOWrapper.reconfigure()。專案用的是
python-build-standalone 3.12,安全;仍以 hasattr 做防禦,
在極舊環境下退化成 no-op 而不是 crash。
"""
for name in ("stdin", "stdout", "stderr"):
stream = getattr(sys, name, None)
if stream is None:
continue # pythonw / 被重導向到 None 的情境
try:
if hasattr(stream, "reconfigure"):
stream.reconfigure(encoding="utf-8", errors="replace")
except (AttributeError, ValueError, OSError):
# 已被換成非 TextIOWrapper例如測試的 StringIO或已關閉。
# 這不是致命錯誤,維持原編碼繼續跑。
pass
_force_utf8_stdio()
def _preload_kneron_dylibs_macos():
"""macOS 專用:用絕對路徑預先 dlopen wheel 內的 libusb + libkplus。
背景:
- KneronPLUS wheel 把 libusb-1.0.0.dylib + libkplus.dylib 放在 kp/lib/。
- macOS dyld 在載入 libkplus 時會去找它的相依 libusb-1.0.0.dylib。
預設搜尋路徑(/usr/local/lib、/usr/lib在 bundled Python 環境下通常
找不到(我們沒有 brew libusb於是 `import kp` 就拋 OSError →
HAS_KP=False → scan 回空陣列。
- macOS hardened runtime 會剝掉 DYLD_LIBRARY_PATH 等環境變數,所以
改從 Go 端注入 env 也不保險;最穩的做法是在 Python 這端用 ctypes
以絕對路徑先載入,後續 `import kp` 時 dyld 會重用已載入的映像。
Windows / Linux 不走這支 — 各自機制已在 Go 端處理Windows 靠 PATH、
Linux 靠 wheel 自帶的 libusb.so.1.0.0 + LD_LIBRARY_PATH
"""
if sys.platform != "darwin":
return
try:
import ctypes
import importlib.util
spec = importlib.util.find_spec("kp")
if spec is None or not spec.submodule_search_locations:
return
kp_dir = spec.submodule_search_locations[0]
lib_dir = os.path.join(kp_dir, "lib")
# 載入順序:先 libusb再 libkpluslibkplus 相依 libusb
for name in ("libusb-1.0.0.dylib", "libkplus.dylib"):
path = os.path.join(lib_dir, name)
if os.path.isfile(path):
try:
ctypes.CDLL(path, mode=ctypes.RTLD_GLOBAL)
except OSError:
pass
except Exception:
pass
_preload_kneron_dylibs_macos()
try:
import kp
HAS_KP = True
except (ImportError, AttributeError, Exception):
HAS_KP = False
try:
import usb.core
HAS_PYUSB = True
except ImportError:
HAS_PYUSB = False
try:
import cv2
HAS_CV2 = True
except ImportError:
HAS_CV2 = False
# ── Global state ──────────────────────────────────────────────────────
_device_group = None
def _clear_device_group():
"""Safely disconnect and clear the global _device_group.
KneronPLUS SDK's DeviceGroup.__del__ calls kp_disconnect_devices on the
native handle, but if the handle is already invalid (failed connect / stale
state) it causes 'OSError: access violation'. By explicitly disconnecting
before setting None, __del__ becomes a no-op on an already-disconnected
handle. All errors are silenced — this is best-effort cleanup.
"""
global _device_group
if _device_group is not None:
try:
kp.core.disconnect_devices(_device_group)
except Exception:
pass
_device_group = None
_model_id = None
_model_nef = None
# _model_nef_path 是當前已載入 model 的 .nef 絕對路徑。
# 需要保留是因為 set_inference_options 在推論期切換解析方式時要重跑
# _detect_model_type —— 該函式用檔名推導 input size沒有原始路徑就只能
# 退回預設 224等於切換解析方式的副作用是偷改 input size。
_model_nef_path = ""
_model_input_size = 224 # updated on model load
# _model_input_width / _model_input_height 是「模型真正的輸入尺寸」,逐軸保存。
#
# 為什麼要逐軸而不是沿用單一 _model_input_size模型的輸入不保證是正方形
# (常見的 w256h192 之類)。舊版全靠檔名的 wNNNhNNN 猜、而且 _parse_size_from_name
# 只取 width 丟掉 height等於「非正方形模型一律被當成正方形」——尺寸錯了
# NPU 不會報錯,只會安靜地給出錯的推論結果。
#
# _model_input_size 保留為「短邊」派生值,維持既有 detection post-process 與
# 對外 JSON 欄位的相容性(見 _set_model_input_size
_model_input_width = 224
_model_input_height = 224
# _model_input_size_source 記錄這次的尺寸「從哪來」,只作 log / 診斷用。
# 值域:見 INPUT_SIZE_SOURCE_*。
_model_input_size_source = "default"
_model_type = "tiny_yolov3" # updated on model load based on model_id / nef name
_firmware_loaded = False
_device_chip = "KL520" # updated on connect from product_id / device_type
# ── Model metadata injected by the Go layer on load_model ────────────
# 這兩個由 Go 端在 load_model payload 帶進來M2。在此之前 bridge 只能靠
# model id / 檔名猜 model type、對「檔名沒有關鍵字」的自訂模型必定猜錯
# (落到 tiny_yolov3 → 走 YOLO 解析 → 回空結果)。外部指定一律優先於猜測。
#
# _task_type_override: "classification" | "object_detection" | None
# _model_labels: list[str]index 對應 class indexNone 表示未注入
_task_type_override = None
_model_labels = None
# _model_declared_input_size 是 models.json / metadata.json 宣告的 (width, height)。
# 保留是為了 set_inference_options —— 推論期切解析方式會重跑 _detect_model_type
# 沒有這份值就只能退回檔名猜測,等於「切解析方式」偷偷改掉了 input size。
_model_declared_input_size = None
# 前端 / models.json / Go 端統一使用的 task type 值域。
# 注意:舊版 bridge 回傳的是 "detection",與 models.json 的 "object_detection"
# 不一致(見 plan §7 R-4此處統一為 object_detection。
TASK_TYPE_CLASSIFICATION = "classification"
TASK_TYPE_OBJECT_DETECTION = "object_detection"
# classification top-K 預設值(可由 load_model / inference 參數覆寫)
DEFAULT_CLASSIFICATION_TOP_K = 5
# 判定「模型輸出是否已是機率分佈」的容差(見 _looks_like_probabilities
#
# 為何是 1e-4 而不是更寬鬆的 0.01:真正經過 softmax 的 float32 輸出,其總和
# 與 1.0 的偏差來自浮點捨入、實測C=2..1000、常態 logits最壞約 4e-7
# 比 1e-4 還小兩個數量級 —— 容差開到 0.01 只會放進誤判、不會救回任何真的
# 機率向量。實測非負 logits 被誤判的機率C=3 時 0.4%0.01)→ 0.003%1e-4
PROB_SUM_TOLERANCE = 1e-4
# COCO 80-class labels
COCO_CLASSES = [
"person", "bicycle", "car", "motorcycle", "airplane", "bus", "train", "truck",
"boat", "traffic light", "fire hydrant", "stop sign", "parking meter", "bench",
"bird", "cat", "dog", "horse", "sheep", "cow", "elephant", "bear", "zebra",
"giraffe", "backpack", "umbrella", "handbag", "tie", "suitcase", "frisbee",
"skis", "snowboard", "sports ball", "kite", "baseball bat", "baseball glove",
"skateboard", "surfboard", "tennis racket", "bottle", "wine glass", "cup",
"fork", "knife", "spoon", "bowl", "banana", "apple", "sandwich", "orange",
"broccoli", "carrot", "hot dog", "pizza", "donut", "cake", "chair", "couch",
"potted plant", "bed", "dining table", "toilet", "tv", "laptop", "mouse",
"remote", "keyboard", "cell phone", "microwave", "oven", "toaster", "sink",
"refrigerator", "book", "clock", "vase", "scissors", "teddy bear", "hair drier",
"toothbrush"
]
# Anchor boxes per model type (each list entry = one output head)
ANCHORS_TINY_YOLOV3 = [
[(81, 82), (135, 169), (344, 319)], # 7×7 head (large objects)
[(10, 14), (23, 27), (37, 58)], # 14×14 head (small objects)
]
# YOLOv5s anchors (Kneron model 20005, no-upsample variant for KL520)
ANCHORS_YOLOV5S = [
[(116, 90), (156, 198), (373, 326)], # P5/32 (large)
[(30, 61), (62, 45), (59, 119)], # P4/16 (medium)
[(10, 13), (16, 30), (33, 23)], # P3/8 (small)
]
CONF_THRESHOLD = 0.25
NMS_IOU_THRESHOLD = 0.45
# Known Kneron model IDs → (model_type, input_size)
KNOWN_MODELS = {
# Tiny YOLO v3 (default KL520 model)
0: ("tiny_yolov3", 224),
# ResNet18 classification (model 20001)
20001: ("resnet18", 224),
# FCOS DarkNet53s detection (model 20004)
20004: ("fcos", 512),
# YOLOv5s no-upsample (model 20005)
20005: ("yolov5s", 640),
}
def _log(msg):
"""Write log messages to stderr (stdout is reserved for JSON-RPC)."""
print(f"[kneron_bridge] {msg}", file=sys.stderr, flush=True)
def _resolve_firmware_paths(chip="KL520"):
"""Resolve firmware paths relative to this script's directory.
Returns (scpu_path, ncpu_path) tuple for backward compat with existing
handle_connect() callers. Use _resolve_firmware_paths_full(chip) to get
loader path additionally (only KL520 has fw_loader.bin in A 階段).
"""
base = os.path.dirname(os.path.abspath(__file__))
fw_dir = os.path.join(base, "firmware", chip)
scpu = os.path.join(fw_dir, "fw_scpu.bin")
ncpu = os.path.join(fw_dir, "fw_ncpu.bin")
if os.path.exists(scpu) and os.path.exists(ncpu):
return scpu, ncpu
# Fallback: check KNERON_FW_DIR env var
fw_dir = os.environ.get("KNERON_FW_DIR", "")
if fw_dir:
scpu = os.path.join(fw_dir, "fw_scpu.bin")
ncpu = os.path.join(fw_dir, "fw_ncpu.bin")
if os.path.exists(scpu) and os.path.exists(ncpu):
return scpu, ncpu
return None, None
_FW_ALLOWED_CHIPS = ("KL520", "KL720") # A 階段範圍、Reviewer m1 雙重防護用
def _resolve_firmware_paths_full(chip="KL520"):
"""Resolve scpu / ncpu / loader paths.
A 階段:只有 KL520 有 fw_loader.bin用於 KDP1 legacy → KDP2 升級的 SDK
loader stage。KL720 不需要 loader不走 SDK loader path、直接 ctypes
呼叫 kp_update_kdp_firmware_from_files 也不需要 loader 檔)。
Reviewer m1對 chip 參數做雙重 allow-list 防護。chip 來自 JSON-RPC stdin、
雖然 caller (handle_firmware_upgrade) 已 enforce allow-list、但這裡再過一道
避免未來 caller 拓寬時破防。額外拒絕含 path separator / 父目錄 / 絕對路徑
的非法輸入、確保 os.path.join 絕不 traverse。
Returns:
dict: {"scpu": <path>, "ncpu": <path>, "loader": <path or None>,
"version": <str or None>}
若 scpu/ncpu 任一缺檔、scpu/ncpu 為 None。
"""
# 雙重 allow-list 防護caller 已過一次、這裡再過一次防 path traversal
if not isinstance(chip, str) or chip not in _FW_ALLOWED_CHIPS:
return {"scpu": None, "ncpu": None, "loader": None, "version": None}
# 額外字元防護(即使 _FW_ALLOWED_CHIPS 拓寬到不安全字串也擋)
if "/" in chip or "\\" in chip or ".." in chip or os.path.isabs(chip):
return {"scpu": None, "ncpu": None, "loader": None, "version": None}
base = os.path.dirname(os.path.abspath(__file__))
fw_dir = os.path.join(base, "firmware", chip)
scpu = os.path.join(fw_dir, "fw_scpu.bin")
ncpu = os.path.join(fw_dir, "fw_ncpu.bin")
loader = os.path.join(fw_dir, "fw_loader.bin")
version_file = os.path.join(fw_dir, "VERSION")
result = {"scpu": None, "ncpu": None, "loader": None, "version": None}
if os.path.exists(scpu) and os.path.exists(ncpu):
result["scpu"] = scpu
result["ncpu"] = ncpu
if os.path.exists(loader):
result["loader"] = loader
if os.path.exists(version_file):
try:
with open(version_file, "r", encoding="utf-8") as f:
result["version"] = f.read().strip()
except Exception:
pass
# Fallback: KNERON_FW_DIR env var
if result["scpu"] is None or result["ncpu"] is None:
env_dir = os.environ.get("KNERON_FW_DIR", "")
if env_dir:
scpu2 = os.path.join(env_dir, "fw_scpu.bin")
ncpu2 = os.path.join(env_dir, "fw_ncpu.bin")
if os.path.exists(scpu2) and os.path.exists(ncpu2):
result["scpu"] = scpu2
result["ncpu"] = ncpu2
loader2 = os.path.join(env_dir, "fw_loader.bin")
if os.path.exists(loader2):
result["loader"] = loader2
return result
# ── Input size 來源標記 ───────────────────────────────────────────────
#
# 依可信度排序(愈前面愈可信)。實機驗收時 log 會印出用的是哪一個 ——
# 這是唯一能一眼看出「尺寸是怎麼來的」的線索,尺寸錯掉時 NPU 不一定報錯,
# 可能只是安靜地給出錯的推論結果。
#
# ⚠️ declared 為什麼**不是**第二可信2026-07 Windows regression 的根因):
# 上傳表單的 inputSize 欄位長期沒有實際作用,使用者是「隨手填」的
# (實際案例:填 640x640模型其實是 224x224。而 KneronPLUS 3.1.2 不再
# 提供 2.0.0 的 shape_onnx 屬性(見 _input_size_from_nefSDK 這層在
# Windows 直接落空 —— 於是垃圾 declared 值成為實際採用值,送進 NPU 得到
# KP_ERROR_INVALID_PARAM_12。相對地檔名的 wNNNhNNN 是模型編譯工具鏈
# 產生的,沒有人為亂填的空間,**明確解析出來時**比 declared 可信。
INPUT_SIZE_SOURCE_SDK = "SDK" # 模型自己宣告的,唯一完全可靠
INPUT_SIZE_SOURCE_FILENAME = "filename" # 檔名 wNNNhNNN 明確解析出的,工具鏈產生
INPUT_SIZE_SOURCE_DECLARED = "declared" # models.json / metadata.json人手填的
INPUT_SIZE_SOURCE_KNOWN_ID = "known-model-id" # 內建 KNOWN_MODELS 表的官方固定值
INPUT_SIZE_SOURCE_DEFAULT = "default" # 什麼都沒有,寫死的預設值
# 可信度排序(僅供 log / 診斷描述用;實際流程由 _resolve_input_size 決定)。
INPUT_SIZE_SOURCE_RANK = {
INPUT_SIZE_SOURCE_SDK: 0,
INPUT_SIZE_SOURCE_FILENAME: 1,
INPUT_SIZE_SOURCE_DECLARED: 2,
INPUT_SIZE_SOURCE_KNOWN_ID: 3,
INPUT_SIZE_SOURCE_DEFAULT: 4,
}
# 合理的 input 邊長範圍。用來擋掉明顯不合理的宣告值 / 解析結果
# (例如 shape 判讀錯把 batch=1 或 channel=3 當成邊長)。
MIN_REASONABLE_INPUT_DIM = 8
MAX_REASONABLE_INPUT_DIM = 8192
def _set_model_input_size(width, height, source):
"""Commit the resolved model input size to the globals.
_model_input_size單一純量沿用「短邊」而非寬邊它在 detection
post-process 裡的用途是 `min_dim`(確保送進 NPU 的影像不小於模型輸入)
以及 anchor / bbox 正規化的除數。取短邊可保證兩軸都達標;取長邊會讓
短的那一軸不足。正方形模型(絕大多數)兩者相同、行為完全不變。
"""
global _model_input_width, _model_input_height
global _model_input_size, _model_input_size_source
_model_input_width = int(width)
_model_input_height = int(height)
_model_input_size = min(_model_input_width, _model_input_height)
_model_input_size_source = source
def _is_reasonable_input_dim(value):
"""True if value looks like a plausible model input edge length."""
return (isinstance(value, int) and not isinstance(value, bool)
and MIN_REASONABLE_INPUT_DIM <= value <= MAX_REASONABLE_INPUT_DIM)
def _input_size_from_nef(nef, model_id=None):
"""Read the **real** input width/height from a loaded NEF descriptor.
這是唯一可靠的來源:模型自己知道它要多大的輸入。檔名可能沒有尺寸資訊、
使用者填的 inputSize 可能是隨手填的,兩者都會靜默地把圖縮到錯的尺寸。
來源是 KneronPLUS 公開 APIkp.ModelNefDescriptor
nef.models[i].input_nodes[j].shape_onnx / .shape_npu → List[int]
shape 判讀規則(優先 shape_onnx因為那是原始 ONNX 的語意):
- 4 維 NCHW (1, 3, H, W) → 取後兩維。這是 KneronPLUS 對影像模型的常態。
- 4 維 NHWC (1, H, W, 3) → 最後一維是通道數時改取中間兩維。
兩者用「哪一個軸看起來像通道數1/3/4」區分。
無法可靠判讀時回 None呼叫端會 fallback**不猜** —— 猜錯等於用錯尺寸
推論,而那正是這個函式要解決的問題。
Args:
nef: kp.ModelNefDescriptorload_model_from_file 的回傳值)
model_id: 指定要取哪個 model 的 inputNone 表示取第一個。
一個 .nef 可以包多個 model。
Returns:
(width, height) tuple或 None。
"""
try:
models = list(nef.models)
except Exception as e:
_log(f"input size from SDK: cannot read nef.models ({type(e).__name__}: {e})")
return None
if not models:
_log("input size from SDK: nef contains no models")
return None
model = None
if model_id is not None:
for candidate in models:
try:
if candidate.id == model_id:
model = candidate
break
except Exception:
continue
if model is None:
model = models[0]
try:
input_nodes = list(model.input_nodes)
except Exception as e:
_log(f"input size from SDK: cannot read input_nodes "
f"({type(e).__name__}: {e})")
return None
if not input_nodes:
_log("input size from SDK: model has no input nodes")
return None
node = input_nodes[0]
candidates = _tensor_shape_candidates(node)
if not candidates:
_log("input size from SDK: input node exposes no readable shape "
"(unsupported KneronPLUS layout?)")
return None
for label, shape in candidates:
size = _input_size_from_shape(shape)
if size is not None:
_log(f"input size from SDK: {label}={shape} -> {size[0]}x{size[1]}")
return size
if shape:
_log(f"input size from SDK: {label}={shape} not interpretable as "
f"an image input shape")
return None
def _tensor_shape_candidates(node):
"""Collect candidate shapes from a TensorDescriptor, newest-API-aware.
KneronPLUS 在 3.x 改了 TensorDescriptor 的結構,**欄位名沒改、但搬了一層**
2.0.0macOS 目前用的版本)
TensorDescriptor.shape_onnx : List[int]
TensorDescriptor.shape_npu : List[int]
3.1.2Windows 目前用的版本)
TensorDescriptor.tensor_shape_info.version : ModelTensorShapeInformationVersion
TensorDescriptor.tensor_shape_info.data : TensorShapeInfoV1 | TensorShapeInfoV2
V1 → .shape_onnx / .shape_npu / .axis_permutation_onnx_to_npu
V2 → .shapedocstring 明寫「ONNX shape of the tensor」
/ .stride_onnx / .stride_npu
3.1.2 的 TensorDescriptor **沒有** shape_onnx 屬性,舊寫法的
`getattr(node, "shape_onnx")` 會拋 AttributeError、被 except 靜默吃掉,
整個 SDK 層等於永遠落空 —— 這就是 Windows「SDK did not report an input
shape」的真正原因不是模型沒帶 shape
回傳 [(label, shape), ...]依可信度排序。label 只作 log 用。
ONNX 語意一律優先於 NPU layout後者可能有對齊 padding拿來當輸入尺寸不準。
"""
candidates = []
def take(label, value):
if value is None:
return
try:
shape = [int(dim) for dim in value]
except (TypeError, ValueError):
return
if shape:
candidates.append((label, shape))
# ── 3.x巢狀 tensor_shape_info先試因為新版沒有平鋪欄位──────
info = getattr(node, "tensor_shape_info", None)
if info is not None:
data = getattr(info, "data", None)
if data is not None:
# V1 與 V2 的欄位名不重疊,直接各自試、不需判斷 version enum
# (少一個跨版本 enum 相依enum 名稱改了也不會整個壞掉)。
take("tensor_shape_info.data.shape_onnx",
getattr(data, "shape_onnx", None))
take("tensor_shape_info.data.shape",
getattr(data, "shape", None))
take("tensor_shape_info.data.shape_npu",
getattr(data, "shape_npu", None))
# ── 2.x平鋪在 TensorDescriptor 上 ────────────────────────────────
take("shape_onnx", getattr(node, "shape_onnx", None))
take("shape_npu", getattr(node, "shape_npu", None))
return candidates
def _input_size_from_shape(shape):
"""Interpret a 4-D tensor shape as (width, height), or None.
只處理 4 維 —— 影像輸入必為 4 維batch + 3 個空間/通道軸)。其他維度
(例如純向量輸入的模型)不是本函式能判讀的,回 None 讓呼叫端 fallback。
"""
if not shape or len(shape) != 4:
return None
_, a, b, c = shape
channel_like = (1, 3, 4)
# NCHW: (N, C, H, W) —— KneronPLUS 對影像模型的常態
if a in channel_like and _is_reasonable_input_dim(b) and _is_reasonable_input_dim(c):
return (c, b) # (width, height)
# NHWC: (N, H, W, C)
if c in channel_like and _is_reasonable_input_dim(a) and _is_reasonable_input_dim(b):
return (b, a) # (width, height)
return None
def _detect_model_type(model_id, nef_path, task_type=None,
nef=None, declared_input_size=None):
"""Detect model type and input size from model ID or .nef filename.
Args:
model_id: model id read from the loaded .nef.
nef_path: path of the .nef file (used for filename heuristics).
task_type: 外部Go 層)指定的 task type。有指定時**一律優先**、
不做任何猜測。這是「自訂 classification 模型跑不出結果」的根因修正:
使用者上傳的模型檔名(如 model.nef / 1784536643_models_520.nef
不含任何關鍵字,舊邏輯會落到 else 被誤判成 tiny_yolov3。
nef: 已載入的 kp.ModelNefDescriptor。有帶就從它取**真正的** input size。
declared_input_size: 外部宣告的 (width, height),來自 models.json /
metadata.json 的 inputSize。只在 SDK 取不到時才用 —— 那是人填的,
使用者可能隨手填(實際案例:填 640x640 但模型是別的尺寸)。
task type 與 input size 是兩件獨立的事task type 決定走哪條 post-process、
input size 決定圖片怎麼縮。外部指定 task type 時仍會照常解析 input size。
"""
global _model_type
# input size 三層來源,與 model type 的判斷完全獨立。
_resolve_input_size(model_id, nef_path, nef=nef,
declared_input_size=declared_input_size)
normalized_task = _normalize_task_type(task_type)
if normalized_task == TASK_TYPE_CLASSIFICATION:
# 先用既有邏輯決定 backbone順帶會寫 input size再強制覆寫 type。
_detect_model_type_by_heuristics(model_id, nef_path)
_model_type = "classification"
_log(f"Model type set by caller: classification "
f"({_describe_input_size()})")
return
if normalized_task == TASK_TYPE_OBJECT_DETECTION:
# detection 有多種 backboneyolo / fcos / ssd…外部只告訴我們
# 「這是 detection」、無法指出是哪一種所以仍需 heuristics 決定
# 用哪個 post-process。但至少可以確定它不是 classification。
_detect_model_type_by_heuristics(model_id, nef_path)
if _model_type == "classification" or _model_type == "resnet18":
_log("Caller declared object_detection but heuristics said "
"classification; falling back to tiny_yolov3 post-process")
_model_type = "tiny_yolov3"
return
_detect_model_type_by_heuristics(model_id, nef_path)
def _describe_input_size():
"""One-line description of the current input size **and where it came from**.
實機驗收要能一眼看出尺寸是怎麼來的 —— 尺寸錯掉時 NPU 不會報錯,
只會安靜地給出錯的推論結果,這行 log 是第一個要看的線索。
"""
desc = (f"input_size={_model_input_width}x{_model_input_height} "
f"(source: {_model_input_size_source}")
if _model_input_size_source == INPUT_SIZE_SOURCE_DECLARED:
desc += ", hand-entered — verify if inference fails"
elif _model_input_size_source == INPUT_SIZE_SOURCE_KNOWN_ID:
desc += ", UNRELIABLE — guessed from a built-in model-id table"
elif _model_input_size_source == INPUT_SIZE_SOURCE_DEFAULT:
desc += ", UNRELIABLE — no size information available anywhere"
return desc + ")"
def _resolve_input_size(model_id, nef_path, nef=None, declared_input_size=None):
"""Resolve the model input size from the most trustworthy source available.
優先序(愈前面愈可信):
1. SDK — 模型自己宣告的 input tensor shape。唯一完全可靠那是
模型編譯進 .nef 的事實,不是任何人的說法。
2. filename — 檔名 wNNNhNNN **明確解析出來**的值。由模型編譯工具鏈
產生kl520_20004_fcos-drk53s_w512h512.nef 這種格式),
沒有人為亂填的空間。注意:只有真的解析到才算這一層,
解析不到不會退化成 backbone 預設值假裝有來源。
3. declared — models.json / metadata.json 的 inputSize**使用者手填**。
4. known id — 內建 KNOWN_MODELS 表Kneron 官方模型的固定尺寸。
5. default — 都沒有backbone 慣用值(多為 224
⚠️ 為什麼 declared 排在 filename **之後**2026-07 Windows regression
上傳表單的 inputSize 欄位過去長期沒有實際作用,既有資料裡的值大多是
使用者隨手填的垃圾(實際案例:填 640x640、模型其實 224x224。把它排在
檔名之前,等於讓最不可信的來源壓過工具鏈產生的事實。
但也不能無腦把 declared 降到最後:使用者若是刻意填對的,它仍然比
「檔名沒有尺寸資訊時的寫死預設值」可信 —— 所以它排在 known id / default
之前,只讓位給 SDK 與檔名明確解析。
known id 排在 declared 之後而非之前KNOWN_MODELS 只涵蓋 Kneron 官方
幾個內建模型,且那些檔名本來就帶 wNNNhNNN會在第 2 層就命中);真正
落到這層的情境是「id 撞上官方編號但檔案已被換過」,此時使用者的宣告
反而比我們的內建表貼近現實。
本函式負責 13 層4 與 5 由 _detect_model_type_by_heuristics 收尾
(它同時要決定 model type順道寫入尺寸
"""
global _model_input_size_source
# 先降級成 default —— 否則上一個模型留下的較可信標記會讓
# _detect_model_type_by_heuristics 誤以為「已有更可信來源」而不敢寫,
# 新模型就會沿用舊模型的尺寸(靜默、且只在換模型時才出現)。
_model_input_size_source = INPUT_SIZE_SOURCE_DEFAULT
declared = _normalize_declared_input_size(declared_input_size)
# ── 1. SDK模型自己說的 ────────────────────────────────────────
if nef is not None:
sdk_size = _input_size_from_nef(nef, model_id=model_id)
if sdk_size is not None:
_set_model_input_size(sdk_size[0], sdk_size[1],
INPUT_SIZE_SOURCE_SDK)
_log(f"Model input size resolved: {_describe_input_size()}")
_warn_if_declared_disagrees(declared)
return
# ── 2. filename工具鏈產生的 wNNNhNNN**明確解析到**才算 ────────
basename = os.path.basename(nef_path).lower() if nef_path else ""
from_name = _size_from_name_or_none(basename)
if from_name is not None:
_set_model_input_size(from_name[0], from_name[1],
INPUT_SIZE_SOURCE_FILENAME)
_log(f"Model input size resolved: {_describe_input_size()} "
f"— SDK did not report an input shape, parsed from the filename")
_warn_if_declared_disagrees(declared)
return
# ── 3. declared外部models.json / metadata.json宣告的 ───────
if declared is not None:
_set_model_input_size(declared[0], declared[1],
INPUT_SIZE_SOURCE_DECLARED)
_log(f"Model input size resolved: {_describe_input_size()} "
f"— no SDK shape and no size in the filename, falling back to "
f"the user-declared value (this value is hand-entered and is a "
f"likely cause if inference fails with KP_ERROR_INVALID_PARAM)")
return
# ── 4 / 5. 交給 known id / 檔名 heuristics它自己會寫 source────
def _warn_if_declared_disagrees(declared):
"""Log when the user-declared size contradicts the source we actually used.
這行 log 是使用者「為什麼我填的尺寸沒有生效」的唯一線索。不改行為 ——
declared 本來就該讓位給更可信的來源,但要讓人看得見它被讓位了。
"""
if declared is None:
return
if (declared[0], declared[1]) == (_model_input_width, _model_input_height):
return
_log(f"Note: declared input size {declared[0]}x{declared[1]} "
f"(hand-entered) disagrees with the resolved "
f"{_model_input_width}x{_model_input_height} "
f"(source: {_model_input_size_source}); the more trustworthy source "
f"wins. Update the model's declared size if the declared one is right.")
def _normalize_declared_input_size(declared):
"""Validate an externally declared input size.
接受 (width, height) tuple/list 或 {"width": w, "height": h} dict。
任一軸不合理就整組拒絕 —— 半套採用(例如寬對高錯)比整組不用更危險,
因為看起來像是有正確來源。
"""
if declared is None:
return None
if isinstance(declared, dict):
width = declared.get("width")
height = declared.get("height")
elif isinstance(declared, (tuple, list)) and len(declared) == 2:
width, height = declared
else:
_log(f"Declared input size ignored: unsupported form {declared!r}")
return None
try:
width = int(width)
height = int(height)
except (TypeError, ValueError):
_log(f"Declared input size ignored: not integers ({declared!r})")
return None
if not _is_reasonable_input_dim(width) or not _is_reasonable_input_dim(height):
_log(f"Declared input size ignored: {width}x{height} outside the "
f"plausible range [{MIN_REASONABLE_INPUT_DIM}, "
f"{MAX_REASONABLE_INPUT_DIM}]")
return None
return (width, height)
def _normalize_task_type(task_type):
"""Normalize an externally supplied task type string.
接受 models.json / 前端使用的 object_detection、以及舊版 bridge 的
detection 別名。無法辨識(含 None / 空字串)回傳 None 代表「未指定」。
"""
if not isinstance(task_type, str):
return None
value = task_type.strip().lower()
if not value:
return None
if value == TASK_TYPE_CLASSIFICATION:
return TASK_TYPE_CLASSIFICATION
if value in (TASK_TYPE_OBJECT_DETECTION, "detection"):
return TASK_TYPE_OBJECT_DETECTION
_log(f"Unknown task_type '{task_type}' from caller, ignoring")
return None
def _detect_model_type_by_heuristics(model_id, nef_path):
"""Guess model type / input size from model ID or .nef filename.
⚠️ Input size 的部分是**最後手段**只有在更可信的來源SDK / 外部宣告)
沒有結果時才會生效。已經由 _resolve_input_size 取到 SDK 或 declared 尺寸
時,這裡只決定 model type走哪條 post-process不動尺寸 —— 否則猜出來
的值會蓋掉模型自己宣告的真值。
"""
global _model_type
# _resolve_input_size 已處理 SDK / filename / declared 三層。只有它什麼
# 都沒找到source 仍是 default這裡的猜測才有機會生效。
keep_size = _model_input_size_source != INPUT_SIZE_SOURCE_DEFAULT
def apply_size(width, height, source):
"""Write the guessed size unless a more trustworthy source already won."""
if keep_size:
return
_set_model_input_size(width, height, source)
# Check known model IDs
if model_id in KNOWN_MODELS:
_model_type, known_size = KNOWN_MODELS[model_id]
# 已知 model id 的尺寸是 Kneron 官方模型的固定值。比寫死的 backbone
# 預設值可信,但不如 SDK / 檔名 / 使用者宣告(見 _resolve_input_size
apply_size(known_size, known_size, INPUT_SIZE_SOURCE_KNOWN_ID)
_log(f"Model type detected by ID {model_id}: {_model_type} "
f"({_describe_input_size()})")
return
# Fallback: try to infer from filename
basename = os.path.basename(nef_path).lower() if nef_path else ""
# 這裡的尺寸一律是「該 backbone 的慣用值」——檔名真的帶 wNNNhNNN 時
# _resolve_input_size 第 2 層早就命中了,走到這裡代表檔名沒有尺寸資訊,
# 所以 source 是 default 而不是 filename不可讓預設值偽裝成解析結果
if "yolov5" in basename:
_model_type = "yolov5s"
apply_size(640, 640, INPUT_SIZE_SOURCE_DEFAULT)
elif "fcos" in basename:
_model_type = "fcos"
apply_size(512, 512, INPUT_SIZE_SOURCE_DEFAULT)
elif "ssd" in basename:
_model_type = "ssd"
apply_size(320, 320, INPUT_SIZE_SOURCE_DEFAULT)
elif "resnet" in basename or "classification" in basename:
_model_type = "resnet18"
apply_size(224, 224, INPUT_SIZE_SOURCE_DEFAULT)
elif "tiny_yolo" in basename or "tinyyolo" in basename:
_model_type = "tiny_yolov3"
apply_size(224, 224, INPUT_SIZE_SOURCE_DEFAULT)
else:
# Default: assume YOLO-like detection.
# 對「外部指定 classification 但檔名沒有型別關鍵字」的自訂模型特別
# 重要 —— input size 與 task type 是兩件獨立的事。
_model_type = "tiny_yolov3"
apply_size(224, 224, INPUT_SIZE_SOURCE_DEFAULT)
_log(f"Model type detected by filename '{basename}': {_model_type} "
f"({_describe_input_size()})")
def _size_from_name_or_none(name):
"""Extract (width, height) from a filename like 'w640h640' / 'w256h192'.
**解析不出來時回 None**,不回任何預設值 —— 這個「有解析到 vs 沒解析到」
的區分是新優先序的前提:真的從檔名讀到尺寸時它比使用者手填的 declared
可信(工具鏈產生、沒有亂填空間);沒讀到時它什麼都不是,絕不能讓
backbone 的慣用預設值224 / 640 …)偽裝成「從檔名來的」而蓋掉 declared。
舊版 _parse_size_from_name 把兩者混在同一個回傳值裡,呼叫端無從分辨。
"""
import re
m = re.search(r'w(\d+)h(\d+)', name or "")
if not m:
return None
width = int(m.group(1))
height = int(m.group(2))
if _is_reasonable_input_dim(width) and _is_reasonable_input_dim(height):
return (width, height)
return None
def _parse_size_from_name(name, default=224):
"""Backward-compatible wrapper: returns (default, default) when unparsed.
保留給既有呼叫端 / 測試。新程式碼請用 _size_from_name_or_none —— 它能
分辨「解析成功」與「用了 default」那是決定優先序所必需的資訊。
"""
parsed = _size_from_name_or_none(name)
if parsed is not None:
return parsed
return (default, default)
# ── Post-processing ──────────────────────────────────────────────────
def _sigmoid(x):
return 1.0 / (1.0 + np.exp(-np.clip(x, -500, 500)))
def _nms(detections, iou_threshold=NMS_IOU_THRESHOLD):
"""Non-Maximum Suppression."""
detections.sort(key=lambda d: d["confidence"], reverse=True)
keep = []
for d in detections:
skip = False
for k in keep:
if d["class_id"] != k["class_id"]:
continue
x1 = max(d["bbox"]["x"], k["bbox"]["x"])
y1 = max(d["bbox"]["y"], k["bbox"]["y"])
x2 = min(d["bbox"]["x"] + d["bbox"]["width"],
k["bbox"]["x"] + k["bbox"]["width"])
y2 = min(d["bbox"]["y"] + d["bbox"]["height"],
k["bbox"]["y"] + k["bbox"]["height"])
inter = max(0, x2 - x1) * max(0, y2 - y1)
a1 = d["bbox"]["width"] * d["bbox"]["height"]
a2 = k["bbox"]["width"] * k["bbox"]["height"]
if inter / (a1 + a2 - inter + 1e-6) > iou_threshold:
skip = True
break
if not skip:
keep.append(d)
return keep
def _get_preproc_info(result):
"""Extract letterbox padding info from the inference result.
Kneron SDK applies letterbox resize (aspect-ratio-preserving + zero padding)
before inference. The hw_pre_proc_info tells us how to reverse it.
Returns (pad_left, pad_top, resize_w, resize_h, model_w, model_h) or None.
"""
try:
info = result.header.hw_pre_proc_info_list[0]
return {
"pad_left": info.pad_left if hasattr(info, 'pad_left') else 0,
"pad_top": info.pad_top if hasattr(info, 'pad_top') else 0,
"resized_w": info.resized_img_width if hasattr(info, 'resized_img_width') else 0,
"resized_h": info.resized_img_height if hasattr(info, 'resized_img_height') else 0,
"model_w": info.model_input_width if hasattr(info, 'model_input_width') else 0,
"model_h": info.model_input_height if hasattr(info, 'model_input_height') else 0,
"img_w": info.img_width if hasattr(info, 'img_width') else 0,
"img_h": info.img_height if hasattr(info, 'img_height') else 0,
}
except Exception:
return None
def _correct_bbox_for_letterbox(x, y, w, h, preproc, model_size):
"""Remove letterbox padding offset from normalized bbox coordinates.
Input (x, y, w, h) is in model-input-space normalized to 0-1.
Output is re-normalized to the original image aspect ratio (still 0-1).
For KP_PADDING_CORNER (default): image is at top-left, padding at bottom/right.
"""
if preproc is None:
return x, y, w, h
model_w = preproc["model_w"] or model_size
model_h = preproc["model_h"] or model_size
pad_left = preproc["pad_left"]
pad_top = preproc["pad_top"]
resized_w = preproc["resized_w"] or model_w
resized_h = preproc["resized_h"] or model_h
# If no padding was applied, skip correction
if pad_left == 0 and pad_top == 0 and resized_w == model_w and resized_h == model_h:
return x, y, w, h
# Convert from normalized (0-1 of model input) to pixel coords in model space
px = x * model_w
py = y * model_h
pw = w * model_w
ph = h * model_h
# Subtract padding offset
px -= pad_left
py -= pad_top
# Re-normalize to the resized (un-padded) image dimensions
nx = px / resized_w
ny = py / resized_h
nw = pw / resized_w
nh = ph / resized_h
# Clip to 0-1
nx = max(0.0, min(1.0, nx))
ny = max(0.0, min(1.0, ny))
nw = min(1.0 - nx, nw)
nh = min(1.0 - ny, nh)
return nx, ny, nw, nh
def _parse_yolo_output(result, anchors, input_size, num_classes=80, labels=None):
"""Parse YOLO (v3/v5) raw output into detection results.
Works for both Tiny YOLOv3 and YOLOv5 — the tensor layout is the same:
(num_anchors * (5 + num_classes), grid_h, grid_w)
The key differences are:
- anchor values
- input_size used for anchor normalization
- number of output heads
Bounding boxes are corrected for letterbox padding so coordinates
are relative to the original image (normalized 0-1).
"""
detections = []
entry_size = 5 + num_classes # 85 for COCO 80 classes
# Get letterbox padding info
preproc = _get_preproc_info(result)
if preproc:
_log(f"Preproc info: pad=({preproc['pad_left']},{preproc['pad_top']}), "
f"resized=({preproc['resized_w']}x{preproc['resized_h']}), "
f"model=({preproc['model_w']}x{preproc['model_h']}), "
f"img=({preproc['img_w']}x{preproc['img_h']})")
for head_idx in range(result.header.num_output_node):
output = kp.inference.generic_inference_retrieve_float_node(
node_idx=head_idx,
generic_raw_result=result,
channels_ordering=kp.ChannelOrdering.KP_CHANNEL_ORDERING_CHW
)
arr = output.ndarray[0] # (C, H, W)
channels, grid_h, grid_w = arr.shape
# Determine number of anchors for this head
num_anchors = channels // entry_size
if num_anchors < 1:
_log(f"Head {head_idx}: unexpected shape {arr.shape}, skipping")
continue
# Use the correct anchor set for this head
if head_idx < len(anchors):
head_anchors = anchors[head_idx]
else:
_log(f"Head {head_idx}: no anchors defined, skipping")
continue
for a_idx in range(min(num_anchors, len(head_anchors))):
off = a_idx * entry_size
for cy in range(grid_h):
for cx in range(grid_w):
obj_conf = _sigmoid(arr[off + 4, cy, cx])
if obj_conf < CONF_THRESHOLD:
continue
cls_scores = _sigmoid(arr[off + 5:off + entry_size, cy, cx])
cls_id = int(np.argmax(cls_scores))
cls_conf = float(cls_scores[cls_id])
conf = float(obj_conf * cls_conf)
if conf < CONF_THRESHOLD:
continue
bx = (_sigmoid(arr[off, cy, cx]) + cx) / grid_w
by = (_sigmoid(arr[off + 1, cy, cx]) + cy) / grid_h
aw, ah = head_anchors[a_idx]
bw = (np.exp(min(float(arr[off + 2, cy, cx]), 10)) * aw) / input_size
bh = (np.exp(min(float(arr[off + 3, cy, cx]), 10)) * ah) / input_size
# Convert center x,y,w,h to corner x,y,w,h (normalized to model input)
x = max(0.0, bx - bw / 2)
y = max(0.0, by - bh / 2)
w = min(1.0, bx + bw / 2) - x
h = min(1.0, by + bh / 2) - y
# Correct for letterbox padding
x, y, w, h = _correct_bbox_for_letterbox(x, y, w, h, preproc, input_size)
label = _resolve_label(cls_id, labels=labels,
fallback_labels=COCO_CLASSES)
detections.append({
"label": label,
"class_id": cls_id,
"confidence": conf,
"bbox": {"x": x, "y": y, "width": w, "height": h},
})
detections = _nms(detections)
# Remove internal class_id before returning
for d in detections:
del d["class_id"]
return detections
def _parse_ssd_output(result, input_size=320, num_classes=2, labels=None):
"""Parse SSD face detection output.
SSD typically outputs two tensors:
- locations: (num_boxes, 4) — bounding box coordinates
- confidences: (num_boxes, num_classes) — class scores
For the KL520 SSD face detection model (kl520_ssd_fd_lm.nef),
the output contains face detections with landmarks.
"""
detections = []
preproc = _get_preproc_info(result)
try:
# Retrieve all output nodes
num_outputs = result.header.num_output_node
outputs = []
for i in range(num_outputs):
output = kp.inference.generic_inference_retrieve_float_node(
node_idx=i,
generic_raw_result=result,
channels_ordering=kp.ChannelOrdering.KP_CHANNEL_ORDERING_CHW
)
outputs.append(output.ndarray[0])
if num_outputs < 2:
_log(f"SSD: expected >=2 output nodes, got {num_outputs}")
return detections
# Heuristic: the larger tensor is locations, smaller is confidences
# Or: first output = locations, second = confidences
locations = outputs[0]
confidences = outputs[1]
# Flatten if needed
if locations.ndim > 2:
locations = locations.reshape(-1, 4)
if confidences.ndim > 2:
confidences = confidences.reshape(-1, confidences.shape[-1])
num_boxes = min(locations.shape[0], confidences.shape[0])
for i in range(num_boxes):
# SSD confidence: class 0 = background, class 1 = face
if confidences.shape[-1] > 1:
conf = float(confidences[i, 1]) # face class
else:
conf = float(_sigmoid(confidences[i, 0]))
if conf < CONF_THRESHOLD:
continue
# SSD outputs are typically [x_min, y_min, x_max, y_max] normalized
x_min = float(np.clip(locations[i, 0], 0.0, 1.0))
y_min = float(np.clip(locations[i, 1], 0.0, 1.0))
x_max = float(np.clip(locations[i, 2], 0.0, 1.0))
y_max = float(np.clip(locations[i, 3], 0.0, 1.0))
w = x_max - x_min
h = y_max - y_min
if w <= 0 or h <= 0:
continue
# Correct for letterbox padding
x_min, y_min, w, h = _correct_bbox_for_letterbox(
x_min, y_min, w, h, preproc, input_size)
# SSD 人臉模型只有單一前景類別;有注入 labels 時用 labels[0]
# 沒有時維持既有的 "face"。
detections.append({
"label": _resolve_label(0, labels=labels,
fallback_labels=["face"]),
"class_id": 0,
"confidence": conf,
"bbox": {"x": x_min, "y": y_min, "width": w, "height": h},
})
detections = _nms(detections)
for d in detections:
del d["class_id"]
except Exception as e:
_log(f"SSD parse error: {e}")
return detections
def _parse_fcos_output(result, input_size=512, num_classes=80, labels=None):
"""Parse FCOS (Fully Convolutional One-Stage) detection output.
FCOS outputs per feature level:
- classification: (num_classes, H, W)
- centerness: (1, H, W)
- regression: (4, H, W) — distances from each pixel to box edges (l, t, r, b)
The outputs come in groups of 3 per feature level.
"""
detections = []
preproc = _get_preproc_info(result)
try:
num_outputs = result.header.num_output_node
outputs = []
for i in range(num_outputs):
output = kp.inference.generic_inference_retrieve_float_node(
node_idx=i,
generic_raw_result=result,
channels_ordering=kp.ChannelOrdering.KP_CHANNEL_ORDERING_CHW
)
outputs.append(output.ndarray[0])
# FCOS typically has 5 feature levels × 3 outputs = 15 output nodes
# Or fewer for simplified models. Group by 3: (cls, centerness, reg)
# If we can't determine the grouping, try a simpler approach.
strides = [8, 16, 32, 64, 128]
num_levels = num_outputs // 3
for level in range(num_levels):
cls_out = outputs[level * 3] # (num_classes, H, W)
cnt_out = outputs[level * 3 + 1] # (1, H, W)
reg_out = outputs[level * 3 + 2] # (4, H, W)
stride = strides[level] if level < len(strides) else (8 * (2 ** level))
h, w = cls_out.shape[1], cls_out.shape[2]
for cy in range(h):
for cx in range(w):
cls_scores = _sigmoid(cls_out[:, cy, cx])
cls_id = int(np.argmax(cls_scores))
cls_conf = float(cls_scores[cls_id])
centerness = float(_sigmoid(cnt_out[0, cy, cx]))
conf = cls_conf * centerness
if conf < CONF_THRESHOLD:
continue
# Regression: distances from pixel center to box edges
px = (cx + 0.5) * stride
py = (cy + 0.5) * stride
l = float(np.exp(min(reg_out[0, cy, cx], 10))) * stride
t = float(np.exp(min(reg_out[1, cy, cx], 10))) * stride
r = float(np.exp(min(reg_out[2, cy, cx], 10))) * stride
b = float(np.exp(min(reg_out[3, cy, cx], 10))) * stride
x_min = max(0.0, (px - l) / input_size)
y_min = max(0.0, (py - t) / input_size)
x_max = min(1.0, (px + r) / input_size)
y_max = min(1.0, (py + b) / input_size)
bw = x_max - x_min
bh = y_max - y_min
if bw <= 0 or bh <= 0:
continue
# Correct for letterbox padding
x_min, y_min, bw, bh = _correct_bbox_for_letterbox(
x_min, y_min, bw, bh, preproc, input_size)
label = _resolve_label(cls_id, labels=labels,
fallback_labels=COCO_CLASSES)
detections.append({
"label": label,
"class_id": cls_id,
"confidence": conf,
"bbox": {"x": x_min, "y": y_min, "width": bw, "height": bh},
})
detections = _nms(detections)
for d in detections:
del d["class_id"]
except Exception as e:
_log(f"FCOS parse error: {e}")
return detections
def _sanitize_labels(raw):
"""Normalize the labels payload coming from the Go layer.
只接受 list/tuple。非字串元素轉成 strNone 轉成空字串(稀疏佔位)。
空 list 視為「未提供」回傳 None讓下游走 class_N / COCO fallback。
"""
if raw is None:
return None
if not isinstance(raw, (list, tuple)):
_log(f"load_model: labels must be a list, got {type(raw).__name__}; ignoring")
return None
labels = []
for item in raw:
if item is None:
labels.append("")
elif isinstance(item, str):
labels.append(item)
else:
labels.append(str(item))
if not any(label.strip() for label in labels):
return None
return labels
def _reset_model_metadata():
"""Clear the injected model metadata (called on disconnect / reset).
不清會讓下一個模型沿用上一個的 labels / task_type —— 例如先載
classification 再載 detectiondetection 會拿到錯的 label 表。
_model_nef_path 一併清掉它描述「_model_id 指向的那個 .nef 在哪」,
model 都沒了還留著路徑set_inference_options 有機會拿舊路徑去推導
input size。實務上 set_inference_options 會先被 _model_id is None
擋掉但讓不變式「path 與 model_id 同生同滅」成立比依賴那道守衛安全。)
input size 三個全域width / height / source也一併回到預設理由相同
它們描述的是「_model_id 那個模型要多大的輸入」,模型沒了就不該留著 ——
留著會讓下一個模型在來源判定時看到上一個模型的 SDK 標記。
"""
global _task_type_override, _model_labels, _model_nef_path
global _model_declared_input_size
_task_type_override = None
_model_labels = None
_model_nef_path = ""
_model_declared_input_size = None
_set_model_input_size(224, 224, INPUT_SIZE_SOURCE_DEFAULT)
def _current_task_type():
"""The task type that will be reported for inference results.
以「實際會執行哪一條 post-process」為準而非外部宣告值確保回傳的
taskType 與 detections / classifications 欄位始終一致。
"""
if _model_type in ("classification", "resnet18"):
return TASK_TYPE_CLASSIFICATION
return TASK_TYPE_OBJECT_DETECTION
def _resolve_label(class_index, labels=None, fallback_labels=None):
"""Resolve a class index to a display label.
優先序:
1. 注入的 labels來自 models.json / 使用者上傳,經 Go 層傳入)
2. fallback_labelsdetection 沿用 COCO_CLASSES 以維持既有行為)
3. f"class_{index}" —— 原始 enum index
「沒有 label」不是錯誤狀態classification 必須照樣能跑,只是顯示成
class_0 / class_1 / … 讓使用者自己對照。labels 長度與實際類別數不符時,
對得到的用 label、對不到的 fallback 回 index不讓整批失敗
"""
if labels:
if 0 <= class_index < len(labels):
label = labels[class_index]
# 稀疏 labels 允許用空字串佔位(見 plan §3.2 A1此時 fallback
if isinstance(label, str) and label.strip():
return label
if fallback_labels and 0 <= class_index < len(fallback_labels):
return fallback_labels[class_index]
return f"class_{class_index}"
def _extract_logits_vector(result):
"""Extract a 1-D score vector from a classification model's output.
自適應處理:不預設任何 output node 數、shape 或類別數。
支援的 shape 變體(取回時已指定 CHW ordering
(1, C) — batch 1、C 類(最常見)
(C,) — 已 squeeze 的一維向量
(1, C, 1, 1) — NCHW、spatial 為 1×1
(C, 1, 1) — CHW、spatial 為 1×1
(1, 1, 1, C) — NHWC 尾端通道
任何 size == C 但其餘維度皆為 1 的張量 → squeeze 後取用
多個 output node 時:挑選 squeeze 後為一維、且元素數最多的那個節點
classification head 通常是唯一的一維輸出)。
無法判定時 **拋出 ValueError**、附上實際 shape讓上層明確報錯。
絕不靜默回傳可能是錯的結果。
Returns:
(scores: np.ndarray 1-D, node_idx: int)
Raises:
ValueError: 沒有任何 output node 可解讀成 classification logits。
"""
num_nodes = getattr(result.header, "num_output_node", 1) or 1
observed = [] # [(node_idx, original_shape)]
candidates = [] # [(node_idx, original_shape, squeezed 1-D ndarray)]
for node_idx in range(num_nodes):
output = kp.inference.generic_inference_retrieve_float_node(
node_idx=node_idx,
generic_raw_result=result,
channels_ordering=kp.ChannelOrdering.KP_CHANNEL_ORDERING_CHW
)
arr = np.asarray(output.ndarray)
observed.append((node_idx, tuple(arr.shape)))
if arr.size == 0:
continue
# squeeze 掉所有長度 1 的維度。classification 的輸出無論是
# (1,C) / (1,C,1,1) / (C,1,1) / (1,1,1,C)squeeze 後都會變成 (C,)。
squeezed = np.squeeze(arr)
if squeezed.ndim == 0:
# 單一純量C == 1視為只有一個類別的合法輸出
squeezed = squeezed.reshape(1)
if squeezed.ndim != 1:
# 仍是多維 → 帶有 spatial 維度,不是 classification head
# (例如 detection 的 (C,H,W))。不猜、直接跳過這個節點。
continue
candidates.append((node_idx, tuple(arr.shape), squeezed))
if not candidates:
shapes = ", ".join(f"node[{i}]={s}" for i, s in observed) or "(no output node)"
raise ValueError(
"classification output not recognizable: expected a tensor that "
"squeezes to 1-D logits, e.g. (1,C) / (C,) / (1,C,1,1); "
f"got {shapes}"
)
# 多個一維候選時取元素數最多的classification head 的類別數通常
# 明顯大於其它輔助輸出)。單一候選時就是它。
node_idx, original_shape, scores = max(candidates, key=lambda c: c[2].size)
if len(candidates) > 1:
_log(f"Classification: {len(candidates)} 1-D output nodes found, "
f"using node[{node_idx}] shape={original_shape} (largest)")
return scores.astype(np.float64, copy=False), node_idx
def _looks_like_probabilities(scores):
"""Heuristic: has this vector already been through softmax?
若模型輸出已是機率、再做一次 softmax 會把分佈壓平3 類會趨近
0.33/0.33/0.33**不報錯但結果無意義、極難 debug**plan §7 R-2
判定條件(兩者皆須成立):
- 所有元素 >= 0logits 幾乎必有負值)
- 總和落在 1.0 ± PROB_SUM_TOLERANCE
單一類別C == 1且值為 1.0 也符合,屬正確判定。
**已知殘留風險(無法用啟發式根除)**:全非負、且總和恰好為 1.0 的
logits 在數學上與機率向量不可區分 —— 例如 C=2 的 [0.4, 0.6]。收緊
PROB_SUM_TOLERANCE 到 1e-4 已把誤判率壓到 C=3 時約 0.003%(見常數處
的實測數據),但無法歸零。因此呼叫端 **必須把判定結果寫進 log**
_parse_classification_output 有做),這是實機驗收時唯一能看出
判斷對錯的線索。要真正根除只能由模型端明示輸出是否已正規化
models.json 加欄位),屬 M4 範圍。
"""
if scores.size == 0:
return False
if not np.all(np.isfinite(scores)):
return False
if float(np.min(scores)) < 0.0:
return False
return abs(float(np.sum(scores)) - 1.0) <= PROB_SUM_TOLERANCE
def _softmax(scores):
"""Numerically stable softmax over a 1-D vector."""
shifted = scores - np.max(scores)
exp_scores = np.exp(shifted)
total = exp_scores.sum()
if not np.isfinite(total) or total <= 0:
# 理論上不會發生shifted 最大值為 0 → exp 至少有一個 1.0
# 但寧可退回均勻分佈也不要回傳 NaN 給前端。
_log("Classification: softmax denominator invalid, falling back to uniform")
return np.full(scores.shape, 1.0 / scores.size, dtype=np.float64)
return exp_scores / total
def _parse_classification_output(result, labels=None, top_k=None):
"""Parse classification model output.
完全自適應:不假設 output node 數、shape 或類別數,類別數一律從實際
張量 shape 取得。無法判定時明確拋錯(由 handle_inference 轉成 error
response不靜默回傳可能是錯的結果。
Args:
result: SDK 的 generic raw result。
labels: 注入的 label 陣列。None / 空 → 輸出原始 enumclass_N
top_k: 取前幾名。None 或 <= 0 → 使用 DEFAULT_CLASSIFICATION_TOP_K。
Returns:
list[dict],依 confidence 降序,每筆含 label / classIndex / confidence。
Raises:
ValueError: output shape 無法解讀成 classification logits。
"""
scores, node_idx = _extract_logits_vector(result)
num_classes = int(scores.size)
if _looks_like_probabilities(scores):
probs = scores
# 這條 log 是實機驗收時唯一能看出「判定為已是機率」對不對的線索
# (見 _looks_like_probabilities 的殘留風險說明)。若信心度數值看起來
# 不合理、先來這裡確認是不是誤判成已正規化而跳過了 softmax。
_log(f"Classification: node[{node_idx}] C={num_classes}, "
f"SKIPPING SOFTMAX — output judged already normalized "
f"(sum={float(np.sum(scores)):.8f}, "
f"range=[{float(np.min(scores)):.4f}, {float(np.max(scores)):.4f}], "
f"tolerance={PROB_SUM_TOLERANCE:g}). "
f"If confidences look wrong, this verdict is the first suspect.")
else:
probs = _softmax(scores)
_log(f"Classification: node[{node_idx}] C={num_classes}, "
f"raw range=[{float(np.min(scores)):.4f}, {float(np.max(scores)):.4f}], "
f"applied softmax")
effective_k = top_k if isinstance(top_k, int) and top_k > 0 else DEFAULT_CLASSIFICATION_TOP_K
effective_k = min(effective_k, num_classes)
if labels and len(labels) != num_classes:
_log(f"Classification: label count ({len(labels)}) != class count "
f"({num_classes}); unmatched indices fall back to class_N")
top_indices = np.argsort(probs)[::-1][:effective_k]
classifications = []
for idx in top_indices:
class_index = int(idx)
classifications.append({
"label": _resolve_label(class_index, labels=labels),
"classIndex": class_index,
"confidence": float(probs[class_index]),
})
return classifications
# ── Command handlers ─────────────────────────────────────────────────
def handle_scan():
"""Scan for connected Kneron devices.
Tries Kneron PLUS SDK first (provides firmware info, kn_number, etc.).
Falls back to pyusb if the SDK is unavailable (e.g. macOS missing .dylib).
"""
if HAS_KP:
try:
descs = kp.core.scan_devices()
devices = []
for i in range(descs.device_descriptor_number):
dev = descs.device_descriptor_list[i]
devices.append({
"port": str(dev.usb_port_id),
"firmware": str(dev.firmware),
"kn_number": f"0x{dev.kn_number:08X}",
"product_id": f"0x{dev.product_id:04X}",
"connectable": dev.is_connectable,
})
return {"devices": devices}
except Exception as e:
_log(f"kp.core.scan_devices failed: {e}, trying pyusb fallback")
# Fallback: use pyusb (same approach as kneron_detect.py)
if HAS_PYUSB:
return _scan_with_pyusb()
return {"devices": [], "error_detail": "neither kp nor pyusb available"}
# Known Kneron product IDs (same as kneron_detect.py)
_KNERON_VENDOR_ID = 0x3231
_KNOWN_PRODUCTS = {
0x0100: "KL520",
0x0200: "KL720",
0x0720: "KL720",
0x0530: "KL530",
0x0630: "KL630",
0x0730: "KL730",
}
def _scan_with_pyusb():
"""Scan for Kneron devices using pyusb (libusb backend)."""
try:
usb_devices = list(usb.core.find(find_all=True, idVendor=_KNERON_VENDOR_ID))
devices = []
for dev in usb_devices:
product_id = f"0x{dev.idProduct:04X}"
chip = _KNOWN_PRODUCTS.get(dev.idProduct, f"Unknown-{product_id}")
# pyusb port_id: bus-address
port = f"{dev.bus}-{dev.address}"
firmware = "unknown"
try:
firmware = dev.product or "unknown"
except Exception:
pass
devices.append({
"port": port,
"firmware": firmware,
"kn_number": "0x00000000",
"product_id": product_id,
"connectable": True,
})
return {"devices": devices}
except Exception as e:
return {"devices": [], "error_detail": f"pyusb scan failed: {e}"}
def handle_connect(params):
"""Connect to a Kneron device and load firmware if needed.
KL520: USB Boot mode — firmware MUST be uploaded every session.
KL720 (KDP2, pid=0x0720): Flash-based — firmware pre-installed.
KL720 (KDP legacy, pid=0x0200): Old firmware — needs connect_without_check
+ firmware load to RAM before normal operation.
"""
global _device_group, _firmware_loaded, _device_chip
if not HAS_KP:
return {"error": "kp module not available"}
try:
port = params.get("port", "")
device_type = params.get("device_type", "")
# Scan to find device
descs = kp.core.scan_devices()
if descs.device_descriptor_number == 0:
return {"error": "no Kneron device found"}
# Find device by port or use first one
target_dev = None
for i in range(descs.device_descriptor_number):
dev = descs.device_descriptor_list[i]
if port and str(dev.usb_port_id) == port:
target_dev = dev
break
if target_dev is None:
target_dev = descs.device_descriptor_list[0]
# Note: KL520 in USB Boot mode has is_connectable=False, which is
# normal — it becomes connectable after firmware is loaded. KL720 KDP
# legacy (pid=0x0200) is also not connectable until firmware load.
# So we do NOT reject is_connectable=False here; instead we attempt
# connection and firmware load as appropriate.
# Determine chip type from device_type param or product_id
pid = target_dev.product_id
if "kl720" in device_type.lower():
_device_chip = "KL720"
elif "kl520" in device_type.lower():
_device_chip = "KL520"
elif pid in (0x0200, 0x0720):
_device_chip = "KL720"
else:
_device_chip = "KL520"
fw_str = str(target_dev.firmware)
is_kdp_legacy = (_device_chip == "KL720" and pid == 0x0200)
_log(f"Chip type: {_device_chip} (product_id=0x{pid:04X}, device_type={device_type}, fw={fw_str})")
# ── KL720 KDP Legacy (pid=0x0200): old firmware, incompatible with SDK ──
if is_kdp_legacy:
_log(f"KL720 has legacy KDP firmware (pid=0x0200). Using connect_devices_without_check...")
_device_group = kp.core.connect_devices_without_check(
usb_port_ids=[target_dev.usb_port_id]
)
kp.core.set_timeout(device_group=_device_group, milliseconds=60000)
# Load KDP2 firmware to RAM so the device can operate with this SDK
scpu_path, ncpu_path = _resolve_firmware_paths("KL720")
if scpu_path and ncpu_path:
_log(f"KL720: Loading KDP2 firmware to RAM: {scpu_path}")
kp.core.load_firmware_from_file(
_device_group, scpu_path, ncpu_path
)
_firmware_loaded = True
_log("KL720: Firmware loaded to RAM, waiting for reboot...")
time.sleep(5)
# Reconnect — device should now be running KDP2 in RAM
descs = kp.core.scan_devices()
reconnected = False
for i in range(descs.device_descriptor_number):
dev = descs.device_descriptor_list[i]
if dev.product_id in (0x0200, 0x0720):
target_dev = dev
reconnected = True
break
if not reconnected:
return {"error": "KL720 not found after firmware load. Unplug and re-plug."}
# Try normal connect first, fallback to without_check
try:
_device_group = kp.core.connect_devices(
usb_port_ids=[target_dev.usb_port_id]
)
except Exception as conn_err:
_log(f"KL720: Normal reconnect failed ({conn_err}), using without_check...")
_device_group = kp.core.connect_devices_without_check(
usb_port_ids=[target_dev.usb_port_id]
)
kp.core.set_timeout(device_group=_device_group, milliseconds=10000)
fw_str = str(target_dev.firmware)
_log(f"KL720: Reconnected after firmware load, pid=0x{target_dev.product_id:04X}, fw={fw_str}")
else:
_log("WARNING: KL720 firmware files not found. Cannot operate with KDP legacy device.")
_clear_device_group()
return {"error": "KL720 has legacy KDP firmware but KDP2 firmware files not found. "
"Run update_kl720_firmware.py to flash KDP2 permanently."}
return {
"status": "connected",
"firmware": fw_str,
"kn_number": f"0x{target_dev.kn_number:08X}",
"chip": _device_chip,
"kdp_legacy": True,
}
# ── Normal connection (KL520 or KL720 KDP2) ──
# Use connect_devices_without_check when:
# - KL720 KDP2: connect_devices() often fails with Error 28
# - KL520 USB Boot: is_connectable=False, connect_devices() rejects it
# In these cases, connect_devices_without_check() works and we can
# still load firmware afterwards.
use_without_check = (_device_chip == "KL720") or (not target_dev.is_connectable)
max_retries = 3
last_err = None
for attempt in range(max_retries):
try:
# Clear any stale device group from previous failed attempt.
_clear_device_group()
if use_without_check:
_log(f"{_device_chip}: connect_devices_without_check(usb_port_id={target_dev.usb_port_id}, connectable={target_dev.is_connectable}) attempt {attempt+1}/{max_retries}...")
_device_group = kp.core.connect_devices_without_check(
usb_port_ids=[target_dev.usb_port_id]
)
else:
_log(f"connect_devices(usb_port_id={target_dev.usb_port_id}) attempt {attempt+1}/{max_retries}...")
_device_group = kp.core.connect_devices(
usb_port_ids=[target_dev.usb_port_id]
)
_log(f"connect succeeded on attempt {attempt+1}")
last_err = None
break
except Exception as conn_err:
_clear_device_group()
last_err = conn_err
_log(f"connect attempt {attempt+1} failed: {conn_err}")
if attempt < max_retries - 1:
time.sleep(2)
# Re-scan to refresh device handle
try:
descs = kp.core.scan_devices()
for i in range(descs.device_descriptor_number):
dev = descs.device_descriptor_list[i]
if port and str(dev.usb_port_id) == port:
target_dev = dev
break
elif not port:
target_dev = descs.device_descriptor_list[0]
break
except Exception:
pass
if last_err is not None:
hint = ""
if sys.platform == "win32":
hint = (" On Windows, ensure the WinUSB driver is installed for this device."
" Re-run the installer or use Zadig (https://zadig.akeo.ie).")
raise RuntimeError(f"Failed to connect after {max_retries} attempts: {last_err}.{hint}")
# KL720 needs longer timeout for large NEF transfers (12MB+ over USB)
_timeout_ms = 60000 if _device_chip == "KL720" else 10000
_log(f"Calling set_timeout(milliseconds={_timeout_ms})...")
kp.core.set_timeout(device_group=_device_group, milliseconds=_timeout_ms)
_log(f"set_timeout succeeded")
# Firmware handling — chip-dependent.
# fresh_firmware_loaded is used by Go driver to decide whether to
# skip the post-connect reset (freshly loaded firmware is already
# in a clean state — reset would just waste 30-60s reloading it).
fresh_firmware_loaded = False
if "Loader" in fw_str:
# Device is in USB Boot (Loader) mode and needs firmware
if _device_chip == "KL720":
_log(f"WARNING: {_device_chip} is in Loader mode (unusual). Attempting firmware load...")
scpu_path, ncpu_path = _resolve_firmware_paths(_device_chip)
if scpu_path and ncpu_path:
_log(f"{_device_chip}: Loading firmware: {scpu_path}")
kp.core.load_firmware_from_file(
_device_group, scpu_path, ncpu_path
)
_firmware_loaded = True
_log("Firmware loaded, waiting for reboot...")
time.sleep(5)
# Reconnect after firmware load (with retry)
_clear_device_group()
for retry in range(3):
try:
descs = kp.core.scan_devices()
target_dev = descs.device_descriptor_list[0]
try:
_device_group = kp.core.connect_devices(
usb_port_ids=[target_dev.usb_port_id]
)
except Exception:
_device_group = kp.core.connect_devices_without_check(
usb_port_ids=[target_dev.usb_port_id]
)
break
except Exception as re_err:
_log(f"Reconnect attempt {retry+1} failed: {re_err}")
if retry < 2:
time.sleep(3)
if _device_group is None:
return {"error": "Device not found after firmware load. Unplug and re-plug the device."}
kp.core.set_timeout(
device_group=_device_group, milliseconds=_timeout_ms
)
fw_str = str(target_dev.firmware)
fresh_firmware_loaded = True
_log(f"Reconnected after firmware load, firmware: {fw_str}")
else:
_log(f"WARNING: {_device_chip} firmware files not found, skipping firmware load")
else:
# Not in Loader mode — firmware already present from a previous
# session. This is the state that triggers Error 15 on inference
# without reset, per observed bug.
_log(f"{_device_chip}: firmware already present (normal). fw={fw_str}")
return {
"status": "connected",
"firmware": fw_str,
"kn_number": f"0x{target_dev.kn_number:08X}",
"chip": _device_chip,
"fresh_firmware_loaded": fresh_firmware_loaded,
}
except Exception as e:
_clear_device_group()
return {"error": str(e)}
def handle_disconnect(params):
"""Disconnect from the current device."""
global _device_group, _model_id, _model_nef, _firmware_loaded
global _model_type, _device_chip
_clear_device_group()
_model_id = None
_model_nef = None
_model_type = "tiny_yolov3"
_firmware_loaded = False
_device_chip = "KL520"
# input size 由 _reset_model_metadata 一併還原(含 width/height/source
# 不在這裡另外寫 _model_input_size —— 兩處各寫一半會讓三個全域不同步。
_reset_model_metadata()
return {"status": "disconnected"}
def handle_reset(params):
"""Reset the device back to USB Boot (Loader) state.
This forces the device to drop its firmware and any loaded models.
After reset the device will re-enumerate on USB, so the caller
must wait and issue a fresh 'connect' command.
"""
global _device_group, _model_id, _model_nef, _firmware_loaded
global _model_type, _device_chip
if _device_group is None:
return {"error": "device not connected"}
try:
_log("Resetting device (kp.core.reset_device KP_RESET_REBOOT)...")
kp.core.reset_device(
device_group=_device_group,
reset_mode=kp.ResetMode.KP_RESET_REBOOT,
)
_log("Device reset command sent successfully")
except Exception as e:
_log(f"reset_device raised: {e}")
# Even if it throws, the device usually does reset.
# Clear all state — the device is gone until it re-enumerates.
_clear_device_group()
_model_id = None
_model_nef = None
_model_type = "tiny_yolov3"
_firmware_loaded = False
_device_chip = "KL520"
# input size 由 _reset_model_metadata 一併還原(含 width/height/source
_reset_model_metadata()
return {"status": "reset"}
def handle_load_model(params):
"""Load a model file onto the device.
KL520 USB Boot mode limitation: only one model can be loaded per
USB session. If error 40 occurs, the error is returned to the Go
driver which handles it by restarting the entire Python bridge.
Accepted params:
path (required) — .nef 檔絕對路徑
task_type (optional) — "classification" | "object_detection"
由 Go 層從 model metadata 帶入。**有指定就不做猜測。**
labels (optional) — list[str]index 對應 class index。
沒帶時 classification 輸出原始 enumclass_N
detection 沿用 COCO_CLASSES。
input_size (optional) — {"width": w, "height": h}models.json /
metadata.json 宣告的 inputSize。**只在 SDK 沒回報
input shape 時才會用**:這個值是人填的,使用者可能
隨手填一個數字,不能當成可信來源。
失敗語意:任一步失敗都 **不會修改任何全域狀態**all-or-nothing
裝置上的舊模型仍可繼續推論,其 _model_id / _model_type / _model_labels
保持互相一致。
"""
global _model_id, _model_nef, _model_nef_path, _task_type_override, _model_labels
global _model_declared_input_size
if _device_group is None:
return {"error": "device not connected"}
path = params.get("path", "")
if not path or not os.path.exists(path):
return {"error": f"model file not found: {path}"}
# 先解析成 local**還不要寫全域**。全域 metadata 必須永遠描述
# 「_model_id 當前指向的那個模型」;載入失敗時舊模型仍在裝置上、
# 仍可能被繼續推論,此時若 labels 已被新模型的值覆蓋,就會用
# 舊模型 + 新 label 表 → 標籤張冠李戴(靜默錯誤、極難 debug
# 所以所有 metadata 一律等到「確定成功」才一次寫入(見下方 commit 段)。
pending_task_type = _normalize_task_type(params.get("task_type"))
pending_labels = _sanitize_labels(params.get("labels"))
pending_declared_size = _normalize_declared_input_size(params.get("input_size"))
_log(f"load_model: task_type={pending_task_type or '(not specified)'}, "
f"labels={len(pending_labels) if pending_labels else 0}, "
f"declared_input_size="
f"{'x'.join(map(str, pending_declared_size)) if pending_declared_size else '(not specified)'}")
try:
nef = kp.core.load_model_from_file(
device_group=_device_group,
file_path=path
)
model = nef.models[0]
model_id = model.id
target_chip = str(nef.target_chip)
except Exception as e:
# 任何一步失敗 → 全域完全不動,維持舊模型的一致狀態。
_log(f"load_model failed, keeping previous model state: {e}")
return {"error": str(e)}
# ── Commit到這裡才確定成功一次寫入全部全域 metadata ──────────
_model_nef = nef
_model_id = model_id
_model_nef_path = path
_task_type_override = pending_task_type
_model_labels = pending_labels
# 留著給 set_inference_options 用:推論期切解析方式會重跑 _detect_model_type
# 那時要能取到同一組 declared 值,否則切換的副作用是偷改 input size。
_model_declared_input_size = pending_declared_size
# Detect model type and input size外部指定的 task_type 優先)。
# nef 帶進去 → input size 由模型自己宣告的 shape 決定,不再靠檔名猜。
_detect_model_type(_model_id, path, task_type=_task_type_override,
nef=nef, declared_input_size=pending_declared_size)
_log(f"Model loaded: id={_model_id}, type={_model_type}, "
f"{_describe_input_size()}, target={target_chip}")
return {
"status": "loaded",
"model_id": _model_id,
"model_type": _model_type,
"task_type": _current_task_type(),
"label_count": len(_model_labels) if _model_labels else 0,
"input_size": _model_input_size,
"input_width": _model_input_width,
"input_height": _model_input_height,
"input_size_source": _model_input_size_source,
"model_path": path,
"target_chip": target_chip,
}
def handle_set_inference_options(params):
"""Update parse mode / labels for the **already loaded** model.
KL520 USB Boot 模式一次只能載一個 model換 model 必須重燒(數十秒)。
但「怎麼解析輸出」與「class index 顯示成什麼名字」都只是 post-process
的事,跟裝置上那份 model 無關 —— 所以這兩件事可以在推論期即時切換,
不需要重新 load_model。
Accepted params兩者皆為 optional但至少要帶一個
task_type (optional) — "classification" | "object_detection"
帶了就覆寫當前的解析方式。
labels (optional) — list[str]index 對應 class index。
帶空 list或全空字串= 清掉 label 表,回到原始 enum
輸出class_N。這是刻意可達的狀態使用者可能上傳錯
label 檔想清掉。
失敗語意(沿用 handle_load_model 的 defer-until-success任一步驗證
失敗 → **零全域變動**。驗證通過才一次 commit。理由相同 —— 全域 metadata
必須永遠與 _model_id 指向的那個模型互相一致,半套寫入會產生「舊解析方式
配新 label 表」這種不報錯的錯誤。
"""
global _task_type_override, _model_labels
if _device_group is None:
return {"error": "device not connected"}
if _model_id is None:
return {"error": "no model loaded — load a model before setting inference options"}
has_task_type = "task_type" in params
has_labels = "labels" in params
if not has_task_type and not has_labels:
return {"error": "set_inference_options requires task_type and/or labels"}
# ── 驗證階段:全部算成 local還不要碰全域 ──────────────────────
pending_task_type = _task_type_override
if has_task_type:
raw_task_type = params.get("task_type")
normalized = _normalize_task_type(raw_task_type)
if normalized is None:
# 這裡與 load_model 不同load_model 的 task_type 是「附帶提示」、
# 無法辨識就 fallback 檔名猜測是合理的。但這個 handler 的**唯一
# 目的**就是套用呼叫端指定的解析方式 —— 靜默忽略等於回 200 但什麼
# 都沒做,使用者會以為切換生效了。所以一律明確報錯。
return {"error": f"invalid task_type: {raw_task_type!r} "
f"(expected {TASK_TYPE_CLASSIFICATION} or "
f"{TASK_TYPE_OBJECT_DETECTION})"}
pending_task_type = normalized
pending_labels = _model_labels
if has_labels:
raw_labels = params.get("labels")
if raw_labels is None:
pending_labels = None
elif not isinstance(raw_labels, (list, tuple)):
return {"error": f"labels must be a list, got {type(raw_labels).__name__}"}
else:
# _sanitize_labels 對空 list / 全空字串回 None正好就是「清掉
# label 表」的語意(下游 _resolve_label 會回 class_N
pending_labels = _sanitize_labels(raw_labels)
# ── Commit驗證全過才一次寫入 ─────────────────────────────────
_task_type_override = pending_task_type
_model_labels = pending_labels
if has_task_type:
# 重跑 model type 偵測讓 _model_type實際決定走哪條 post-process
# 跟著新的 task_type 走。
#
# ⚠️ 三個 input size 來源都要一起帶回去_model_nef / declared / 路徑),
# 否則這次重跑會退回較低可信度的來源,等於「切解析方式」這個動作
# 偷偷改掉了 input size —— 而 input size 錯了 NPU 不會報錯。
_detect_model_type(_model_id, _model_nef_path,
task_type=_task_type_override,
nef=_model_nef,
declared_input_size=_model_declared_input_size)
_log(f"set_inference_options: task_type={_task_type_override or '(unchanged/none)'}, "
f"labels={len(_model_labels) if _model_labels else 0}, "
f"model_type={_model_type}, {_describe_input_size()}")
return {
"status": "updated",
"model_id": _model_id,
"model_type": _model_type,
"task_type": _current_task_type(),
"label_count": len(_model_labels) if _model_labels else 0,
"input_size": _model_input_size,
"input_width": _model_input_width,
"input_height": _model_input_height,
"input_size_source": _model_input_size_source,
}
def handle_inference(params):
"""Run inference on the provided image data."""
if _device_group is None:
return {"error": "device not connected"}
if _model_id is None:
return {"error": "no model loaded"}
image_b64 = params.get("image_base64", "")
try:
t0 = time.time()
if image_b64:
# Decode base64 image
img_bytes = base64.b64decode(image_b64)
if HAS_CV2:
# Decode image with OpenCV
img_array = np.frombuffer(img_bytes, dtype=np.uint8)
img = cv2.imdecode(img_array, cv2.IMREAD_COLOR)
if img is None:
return {"error": "failed to decode image"}
h, w = img.shape[:2]
# KL520 NPU requires input image dimensions >= model input size
# and both width/height must be even numbers.
#
# 逐軸比較(而非拿單一 min_dim 比兩軸):非正方形模型用單一值
# 會讓較長的那一軸不足 —— 例如模型要 256x192、影像 200x200
# 用 min_dim=192 會判定「已達標」而不放大,但寬度其實差 56px。
min_w = _model_input_width
min_h = _model_input_height
if w < min_w or h < min_h or w % 2 != 0 or h % 2 != 0:
if w < min_w or h < min_h:
# 等比例放大到兩軸都達標(取較大的縮放倍率)。
scale = max(min_w / w, min_h / h)
new_w = int(w * scale)
new_h = int(h * scale)
else:
new_w, new_h = w, h
new_w = (new_w + 1) & ~1
new_h = (new_h + 1) & ~1
img = cv2.resize(img, (new_w, new_h), interpolation=cv2.INTER_LINEAR)
_log(f"Inference image resized: {w}x{h} -> {new_w}x{new_h} "
f"(model needs >= {min_w}x{min_h})")
# Convert BGR to BGR565
img_bgr565 = cv2.cvtColor(src=img, code=cv2.COLOR_BGR2BGR565)
else:
img_bgr565 = np.frombuffer(img_bytes, dtype=np.uint8)
else:
return {"error": "no image data provided"}
# Create inference config (original: pass numpy ndarray, SDK reads shape)
inf_config = kp.GenericImageInferenceDescriptor(
model_id=_model_id,
inference_number=0,
input_node_image_list=[
kp.GenericInputNodeImage(
image=img_bgr565,
image_format=kp.ImageFormat.KP_IMAGE_FORMAT_RGB565,
)
]
)
# Send and receive
_log(f"Inference: sending to NPU (model_type={_model_type}, "
f"{_describe_input_size()})")
kp.inference.generic_image_inference_send(_device_group, inf_config)
result = kp.inference.generic_image_inference_receive(_device_group)
_log(f"Inference: receive complete, parsing...")
elapsed_ms = (time.time() - t0) * 1000
# Parse output based on model type
detections = []
classifications = []
task_type = _current_task_type()
if task_type == TASK_TYPE_CLASSIFICATION:
# 解析失敗會拋 ValueError附實際 shape由外層 except 轉成
# error response —— 不靜默回傳可能是錯的空結果。
classifications = _parse_classification_output(
result,
labels=_model_labels,
top_k=params.get("top_k"),
)
elif _model_type == "ssd":
detections = _parse_ssd_output(
result, input_size=_model_input_size, labels=_model_labels)
elif _model_type == "fcos":
detections = _parse_fcos_output(
result, input_size=_model_input_size, labels=_model_labels)
elif _model_type == "yolov5s":
detections = _parse_yolo_output(
result,
anchors=ANCHORS_YOLOV5S,
input_size=_model_input_size,
labels=_model_labels,
)
else:
# Default: Tiny YOLOv3
detections = _parse_yolo_output(
result,
anchors=ANCHORS_TINY_YOLOV3,
input_size=_model_input_size,
labels=_model_labels,
)
_log(f"Inference: parse done, detections={len(detections)}, classifications={len(classifications)}, elapsed={elapsed_ms:.1f}ms")
return {
"taskType": task_type,
"timestamp": int(time.time() * 1000),
"latencyMs": round(elapsed_ms, 1),
"detections": detections,
"classifications": classifications,
}
except Exception as e:
import traceback
_log(f"Inference EXCEPTION: {type(e).__name__}: {e}\n{traceback.format_exc()}")
return {"error": _annotate_inference_error(e)}
# KP_ERROR_INVALID_PARAM_12 幾乎都是「送進去的影像尺寸不是模型要的」。
# SDK 只回一個對使用者毫無意義的錯誤碼,不會說是哪個參數不對。
_INVALID_PARAM_MARKERS = ("KP_ERROR_INVALID_PARAM", "Error code: 12")
def _annotate_inference_error(exc):
"""Attach input-size context to errors that are most likely size mismatches.
為什麼要做這件事KP_ERROR_INVALID_PARAM_12 對使用者完全無法理解,而
尺寸不符是它最常見的成因。把「當前用的尺寸 + 尺寸從哪來」直接寫進錯誤
訊息,下次遇到就不必再從 log 反推來源2026-07 那次繞了一大圈)。
"""
message = str(exc)
if not any(marker in message for marker in _INVALID_PARAM_MARKERS):
return message
hint = (f"推論失敗KP_ERROR_INVALID_PARAM"
f"當前 input_size={_model_input_width}x{_model_input_height} "
f"(source: {_model_input_size_source})。"
f"若模型實際尺寸不同,這是最可能的原因。")
if _model_input_size_source == INPUT_SIZE_SOURCE_DECLARED:
hint += ("此尺寸來自上傳時手動填寫的欄位(非模型自述),"
"請優先確認它是否填錯。")
elif _model_input_size_source in (INPUT_SIZE_SOURCE_KNOWN_ID,
INPUT_SIZE_SOURCE_DEFAULT):
hint += ("此尺寸是在沒有任何可靠來源時推測的,"
"請於模型設定中填入正確的輸入尺寸。")
_log(hint)
return f"{hint} 原始錯誤:{message}"
# ── Firmware upgrade (A 階段 M9-1) ───────────────────────────────────
#
# 對應 TDD v2/firmware-management.md §5.1 / §6.1
# - 自動升級 KDP1 legacy → KDP2含 KL520USB Boot mode + loader stage
# 與 KL720含 KDP legacy pid=0x0200
# - Stage 命名採 Designpreparing / loading / flashing / verifying / done / error
# TDD §4.3 為 source of truth
# - 失敗 reason enumTDD §3.4scan_not_found / connect_failed /
# loader_write_failed / upgrade_mid_failed / disconnect_during_op /
# timeout / verify_mismatch / verify_not_found。
#
# 為什麼走 ctypesKneronPLUS Python wrapper 沒 export
# `kp_update_kdp_firmware_from_files`(見 research-kl520-fw-management/
# 56-m9-6-strong-validation-result.md 附帶發現 1warrenchen reference
# 實作 `LocalAPI/legacy_plus121_runner.py` 直接 ctypes 打 C symbol本檔
# 沿用該模式。
KDP_MAGIC_CONNECTION_PASS = 536173391 # 與 warrenchen reference 一致
KP_SUCCESS = 0
USB_WAIT_AFTER_REBOOT_MS = 2000 # SDK loader 階段 reboot 等待
USB_WAIT_AFTER_UPGRADE_MS = 5000 # AC-FW-1.6:升級後 5-8s USB stable
USB_WAIT_RETRY_CONNECT_MS = 200
MAX_RECONNECT_RETRIES = 15 # 5s sleep + 15 * 200ms = 8s 上界
KL520_UPGRADE_TIMEOUT_S = 60 # AC-FW-1.7
KL720_UPGRADE_TIMEOUT_S = 200 # AC-FW-1.7
# 進度事件 stage % 對照TDD §4.3
_FW_STAGE_PERCENT = {
"preparing": 5,
"loading": 20,
"flashing": 50,
"verifying": 90,
"done": 100,
"error": -1,
}
# 升級進行中旗標SIGTERM handler 用、AC-FW-1.9 graceful shutdown 拒絕)
# Reviewer m4原本還有 _firmware_upgrade_start_ts 全域變數、與 SIGTERM handler
# closure capture 的 start_ts 重複、容易未來 desync → 砍掉、單一 source of truth
# 走 closure。
_firmware_upgrade_in_progress = False
def _fw_normalize_code(code):
"""Convert int8-like unsigned (e.g. 253 for -3) to signed.
與 warrenchen reference 一致:某些 legacy 路徑回 unsigned int8 值。
"""
try:
c = int(code)
except Exception:
return code
if c > 127:
return c - 256
return c
def _fw_emit_progress(stage, message="", elapsed_ms=0, eta_ms=0, extra=None):
"""Push a progress event to stderr as a JSON-RPC notification line.
Go driver 抓 stderr line-by-line、轉成 WebSocket FirmwareProgress 給前端。
Schema 對齊 TDD §4.2 `FirmwareProgress`
{"event": "firmware_progress", "percent": int, "stage": str,
"message": str, "elapsed_ms": int, "eta_ms": int, ...}
Stage `error` 時 caller 應 push 額外 reason / raw_error / before_version
透過 extra dict。
"""
payload = {
"event": "firmware_progress",
"percent": _FW_STAGE_PERCENT.get(stage, 0),
"stage": stage,
"message": message,
"elapsed_ms": int(elapsed_ms),
"eta_ms": int(eta_ms),
}
if extra:
payload.update(extra)
try:
# 寫到 stderr、與既有 _log() 同 fd、但用 JSON 格式(不加 [kneron_bridge] prefix
# 方便 Go driver 區分「progress event JSON」vs「自由文字 log」。
print(json.dumps(payload), file=sys.stderr, flush=True)
except Exception:
# progress emit 失敗不該影響升級流程本身
pass
def _fw_load_libkplus():
"""Load libkplus shared library via ctypes、bind needed C symbol signatures.
跨平台macOS .dylib / Linux .so / Windows .dll。優先用 `kp` module 已載
入的 lib path避免重複載入造成 mismatchfallback 到 wheel 內 lib/ 目錄。
Raises:
RuntimeError: 若 libkplus 找不到或符號 binding 失敗。
"""
import ctypes
import importlib.util
spec = importlib.util.find_spec("kp")
if spec is None or not spec.submodule_search_locations:
raise RuntimeError("kp module spec not found")
kp_dir = spec.submodule_search_locations[0]
lib_dir = os.path.join(kp_dir, "lib")
# 平台對應的 lib filename
if sys.platform == "darwin":
lib_name = "libkplus.dylib"
elif sys.platform == "win32":
lib_name = "libkplus.dll"
else:
lib_name = "libkplus.so"
lib_path = os.path.join(lib_dir, lib_name)
if not os.path.isfile(lib_path):
# Windows 可能用其他命名warrenchen reference 是 libkplus.dll
# 嘗試找任何 libkplus* 檔案
# Reviewer m2sort() 確保 deterministic 順序、不依賴 os.listdir 回傳次序
candidates = sorted(
f for f in os.listdir(lib_dir) if f.startswith("libkplus")
)
if not candidates:
raise RuntimeError(f"libkplus not found in {lib_dir}")
lib_path = os.path.join(lib_dir, candidates[0])
_log(f"WARNING: libkplus fallback using {candidates[0]} (primary {lib_name} not found)")
# Windows: add_dll_directory 確保相依 dll 可解析
if sys.platform == "win32" and hasattr(os, "add_dll_directory"):
try:
os.add_dll_directory(lib_dir)
except Exception:
pass
lib = ctypes.CDLL(lib_path)
# Bind C symbol signatures與 warrenchen reference 完全一致)
lib.kp_connect_devices.argtypes = [
ctypes.c_int, # num_devices
ctypes.POINTER(ctypes.c_int), # usb_port_ids
ctypes.POINTER(ctypes.c_int), # status_out
]
lib.kp_connect_devices.restype = ctypes.c_void_p # device_group handle
lib.kp_set_timeout.argtypes = [ctypes.c_void_p, ctypes.c_int]
lib.kp_set_timeout.restype = None
lib.kp_load_firmware_from_file.argtypes = [
ctypes.c_void_p, ctypes.c_char_p, ctypes.c_char_p
]
lib.kp_load_firmware_from_file.restype = ctypes.c_int
lib.kp_update_kdp_firmware_from_files.argtypes = [
ctypes.c_void_p, # device_group
ctypes.c_char_p, # scpu_or_loader path
ctypes.c_char_p, # ncpu path or NULL
ctypes.c_bool, # auto_reboot
]
lib.kp_update_kdp_firmware_from_files.restype = ctypes.c_int
lib.kp_disconnect_devices.argtypes = [ctypes.c_void_p]
lib.kp_disconnect_devices.restype = ctypes.c_int
if hasattr(lib, "kp_error_string"):
lib.kp_error_string.argtypes = [ctypes.c_int]
lib.kp_error_string.restype = ctypes.c_char_p
return lib
def _fw_errstr(lib, code):
"""Decode kp error code → string via kp_error_string()。
與 warrenchen 一致:先試 raw code、若無回應再試 signed normalize 後值。
"""
signed = _fw_normalize_code(code)
if hasattr(lib, "kp_error_string"):
try:
msg = lib.kp_error_string(int(code))
if not msg and signed != code:
msg = lib.kp_error_string(int(signed))
if msg:
return msg.decode("utf-8", errors="replace")
except Exception:
pass
return f"code={code}"
def _fw_connect_with_magic(lib, port_id):
"""Connect with magic pass = 536173391 (允許 KDP1 legacy device 連線)。
Returns:
device_group handle (c_void_p int).
Raises:
RuntimeError("connect_failed: ...") on failure.
"""
import ctypes
port_ids = (ctypes.c_int * 1)(int(port_id))
status = ctypes.c_int(KDP_MAGIC_CONNECTION_PASS)
dg = lib.kp_connect_devices(1, port_ids, ctypes.byref(status))
if not dg or status.value != KP_SUCCESS:
signed = _fw_normalize_code(status.value)
raise RuntimeError(
f"connect_failed: raw_code={status.value}, signed={signed}, "
f"msg={_fw_errstr(lib, status.value)}"
)
return dg
def _fw_scan_target(port):
"""Scan devices via kp.core.scan_devices() and find target by usb_port_id.
Returns:
descriptor or None.
"""
try:
descs = kp.core.scan_devices()
except Exception as e:
_log(f"fw_scan_target: scan_devices failed: {e}")
return None
if descs.device_descriptor_number == 0:
return None
for i in range(descs.device_descriptor_number):
dev = descs.device_descriptor_list[i]
if port and str(dev.usb_port_id) == str(port):
return dev
return None
def _fw_rescan_and_wait(port, max_wait_s=8.0, initial_sleep_s=5.0):
"""等 USB re-enumerate stable → rescan 找回 target by port (AC-FW-1.6)。
Args:
port: 原 usb_port_id升級後 re-enumerate 通常保留同 port
max_wait_s: 從 initial_sleep_s 過後再加 max_wait_s - initial_sleep_s
秒輪詢上界。實測 5 秒已穩、保留上界 8 秒AC-FW-1.6)。
initial_sleep_s: 第一次 rescan 前固定等的秒數。
Returns:
(descriptor or None, total_wait_s).
"""
time.sleep(initial_sleep_s)
waited = initial_sleep_s
target = _fw_scan_target(port)
if target is not None:
return target, waited
# 多輪 short-poll
poll_step = 0.5
while waited < max_wait_s:
time.sleep(poll_step)
waited += poll_step
target = _fw_scan_target(port)
if target is not None:
return target, waited
return None, waited
def _fw_classify_legacy(firmware_str, product_id):
"""判斷 device 是否為 KDP1 legacy state需走 loader stage
KL520 legacy 訊號firmware 字串為 "KDP""KDP1""KDP1.x""USB Boot"
"USB Boot Loader""LOADER" 等 legacy state、或空字串
(某些 USB Boot state 不回 firmware string
KL720 legacy 訊號product_id == 0x0200 (KP_DEVICE_KL720_LEGACY)。
Reviewer M3 + s3原本只用 substring match `"KDP" in fw and "KDP2" not in fw`
對 KDP3未來 firmware會誤判 legacy → 改用顯式 prefix 比對表 + 已知字串
enumeration、確保覆蓋 KDP1 各種 firmware 字串變體、forward-compat KDP3+。
Returns True if needs SDK loader stage、False if can short-circuit to flashing.
"""
if product_id == 0x0200:
return True # KL720 KDP1 legacypid 明示、不靠 firmware 字串)
fw = (firmware_str or "").strip().upper()
# 已知 KDP1 legacy firmware 字串完整列舉(明示比對、不靠 substring
legacy_exact = {
"", # 某些 USB Boot state 不回 firmware string
"KDP",
"KDP1",
"USB BOOT",
"USB BOOT LOADER",
"LOADER",
"BOOTLOADER",
}
if fw in legacy_exact:
return True
# KDP1.xKDP1.0 / KDP1.5 等版本字串)
if fw.startswith("KDP1.") or fw.startswith("KDP1 "):
return True
# 明示放行 KDP2 / KDP3+forward-compat、避免 substring match 對未來 firmware 誤判)
# KDP2.x / KDP3.x / KDP4.x ... 皆為 modern firmware、不需走 loader
for prefix in ("KDP2", "KDP3", "KDP4", "KDP5", "KDP6", "KDP7", "KDP8", "KDP9"):
if fw.startswith(prefix):
return False
# 未知 firmware 字串:保守 default = 不走 loader避免誤觸 loader stage brick device
# 例:未來 firmware 用全新命名("NEF"、"K3"、等)→ 假設是 modern firmware
# 若這判斷錯了、verify 階段會 detect verify_mismatch、不致 brick
return False
def _fw_eta_ms(chip, current_stage):
"""估算剩餘 ms給前端顯示 ~X 秒、非精確)。
依 TDD §4.2UI 顯示「~X 秒 remaining」、精度低可接受。
"""
# 各 stage 預估完成時刻(以升級開始為 0
if chip == "KL520":
total_ms = 30000 # AC-FW-1.7 預估 30s
cum = {"preparing": 2000, "loading": 8000, "flashing": 22000, "verifying": 28000}
else: # KL720
total_ms = 180000 # AC-FW-1.7 預估 180s
cum = {"preparing": 5000, "loading": 30000, "flashing": 160000, "verifying": 175000}
done_at = cum.get(current_stage, total_ms)
return max(0, total_ms - done_at)
# ── Firmware upgrade exceptions + failure handler ────────────────────
#
# Reviewer M1原本 _FwError / _FwTimeoutError / _fw_handle_failure 宣告位於
# handle_firmware_upgrade **之後**(語法上 Python module load 時會先掃完整個檔
# 才走 handler、所以 happy-path 不會炸 NameError、但 readability 差、且若有人
# 在 handler 中間插入 module-level code 觸發呼叫就會炸)。
# 移到 handler 之前、讓讀者從上而下能理解 error flow。
class _FwError(Exception):
"""Internal exception carrying (stage, reason, message) for firmware ops."""
def __init__(self, stage, reason, message):
super().__init__(message)
self.stage = stage
self.reason = reason
self.message = message
class _FwTimeoutError(Exception):
"""Raised when total upgrade duration exceeds chip timeout."""
def __init__(self, stage):
super().__init__(f"timeout at stage={stage}")
self.stage = stage
def _fw_handle_failure(stage, reason, message, before_fw, start_ts, dg, lib, raw=""):
"""彙整失敗 progress event + return 給 caller 的 error dict。
對齊 TDD §6.1 失敗回傳格式:
{"error":<str>, "stage":<str>, "reason":<str>, "raw_error":<str>}
Reviewer m3原本此 helper 內 disconnect、caller 的 finally 也 disconnect、
雙重 disconnect 對 SDK 行為未定。改成「single owner of disconnect」原則
本 helper 不再 disconnect、由 caller 的 finally 統一處理。本函式只負責 emit
progress event + 組裝 error dict。
"""
elapsed = int((time.monotonic() - start_ts) * 1000)
_log(f"firmware_upgrade FAILED: stage={stage}, reason={reason}, "
f"message={message}, elapsed_ms={elapsed}")
_fw_emit_progress(
"error",
message=message,
elapsed_ms=elapsed,
eta_ms=0,
extra={
"error": message,
"reason": reason,
"raw_error": raw or message,
"before_version": before_fw,
},
)
return {
"error": message,
"stage": stage,
"reason": reason,
"raw_error": raw or message,
}
def handle_firmware_upgrade(params):
"""A 階段 M9-1自動升級 KDP1 → KDP2、KL520 與 KL720。
對應 TDD §6.1 表 + §5.1 流程:
Input: {"port": "<usb_port_id>", "chip": "KL520" | "KL720"}
Output (success):
{"status":"upgraded", "before_firmware":<str>, "after_firmware":<str>,
"method":"ctypes_kp_update_kdp_firmware_from_files",
"duration_ms":<int>}
Output (failure):
{"error":<str>, "stage":<preparing|loading|flashing|verifying>,
"reason":<scan_not_found|connect_failed|loader_write_failed|
upgrade_mid_failed|disconnect_during_op|timeout|
verify_mismatch|verify_not_found>,
"raw_error":<str>}
每進入一個 stage 透過 _fw_emit_progress() 推 progress event 到 stderr
Go driver 抓 stderr line-by-line 轉成 WebSocket FirmwareProgress 給前端。
"""
global _firmware_upgrade_in_progress
if not HAS_KP:
return {"error": "kp module not available", "stage": "preparing",
"reason": "scan_not_found", "raw_error": "kp not available"}
chip = params.get("chip", "KL520")
port = str(params.get("port", ""))
if chip not in ("KL520", "KL720"):
return {"error": f"unsupported chip for A 階段: {chip}",
"stage": "preparing", "reason": "scan_not_found",
"raw_error": f"chip={chip} not in (KL520, KL720)"}
timeout_s = KL520_UPGRADE_TIMEOUT_S if chip == "KL520" else KL720_UPGRADE_TIMEOUT_S
start_ts = time.monotonic()
def elapsed_ms():
return int((time.monotonic() - start_ts) * 1000)
def check_timeout(current_stage):
if (time.monotonic() - start_ts) > timeout_s:
raise _FwTimeoutError(current_stage)
# ── AC-FW-1.9 graceful shutdown 拒絕:標記升級進行中 ──
# Reviewer m4原本還寫 _firmware_upgrade_start_ts 全域、與 SIGTERM handler
# closure 重複、已移除、改由 closure capture start_ts 為 single source。
_firmware_upgrade_in_progress = True
# 在升降版進入 critical section 期間註冊 SIGTERM handler
# (收 SIGTERM 不立即退、改 log warning event實際 server 端 lock
# 由 M9-2 Go driver / M9-3 service 實作、bridge.py 只負責「正在跑時
# 拒絕被 kill」
_fw_register_sigterm_handler(start_ts)
method = "ctypes_kp_update_kdp_firmware_from_files"
before_fw = ""
lib = None
dg = None
try:
# ── preparingscan + connect ────────────────────────────────
_fw_emit_progress(
"preparing",
message=f"scanning {chip} on port {port}",
elapsed_ms=elapsed_ms(),
eta_ms=_fw_eta_ms(chip, "preparing"),
)
check_timeout("preparing")
# 先 disconnect 既有 _device_group若有、避免 handle 衝突
_clear_device_group()
target = _fw_scan_target(port)
if target is None:
raise _FwError(
"preparing", "scan_not_found",
f"device with port_id={port} not found in scan",
)
before_fw = str(target.firmware)
target_port_id = int(target.usb_port_id)
target_pid = int(target.product_id)
_log(f"firmware_upgrade: chip={chip}, port={target_port_id}, "
f"pid=0x{target_pid:04X}, firmware='{before_fw}'")
# ── 解析 firmware 檔路徑 ─────────────────────────────────────
fw_paths = _resolve_firmware_paths_full(chip)
if fw_paths["scpu"] is None or fw_paths["ncpu"] is None:
raise _FwError(
"preparing", "scan_not_found",
f"firmware files not found for {chip} "
f"(scpu/ncpu missing in server/scripts/firmware/{chip}/)",
)
# ── 載入 libkplus + ctypes binding ──────────────────────────
try:
lib = _fw_load_libkplus()
except Exception as e:
raise _FwError(
"preparing", "connect_failed",
f"libkplus load failed: {e}",
)
# ── connect with magicallow KDP1 legacy device───────────
try:
dg = _fw_connect_with_magic(lib, target_port_id)
except RuntimeError as e:
raise _FwError("preparing", "connect_failed", str(e))
# set timeout for SDK operations注意不是整體 upgrade timeout、
# 是單一 SDK call 的 timeout、避免單個 kp_load/update call 卡住)
lib.kp_set_timeout(dg, int(timeout_s * 1000))
# ── 判斷是否走 SDK loader stage ──────────────────────────────
# Reviewer M2原本控制流隱式`if needs_loader: if loader_path is None: ...`
# nested、讀者不易看清「實際會跑 loading stage」的條件。改為三個顯式 bool
#
# needs_loader = device 處於 KDP1 legacy state_fw_classify_legacy
# should_run_loader_stage = 實際會跑 loading stageloader.bin 存在 + needs_loader
# loader_required_but_missing = KL520 KDP1 legacy 但缺 loader.bin必失敗
#
# 三個情境的流程:
# 1. KL520 KDP1 legacy + loader.bin 存在 → loading → flashing(SDK load)
# → verifying → done (should_run_loader_stage=True)
# 2. KL520 KDP1 legacy + loader.bin 缺 → fail at loading (loader_write_failed)
# 3. KL720 KDP1 legacy + loader.bin 缺 → skip loading、直接 flashing(warrenchen 模式)
# → verifying → done (should_run_loader_stage=False)
# 4. already KDP2KL520/KL720→ skip loading、直接 flashing(warrenchen 模式)
# → verifying → done (should_run_loader_stage=False)
needs_loader = _fw_classify_legacy(before_fw, target_pid)
loader_path = fw_paths["loader"]
should_run_loader_stage = needs_loader and loader_path is not None
loader_required_but_missing = (
needs_loader and loader_path is None and chip == "KL520"
)
_log(f"firmware_upgrade: needs_loader={needs_loader}, "
f"should_run_loader_stage={should_run_loader_stage}, "
f"loader_required_but_missing={loader_required_but_missing}, "
f"legacy={'yes' if needs_loader else 'no'}")
# ── 情境 2KL520 KDP1 legacy 但缺 loader.bin → 直接失敗 ─────
if loader_required_but_missing:
check_timeout("loading")
raise _FwError(
"loading", "loader_write_failed",
f"fw_loader.bin not found for {chip} but device is in "
f"KDP1 legacy state (firmware='{before_fw}')",
)
# ── 情境 1跑 loading stageKL520 KDP1 legacy + loader.bin──
if should_run_loader_stage:
check_timeout("loading")
_fw_emit_progress(
"loading",
message="writing USB Boot loader firmware",
elapsed_ms=elapsed_ms(),
eta_ms=_fw_eta_ms(chip, "loading"),
)
ret = lib.kp_update_kdp_firmware_from_files(
dg,
loader_path.encode("utf-8"),
None, # loader stage: ncpu = NULL
True, # auto_reboot
)
if ret != KP_SUCCESS:
raise _FwError(
"loading", "loader_write_failed",
f"kp_update_kdp_firmware_from_files(loader) ret={ret} "
f"({_fw_errstr(lib, ret)})",
)
# auto_reboot 後 disconnect 可能失敗USB re-enumerate容忍
try:
lib.kp_disconnect_devices(dg)
except Exception:
pass
# disconnect 完設 dg=None、避免 finally double-disconnect 已 freed handle
dg = None
# 等 device reboot 完進 USB Boot modeLoader firmware loaded
time.sleep(USB_WAIT_AFTER_REBOOT_MS / 1000.0)
# rescan + reconnect with magic
target = _fw_scan_target(port)
if target is None:
raise _FwError(
"loading", "disconnect_during_op",
f"device disappeared after loader write, port={port}",
)
try:
dg = _fw_connect_with_magic(lib, int(target.usb_port_id))
except RuntimeError as e:
raise _FwError(
"loading", "connect_failed",
f"reconnect after loader failed: {e}",
)
lib.kp_set_timeout(dg, int(timeout_s * 1000))
elif needs_loader:
# 情境 3KL720 KDP1 legacy 沒 loader.bin → 跳過 loading、直接 flashing
# warrenchen 模式kp_update_kdp_firmware_from_files(scpu, ncpu, True) 一次寫
_log(f"firmware_upgrade: {chip} legacy without loader.bin、"
f"skipping loading stage, will go directly to flashing")
# ── flashing寫入 KDP2 firmwarescpu + ncpu─────────────
check_timeout("flashing")
_fw_emit_progress(
"flashing",
message="writing KDP2 firmware (scpu + ncpu)",
elapsed_ms=elapsed_ms(),
eta_ms=_fw_eta_ms(chip, "flashing"),
)
if should_run_loader_stage:
# 情境 1device 已透過 loader stage 進 Loader mode、用
# kp_load_firmware_from_file 載 scpu + ncpu 到 RAM
ret = lib.kp_load_firmware_from_file(
dg,
fw_paths["scpu"].encode("utf-8"),
fw_paths["ncpu"].encode("utf-8"),
)
if ret != KP_SUCCESS:
raise _FwError(
"flashing", "upgrade_mid_failed",
f"kp_load_firmware_from_file ret={ret} "
f"({_fw_errstr(lib, ret)})",
)
else:
# 情境 3 / 4沒走 loader stageKL720 legacy without loader.bin、
# 或 already KDP2→ warrenchen 模式:直接
# kp_update_kdp_firmware_from_files(scpu, ncpu, True) 一次寫
ret = lib.kp_update_kdp_firmware_from_files(
dg,
fw_paths["scpu"].encode("utf-8"),
fw_paths["ncpu"].encode("utf-8"),
True, # auto_reboot
)
if ret != KP_SUCCESS:
raise _FwError(
"flashing", "upgrade_mid_failed",
f"kp_update_kdp_firmware_from_files ret={ret} "
f"({_fw_errstr(lib, ret)})",
)
# disconnect after upgradeauto_reboot 後 disconnect 失敗預期、容忍
try:
lib.kp_disconnect_devices(dg)
except Exception:
pass
dg = None
# ── verifying等 USB re-enumerate → rescan → 驗 firmware 字串 ──
check_timeout("verifying")
_fw_emit_progress(
"verifying",
message="waiting USB re-enumerate and verifying firmware version",
elapsed_ms=elapsed_ms(),
eta_ms=_fw_eta_ms(chip, "verifying"),
)
# AC-FW-1.6: 等 5-8 秒 USB stable
target_after, waited = _fw_rescan_and_wait(
port,
max_wait_s=USB_WAIT_AFTER_UPGRADE_MS / 1000.0 + 3.0, # 5 + 3 = 8s 上界
initial_sleep_s=USB_WAIT_AFTER_UPGRADE_MS / 1000.0,
)
if target_after is None:
raise _FwError(
"verifying", "verify_not_found",
f"device not found after upgrade (waited {waited:.1f}s)、"
f"USB may still be re-enumerating, please re-plug",
)
after_fw = str(target_after.firmware)
after_pid = int(target_after.product_id)
# 驗證 firmware 字串已升到 KDP2不再是 KDP1 legacy
if _fw_classify_legacy(after_fw, after_pid):
raise _FwError(
"verifying", "verify_mismatch",
f"firmware after upgrade still appears legacy: "
f"firmware='{after_fw}', pid=0x{after_pid:04X}",
)
# ── done ──
duration_ms = elapsed_ms()
_fw_emit_progress(
"done",
message=f"upgraded from '{before_fw}' to '{after_fw}'",
elapsed_ms=duration_ms,
eta_ms=0,
)
return {
"status": "upgraded",
"before_firmware": before_fw,
"after_firmware": after_fw,
"method": method,
"duration_ms": duration_ms,
}
except _FwTimeoutError as e:
return _fw_handle_failure(
e.stage, "timeout",
f"upgrade exceeded {timeout_s}s timeout at stage={e.stage}",
before_fw, start_ts, dg, lib, raw=str(e),
)
except _FwError as e:
return _fw_handle_failure(
e.stage, e.reason, e.message, before_fw, start_ts, dg, lib, raw=str(e),
)
except Exception as e:
import traceback
tb = traceback.format_exc()
_log(f"firmware_upgrade UNEXPECTED EXCEPTION: {type(e).__name__}: {e}\n{tb}")
return _fw_handle_failure(
"flashing", "upgrade_mid_failed",
f"unexpected: {type(e).__name__}: {e}",
before_fw, start_ts, dg, lib, raw=tb,
)
finally:
_firmware_upgrade_in_progress = False
# Reviewer m3disconnect 的 single owner = 此 finally block。
# _fw_handle_failure 已改為「不在裡面 disconnect」、避免 double-disconnect。
# success path 在 1810 行已 disconnect 並設 dg=None、此處 if dg is not None
# 會 short-circuit 跳過、不會 double。
# fail pathdg 可能還持有 handle、由本 finally 統一收尾。
if dg is not None and lib is not None:
try:
lib.kp_disconnect_devices(dg)
except Exception:
pass
dg = None # 確保不會被外部誤用
_fw_unregister_sigterm_handler()
# ── SIGTERM handler (AC-FW-1.9 graceful shutdown rejection) ──────────
#
# 升級進行中收到 SIGTERM 時,不立即退出、改在 stderr push warning event。
# 實際的 server-side lock 機制由 M9-2 / M9-3 實作progress.md「未解決問題」
# 註記為依賴)。本處 bridge.py 端的責任:「正在跑時拒絕被 kill」。
#
# Windows 沒有 SIGTERM 概念、改用 atexit。Linux/macOS 用 signal handler。
_fw_original_sigterm_handler = None
def _fw_register_sigterm_handler(start_ts):
"""註冊 SIGTERM handler升級進行中時拒絕並 log warning。"""
global _fw_original_sigterm_handler
if sys.platform == "win32":
return # Windows 沒 SIGTERM
try:
import signal
def handler(signum, frame):
if _firmware_upgrade_in_progress:
elapsed = int((time.monotonic() - start_ts) * 1000)
try:
print(
json.dumps({
"event": "shutdown_rejected",
"reason": "firmware_upgrade_in_progress",
"task": "firmware_upgrade",
"elapsed_ms": elapsed,
}),
file=sys.stderr,
flush=True,
)
except Exception:
pass
# 拒絕 SIGTERM不呼叫 sys.exit、不 raise、繼續執行升級
return
# 沒升級進行中、走預設行為
if callable(_fw_original_sigterm_handler):
_fw_original_sigterm_handler(signum, frame)
else:
sys.exit(0)
_fw_original_sigterm_handler = signal.signal(signal.SIGTERM, handler)
except Exception as e:
_log(f"SIGTERM handler registration failed: {e}")
def _fw_unregister_sigterm_handler():
"""還原 SIGTERM handler 為 install 前狀態。"""
global _fw_original_sigterm_handler
if sys.platform == "win32":
return
try:
import signal
if _fw_original_sigterm_handler is not None:
signal.signal(signal.SIGTERM, _fw_original_sigterm_handler)
_fw_original_sigterm_handler = None
else:
signal.signal(signal.SIGTERM, signal.SIG_DFL)
except Exception:
pass
# ── Main loop ────────────────────────────────────────────────────────
def main():
"""Main loop: read JSON commands from stdin, write responses to stdout."""
# The Kneron C SDK may write ANSI-colored warnings directly to fd 1
# (stdout), which corrupts our JSON-RPC protocol. To prevent this we
# dup the real stdout fd, then redirect fd 1 to stderr so any C-level
# writes go to stderr. Our JSON responses use the duped fd.
_real_stdout_fd = os.dup(1) # duplicate fd 1
os.dup2(2, 1) # fd 1 now points to stderr
# encoding 必須顯式指定 —— os.fdopen() 不帶 encoding 時會用
# locale.getpreferredencoding(),在繁中 Windows 上是 cp950。這個檔案
# 物件是 JSON-RPC 的**回應通道**,且它是全新建立的,不受
# _force_utf8_stdio() 對 sys.stdout 的 reconfigure 影響,所以必須在
# 這裡各自綁一次。errors 理由同 _force_utf8_stdio():寧可壞一個字元
# 也不要讓 bridge 崩潰而導致裝置失聯。
_real_stdout = os.fdopen(
_real_stdout_fd, "w", encoding="utf-8", errors="replace")
sys.stdout = sys.stderr # Python-level redirect too
def _respond(obj):
"""Write a JSON response to the real stdout (not stderr)."""
_real_stdout.write(json.dumps(obj) + "\n")
_real_stdout.flush()
# Signal readiness
_respond({"status": "ready"})
_log(f"Bridge started (kp={'yes' if HAS_KP else 'no'}, pyusb={'yes' if HAS_PYUSB else 'no'}, cv2={'yes' if HAS_CV2 else 'no'})")
for line in sys.stdin:
line = line.strip()
if not line:
continue
try:
cmd = json.loads(line)
action = cmd.get("cmd", "")
if action == "scan":
result = handle_scan()
elif action == "connect":
result = handle_connect(cmd)
elif action == "disconnect":
result = handle_disconnect(cmd)
elif action == "reset":
result = handle_reset(cmd)
elif action == "load_model":
result = handle_load_model(cmd)
elif action == "set_inference_options":
result = handle_set_inference_options(cmd)
elif action == "inference":
result = handle_inference(cmd)
elif action == "firmware_upgrade":
result = handle_firmware_upgrade(cmd)
else:
result = {"error": f"unknown command: {action}"}
_respond(result)
except Exception as e:
_respond({"error": str(e)})
def _cleanup():
"""Explicitly disconnect and clear _device_group before Python GC runs.
KneronPLUS SDK's DeviceGroup.__del__ calls kp_disconnect_devices on a
native handle that may already be freed when the interpreter is shutting
down, causing 'OSError: access violation reading 0x00...'. By doing a
clean disconnect + setting the global to None here, __del__ becomes a
no-op (None has no __del__).
"""
global _device_group
if _device_group is not None:
try:
kp.core.disconnect_devices(_device_group)
except Exception:
pass
_device_group = None
if __name__ == "__main__":
import atexit
atexit.register(_cleanup)
main()
_cleanup() # also call synchronously in case atexit doesn't fire