1.NCNN介绍
NCNN 是腾讯优图实验室开源的高性能神经网络推理框架,专为移动端和嵌入式设备设计。 NCNN是通用型推理框架,主打跨平台和轻量。类似对比这个瑞芯微专用的(RKNN是瑞芯微专用推理框架,深度绑定自家芯片, 靠NPU硬件加速换性能)。 对于需要在PC或者非瑞芯微的ARM上,NCNN是通用框架,支持ARM、x86、MIPS等多种架构,主要靠CPU(NEON指令集)或Vulkan GPU加速,没有NPU也能运行, 当然代价就是NCNN 在 CPU上运行速度,对比RKNN差异巨大。我们本次wiki记录主要就是在PC上CPU运行加载 这个yolov5训练的模型,进行测试。
2.主要流程步骤
2.1 模型转换
需要将这个yolov5 训练出的模型,转换成 NCNN能识别的。NCNN里面默认带有转换python,但是存在一个 很大的坑。
现代 yolov5 export.py 默认把 Detect 的 anchor 解码烙进图里,输出是单个已解码
的 output0[1,25200,6],但它依赖 5D 的 grid/anchor 常量 [1,3,H,W,2],
onnx2ncnn 只支持 4D 常量,转换后 MemoryData 没有形状 -> 解码层崩。
所以必须让 Detect 输出「未解码」的 3 个原始头(和官方 examples/yolov5.cpp 一致)
处理解决方法:
加载 pt 后把 model.model[-1].training 置 True,Detect.forward 就会走
`return x if self.training` 分支,输出 3 个原始头 (1, 3, ny, nx, no)。
no = nc + 5 = 6(simcard 是 1 类)。
2.2 加载图片方案
对于PC或者方便移植到嵌入式平台,我们没有使用opencv这类库,使用stb_image, 这个 主要就是将jpg图片等,变换成RGB格式 (主要opencv的MAT 颜色通道不一样,BGR)。
2.3 加载模型识别
我们还是以simcard检查这个模型做实例
// 检测主函数:stb 读出的 RGB 像素 -> 预处理 -> 推理 -> 后处理
// rgb_data : RGB 像素(3 通道连续内存),w/h : 原图宽高
static int detect_simcard(const unsigned char* rgb_data, int w, int h, std::vector<Object>& objects)
{
ncnn::Net net;
// 本机库没编 Vulkan,走 CPU;要 GPU 可开下面这行(需重新编译带 Vulkan 的库)
// net.opt.use_vulkan_compute = true;
auto t_load = std::chrono::steady_clock::now();
// 加载模型(先用文件路径;嵌入式上可改 load_param/load_model(const unsigned char* mem))
if (net.load_param(MODEL_PARAM) != 0)
{
fprintf(stderr, "load_param %s failed\n", MODEL_PARAM);
return -1;
}
if (net.load_model(MODEL_BIN) != 0)
{
fprintf(stderr, "load_model %s failed\n", MODEL_BIN);
return -1;
}
fprintf(stdout, "[time] load param+bin : %8.2f ms\n", elapsed_ms(t_load));
const int target_size = 640;
const float prob_threshold = 0.25f;
const float nms_threshold = 0.45f;
auto t_pre = std::chrono::steady_clock::now();
// letterbox 等比缩放:长边缩到 640,短边按同比例缩放
int img_w = w;
int img_h = h;
float scale = 1.f;
int rw, rh;
if (img_w > img_h)
{
scale = (float)target_size / img_w;
rw = target_size;
rh = (int)(img_h * scale);
}
else
{
scale = (float)target_size / img_h;
rh = target_size;
rw = (int)(img_w * scale);
}
// stb 输出是 RGB,格式直接填 PIXEL_RGB
ncnn::Mat in = ncnn::Mat::from_pixels_resize(rgb_data, ncnn::Mat::PIXEL_RGB, img_w, img_h, rw, rh);
// 补边到训练尺寸 640x640(YOLOv5 灰边 114)。注意:不是补到 64 的倍数,
// 是补满到 640x640。若补到 512x640 这种异形尺寸,模型输出网格就错了。
int wpad = target_size - rw;
int hpad = target_size - rh;
ncnn::Mat in_pad;
ncnn::copy_make_border(in, in_pad, hpad / 2, hpad - hpad / 2, wpad / 2, wpad - wpad / 2,
ncnn::BORDER_CONSTANT, 114.f);
// 归一化:/255(YOLOv5 无减均值,mean 传 0)
const float norm_vals[3] = {1 / 255.f, 1 / 255.f, 1 / 255.f};
in_pad.substract_mean_normalize(0, norm_vals);
fprintf(stdout, "[time] preprocess : %8.2f ms\n", elapsed_ms(t_pre));
// 推理
auto t_infer = std::chrono::steady_clock::now();
ncnn::Extractor ex = net.create_extractor();
ex.input(INPUT_NAME, in_pad);
std::vector<Object> proposals;
// stride 8 输出(小目标)
{
ncnn::Mat out;
ex.extract(OUT_STRIDE8, out);
ncnn::Mat anchors(6);
anchors[0] = 10.f; anchors[1] = 13.f;
anchors[2] = 16.f; anchors[3] = 30.f;
anchors[4] = 33.f; anchors[5] = 23.f;
std::vector<Object> objs8;
generate_proposals(anchors, 8, in_pad, out, prob_threshold, objs8);
proposals.insert(proposals.end(), objs8.begin(), objs8.end());
}
// stride 16 输出(中目标)
{
ncnn::Mat out;
ex.extract(OUT_STRIDE16, out);
ncnn::Mat anchors(6);
anchors[0] = 30.f; anchors[1] = 61.f;
anchors[2] = 62.f; anchors[3] = 45.f;
anchors[4] = 59.f; anchors[5] = 119.f;
std::vector<Object> objs16;
generate_proposals(anchors, 16, in_pad, out, prob_threshold, objs16);
proposals.insert(proposals.end(), objs16.begin(), objs16.end());
}
// stride 32 输出(大目标)
{
ncnn::Mat out;
ex.extract(OUT_STRIDE32, out);
ncnn::Mat anchors(6);
anchors[0] = 116.f; anchors[1] = 90.f;
anchors[2] = 156.f; anchors[3] = 198.f;
anchors[4] = 373.f; anchors[5] = 326.f;
std::vector<Object> objs32;
generate_proposals(anchors, 32, in_pad, out, prob_threshold, objs32);
proposals.insert(proposals.end(), objs32.begin(), objs32.end());
}
fprintf(stdout, "[time] inference : %8.2f ms\n", elapsed_ms(t_infer));
// 排序 + NMS
auto t_post = std::chrono::steady_clock::now();
qsort_descent_inplace(proposals);
std::vector<int> picked;
nms_sorted_bboxes(proposals, picked, nms_threshold);
// 把框坐标从「补边+缩放后的图」还原回原图
int count = (int)picked.size();
objects.resize(count);
for (int i = 0; i < count; i++)
{
const Object& p = proposals[picked[i]];
float x0 = (p.x - wpad / 2.f) / scale;
float y0 = (p.y - hpad / 2.f) / scale;
float x1 = (p.x + p.w - wpad / 2.f) / scale;
float y1 = (p.y + p.h - hpad / 2.f) / scale;
// 裁剪到图内
x0 = std::max(0.f, std::min(x0, (float)(img_w - 1)));
y0 = std::max(0.f, std::min(y0, (float)(img_h - 1)));
x1 = std::max(0.f, std::min(x1, (float)(img_w - 1)));
y1 = std::max(0.f, std::min(y1, (float)(img_h - 1)));
objects[i].x = x0;
objects[i].y = y0;
objects[i].w = x1 - x0;
objects[i].h = y1 - y0;
objects[i].label = p.label;
objects[i].prob = p.prob;
}
fprintf(stdout, "[time] postprocess : %8.2f ms\n", elapsed_ms(t_post));
return 0;
}
3.测试效果层面
我们在实际pc上测试一下这个效果,根据每次消耗时间看一下主要瓶颈:
dong@dong:/dong/app-workspace/gs_workspace/ncnn/demo/build$ ./simcard_demo ../simcard-test.jpg
image ../simcard-test.jpg loaded: 3456x4608 (orig channels=3)
[time] load param+bin : 371.88 ms
[time] preprocess : 8.73 ms
[time] inference : 287.64 ms
[time] postprocess : 0.01 ms
detected 2 objects:
[0] simcard 81.52% x=1561.86 y=2666.96 w=895.51 h=735.07
[0] simcard 54.27% x=3014.08 y=382.85 w=440.92 h=452.41
[time] total : 1130.23 ms
实际测试过程中 load parma+bin 较耗时,优化时,对于多轮检测,这个过程建议放到初始化就完成。 对于检测的置信度,里面可以设置阈值。
您还没有登录,请您登录后发表评论。