Yolov5 ncnn本地识别

1.NCNN介绍

NCNN 是腾讯优图实验室开源的高性能神经网络推理框架,专为移动端和嵌入式设备设计。 NCNN是通用型推理框架,主打跨平台和轻量。类似对比这个瑞芯微专用的(RKNN是瑞芯微专用推理框架,深度绑定自家芯片, 靠NPU硬件加速换性能)。 对于需要在PC或者非瑞芯微的ARM上,NCNN是通用框架,支持ARM、x86、MIPS等多种架构,主要靠CPU(NEON指令集)或Vulkan GPU加速,没有NPU也能运行, 当然代价就是NCNN 在 CPU上运行速度,对比RKNN差异巨大。我们本次wiki记录主要就是在PC上CPU运行加载 这个yolov5训练的模型,进行测试。

2.主要流程步骤

2.1 模型转换

需要将这个yolov5 训练出的模型,转换成 NCNN能识别的。NCNN里面默认带有转换python,但是存在一个 很大的坑。

现代 yolov5 export.py 默认把 Detect 的 anchor 解码烙进图里,输出是单个已解码
    的 output0[1,25200,6],但它依赖 5D 的 grid/anchor 常量 [1,3,H,W,2],
    onnx2ncnn 只支持 4D 常量,转换后 MemoryData 没有形状 -> 解码层崩。
    所以必须让 Detect 输出「未解码」的 3 个原始头(和官方 examples/yolov5.cpp 一致)

处理解决方法:

加载 pt 后把 model.model[-1].training 置 True,Detect.forward 就会走
    `return x if self.training` 分支,输出 3 个原始头 (1, 3, ny, nx, no)。
    no = nc + 5 = 6(simcard 是 1 类)。

2.2 加载图片方案

对于PC或者方便移植到嵌入式平台,我们没有使用opencv这类库,使用stb_image, 这个 主要就是将jpg图片等,变换成RGB格式 (主要opencv的MAT 颜色通道不一样,BGR)。

2.3 加载模型识别

我们还是以simcard检查这个模型做实例

// 检测主函数:stb 读出的 RGB 像素 -> 预处理 -> 推理 -> 后处理
// rgb_data : RGB 像素(3 通道连续内存),w/h : 原图宽高
static int detect_simcard(const unsigned char* rgb_data, int w, int h, std::vector<Object>& objects)
{
    ncnn::Net net;

    // 本机库没编 Vulkan,走 CPU;要 GPU 可开下面这行(需重新编译带 Vulkan 的库)
    // net.opt.use_vulkan_compute = true;

    auto t_load = std::chrono::steady_clock::now();

    // 加载模型(先用文件路径;嵌入式上可改 load_param/load_model(const unsigned char* mem))
    if (net.load_param(MODEL_PARAM) != 0)
    {
        fprintf(stderr, "load_param %s failed\n", MODEL_PARAM);
        return -1;
    }
    if (net.load_model(MODEL_BIN) != 0)
    {
        fprintf(stderr, "load_model %s failed\n", MODEL_BIN);
        return -1;
    }
    fprintf(stdout, "[time] load param+bin   : %8.2f ms\n", elapsed_ms(t_load));

    const int target_size = 640;
    const float prob_threshold = 0.25f;
    const float nms_threshold = 0.45f;

    auto t_pre = std::chrono::steady_clock::now();

    // letterbox 等比缩放:长边缩到 640,短边按同比例缩放
    int img_w = w;
    int img_h = h;
    float scale = 1.f;
    int rw, rh;
    if (img_w > img_h)
    {
        scale = (float)target_size / img_w;
        rw = target_size;
        rh = (int)(img_h * scale);
    }
    else
    {
        scale = (float)target_size / img_h;
        rh = target_size;
        rw = (int)(img_w * scale);
    }

    // stb 输出是 RGB,格式直接填 PIXEL_RGB
    ncnn::Mat in = ncnn::Mat::from_pixels_resize(rgb_data, ncnn::Mat::PIXEL_RGB, img_w, img_h, rw, rh);

    // 补边到训练尺寸 640x640(YOLOv5 灰边 114)。注意:不是补到 64 的倍数,
    // 是补满到 640x640。若补到 512x640 这种异形尺寸,模型输出网格就错了。
    int wpad = target_size - rw;
    int hpad = target_size - rh;
    ncnn::Mat in_pad;
    ncnn::copy_make_border(in, in_pad, hpad / 2, hpad - hpad / 2, wpad / 2, wpad - wpad / 2,
                           ncnn::BORDER_CONSTANT, 114.f);

    // 归一化:/255(YOLOv5 无减均值,mean 传 0)
    const float norm_vals[3] = {1 / 255.f, 1 / 255.f, 1 / 255.f};
    in_pad.substract_mean_normalize(0, norm_vals);
    fprintf(stdout, "[time] preprocess       : %8.2f ms\n", elapsed_ms(t_pre));

    // 推理
    auto t_infer = std::chrono::steady_clock::now();
    ncnn::Extractor ex = net.create_extractor();
    ex.input(INPUT_NAME, in_pad);

    std::vector<Object> proposals;

    // stride 8 输出(小目标)
    {
        ncnn::Mat out;
        ex.extract(OUT_STRIDE8, out);

        ncnn::Mat anchors(6);
        anchors[0] = 10.f; anchors[1] = 13.f;
        anchors[2] = 16.f; anchors[3] = 30.f;
        anchors[4] = 33.f; anchors[5] = 23.f;

        std::vector<Object> objs8;
        generate_proposals(anchors, 8, in_pad, out, prob_threshold, objs8);
        proposals.insert(proposals.end(), objs8.begin(), objs8.end());
    }

    // stride 16 输出(中目标)
    {
        ncnn::Mat out;
        ex.extract(OUT_STRIDE16, out);

        ncnn::Mat anchors(6);
        anchors[0] = 30.f; anchors[1] = 61.f;
        anchors[2] = 62.f; anchors[3] = 45.f;
        anchors[4] = 59.f; anchors[5] = 119.f;

        std::vector<Object> objs16;
        generate_proposals(anchors, 16, in_pad, out, prob_threshold, objs16);
        proposals.insert(proposals.end(), objs16.begin(), objs16.end());
    }

    // stride 32 输出(大目标)
    {
        ncnn::Mat out;
        ex.extract(OUT_STRIDE32, out);

        ncnn::Mat anchors(6);
        anchors[0] = 116.f; anchors[1] = 90.f;
        anchors[2] = 156.f; anchors[3] = 198.f;
        anchors[4] = 373.f; anchors[5] = 326.f;

        std::vector<Object> objs32;
        generate_proposals(anchors, 32, in_pad, out, prob_threshold, objs32);
        proposals.insert(proposals.end(), objs32.begin(), objs32.end());
    }
    fprintf(stdout, "[time] inference        : %8.2f ms\n", elapsed_ms(t_infer));

    // 排序 + NMS
    auto t_post = std::chrono::steady_clock::now();
    qsort_descent_inplace(proposals);

    std::vector<int> picked;
    nms_sorted_bboxes(proposals, picked, nms_threshold);

    // 把框坐标从「补边+缩放后的图」还原回原图
    int count = (int)picked.size();
    objects.resize(count);
    for (int i = 0; i < count; i++)
    {
        const Object& p = proposals[picked[i]];

        float x0 = (p.x - wpad / 2.f) / scale;
        float y0 = (p.y - hpad / 2.f) / scale;
        float x1 = (p.x + p.w - wpad / 2.f) / scale;
        float y1 = (p.y + p.h - hpad / 2.f) / scale;

        // 裁剪到图内
        x0 = std::max(0.f, std::min(x0, (float)(img_w - 1)));
        y0 = std::max(0.f, std::min(y0, (float)(img_h - 1)));
        x1 = std::max(0.f, std::min(x1, (float)(img_w - 1)));
        y1 = std::max(0.f, std::min(y1, (float)(img_h - 1)));

        objects[i].x = x0;
        objects[i].y = y0;
        objects[i].w = x1 - x0;
        objects[i].h = y1 - y0;
        objects[i].label = p.label;
        objects[i].prob = p.prob;
    }
    fprintf(stdout, "[time] postprocess      : %8.2f ms\n", elapsed_ms(t_post));

    return 0;
}

3.测试效果层面

我们在实际pc上测试一下这个效果,根据每次消耗时间看一下主要瓶颈:

dong@dong:/dong/app-workspace/gs_workspace/ncnn/demo/build$ ./simcard_demo ../simcard-test.jpg
image ../simcard-test.jpg loaded: 3456x4608 (orig channels=3)
[time] load param+bin   :   371.88 ms
[time] preprocess       :     8.73 ms
[time] inference        :   287.64 ms
[time] postprocess      :     0.01 ms
detected 2 objects:
  [0] simcard    81.52%  x=1561.86 y=2666.96 w=895.51 h=735.07
  [0] simcard    54.27%  x=3014.08 y=382.85 w=440.92 h=452.41
[time] total            :  1130.23 ms

实际测试过程中 load parma+bin 较耗时,优化时,对于多轮检测,这个过程建议放到初始化就完成。 对于检测的置信度,里面可以设置阈值。