HOME
Shop
  • English
  • 简体中文
HOME
Shop
  • English
  • 简体中文
  • Product Series

    • FPGA+ARM

      • GM-3568JHF

        • Introduction

          • GM-3568JHF Introduction
        • Quick Start

          • Preface
          • Environment Setup
          • Compilation Notes
          • Flashing Guide
          • Debugging Tools
          • Software Update
          • Viewing System Information
          • Test Commands
          • Application Compilation
          • Source Code Access
        • Peripherals & Interfaces

          • USB
          • Display and Touch
          • Ethernet
          • WIFI
          • Bluetooth
          • TF-Card
          • Audio
          • Serial Port
          • CAN
          • RTC
        • Application Development

          • UART Read/Write Demo
          • Key Detection Demo
          • LED Blink Demo
          • MIPI Screen Detection Demo
          • Read USB Device Information Demo
          • FAN Detection Demo
          • FPGA FSPI Communication Demo
          • FPGA DMA Read/Write Demo
          • GPS Debugging Demo
          • Ethernet Test Demo
          • RS485 Read/Write Demo
          • FPGA I2C Read/Write Demo
          • PN532 NFC Card-Reading Demo
          • TF Card Read/Write Demo
        • QT Development

          • ARM64 Cross-Compiler Environment Setup
          • Adding a QT Program to Boot Auto-Start
        • RKNN_NPU Development

          • RK3568 NPU Overview
          • Development Environment Setup
          • Run the Official YOLOv5 Example
        • FPGA Development

          • ARM and FPGA Communication
          • FPGA Development Manual
        • Others

          • Modifying the Root Filesystem
          • System Auto-Start Services
        • Downloads

          • Downloads
      • MB-E30P

        • Introduction

          • MB-E30P Introduction
        • Quick Start

          • Preface
          • Environment Setup
          • Compilation Instructions
          • Flashing Guide
          • Debugging Tools
          • Software Update
          • Viewing Information
          • Test Commands
          • Application Compilation
          • Source Code Acquisition
        • Peripherals & Interfaces

          • USB
          • Display and Touch
          • Ethernet
          • WIFI
          • Bluetooth
          • TF-Card
          • Audio
          • RTC
        • Application Development

          • Key Detection Demo
          • LED Blink Demo
          • MIPI Screen Detection Demo
          • Read USB Device Information Demo
          • FAN Detection Demo
          • FPGA FSPI Communication Demo
          • FPGA DMA Read/Write Demo
          • Ethernet Test Demo
          • FPGA IIC Read/Write Demo
          • PN532 NFC Card Reading Demo
          • TF Card Read/Write Demo
        • QT Development

          • ARM64 Cross-Compiler Environment Setup
          • Adding a QT Program to the Boot Auto-Start Service
        • RKNN_NPU Development

          • RK3568 NPU Overview
          • Development Environment Setup
          • Run the Official YOLOv5 Example
          • Model Conversion In Detail
          • Run Custom Models on the Board
        • FPGA Development

          • ARM and FPGA Communication
          • FPGA Development Manual
        • Others

          • Modifying the Root Filesystem
          • System Auto-Start Service
        • Downloads

          • Downloads
    • ShimetaPi

      • M4-R1

        • Introduction

          • M4-R1 Introduction
        • Quick Start

          • OpenHarmony Overview
          • Image Burning
          • Application Development Quick Start
          • Device Development Quick Start
        • Application Development

          • ArkUI

            • ArkTS Language Overview
            • UI Components - Row Container Introduction
            • UI Components - Column Container Introduction
            • UI Components - Text Component
            • UI Components - Toggle Component
            • UI Components - Slider Component
            • UI Components - Animation Component & Transition Component
          • Documentation

            • OpenHarmony Official Materials
          • Development Notes

            • Full-SDK Replacement Tutorial
            • Introducing and Using Third-Party Libraries
            • HDC Debugging
            • Restore Factory Mode via Command Line
            • Upgrade App to System Permission
          • First App

            • Build Your First ArkTS Application - HelloWorld
          • Demos

            • Serial-Debug-Assistant Application Demo
            • Writing-Board Application Demo
            • Digital Clock Application Demo
            • Wi-Fi Information Acquisition Application Demo
        • Device Development

          • Ubuntu Development

            • Environment Setup
            • Download Source Code
            • Compile Source Code
          • DevEco Device Tool

            • Tool Introduction
            • Development Environment Construction
            • Import the SDK
            • HUAWEI DevEco Tool Function Introduction
        • Kernel Peripherals & Interfaces

          • Guide
          • Device Tree Introduction
          • NAPI Introduction
          • ArkTS Introduction
          • NAPI Development Hands-on Demo
          • GPIO Introduction
          • I2C Communication
          • SPI Communication
          • PWM Control
          • UART Communication
          • TF Card (MicroSD)
          • Screen (Display)
          • Touch
          • Ethernet
          • M.2 SSD
          • Audio
          • WIFI & BT
          • Camera
        • Downloads

          • Downloads
      • M5-R1

        • Introduction

          • M5-R1 Development Docs
        • Quick Start

          • Image Burning
          • Environment Setup
          • Download Source Code
        • Peripherals & Interfaces

          • Raspberry Pi Interfaces
          • GPIO Interface
          • I2C Interface
          • SPI Communication
          • PWM Control
          • Serial Port Communication
          • TF Card
          • Display
          • Touch
          • Audio
          • RTC
          • Ethernet
          • M.2
          • MINI-PCIE
          • Camera
          • WIFI & BT
        • Downloads

          • Downloads
      • Pico-G1

        • Product Overview

          • Product Introduction
          • SDK Version Information
        • Quick Start

          • Development Environment Setup
          • Image Build
          • Image Flashing
          • System Login
          • Network Configuration
          • File Transfer
          • SDK Directory Structure
          • Deploying Your First Application
          • Deploying Your First Driver
          • Mounting an SD Card
        • Peripherals & Interfaces

          • GPIO Control
          • UART Serial Communication
          • I2C Communication
          • SPI Communication
        • MPP Media Development

          • MPP Media Processing Software
          • Image Processing Chain
          • Video Input
          • Image Encoding
        • NPU & AI

          • NPU Driver and Runtime Library Architecture
          • .xmm Model Loading
          • SVP Video Processing
          • AI Noise Reduction (AI_NR)
        • Application Samples

          • Encryption/Decryption Application
          • ADC Acquisition Application
          • Low-Power Application
          • Audio Processing Application
          • Video Encoding Application
          • Video Input Application
          • Video Graphics Subsystem (VGS) Application
          • 08 Region Overlay Application
          • 09 Intelligent Video Engine Application
          • 10 UVC Webcam Application
          • 11 All-in-One Quickstart Application
          • 12 FPN Correction Application
          • 13 Regional Motion Detection Application
          • 14 MTCNN Face Detection Application
        • Expansion Board Peripheral Examples

          • 00 - Pico Expansion Board Peripheral Examples Overview
          • 01 - OLED Display Application
          • 02 - TFT Display Application
          • 03 - MPU6050 Gyroscope Application
          • 04 - ADC Acquisition Application
          • 05 - Passive Buzzer Application
          • 06 - MQ Gas Sensor Application
          • 07 - GPS Positioning Application
          • 08 - SHT20 Temperature & Humidity Application
          • 09 - Ultrasonic Ranging Application
          • 10 - SpO2 Sensor Application
          • 11 - DC Motor Control Application
          • 12 - Servo Control Application
    • OpenHarmony

      • SC-3568HA

        • Introduction

          • SC-3568HA Overview
        • Quick Start Guide

          • OpenHarmony Overview
          • Image Flashing
          • Setting Up the Development Environment
          • Hello World Application and Deployment
        • Application Development

          • ArkUI

            • Introduction to ArkTS Language
            • Introduction to UI Components and Practical Applications (Part 1)
            • Introduction to UI Components and Practical Applications (Part 2)
            • Introduction to UI Components and Practical Applications (Part 3)
          • Expand

            • Getting Started Guide
            • Referencing and Using Third-Party Libraries
            • Application Compilation and Deployment
            • Command-Line Factory Reset
            • System Debugging -- HDC Debugging
            • APP Stability Testing
            • Chapter 7 Application Testing
        • Device Development

          • Environment Setup
          • Download Source Code
          • Compiling Source Code
        • Peripheral And Interface

          • Raspberry Pi interface
          • GPIO Interface
          • I2C Interface
          • SPI communication
          • PWM (Pulse Width Modulation) control
          • Serial port communication
          • TF Card
          • Display Screen
          • Touch
          • Audio
          • RTC
          • Ethernet
          • M.2
          • MINI-PCIE
          • Camera
          • WIFI&BT
          • Raspberry Pi expansion board
        • Downloads

          • Downloads
      • M-K1HSE

        • Introduction

          • M-K1HSE Introduction
        • Quick Start

          • Development environment construction
          • Source code acquisition
          • Compilation Notes
          • Burning Guide
        • Application Development

          • Application Development Environment Setup
          • First Application - Hello World
        • Peripherals and interfaces

          • 01 Audio
          • 02 RS485
          • 03 Display
        • System customization development

          • System transplant
          • System customization
          • Driver Development
          • System Debugging
          • OTA Update
        • Downloads

          • Downloads
    • HVS Camera

      • Quick Start

        • SDK Overview
        • Downloads
        • Your First C++ Program
        • Python Data Analysis
        • MultiVision Studio
      • Development

        • Programming Guides

          • Open Camera
          • Read Events
          • Recording & Replay
          • Event Processing (Denoising)
          • Display & Visualization
          • Tuning
          • Capture APS Image
        • Toolkit SDK

          • Hybrid Vision Toolkit
          • Quick Start
          • C++ API
          • Python API
        • Algorithm

          • Hybrid Vision Algo
          • Hybrid Vision Algo API
          • Windows Algo SDK
        • Samples Overview
        • Applications
      • Fundamentals

        • Event Camera Fundamentals
        • HVS Hybrid Vision
        • Event Visualization
        • Data Formats Reference
        • Glossary
        • Bias & Tuning
        • Video Tutorials
      • USB Cameras

        • HVS Camera Quick Start
        • Networking Capabilities

          • HVS Camera System Architecture
          • EVS Network Server
          • EVS Time Sync
          • Web Window
        • HVS Camera Compatibility Matrix
        • FAQ & Troubleshooting Guide
        • Products

          • CF-NRS1 (Lingguang No.1 Hybrid Vision Camera)
      • MIPI Modules

        • MIPI Module Quick Start
        • Carrier Boards

          • RDK X5 Carrier Board Adaptation
          • Raspberry Pi Carrier Board Adaptation
          • Digua Pi Carrier Board Adaptation
          • ShimeTai Board Carrier Board Adaptation
        • MIPI Module Compatibility Matrix
        • Products

          • EVS_003 Sensor Module
    • AI-model

      • 1684XB-32T

        • Introduction

          • AIBOX-1684XB-32 Introduction
        • Quick Start

          • First Use
          • Network Configuration
          • Disk Usage
          • Memory Allocation
          • Fan Control Strategy
          • Firmware Upgrade
          • Cross Compilation
          • Model Quantization
        • Application Development

          • Development Overview

            • Sophgo SDK Development
            • Sophgo Demo Introduction
          • Large Language Models

            • Deploying Llama3 Example
            • Sophon LLM_api_server Development
            • Deploying MiniCPM-V-2_6
            • Qwen-2-5-VL Image and Video Recognition Demo
            • Qwen3-chat Demo
            • Qwen3-Qwen Agent-MCP Development
            • Qwen3-langchain-AI Agent
          • Deep Learning

            • ResNet (Image Classification)
            • LPRNet (License Plate Recognition)
            • SAM (General Image Segmentation Foundation Model)
            • YOLOv5 (Object Detection)
            • OpenPose (Human Keypoint Detection)
            • PP-OCR (Optical Character Recognition)
        • Downloads

          • Downloads
      • 1684X-416T

        • Introduction

          • AIBOX-1684X-416 Introduction
        • Demo Quick Guide

          • ShimeTai Intelligent Monitoring Demo Quick Usage Guide
      • RDK-X5

        • Introduction

          • RDK-X5 Hardware Introduction
        • Quick Start

          • RDK-X5 Quick Start
        • Application Development

          • AI Online Model Development

            • Experiment 01 - Access Volcengine Doubao AI
            • Experiment 02 - Image Analysis
            • Experiment 03 - Multimodal Visual Analysis & Localization
            • Experiment 04 - Multimodal Image-Text Comparison
            • Experiment 05 - Multimodal Document/Table Analysis
            • Experiment 06 - Camera-based AI Visual Analysis
          • Large Language Models

            • Experiment 01 - Speech Recognition
            • Experiment 02 - Voice Conversation
            • Experiment 03 - Multimodal Image Analysis - Voice
            • Experiment 04 - Multimodal Image Comparison - Voice
            • Experiment 05 - Multimodal Document Analysis - Voice
            • Experiment 06 - Multimodal Vision Application - Voice
          • ROS2 Basics

            • Experiment 01 - Environment Setup
            • Experiment 02 - Create & Build a Workspace Package
            • Experiment 03 - Run ROS2 Topic Communication Node
            • Experiment 04 - ROS2 Camera Application
          • 40-pin IO Development

            • Experiment 01 - GPIO Output (LED Blink)
            • Experiment 02 - GPIO Input
            • Experiment 03 - Button-controlled LED
            • Experiment 04 - PWM Output
            • Experiment 05 - Serial Output
            • Experiment 06 - I2C Experiment
            • Experiment 07 - SPI Experiment
          • USB Module Usage

            • Experiment 01 - USB Voice Module Usage
            • Experiment 02 - Sound Source Localization Module
          • Machine Vision Practice

            • Experiment 01 - Open USB Camera
            • Experiment 02 - Color Recognition
            • Experiment 03 - Gesture Recognition
            • Experiment 04 - YOLOv5 Object Detection
      • RDK-S100

        • Introduction

          • RDK-S100 Hardware Introduction
        • Quick Start

          • RDK-S100 Quick Start
        • Application Development

          • AI Online Model Development

            • Experiment 01 - Access Volcengine Doubao AI
            • Experiment 02 - Image Analysis
            • Experiment 03 - Multimodal Visual Analysis & Localization
            • Experiment 04 - Multimodal Image-Text Comparison
            • Experiment 05 - Multimodal Document/Table Analysis
            • Experiment 06 - Camera-based AI Visual Analysis
          • Large Language Models

            • Experiment 01 - Speech Recognition
            • Experiment 02 - Voice Conversation
            • Experiment 03 - Multimodal Image Analysis - Voice
            • Experiment 04 - Multimodal Image Comparison - Voice
            • Experiment 05 - Multimodal Document Analysis - Voice
            • Experiment 06 - Multimodal Vision Application - Voice
          • ROS2 Basics

            • Experiment 01 - Environment Setup
            • Experiment 02 - Create & Build a Workspace Package
            • Experiment 03 - Run ROS2 Topic Communication Node
            • Experiment 04 - ROS2 Camera Application
          • 40-pin IO Development

            • Experiment 01 - GPIO Output (LED Blink)
            • Experiment 02 - GPIO Input
            • Experiment 03 - Button-controlled LED
            • Experiment 04 - PWM Output
            • Experiment 05 - Serial Output
            • Experiment 06 - I2C Experiment
            • Experiment 07 - SPI Experiment
          • USB Module Usage

            • Experiment 01 - USB Voice Module Usage
            • Experiment 02 - Sound Source Localization Module
          • Machine Vision Practice

            • Experiment 01 - Open USB Camera
            • Experiment 02 - Image Processing Basics
            • Experiment 03 - Object Detection
            • Experiment 04 - Image Segmentation
      • RK1828

        • Introduction

          • M5-182X-A1 AI Edge Box - Product Introduction
          • M5-182X-A1 Hardware Specifications
          • M5-182X-A1 Usage & Safety
        • Quick Start

          • M5-182X-A1 Image Flashing
          • RK182X Hardware Installation & Verification
          • RK182X Development Environment Quick Setup
          • RK182X SDK Overview
          • RK182X Environment Setup in Detail
          • RK182X Quick Start
          • Vendor SDK Data Extraction Record
        • Development Guide

          • ClawChips Architecture and Principles
          • SKILL User Manual
          • RK182X Series LLM Inference (RK1828 Model)
          • RK182X Series CNN Inference (RK1828 Model)
          • Model Conversion
          • RK182X AI Agent Application Development Guide
          • RK182X Industrial Anomaly Detection Application
        • SDK Reference

          • RKNN3-SDK Overview

            • RKNN3 SDK Overview
          • RKNN3-Toolkit

            • RKNN3 Toolkit Installation and Usage
          • RKLLM

            • RKLLM On-Device LLM Inference
          • RK182X Series NPU Overview and Architecture (RK1828 Model)
          • RK182X INT8 Quantized Inference Deployment
          • RK182X MPP Multimedia Framework
          • MPP Details

            • RK182X Video Decoding
            • RK182X Video Encoding
          • NPU Details

            • RKNN Model Conversion
            • RK182X NPU INT8 Quantized Inference
            • RK182X Multi-Model Parallel Inference
          • RGA Details

            • RK182X RGA 2D Graphics Acceleration
          • VPU Details

            • RK182X VPU Codec
        • Hardware Reference

          • RK182X Series Hardware Architecture Overview (RK1828 Model)
          • RK182X Pin Definitions and Multiplexing Configuration
          • RK182X Pin Definitions
          • RK182X Power Management
          • RK182X Clock and PLL Configuration
          • RK182X Clock and Frequency Configuration
        • Tutorials

          • Hello World
          • Hello RK1828 - The First Program
          • RTSP Streaming
          • RTSP Streaming + AI Analysis
          • ShiMetaPi AI Lobster One-Click Deployment
          • PaddleOCR-VL Text Recognition
          • Qwen3-1.7B LLM Text Chat
          • AI Multi-View Inspection (Qwen3-VL Wrapper)
          • YOLOv5 Object Detection
        • Downloads

          • Downloads
        • FAQ

          • FAQ
    • Core-Board

      • C-3568BQ

        • Introduction

          • C-3568BQ Overview
      • C-3588LQ

        • Introduction

          • C-3588LQ Overview
      • GC-3568JBAF

        • Introduction

          • GC-3568JBAF Overview
      • C-K1BA

        • Introduction

          • C-K1BA Overview
    • Software Platform

      • ShiMetaPi Workbench

        • Introduction

          • Product Overview
          • Core Architecture
          • Feature Entries
          • Supported Hardware
          • Release Notes
        • Quick Start

          • Install & Login
          • Connect the Device
          • Set Up the Environment
          • Connect to AIHub
          • First Inference
        • User Guide

          • Workspace Overview
          • Device Manager
          • Model Market
          • One-Click Deploy
          • Vision — SVP
          • Vision - Custom Models
          • shimeta-py IDE
          • Terminal
          • Agent Debug Assistant
          • Settings and Resources
        • FAQ

          • Installation & Login
          • Device Connection
          • Models & Deployment
          • Vision & Runtime
          • Settings & Other
      • ShimetaPi Repository

        • Introduction

          • ShimetaPi Software Repository
        • Pico G1 (GK7206)

          • Quick Start

            • Installation & First Inference
            • shimeta_infer — Image Inference
            • shimeta_camera — Real-time Camera Inference
            • SVP Scene Detection
            • File Transfer & Built-in Model Reference
            • FAQ
          • HTTP API & Python SDK

            • HTTP API Reference
      • Model Fine-tuning Platform

        • Introduction

          • Model Training Platform
        • Quick Start

          • Register & Login
          • Create Your First Model (30-Minute Quick Experience)
        • Training Guide

          • Data Preparation & Annotation
          • Training Parameter Configuration
          • Start & Monitor Training
          • Model Evaluation & Testing
        • Model Deployment

          • Export Model
          • Deploy to Edge Device

Model Conversion In Detail

4.1 Basic Concepts of Model Conversion

What is Model Conversion

Definition of Model Conversion

Model conversion is the process of converting a trained deep-learning model from one format to another. In RK3568 NPU development, the main task is converting models from common deep-learning frameworks (such as PyTorch, TensorFlow, ONNX, etc.) into the RKNN format, so they can run efficiently on the Rockchip NPU.

Necessity of Conversion

原始模型 (PyTorch/TensorFlow/ONNX)
    ↓
模型优化 (图优化、算子融合)
    ↓
量化处理 (FP32 → INT8/INT16)
    ↓
硬件适配 (NPU指令集映射)
    ↓
RKNN模型 (可在RK3568 NPU运行)

Main purposes of conversion:

  • Hardware adaptation: Adapt a general model to specific NPU hardware
  • Performance optimization: Improve inference speed through graph optimization and operator fusion
  • Memory optimization: Reduce model size and runtime memory usage
  • Quantization acceleration: Quantize an FP32 model to INT8 to speed up inference

Supported Model Formats

Input Format Support

FrameworkFormatSupported versionsNotes
ONNX.onnx1.6-1.12Recommended format, best compatibility
TensorFlow.pb1.x, 2.xRequires frozen graph
TensorFlow Lite.tflite2.xLightweight model
Caffe.prototxt + .caffemodel1.0Classic framework
DarkNet.cfg + .weights-YOLO series models

Recommended Conversion Paths

PyTorch → ONNX → RKNN (推荐)
TensorFlow → ONNX → RKNN (推荐)
TensorFlow → TensorFlow Lite → RKNN
Caffe → RKNN (直接转换)

Quantization Techniques In Detail

Quantization Type Comparison

Quantization typePrecisionSpeedModel sizeApplicable scenarios
FP32HighestSlowLargeVery high precision requirements
FP16HighMediumMediumBalance precision and performance
INT8MediumFastSmallMost application scenarios
Mixed precisionHighFastSmallKeep key layers at high precision

Quantization Strategies

# 对称量化 vs 非对称量化
symmetric_quantization = {
    "range": "[-127, 127]",
    "zero_point": 0,
    "advantages": "计算简单,硬件友好",
    "disadvantages": "可能浪费量化范围"
}

asymmetric_quantization = {
    "range": "[0, 255] 或 [-128, 127]",
    "zero_point": "非零",
    "advantages": "充分利用量化范围",
    "disadvantages": "计算复杂度稍高"
}

Conversion Workflow Overview

Complete Conversion Flow

graph TD
    A[原始模型] --> B[模型验证]
    B --> C[预处理配置]
    C --> D[量化数据准备]
    D --> E[模型转换]
    E --> F[精度验证]
    F --> G[性能测试]
    G --> H[模型优化]
    H --> I[最终部署]

Key Step Descriptions

  1. Model validation: Ensure the original model can inference normally
  2. Preprocessing configuration: Set preprocessing parameters for the input data
  3. Quantization data preparation: Prepare a representative dataset for quantization calibration
  4. Model conversion: Execute the actual conversion process
  5. Accuracy validation: Compare accuracy differences before and after conversion
  6. Performance testing: Test the inference performance of the converted model
  7. Model optimization: Perform further optimization based on test results

4.2 Prepare the Model to Convert

Get Pre-trained Models

Get ONNX Models from Official Sources

#!/usr/bin/env python3
# download_models.py

import torch
import torchvision.models as models
import requests
import os

def download_classification_models():
    """下载分类模型"""
    models_info = {
        'resnet18': 'https://download.pytorch.org/models/resnet18-5c106cde.pth',
        'resnet50': 'https://download.pytorch.org/models/resnet50-19c8e357.pth',
        'mobilenet_v2': 'https://download.pytorch.org/models/mobilenet_v2-b0353104.pth',
        'efficientnet_b0': 'https://download.pytorch.org/models/efficientnet_b0_rwightman-3dd342df.pth'
    }

    os.makedirs('models/classification', exist_ok=True)

    for model_name, url in models_info.items():
        print(f"下载 {model_name}...")

        # 加载预训练模型
        if model_name == 'resnet18':
            model = models.resnet18(pretrained=True)
        elif model_name == 'resnet50':
            model = models.resnet50(pretrained=True)
        elif model_name == 'mobilenet_v2':
            model = models.mobilenet_v2(pretrained=True)

        model.eval()

        # 导出为ONNX
        dummy_input = torch.randn(1, 3, 224, 224)
        onnx_path = f'models/classification/{model_name}.onnx'

        torch.onnx.export(
            model,
            dummy_input,
            onnx_path,
            export_params=True,
            opset_version=11,
            do_constant_folding=True,
            input_names=['input'],
            output_names=['output'],
            dynamic_axes={
                'input': {0: 'batch_size'},
                'output': {0: 'batch_size'}
            }
        )

        print(f"ONNX模型保存到: {onnx_path}")

def download_yolo_models():
    """下载YOLO模型"""
    import ultralytics

    os.makedirs('models/detection', exist_ok=True)

    # YOLOv5模型
    yolo_models = ['yolov5s', 'yolov5m', 'yolov5l']

    for model_name in yolo_models:
        print(f"下载 {model_name}...")

        # 加载模型
        model = torch.hub.load('ultralytics/yolov5', model_name, pretrained=True)
        model.eval()

        # 导出ONNX
        dummy_input = torch.randn(1, 3, 640, 640)
        onnx_path = f'models/detection/{model_name}.onnx'

        torch.onnx.export(
            model,
            dummy_input,
            onnx_path,
            export_params=True,
            opset_version=11,
            do_constant_folding=True,
            input_names=['images'],
            output_names=['output'],
            dynamic_axes={
                'images': {0: 'batch_size'},
                'output': {0: 'batch_size'}
            }
        )

        print(f"ONNX模型保存到: {onnx_path}")

if __name__ == "__main__":
    download_classification_models()
    download_yolo_models()

Get Models from Hugging Face

#!/usr/bin/env python3
# download_huggingface_models.py

from transformers import AutoModel, AutoTokenizer
import torch
import os

def download_transformer_models():
    """下载Transformer模型"""
    models_info = {
        'bert-base-uncased': 'bert-base-uncased',
        'distilbert-base-uncased': 'distilbert-base-uncased',
        'roberta-base': 'roberta-base'
    }

    os.makedirs('models/nlp', exist_ok=True)

    for model_name, model_id in models_info.items():
        print(f"下载 {model_name}...")

        # 下载模型和tokenizer
        model = AutoModel.from_pretrained(model_id)
        tokenizer = AutoTokenizer.from_pretrained(model_id)

        # 保存模型
        model_dir = f'models/nlp/{model_name}'
        model.save_pretrained(model_dir)
        tokenizer.save_pretrained(model_dir)

        # 导出ONNX (示例)
        model.eval()
        dummy_input = torch.randint(0, 1000, (1, 128))  # 序列长度128

        onnx_path = f'{model_dir}/{model_name}.onnx'
        torch.onnx.export(
            model,
            dummy_input,
            onnx_path,
            export_params=True,
            opset_version=11,
            input_names=['input_ids'],
            output_names=['last_hidden_state'],
            dynamic_axes={
                'input_ids': {0: 'batch_size', 1: 'sequence'},
                'last_hidden_state': {0: 'batch_size', 1: 'sequence'}
            }
        )

        print(f"模型保存到: {model_dir}")

if __name__ == "__main__":
    download_transformer_models()

Model Validation and Preprocessing

Model Integrity Check

#!/usr/bin/env python3
# model_validation.py

import onnx
import onnxruntime as ort
import numpy as np
import cv2

def validate_onnx_model(model_path):
    """验证ONNX模型的完整性"""
    try:
        # 加载模型
        model = onnx.load(model_path)

        # 检查模型
        onnx.checker.check_model(model)
        print(f"✓ 模型 {model_path} 验证通过")

        # 打印模型信息
        print(f"模型版本: {model.ir_version}")
        print(f"生产者: {model.producer_name}")
        print(f"操作集版本: {[opset.version for opset in model.opset_import]}")

        # 打印输入输出信息
        print("\n输入信息:")
        for input_tensor in model.graph.input:
            print(f"  名称: {input_tensor.name}")
            print(f"  形状: {[dim.dim_value for dim in input_tensor.type.tensor_type.shape.dim]}")
            print(f"  类型: {input_tensor.type.tensor_type.elem_type}")

        print("\n输出信息:")
        for output_tensor in model.graph.output:
            print(f"  名称: {output_tensor.name}")
            print(f"  形状: {[dim.dim_value for dim in output_tensor.type.tensor_type.shape.dim]}")
            print(f"  类型: {output_tensor.type.tensor_type.elem_type}")

        return True

    except Exception as e:
        print(f"✗ 模型验证失败: {e}")
        return False

def test_onnx_inference(model_path, input_shape):
    """测试ONNX模型推理"""
    try:
        # 创建推理会话
        session = ort.InferenceSession(model_path)

        # 获取输入输出名称
        input_name = session.get_inputs()[0].name
        output_name = session.get_outputs()[0].name

        # 创建随机输入
        dummy_input = np.random.randn(*input_shape).astype(np.float32)

        # 执行推理
        result = session.run([output_name], {input_name: dummy_input})

        print(f"✓ 推理测试成功")
        print(f"输入形状: {dummy_input.shape}")
        print(f"输出形状: {result[0].shape}")

        return True

    except Exception as e:
        print(f"✗ 推理测试失败: {e}")
        return False

def analyze_model_complexity(model_path):
    """分析模型复杂度"""
    model = onnx.load(model_path)

    # 统计节点类型
    node_types = {}
    for node in model.graph.node:
        op_type = node.op_type
        node_types[op_type] = node_types.get(op_type, 0) + 1

    print(f"\n模型复杂度分析:")
    print(f"总节点数: {len(model.graph.node)}")
    print(f"节点类型分布:")
    for op_type, count in sorted(node_types.items()):
        print(f"  {op_type}: {count}")

    # 估算参数量
    total_params = 0
    for initializer in model.graph.initializer:
        param_size = 1
        for dim in initializer.dims:
            param_size *= dim
        total_params += param_size

    print(f"估算参数量: {total_params:,}")
    print(f"估算模型大小: {total_params * 4 / 1024 / 1024:.2f} MB (FP32)")

if __name__ == "__main__":
    # 测试示例
    model_path = "models/classification/resnet18.onnx"

    if validate_onnx_model(model_path):
        test_onnx_inference(model_path, (1, 3, 224, 224))
        analyze_model_complexity(model_path)

Prepare the Quantization Dataset

Create a Quantization Calibration Dataset

#!/usr/bin/env python3
# prepare_calibration_dataset.py

import os
import cv2
import numpy as np
import random
from pathlib import Path

class CalibrationDataset:
    """量化校准数据集"""

    def __init__(self, data_dir, input_size=(224, 224), num_samples=100):
        self.data_dir = Path(data_dir)
        self.input_size = input_size
        self.num_samples = num_samples
        self.image_paths = self._collect_images()

    def _collect_images(self):
        """收集图片路径"""
        extensions = ['.jpg', '.jpeg', '.png', '.bmp']
        image_paths = []

        for ext in extensions:
            image_paths.extend(self.data_dir.glob(f"**/*{ext}"))
            image_paths.extend(self.data_dir.glob(f"**/*{ext.upper()}"))

        # 随机采样
        if len(image_paths) > self.num_samples:
            image_paths = random.sample(image_paths, self.num_samples)

        print(f"收集到 {len(image_paths)} 张校准图片")
        return image_paths

    def preprocess_image(self, image_path):
        """图像预处理"""
        # 读取图像
        image = cv2.imread(str(image_path))
        if image is None:
            return None

        # 转换颜色空间
        image = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)

        # Resize
        image = cv2.resize(image, self.input_size)

        # 归一化
        image = image.astype(np.float32) / 255.0

        # ImageNet标准化
        mean = np.array([0.485, 0.456, 0.406])
        std = np.array([0.229, 0.224, 0.225])
        image = (image - mean) / std

        # 转换为NCHW格式
        image = np.transpose(image, (2, 0, 1))

        return image

    def generate_calibration_data(self, output_file):
        """生成校准数据"""
        calibration_data = []

        print("生成校准数据...")
        for i, image_path in enumerate(self.image_paths):
            processed_image = self.preprocess_image(image_path)
            if processed_image is not None:
                calibration_data.append(processed_image)

            if (i + 1) % 20 == 0:
                print(f"处理进度: {i + 1}/{len(self.image_paths)}")

        # 转换为numpy数组
        calibration_data = np.array(calibration_data)

        # 保存
        np.save(output_file, calibration_data)
        print(f"校准数据保存到: {output_file}")
        print(f"数据形状: {calibration_data.shape}")

        return calibration_data

def download_imagenet_samples():
    """下载ImageNet样本数据"""
    import urllib.request

    # ImageNet验证集样本URL (示例)
    sample_urls = [
        "https://github.com/pytorch/hub/raw/master/images/dog.jpg",
        "https://github.com/pytorch/hub/raw/master/images/deeplab1.png",
        # 添加更多样本URL
    ]

    os.makedirs("calibration_data/imagenet_samples", exist_ok=True)

    for i, url in enumerate(sample_urls):
        try:
            filename = f"calibration_data/imagenet_samples/sample_{i:03d}.jpg"
            urllib.request.urlretrieve(url, filename)
            print(f"下载: {filename}")
        except Exception as e:
            print(f"下载失败 {url}: {e}")

def create_synthetic_dataset(output_dir, num_samples=100, input_size=(224, 224)):
    """创建合成数据集 (用于测试)"""
    os.makedirs(output_dir, exist_ok=True)

    print(f"创建合成数据集: {num_samples} 张图片")

    for i in range(num_samples):
        # 生成随机图像
        image = np.random.randint(0, 256, (input_size[1], input_size[0], 3), dtype=np.uint8)

        # 添加一些结构
        cv2.rectangle(image, (50, 50), (150, 150), (255, 0, 0), -1)
        cv2.circle(image, (100, 100), 30, (0, 255, 0), -1)

        # 保存图像
        filename = f"{output_dir}/synthetic_{i:03d}.jpg"
        cv2.imwrite(filename, image)

    print(f"合成数据集创建完成: {output_dir}")

if __name__ == "__main__":
    # 创建校准数据集

    # 方法1: 使用现有图片目录
    if os.path.exists("path/to/your/images"):
        dataset = CalibrationDataset("path/to/your/images")
        dataset.generate_calibration_data("calibration_data.npy")

    # 方法2: 下载样本数据
    download_imagenet_samples()

    # 方法3: 创建合成数据集
    create_synthetic_dataset("calibration_data/synthetic", num_samples=50)

    # 使用合成数据集
    dataset = CalibrationDataset("calibration_data/synthetic")
    dataset.generate_calibration_data("calibration_data_synthetic.npy")

4.3 RKNN-Toolkit2 Conversion API Explanation

Core API Introduction

Basic Usage of the RKNN Class

#!/usr/bin/env python3
# rknn_api_tutorial.py

from rknn.api import RKNN
import numpy as np

class RKNNConverter:
    """RKNN转换器封装类"""

    def __init__(self, verbose=True):
        self.rknn = RKNN(verbose=verbose)
        self.model_loaded = False
        self.model_built = False

    def configure_model(self, target_platform='rk3568', **kwargs):
        """配置模型转换参数"""
        config_params = {
            'target_platform': target_platform,
            'quantized_dtype': kwargs.get('quantized_dtype', 'asymmetric_quantized-u8'),
            'optimization_level': kwargs.get('optimization_level', 3),
            'output_optimize': kwargs.get('output_optimize', 1),
            'compress_weight': kwargs.get('compress_weight', False),
            'single_core_mode': kwargs.get('single_core_mode', False),
            'model_pruning': kwargs.get('model_pruning', False)
        }

        print("配置转换参数:")
        for key, value in config_params.items():
            print(f"  {key}: {value}")

        ret = self.rknn.config(**config_params)
        if ret != 0:
            raise Exception("模型配置失败!")

        return ret

    def load_model(self, model_path, model_type='onnx'):
        """加载模型"""
        print(f"加载 {model_type.upper()} 模型: {model_path}")

        if model_type.lower() == 'onnx':
            ret = self.rknn.load_onnx(model=model_path)
        elif model_type.lower() == 'tensorflow':
            ret = self.rknn.load_tensorflow(
                tf_pb=model_path,
                inputs=['input'],
                outputs=['output'],
                input_size_list=[[1, 224, 224, 3]]
            )
        elif model_type.lower() == 'tflite':
            ret = self.rknn.load_tflite(model=model_path)
        elif model_type.lower() == 'caffe':
            ret = self.rknn.load_caffe(
                model=model_path + '.prototxt',
                blobs=model_path + '.caffemodel'
            )
        else:
            raise ValueError(f"不支持的模型类型: {model_type}")

        if ret != 0:
            raise Exception(f"{model_type.upper()} 模型加载失败!")

        self.model_loaded = True
        print("模型加载成功!")
        return ret

    def build_model(self, do_quantization=True, dataset=None):
        """构建模型"""
        if not self.model_loaded:
            raise Exception("请先加载模型!")

        print("开始构建模型...")

        build_params = {'do_quantization': do_quantization}

        if do_quantization and dataset is not None:
            print("使用自定义数据集进行量化...")
            build_params['dataset'] = dataset

        ret = self.rknn.build(**build_params)
        if ret != 0:
            raise Exception("模型构建失败!")

        self.model_built = True
        print("模型构建成功!")
        return ret

    def export_model(self, export_path):
        """导出RKNN模型"""
        if not self.model_built:
            raise Exception("请先构建模型!")

        print(f"导出模型到: {export_path}")
        ret = self.rknn.export_rknn(export_path)
        if ret != 0:
            raise Exception("模型导出失败!")

        print("模型导出成功!")
        return ret

    def init_runtime(self, target='rk3568'):
        """初始化运行时"""
        print(f"初始化运行时 (目标: {target})...")
        ret = self.rknn.init_runtime(target=target)
        if ret != 0:
            raise Exception("运行时初始化失败!")

        print("运行时初始化成功!")
        return ret

    def inference(self, inputs):
        """执行推理"""
        return self.rknn.inference(inputs=inputs)

    def release(self):
        """释放资源"""
        if self.rknn:
            self.rknn.release()
            print("资源释放完成")

# 使用示例
def basic_conversion_example():
    """基本转换示例"""
    converter = RKNNConverter(verbose=True)

    try:
        # 1. 配置参数
        converter.configure_model(
            target_platform='rk3568',
            quantized_dtype='asymmetric_quantized-u8',
            optimization_level=3
        )

        # 2. 加载模型
        converter.load_model('models/classification/resnet18.onnx', 'onnx')

        # 3. 构建模型
        converter.build_model(do_quantization=True)

        # 4. 导出模型
        converter.export_model('resnet18_rk3568.rknn')

        # 5. 测试推理 (可选)
        converter.init_runtime()
        dummy_input = np.random.randn(1, 3, 224, 224).astype(np.float32)
        outputs = converter.inference([dummy_input])
        print(f"推理输出形状: {outputs[0].shape}")

    finally:
        converter.release()

if __name__ == "__main__":
    basic_conversion_example()

Advanced Configuration Options

Detailed Quantization Configuration

#!/usr/bin/env python3
# advanced_quantization.py

from rknn.api import RKNN
import numpy as np

def configure_quantization_options():
    """配置量化选项"""

    # 量化数据类型选项
    quantization_types = {
        'asymmetric_quantized-u8': {
            'description': '非对称8位无符号整数量化',
            'range': '[0, 255]',
            'precision': '中等',
            'speed': '快',
            'recommended': True
        },
        'asymmetric_quantized-i8': {
            'description': '非对称8位有符号整数量化',
            'range': '[-128, 127]',
            'precision': '中等',
            'speed': '快',
            'recommended': False
        },
        'symmetric_quantized-u8': {
            'description': '对称8位无符号整数量化',
            'range': '[0, 255]',
            'precision': '中等',
            'speed': '快',
            'recommended': False
        },
        'dynamic_fixed_point-i8': {
            'description': '动态定点8位量化',
            'range': '[-128, 127]',
            'precision': '高',
            'speed': '中等',
            'recommended': False
        },
        'dynamic_fixed_point-i16': {
            'description': '动态定点16位量化',
            'range': '[-32768, 32767]',
            'precision': '很高',
            'speed': '慢',
            'recommended': False
        }
    }

    print("支持的量化类型:")
    for qtype, info in quantization_types.items():
        print(f"\n{qtype}:")
        for key, value in info.items():
            print(f"  {key}: {value}")

    return quantization_types

def advanced_quantization_config():
    """高级量化配置示例"""
    rknn = RKNN(verbose=True)

    # 高级配置选项
    advanced_config = {
        'target_platform': 'rk3568',
        'quantized_dtype': 'asymmetric_quantized-u8',
        'optimization_level': 3,  # 0-3, 3为最高优化级别
        'output_optimize': 1,     # 输出优化
        'compress_weight': True,  # 权重压缩
        'single_core_mode': False, # 单核模式
        'model_pruning': False,   # 模型剪枝
        'quantized_algorithm': 'normal',  # 量化算法
        'quantized_method': 'channel',    # 量化方法
        'float_dtype': 'float16'  # 浮点数据类型
    }

    print("高级配置参数:")
    for key, value in advanced_config.items():
        print(f"  {key}: {value}")

    ret = rknn.config(**advanced_config)
    rknn.release()

    return ret

def mixed_precision_quantization():
    """混合精度量化示例"""
    rknn = RKNN(verbose=True)

    # 混合精度配置
    # 某些层保持高精度,其他层使用低精度
    mixed_precision_config = {
        'target_platform': 'rk3568',
        'quantized_dtype': 'asymmetric_quantized-u8',
        'optimization_level': 3,
        # 指定特定层的量化类型
        'quantize_input_node': False,  # 输入节点不量化
        'quantize_output_node': False, # 输出节点不量化
    }

    ret = rknn.config(**mixed_precision_config)
    rknn.release()

    return ret

if __name__ == "__main__":
    configure_quantization_options()
    advanced_quantization_config()
    mixed_precision_quantization()

Optimization Levels In Detail

#!/usr/bin/env python3
# optimization_levels.py

def explain_optimization_levels():
    """解释优化级别"""

    optimization_levels = {
        0: {
            'name': '无优化',
            'description': '保持原始模型结构,不进行任何优化',
            'speed': '慢',
            'accuracy': '最高',
            'model_size': '大',
            'use_case': '调试和精度对比'
        },
        1: {
            'name': '基础优化',
            'description': '基本的图优化,如常量折叠',
            'speed': '中等',
            'accuracy': '高',
            'model_size': '中等',
            'use_case': '平衡精度和性能'
        },
        2: {
            'name': '标准优化',
            'description': '包含算子融合和内存优化',
            'speed': '快',
            'accuracy': '中等',
            'model_size': '小',
            'use_case': '大多数应用场景'
        },
        3: {
            'name': '激进优化',
            'description': '最大程度的优化,可能影响精度',
            'speed': '最快',
            'accuracy': '中等偏低',
            'model_size': '最小',
            'use_case': '性能要求极高的场景'
        }
    }

    print("RKNN优化级别详解:")
    for level, info in optimization_levels.items():
        print(f"\n级别 {level} - {info['name']}:")
        for key, value in info.items():
            if key != 'name':
                print(f"  {key}: {value}")

    return optimization_levels

def benchmark_optimization_levels(model_path):
    """测试不同优化级别的效果"""
    from rknn.api import RKNN
    import time
    import os

    results = {}

    for opt_level in range(4):
        print(f"\n测试优化级别 {opt_level}...")

        rknn = RKNN(verbose=False)

        try:
            # 配置
            rknn.config(
                target_platform='rk3568',
                quantized_dtype='asymmetric_quantized-u8',
                optimization_level=opt_level
            )

            # 加载和构建
            start_time = time.time()
            rknn.load_onnx(model=model_path)
            rknn.build(do_quantization=True)
            build_time = time.time() - start_time

            # 导出
            output_path = f'model_opt_{opt_level}.rknn'
            rknn.export_rknn(output_path)

            # 获取文件大小
            model_size = os.path.getsize(output_path) / 1024 / 1024  # MB

            # 测试推理速度
            rknn.init_runtime()
            dummy_input = np.random.randn(1, 3, 224, 224).astype(np.float32)

            # 预热
            for _ in range(10):
                rknn.inference(inputs=[dummy_input])

            # 测试
            inference_times = []
            for _ in range(100):
                start_time = time.time()
                rknn.inference(inputs=[dummy_input])
                inference_times.append(time.time() - start_time)

            avg_inference_time = np.mean(inference_times)

            results[opt_level] = {
                'build_time': build_time,
                'model_size_mb': model_size,
                'avg_inference_time_ms': avg_inference_time * 1000,
                'fps': 1.0 / avg_inference_time
            }

            print(f"  构建时间: {build_time:.2f}s")
            print(f"  模型大小: {model_size:.2f}MB")
            print(f"  推理时间: {avg_inference_time*1000:.2f}ms")
            print(f"  FPS: {1.0/avg_inference_time:.2f}")

            # 清理
            os.remove(output_path)

        except Exception as e:
            print(f"  优化级别 {opt_level} 测试失败: {e}")
            results[opt_level] = None

        finally:
            rknn.release()

    return results

if __name__ == "__main__":
    explain_optimization_levels()

    # 如果有模型文件,可以测试不同优化级别
    model_path = "models/classification/resnet18.onnx"
    if os.path.exists(model_path):
        results = benchmark_optimization_levels(model_path)
        print("\n优化级别对比结果:")
        for level, result in results.items():
            if result:
                print(f"级别 {level}: {result}")

4.4 Conversion In Practice: ONNX Models

ResNet Classification Model Conversion

Complete ResNet Conversion Flow

#!/usr/bin/env python3
# resnet_conversion.py

import os
import cv2
import numpy as np
import time
from rknn.api import RKNN

class ResNetConverter:
    """ResNet模型转换器"""

    def __init__(self, model_path, output_path):
        self.model_path = model_path
        self.output_path = output_path
        self.rknn = RKNN(verbose=True)

    def prepare_calibration_data(self, num_samples=50):
        """准备校准数据"""
        print("准备校准数据...")

        # 创建合成数据 (实际使用中应该用真实数据)
        calibration_data = []

        for i in range(num_samples):
            # 生成随机图像
            image = np.random.randint(0, 256, (224, 224, 3), dtype=np.uint8)

            # 预处理
            image = image.astype(np.float32) / 255.0

            # ImageNet标准化
            mean = np.array([0.485, 0.456, 0.406])
            std = np.array([0.229, 0.224, 0.225])
            image = (image - mean) / std

            # 转换为NCHW格式
            image = np.transpose(image, (2, 0, 1))
            calibration_data.append(image)

        calibration_data = np.array(calibration_data)
        print(f"校准数据形状: {calibration_data.shape}")

        return calibration_data

    def convert_model(self, use_custom_dataset=True):
        """转换模型"""
        try:
            # 1. 配置转换参数
            print("配置转换参数...")
            ret = self.rknn.config(
                target_platform='rk3568',
                quantized_dtype='asymmetric_quantized-u8',
                optimization_level=3,
                output_optimize=1,
                compress_weight=True
            )
            if ret != 0:
                raise Exception("配置失败!")

            # 2. 加载ONNX模型
            print(f"加载ONNX模型: {self.model_path}")
            ret = self.rknn.load_onnx(model=self.model_path)
            if ret != 0:
                raise Exception("模型加载失败!")

            # 3. 构建模型
            print("构建模型...")
            if use_custom_dataset:
                calibration_data = self.prepare_calibration_data()
                ret = self.rknn.build(do_quantization=True, dataset=calibration_data)
            else:
                ret = self.rknn.build(do_quantization=True)

            if ret != 0:
                raise Exception("模型构建失败!")

            # 4. 导出RKNN模型
            print(f"导出RKNN模型: {self.output_path}")
            ret = self.rknn.export_rknn(self.output_path)
            if ret != 0:
                raise Exception("模型导出失败!")

            print("✓ 模型转换成功!")
            return True

        except Exception as e:
            print(f"✗ 模型转换失败: {e}")
            return False

        finally:
            self.rknn.release()

    def test_converted_model(self):
        """测试转换后的模型"""
        print("测试转换后的模型...")

        rknn = RKNN(verbose=False)

        try:
            # 加载RKNN模型
            ret = rknn.load_rknn(self.output_path)
            if ret != 0:
                raise Exception("RKNN模型加载失败!")

            # 初始化运行时
            ret = rknn.init_runtime()
            if ret != 0:
                raise Exception("运行时初始化失败!")

            # 准备测试数据
            test_image = np.random.randn(1, 3, 224, 224).astype(np.float32)

            # 推理测试
            print("执行推理测试...")
            start_time = time.time()
            outputs = rknn.inference(inputs=[test_image])
            inference_time = time.time() - start_time

            print(f"✓ 推理成功!")
            print(f"输入形状: {test_image.shape}")
            print(f"输出形状: {outputs[0].shape}")
            print(f"推理时间: {inference_time*1000:.2f}ms")
            print(f"FPS: {1/inference_time:.2f}")

            # 分析输出
            output = outputs[0][0]  # 移除batch维度
            top5_indices = np.argsort(output)[-5:][::-1]

            print("Top-5 预测结果:")
            for i, idx in enumerate(top5_indices):
                print(f"  {i+1}. 类别 {idx}: {output[idx]:.4f}")

            return True

        except Exception as e:
            print(f"✗ 模型测试失败: {e}")
            return False

        finally:
            rknn.release()

def convert_resnet_models():
    """转换多个ResNet模型"""
    models = {
        'resnet18': 'models/classification/resnet18.onnx',
        'resnet50': 'models/classification/resnet50.onnx',
        'mobilenet_v2': 'models/classification/mobilenet_v2.onnx'
    }

    for model_name, model_path in models.items():
        if not os.path.exists(model_path):
            print(f"跳过不存在的模型: {model_path}")
            continue

        print(f"\n{'='*50}")
        print(f"转换模型: {model_name}")
        print(f"{'='*50}")

        output_path = f"converted_models/{model_name}_rk3568.rknn"
        os.makedirs("converted_models", exist_ok=True)

        converter = ResNetConverter(model_path, output_path)

        # 转换模型
        if converter.convert_model():
            # 测试模型
            converter.test_converted_model()

        print(f"模型 {model_name} 处理完成")

if __name__ == "__main__":
    convert_resnet_models()

YOLOv5 Object Detection Model Conversion

YOLOv5-Specific Conversion Script

#!/usr/bin/env python3
# yolov5_conversion.py

import os
import cv2
import numpy as np
import time
from rknn.api import RKNN

class YOLOv5Converter:
    """YOLOv5模型转换器"""

    def __init__(self, model_path, output_path, input_size=(640, 640)):
        self.model_path = model_path
        self.output_path = output_path
        self.input_size = input_size
        self.rknn = RKNN(verbose=True)

    def prepare_yolo_calibration_data(self, num_samples=100):
        """准备YOLO专用校准数据"""
        print("准备YOLO校准数据...")

        calibration_data = []

        # 创建多样化的测试图像
        for i in range(num_samples):
            if i % 20 == 0:
                print(f"生成校准数据: {i+1}/{num_samples}")

            # 生成不同类型的图像
            if i % 4 == 0:
                # 随机噪声图像
                image = np.random.randint(0, 256, (*self.input_size, 3), dtype=np.uint8)
            elif i % 4 == 1:
                # 几何图形图像
                image = np.zeros((*self.input_size, 3), dtype=np.uint8)
                cv2.rectangle(image, (100, 100), (300, 300), (255, 0, 0), -1)
                cv2.circle(image, (400, 400), 80, (0, 255, 0), -1)
            elif i % 4 == 2:
                # 渐变图像
                image = np.zeros((*self.input_size, 3), dtype=np.uint8)
                for y in range(self.input_size[1]):
                    image[y, :, :] = int(255 * y / self.input_size[1])
            else:
                # 棋盘图像
                image = np.zeros((*self.input_size, 3), dtype=np.uint8)
                square_size = 40
                for y in range(0, self.input_size[1], square_size):
                    for x in range(0, self.input_size[0], square_size):
                        if (x // square_size + y // square_size) % 2 == 0:
                            image[y:y+square_size, x:x+square_size] = 255

            # YOLO预处理
            processed_image = self.preprocess_yolo_image(image)
            calibration_data.append(processed_image)

        calibration_data = np.array(calibration_data)
        print(f"YOLO校准数据形状: {calibration_data.shape}")

        return calibration_data

    def preprocess_yolo_image(self, image):
        """YOLO图像预处理"""
        # 确保输入是正确的尺寸
        if image.shape[:2] != self.input_size:
            image = cv2.resize(image, self.input_size)

        # 转换为RGB
        if len(image.shape) == 3:
            image = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)

        # 归一化到[0,1]
        image = image.astype(np.float32) / 255.0

        # 转换为NCHW格式
        image = np.transpose(image, (2, 0, 1))

        return image

    def convert_yolo_model(self):
        """转换YOLO模型"""
        try:
            # 1. 配置转换参数 (YOLO专用配置)
            print("配置YOLO转换参数...")
            ret = self.rknn.config(
                target_platform='rk3568',
                quantized_dtype='asymmetric_quantized-u8',
                optimization_level=3,
                output_optimize=1,
                compress_weight=True,
                # YOLO特定配置
                quantized_algorithm='normal',
                quantized_method='channel'
            )
            if ret != 0:
                raise Exception("配置失败!")

            # 2. 加载ONNX模型
            print(f"加载YOLOv5 ONNX模型: {self.model_path}")
            ret = self.rknn.load_onnx(model=self.model_path)
            if ret != 0:
                raise Exception("模型加载失败!")

            # 3. 构建模型
            print("构建YOLOv5模型...")
            calibration_data = self.prepare_yolo_calibration_data()
            ret = self.rknn.build(do_quantization=True, dataset=calibration_data)
            if ret != 0:
                raise Exception("模型构建失败!")

            # 4. 导出RKNN模型
            print(f"导出YOLOv5 RKNN模型: {self.output_path}")
            ret = self.rknn.export_rknn(self.output_path)
            if ret != 0:
                raise Exception("模型导出失败!")

            print("✓ YOLOv5模型转换成功!")
            return True

        except Exception as e:
            print(f"✗ YOLOv5模型转换失败: {e}")
            return False

        finally:
            self.rknn.release()

    def test_yolo_model(self):
        """测试转换后的YOLO模型"""
        print("测试转换后的YOLOv5模型...")

        rknn = RKNN(verbose=False)

        try:
            # 加载RKNN模型
            ret = rknn.load_rknn(self.output_path)
            if ret != 0:
                raise Exception("RKNN模型加载失败!")

            # 初始化运行时
            ret = rknn.init_runtime()
            if ret != 0:
                raise Exception("运行时初始化失败!")

            # 准备测试数据
            test_image = np.random.randint(0, 256, (*self.input_size, 3), dtype=np.uint8)
            processed_image = self.preprocess_yolo_image(test_image)
            test_input = np.expand_dims(processed_image, axis=0)

            # 推理测试
            print("执行YOLOv5推理测试...")
            start_time = time.time()
            outputs = rknn.inference(inputs=[test_input])
            inference_time = time.time() - start_time

            print(f"✓ YOLOv5推理成功!")
            print(f"输入形状: {test_input.shape}")
            print(f"输出数量: {len(outputs)}")
            for i, output in enumerate(outputs):
                print(f"输出 {i} 形状: {output.shape}")
            print(f"推理时间: {inference_time*1000:.2f}ms")
            print(f"FPS: {1/inference_time:.2f}")

            # 分析YOLO输出
            self.analyze_yolo_output(outputs[0])

            return True

        except Exception as e:
            print(f"✗ YOLOv5模型测试失败: {e}")
            return False

        finally:
            rknn.release()

    def analyze_yolo_output(self, output):
        """分析YOLO输出"""
        print("\nYOLO输出分析:")

        # YOLOv5输出格式: [batch, num_detections, 85]
        # 85 = 4(bbox) + 1(confidence) + 80(classes)

        batch_size, num_detections, features = output.shape
        print(f"批次大小: {batch_size}")
        print(f"检测数量: {num_detections}")
        print(f"特征维度: {features}")

        # 分析第一个批次的检测结果
        detections = output[0]  # [num_detections, 85]

        # 提取置信度
        confidences = detections[:, 4]
        max_conf = np.max(confidences)
        min_conf = np.min(confidences)
        mean_conf = np.mean(confidences)

        print(f"置信度统计:")
        print(f"  最大值: {max_conf:.4f}")
        print(f"  最小值: {min_conf:.4f}")
        print(f"  平均值: {mean_conf:.4f}")

        # 统计高置信度检测
        high_conf_count = np.sum(confidences > 0.5)
        print(f"高置信度检测 (>0.5): {high_conf_count}")

def convert_yolo_models():
    """转换多个YOLO模型"""
    yolo_models = {
        'yolov5s': {
            'path': 'models/detection/yolov5s.onnx',
            'input_size': (640, 640)
        },
        'yolov5m': {
            'path': 'models/detection/yolov5m.onnx',
            'input_size': (640, 640)
        },
        'yolov5l': {
            'path': 'models/detection/yolov5l.onnx',
            'input_size': (640, 640)
        }
    }

    os.makedirs("converted_models", exist_ok=True)

    for model_name, model_info in yolo_models.items():
        model_path = model_info['path']
        input_size = model_info['input_size']

        if not os.path.exists(model_path):
            print(f"跳过不存在的模型: {model_path}")
            continue

        print(f"\n{'='*60}")
        print(f"转换YOLOv5模型: {model_name}")
        print(f"{'='*60}")

        output_path = f"converted_models/{model_name}_rk3568.rknn"

        converter = YOLOv5Converter(model_path, output_path, input_size)

        # 转换模型
        if converter.convert_yolo_model():
            # 测试模型
            converter.test_yolo_model()

        print(f"YOLOv5模型 {model_name} 处理完成")

if __name__ == "__main__":
    convert_yolo_models()

Accuracy Validation and Comparison

Accuracy Comparison Before and After Conversion

#!/usr/bin/env python3
# accuracy_validation.py

import os
import cv2
import numpy as np
import onnxruntime as ort
from rknn.api import RKNN
import matplotlib.pyplot as plt

class AccuracyValidator:
    """精度验证器"""

    def __init__(self, onnx_path, rknn_path):
        self.onnx_path = onnx_path
        self.rknn_path = rknn_path
        self.onnx_session = None
        self.rknn = None

    def load_models(self):
        """加载模型"""
        # 加载ONNX模型
        self.onnx_session = ort.InferenceSession(self.onnx_path)
        print(f"✓ ONNX模型加载成功: {self.onnx_path}")

        # 加载RKNN模型
        self.rknn = RKNN(verbose=False)
        ret = self.rknn.load_rknn(self.rknn_path)
        if ret != 0:
            raise Exception("RKNN模型加载失败!")

        ret = self.rknn.init_runtime()
        if ret != 0:
            raise Exception("RKNN运行时初始化失败!")

        print(f"✓ RKNN模型加载成功: {self.rknn_path}")

    def prepare_test_data(self, num_samples=50):
        """准备测试数据"""
        test_data = []

        for i in range(num_samples):
            # 生成随机测试图像
            image = np.random.randint(0, 256, (224, 224, 3), dtype=np.uint8)

            # 预处理
            processed = self.preprocess_image(image)
            test_data.append(processed)

        return np.array(test_data)

    def preprocess_image(self, image):
        """图像预处理"""
        # 转换为RGB
        image = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)

        # 归一化
        image = image.astype(np.float32) / 255.0

        # ImageNet标准化
        mean = np.array([0.485, 0.456, 0.406])
        std = np.array([0.229, 0.224, 0.225])
        image = (image - mean) / std

        # 转换为NCHW格式
        image = np.transpose(image, (2, 0, 1))

        return image

    def compare_outputs(self, test_data):
        """比较输出结果"""
        onnx_outputs = []
        rknn_outputs = []

        print("执行精度对比测试...")

        for i, data in enumerate(test_data):
            if (i + 1) % 10 == 0:
                print(f"处理进度: {i+1}/{len(test_data)}")

            # ONNX推理
            input_name = self.onnx_session.get_inputs()[0].name
            onnx_result = self.onnx_session.run(None, {input_name: np.expand_dims(data, 0)})
            onnx_outputs.append(onnx_result[0][0])  # 移除batch维度

            # RKNN推理
            rknn_result = self.rknn.inference(inputs=[np.expand_dims(data, 0)])
            rknn_outputs.append(rknn_result[0][0])  # 移除batch维度

        onnx_outputs = np.array(onnx_outputs)
        rknn_outputs = np.array(rknn_outputs)

        return onnx_outputs, rknn_outputs

    def calculate_metrics(self, onnx_outputs, rknn_outputs):
        """计算精度指标"""
        # 均方误差 (MSE)
        mse = np.mean((onnx_outputs - rknn_outputs) ** 2)

        # 平均绝对误差 (MAE)
        mae = np.mean(np.abs(onnx_outputs - rknn_outputs))

        # 余弦相似度
        def cosine_similarity(a, b):
            return np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b))

        cos_similarities = []
        for onnx_out, rknn_out in zip(onnx_outputs, rknn_outputs):
            cos_sim = cosine_similarity(onnx_out.flatten(), rknn_out.flatten())
            cos_similarities.append(cos_sim)

        avg_cos_similarity = np.mean(cos_similarities)

        # Top-1准确率对比
        onnx_top1 = np.argmax(onnx_outputs, axis=1)
        rknn_top1 = np.argmax(rknn_outputs, axis=1)
        top1_accuracy = np.mean(onnx_top1 == rknn_top1)

        # Top-5准确率对比
        def top5_accuracy(onnx_out, rknn_out):
            onnx_top5 = np.argsort(onnx_out, axis=1)[:, -5:]
            rknn_top5 = np.argsort(rknn_out, axis=1)[:, -5:]

            matches = 0
            for i in range(len(onnx_out)):
                if np.any(np.isin(onnx_top5[i], rknn_top5[i])):
                    matches += 1

            return matches / len(onnx_out)

        top5_acc = top5_accuracy(onnx_outputs, rknn_outputs)

        metrics = {
            'mse': mse,
            'mae': mae,
            'cosine_similarity': avg_cos_similarity,
            'top1_accuracy': top1_accuracy,
            'top5_accuracy': top5_acc
        }

        return metrics

    def visualize_comparison(self, onnx_outputs, rknn_outputs, save_path='accuracy_comparison.png'):
        """可视化对比结果"""
        fig, axes = plt.subplots(2, 2, figsize=(12, 10))

        # 输出分布对比
        axes[0, 0].hist(onnx_outputs.flatten(), bins=50, alpha=0.7, label='ONNX', density=True)
        axes[0, 0].hist(rknn_outputs.flatten(), bins=50, alpha=0.7, label='RKNN', density=True)
        axes[0, 0].set_title('输出分布对比')
        axes[0, 0].legend()

        # 散点图对比
        sample_indices = np.random.choice(len(onnx_outputs.flatten()), 1000, replace=False)
        onnx_sample = onnx_outputs.flatten()[sample_indices]
        rknn_sample = rknn_outputs.flatten()[sample_indices]

        axes[0, 1].scatter(onnx_sample, rknn_sample, alpha=0.5)
        axes[0, 1].plot([onnx_sample.min(), onnx_sample.max()],
                       [onnx_sample.min(), onnx_sample.max()], 'r--')
        axes[0, 1].set_xlabel('ONNX输出')
        axes[0, 1].set_ylabel('RKNN输出')
        axes[0, 1].set_title('输出相关性')

        # 误差分布
        errors = np.abs(onnx_outputs - rknn_outputs)
        axes[1, 0].hist(errors.flatten(), bins=50)
        axes[1, 0].set_title('绝对误差分布')
        axes[1, 0].set_xlabel('绝对误差')

        # Top-1预测对比
        onnx_top1 = np.argmax(onnx_outputs, axis=1)
        rknn_top1 = np.argmax(rknn_outputs, axis=1)

        axes[1, 1].scatter(onnx_top1, rknn_top1, alpha=0.6)
        axes[1, 1].plot([0, max(onnx_top1.max(), rknn_top1.max())],
                       [0, max(onnx_top1.max(), rknn_top1.max())], 'r--')
        axes[1, 1].set_xlabel('ONNX Top-1预测')
        axes[1, 1].set_ylabel('RKNN Top-1预测')
        axes[1, 1].set_title('Top-1预测对比')

        plt.tight_layout()
        plt.savefig(save_path, dpi=300, bbox_inches='tight')
        plt.close()

        print(f"对比图表保存到: {save_path}")

    def run_validation(self):
        """运行完整的精度验证"""
        print("开始精度验证...")

        # 加载模型
        self.load_models()

        # 准备测试数据
        test_data = self.prepare_test_data(num_samples=100)

        # 比较输出
        onnx_outputs, rknn_outputs = self.compare_outputs(test_data)

        # 计算指标
        metrics = self.calculate_metrics(onnx_outputs, rknn_outputs)

        # 打印结果
        print("\n精度验证结果:")
        print(f"均方误差 (MSE): {metrics['mse']:.6f}")
        print(f"平均绝对误差 (MAE): {metrics['mae']:.6f}")
        print(f"余弦相似度: {metrics['cosine_similarity']:.6f}")
        print(f"Top-1准确率: {metrics['top1_accuracy']:.4f}")
        print(f"Top-5准确率: {metrics['top5_accuracy']:.4f}")

        # 可视化
        self.visualize_comparison(onnx_outputs, rknn_outputs)

        # 释放资源
        if self.rknn:
            self.rknn.release()

        return metrics

def batch_accuracy_validation():
    """批量精度验证"""
    model_pairs = [
        {
            'name': 'ResNet18',
            'onnx': 'models/classification/resnet18.onnx',
            'rknn': 'converted_models/resnet18_rk3568.rknn'
        },
        {
            'name': 'MobileNetV2',
            'onnx': 'models/classification/mobilenet_v2.onnx',
            'rknn': 'converted_models/mobilenet_v2_rk3568.rknn'
        }
    ]

    results = {}

    for model_info in model_pairs:
        model_name = model_info['name']
        onnx_path = model_info['onnx']
        rknn_path = model_info['rknn']

        if not (os.path.exists(onnx_path) and os.path.exists(rknn_path)):
            print(f"跳过 {model_name}: 模型文件不存在")
            continue

        print(f"\n{'='*50}")
        print(f"验证模型: {model_name}")
        print(f"{'='*50}")

        validator = AccuracyValidator(onnx_path, rknn_path)
        metrics = validator.run_validation()
        results[model_name] = metrics

    # 汇总结果
    print(f"\n{'='*60}")
    print("精度验证汇总")
    print(f"{'='*60}")

    for model_name, metrics in results.items():
        print(f"\n{model_name}:")
        for metric_name, value in metrics.items():
            print(f"  {metric_name}: {value:.6f}")

if __name__ == "__main__":
    batch_accuracy_validation()

Common Problems and Solutions

Conversion Problem Troubleshooting

# 1. 检查ONNX模型
python3 -c "
import onnx
model = onnx.load('model.onnx')
onnx.checker.check_model(model)
print('ONNX模型验证通过')
"

# 2. 检查RKNN-Toolkit2版本
pip show rknn-toolkit2

# 3. 检查支持的算子
python3 -c "
from rknn.api import RKNN
rknn = RKNN()
print('支持的算子:', rknn.list_ops())
"

# 4. 内存使用监控
free -h
top -p $(pgrep python3)

Performance Optimization Suggestions

# 1. 量化数据集优化
# 使用真实数据而非随机数据
# 确保数据分布代表性

# 2. 模型结构优化
# 减少动态形状
# 合并小算子
# 使用硬件友好的算子

# 3. 转换参数调优
optimization_levels = [0, 1, 2, 3]  # 测试不同优化级别
quantization_types = [
    'asymmetric_quantized-u8',
    'dynamic_fixed_point-i8',
    'dynamic_fixed_point-i16'
]

# 4. 后处理优化
# 在NPU上执行尽可能多的计算
# 减少数据传输

Summary

Through this chapter, you have mastered:

  1. Basic concepts of model conversion: understood why model conversion is needed and the basic conversion flow
  2. Model preparation: learned how to obtain, validate, and prepare models for conversion
  3. RKNN-Toolkit2 API: mastered the use of the core conversion API and advanced configuration options
  4. Conversion in practice: completed hands-on conversion of a ResNet classification model and a YOLOv5 detection model
  5. Accuracy validation: learned how to validate the accuracy and performance of converted models

The next chapter describes how to run these converted custom models on the board, including the use of the Python and C++ APIs.

Edit this page on GitHub
Prev
Run the Official YOLOv5 Example
Next
Run Custom Models on the Board