HOME
Shop
  • English
  • 简体中文
HOME
Shop
  • English
  • 简体中文
  • Product Series

    • FPGA+ARM

      • GM-3568JHF

        • Introduction

          • GM-3568JHF Introduction
        • Quick Start

          • Preface
          • Environment Setup
          • Compilation Notes
          • Flashing Guide
          • Debugging Tools
          • Software Update
          • Viewing System Information
          • Test Commands
          • Application Compilation
          • Source Code Access
        • Peripherals & Interfaces

          • USB
          • Display and Touch
          • Ethernet
          • WIFI
          • Bluetooth
          • TF-Card
          • Audio
          • Serial Port
          • CAN
          • RTC
        • Application Development

          • UART Read/Write Demo
          • Key Detection Demo
          • LED Blink Demo
          • MIPI Screen Detection Demo
          • Read USB Device Information Demo
          • FAN Detection Demo
          • FPGA FSPI Communication Demo
          • FPGA DMA Read/Write Demo
          • GPS Debugging Demo
          • Ethernet Test Demo
          • RS485 Read/Write Demo
          • FPGA I2C Read/Write Demo
          • PN532 NFC Card-Reading Demo
          • TF Card Read/Write Demo
        • QT Development

          • ARM64 Cross-Compiler Environment Setup
          • Adding a QT Program to Boot Auto-Start
        • RKNN_NPU Development

          • RK3568 NPU Overview
          • Development Environment Setup
          • Run the Official YOLOv5 Example
        • FPGA Development

          • ARM and FPGA Communication
          • FPGA Development Manual
        • Others

          • Modifying the Root Filesystem
          • System Auto-Start Services
        • Downloads

          • Downloads
      • MB-E30P

        • Introduction

          • MB-E30P Introduction
        • Quick Start

          • Preface
          • Environment Setup
          • Compilation Instructions
          • Flashing Guide
          • Debugging Tools
          • Software Update
          • Viewing Information
          • Test Commands
          • Application Compilation
          • Source Code Acquisition
        • Peripherals & Interfaces

          • USB
          • Display and Touch
          • Ethernet
          • WIFI
          • Bluetooth
          • TF-Card
          • Audio
          • RTC
        • Application Development

          • Key Detection Demo
          • LED Blink Demo
          • MIPI Screen Detection Demo
          • Read USB Device Information Demo
          • FAN Detection Demo
          • FPGA FSPI Communication Demo
          • FPGA DMA Read/Write Demo
          • Ethernet Test Demo
          • FPGA IIC Read/Write Demo
          • PN532 NFC Card Reading Demo
          • TF Card Read/Write Demo
        • QT Development

          • ARM64 Cross-Compiler Environment Setup
          • Adding a QT Program to the Boot Auto-Start Service
        • RKNN_NPU Development

          • RK3568 NPU Overview
          • Development Environment Setup
          • Run the Official YOLOv5 Example
          • Model Conversion In Detail
          • Run Custom Models on the Board
        • FPGA Development

          • ARM and FPGA Communication
          • FPGA Development Manual
        • Others

          • Modifying the Root Filesystem
          • System Auto-Start Service
        • Downloads

          • Downloads
    • ShimetaPi

      • M4-R1

        • Introduction

          • M4-R1 Introduction
        • Quick Start

          • OpenHarmony Overview
          • Image Burning
          • Application Development Quick Start
          • Device Development Quick Start
        • Application Development

          • ArkUI

            • ArkTS Language Overview
            • UI Components - Row Container Introduction
            • UI Components - Column Container Introduction
            • UI Components - Text Component
            • UI Components - Toggle Component
            • UI Components - Slider Component
            • UI Components - Animation Component & Transition Component
          • Documentation

            • OpenHarmony Official Materials
          • Development Notes

            • Full-SDK Replacement Tutorial
            • Introducing and Using Third-Party Libraries
            • HDC Debugging
            • Restore Factory Mode via Command Line
            • Upgrade App to System Permission
          • First App

            • Build Your First ArkTS Application - HelloWorld
          • Demos

            • Serial-Debug-Assistant Application Demo
            • Writing-Board Application Demo
            • Digital Clock Application Demo
            • Wi-Fi Information Acquisition Application Demo
        • Device Development

          • Ubuntu Development

            • Environment Setup
            • Download Source Code
            • Compile Source Code
          • DevEco Device Tool

            • Tool Introduction
            • Development Environment Construction
            • Import the SDK
            • HUAWEI DevEco Tool Function Introduction
        • Kernel Peripherals & Interfaces

          • Guide
          • Device Tree Introduction
          • NAPI Introduction
          • ArkTS Introduction
          • NAPI Development Hands-on Demo
          • GPIO Introduction
          • I2C Communication
          • SPI Communication
          • PWM Control
          • UART Communication
          • TF Card (MicroSD)
          • Screen (Display)
          • Touch
          • Ethernet
          • M.2 SSD
          • Audio
          • WIFI & BT
          • Camera
        • Downloads

          • Downloads
      • M5-R1

        • Introduction

          • M5-R1 Development Docs
        • Quick Start

          • Image Burning
          • Environment Setup
          • Download Source Code
        • Peripherals & Interfaces

          • Raspberry Pi Interfaces
          • GPIO Interface
          • I2C Interface
          • SPI Communication
          • PWM Control
          • Serial Port Communication
          • TF Card
          • Display
          • Touch
          • Audio
          • RTC
          • Ethernet
          • M.2
          • MINI-PCIE
          • Camera
          • WIFI & BT
        • Downloads

          • Downloads
      • Pico-G1

        • Product Overview

          • Product Introduction
          • SDK Version Information
        • Quick Start

          • Development Environment Setup
          • Image Build
          • Image Flashing
          • System Login
          • Network Configuration
          • File Transfer
          • SDK Directory Structure
          • Deploying Your First Application
          • Deploying Your First Driver
          • Mounting an SD Card
        • Peripherals & Interfaces

          • GPIO Control
          • UART Serial Communication
          • I2C Communication
          • SPI Communication
        • MPP Media Development

          • MPP Media Processing Software
          • Image Processing Chain
          • Video Input
          • Image Encoding
        • NPU & AI

          • NPU Driver and Runtime Library Architecture
          • .xmm Model Loading
          • SVP Video Processing
          • AI Noise Reduction (AI_NR)
        • Application Samples

          • Encryption/Decryption Application
          • ADC Acquisition Application
          • Low-Power Application
          • Audio Processing Application
          • Video Encoding Application
          • Video Input Application
          • Video Graphics Subsystem (VGS) Application
          • 08 Region Overlay Application
          • 09 Intelligent Video Engine Application
          • 10 UVC Webcam Application
          • 11 All-in-One Quickstart Application
          • 12 FPN Correction Application
          • 13 Regional Motion Detection Application
          • 14 MTCNN Face Detection Application
        • Expansion Board Peripheral Examples

          • 00 - Pico Expansion Board Peripheral Examples Overview
          • 01 - OLED Display Application
          • 02 - TFT Display Application
          • 03 - MPU6050 Gyroscope Application
          • 04 - ADC Acquisition Application
          • 05 - Passive Buzzer Application
          • 06 - MQ Gas Sensor Application
          • 07 - GPS Positioning Application
          • 08 - SHT20 Temperature & Humidity Application
          • 09 - Ultrasonic Ranging Application
          • 10 - SpO2 Sensor Application
          • 11 - DC Motor Control Application
          • 12 - Servo Control Application
    • OpenHarmony

      • SC-3568HA

        • Introduction

          • SC-3568HA Overview
        • Quick Start Guide

          • OpenHarmony Overview
          • Image Flashing
          • Setting Up the Development Environment
          • Hello World Application and Deployment
        • Application Development

          • ArkUI

            • Introduction to ArkTS Language
            • Introduction to UI Components and Practical Applications (Part 1)
            • Introduction to UI Components and Practical Applications (Part 2)
            • Introduction to UI Components and Practical Applications (Part 3)
          • Expand

            • Getting Started Guide
            • Referencing and Using Third-Party Libraries
            • Application Compilation and Deployment
            • Command-Line Factory Reset
            • System Debugging -- HDC Debugging
            • APP Stability Testing
            • Chapter 7 Application Testing
        • Device Development

          • Environment Setup
          • Download Source Code
          • Compiling Source Code
        • Peripheral And Interface

          • Raspberry Pi interface
          • GPIO Interface
          • I2C Interface
          • SPI communication
          • PWM (Pulse Width Modulation) control
          • Serial port communication
          • TF Card
          • Display Screen
          • Touch
          • Audio
          • RTC
          • Ethernet
          • M.2
          • MINI-PCIE
          • Camera
          • WIFI&BT
          • Raspberry Pi expansion board
        • Downloads

          • Downloads
      • M-K1HSE

        • Introduction

          • M-K1HSE Introduction
        • Quick Start

          • Development environment construction
          • Source code acquisition
          • Compilation Notes
          • Burning Guide
        • Application Development

          • Application Development Environment Setup
          • First Application - Hello World
        • Peripherals and interfaces

          • 01 Audio
          • 02 RS485
          • 03 Display
        • System customization development

          • System transplant
          • System customization
          • Driver Development
          • System Debugging
          • OTA Update
        • Downloads

          • Downloads
    • HVS Camera

      • Quick Start

        • SDK Overview
        • Downloads
        • Your First C++ Program
        • Python Data Analysis
        • MultiVision Studio
      • Development

        • Programming Guides

          • Open Camera
          • Read Events
          • Recording & Replay
          • Event Processing (Denoising)
          • Display & Visualization
          • Tuning
          • Capture APS Image
        • Toolkit SDK

          • Hybrid Vision Toolkit
          • Quick Start
          • C++ API
          • Python API
        • Algorithm

          • Hybrid Vision Algo
          • Hybrid Vision Algo API
          • Windows Algo SDK
        • Samples Overview
        • Applications
      • Fundamentals

        • Event Camera Fundamentals
        • HVS Hybrid Vision
        • Event Visualization
        • Data Formats Reference
        • Glossary
        • Bias & Tuning
        • Video Tutorials
      • USB Cameras

        • HVS Camera Quick Start
        • Networking Capabilities

          • HVS Camera System Architecture
          • EVS Network Server
          • EVS Time Sync
          • Web Window
        • HVS Camera Compatibility Matrix
        • FAQ & Troubleshooting Guide
        • Products

          • CF-NRS1 (Lingguang No.1 Hybrid Vision Camera)
      • MIPI Modules

        • MIPI Module Quick Start
        • Carrier Boards

          • RDK X5 Carrier Board Adaptation
          • Raspberry Pi Carrier Board Adaptation
          • Digua Pi Carrier Board Adaptation
          • ShimeTai Board Carrier Board Adaptation
        • MIPI Module Compatibility Matrix
        • Products

          • EVS_003 Sensor Module
    • AI-model

      • 1684XB-32T

        • Introduction

          • AIBOX-1684XB-32 Introduction
        • Quick Start

          • First Use
          • Network Configuration
          • Disk Usage
          • Memory Allocation
          • Fan Control Strategy
          • Firmware Upgrade
          • Cross Compilation
          • Model Quantization
        • Application Development

          • Development Overview

            • Sophgo SDK Development
            • Sophgo Demo Introduction
          • Large Language Models

            • Deploying Llama3 Example
            • Sophon LLM_api_server Development
            • Deploying MiniCPM-V-2_6
            • Qwen-2-5-VL Image and Video Recognition Demo
            • Qwen3-chat Demo
            • Qwen3-Qwen Agent-MCP Development
            • Qwen3-langchain-AI Agent
          • Deep Learning

            • ResNet (Image Classification)
            • LPRNet (License Plate Recognition)
            • SAM (General Image Segmentation Foundation Model)
            • YOLOv5 (Object Detection)
            • OpenPose (Human Keypoint Detection)
            • PP-OCR (Optical Character Recognition)
        • Downloads

          • Downloads
      • 1684X-416T

        • Introduction

          • AIBOX-1684X-416 Introduction
        • Demo Quick Guide

          • ShimeTai Intelligent Monitoring Demo Quick Usage Guide
      • RDK-X5

        • Introduction

          • RDK-X5 Hardware Introduction
        • Quick Start

          • RDK-X5 Quick Start
        • Application Development

          • AI Online Model Development

            • Experiment 01 - Access Volcengine Doubao AI
            • Experiment 02 - Image Analysis
            • Experiment 03 - Multimodal Visual Analysis & Localization
            • Experiment 04 - Multimodal Image-Text Comparison
            • Experiment 05 - Multimodal Document/Table Analysis
            • Experiment 06 - Camera-based AI Visual Analysis
          • Large Language Models

            • Experiment 01 - Speech Recognition
            • Experiment 02 - Voice Conversation
            • Experiment 03 - Multimodal Image Analysis - Voice
            • Experiment 04 - Multimodal Image Comparison - Voice
            • Experiment 05 - Multimodal Document Analysis - Voice
            • Experiment 06 - Multimodal Vision Application - Voice
          • ROS2 Basics

            • Experiment 01 - Environment Setup
            • Experiment 02 - Create & Build a Workspace Package
            • Experiment 03 - Run ROS2 Topic Communication Node
            • Experiment 04 - ROS2 Camera Application
          • 40-pin IO Development

            • Experiment 01 - GPIO Output (LED Blink)
            • Experiment 02 - GPIO Input
            • Experiment 03 - Button-controlled LED
            • Experiment 04 - PWM Output
            • Experiment 05 - Serial Output
            • Experiment 06 - I2C Experiment
            • Experiment 07 - SPI Experiment
          • USB Module Usage

            • Experiment 01 - USB Voice Module Usage
            • Experiment 02 - Sound Source Localization Module
          • Machine Vision Practice

            • Experiment 01 - Open USB Camera
            • Experiment 02 - Color Recognition
            • Experiment 03 - Gesture Recognition
            • Experiment 04 - YOLOv5 Object Detection
      • RDK-S100

        • Introduction

          • RDK-S100 Hardware Introduction
        • Quick Start

          • RDK-S100 Quick Start
        • Application Development

          • AI Online Model Development

            • Experiment 01 - Access Volcengine Doubao AI
            • Experiment 02 - Image Analysis
            • Experiment 03 - Multimodal Visual Analysis & Localization
            • Experiment 04 - Multimodal Image-Text Comparison
            • Experiment 05 - Multimodal Document/Table Analysis
            • Experiment 06 - Camera-based AI Visual Analysis
          • Large Language Models

            • Experiment 01 - Speech Recognition
            • Experiment 02 - Voice Conversation
            • Experiment 03 - Multimodal Image Analysis - Voice
            • Experiment 04 - Multimodal Image Comparison - Voice
            • Experiment 05 - Multimodal Document Analysis - Voice
            • Experiment 06 - Multimodal Vision Application - Voice
          • ROS2 Basics

            • Experiment 01 - Environment Setup
            • Experiment 02 - Create & Build a Workspace Package
            • Experiment 03 - Run ROS2 Topic Communication Node
            • Experiment 04 - ROS2 Camera Application
          • 40-pin IO Development

            • Experiment 01 - GPIO Output (LED Blink)
            • Experiment 02 - GPIO Input
            • Experiment 03 - Button-controlled LED
            • Experiment 04 - PWM Output
            • Experiment 05 - Serial Output
            • Experiment 06 - I2C Experiment
            • Experiment 07 - SPI Experiment
          • USB Module Usage

            • Experiment 01 - USB Voice Module Usage
            • Experiment 02 - Sound Source Localization Module
          • Machine Vision Practice

            • Experiment 01 - Open USB Camera
            • Experiment 02 - Image Processing Basics
            • Experiment 03 - Object Detection
            • Experiment 04 - Image Segmentation
      • RK1828

        • Introduction

          • M5-182X-A1 AI Edge Box - Product Introduction
          • M5-182X-A1 Hardware Specifications
          • M5-182X-A1 Usage & Safety
        • Quick Start

          • M5-182X-A1 Image Flashing
          • RK182X Hardware Installation & Verification
          • RK182X Development Environment Quick Setup
          • RK182X SDK Overview
          • RK182X Environment Setup in Detail
          • RK182X Quick Start
          • Vendor SDK Data Extraction Record
        • Development Guide

          • ClawChips Architecture and Principles
          • SKILL User Manual
          • RK182X Series LLM Inference (RK1828 Model)
          • RK182X Series CNN Inference (RK1828 Model)
          • Model Conversion
          • RK182X AI Agent Application Development Guide
          • RK182X Industrial Anomaly Detection Application
        • SDK Reference

          • RKNN3-SDK Overview

            • RKNN3 SDK Overview
          • RKNN3-Toolkit

            • RKNN3 Toolkit Installation and Usage
          • RKLLM

            • RKLLM On-Device LLM Inference
          • RK182X Series NPU Overview and Architecture (RK1828 Model)
          • RK182X INT8 Quantized Inference Deployment
          • RK182X MPP Multimedia Framework
          • MPP Details

            • RK182X Video Decoding
            • RK182X Video Encoding
          • NPU Details

            • RKNN Model Conversion
            • RK182X NPU INT8 Quantized Inference
            • RK182X Multi-Model Parallel Inference
          • RGA Details

            • RK182X RGA 2D Graphics Acceleration
          • VPU Details

            • RK182X VPU Codec
        • Hardware Reference

          • RK182X Series Hardware Architecture Overview (RK1828 Model)
          • RK182X Pin Definitions and Multiplexing Configuration
          • RK182X Pin Definitions
          • RK182X Power Management
          • RK182X Clock and PLL Configuration
          • RK182X Clock and Frequency Configuration
        • Tutorials

          • Hello World
          • Hello RK1828 - The First Program
          • RTSP Streaming
          • RTSP Streaming + AI Analysis
          • ShiMetaPi AI Lobster One-Click Deployment
          • PaddleOCR-VL Text Recognition
          • Qwen3-1.7B LLM Text Chat
          • AI Multi-View Inspection (Qwen3-VL Wrapper)
          • YOLOv5 Object Detection
        • Downloads

          • Downloads
        • FAQ

          • FAQ
    • Core-Board

      • C-3568BQ

        • Introduction

          • C-3568BQ Overview
      • C-3588LQ

        • Introduction

          • C-3588LQ Overview
      • GC-3568JBAF

        • Introduction

          • GC-3568JBAF Overview
      • C-K1BA

        • Introduction

          • C-K1BA Overview
    • Software Platform

      • ShiMetaPi Workbench

        • Introduction

          • Product Overview
          • Core Architecture
          • Feature Entries
          • Supported Hardware
          • Release Notes
        • Quick Start

          • Install & Login
          • Connect the Device
          • Set Up the Environment
          • Connect to AIHub
          • First Inference
        • User Guide

          • Workspace Overview
          • Device Manager
          • Model Market
          • One-Click Deploy
          • Vision — SVP
          • Vision - Custom Models
          • shimeta-py IDE
          • Terminal
          • Agent Debug Assistant
          • Settings and Resources
        • FAQ

          • Installation & Login
          • Device Connection
          • Models & Deployment
          • Vision & Runtime
          • Settings & Other
      • ShimetaPi Repository

        • Introduction

          • ShimetaPi Software Repository
        • Pico G1 (GK7206)

          • Quick Start

            • Installation & First Inference
            • shimeta_infer — Image Inference
            • shimeta_camera — Real-time Camera Inference
            • SVP Scene Detection
            • File Transfer & Built-in Model Reference
            • FAQ
          • HTTP API & Python SDK

            • HTTP API Reference
      • Model Fine-tuning Platform

        • Introduction

          • Model Training Platform
        • Quick Start

          • Register & Login
          • Create Your First Model (30-Minute Quick Experience)
        • Training Guide

          • Data Preparation & Annotation
          • Training Parameter Configuration
          • Start & Monitor Training
          • Model Evaluation & Testing
        • Model Deployment

          • Export Model
          • Deploy to Edge Device

14 MTCNN Face Detection Application

This chapter describes a complete face detection application example based on the GK7206 NPU — face_recognize. The application uses the MTCNN (Multi-task Cascaded Convolutional Networks) three-stage cascaded network to perform real-time face detection and landmark localization on the board, and provides a browser visualization UI through an embedded web server.

The application source code is located in the SDK directory app_sample/face_recognize/. It is a standalone, self-contained example project covering the complete chain of video capture, NPU inference, and web display, suitable as a reference template for developing custom AI vision applications.

1 Application Overview

1.1 Features

  • MTCNN three-stage cascaded detection: P-Net (coarse screening) → R-Net (fine screening) → O-Net (landmarks), filtering face candidate boxes stage by stage
  • 5-point facial landmark localization: left eye, right eye, nose tip, left mouth corner, right mouth corner
  • Real-time web visualization: embedded HTTP server; view the MJPEG video stream and detection overlay directly in a browser
  • Inter-frame tracking stability: IoU-based inter-frame face tracking + EMA smoothing, eliminating detection box jitter
  • Image pyramid multi-scale detection: supports face detection at different scales

1.2 Technical Parameters

ParameterValue
Sensor resolution2560 × 1440 (SC465SL)
Detection input resolution320 × 180 (16:9, matching the sensor aspect ratio)
NPU inference frame rate~8 FPS (limited by the P-Net sliding window count)
Web service port80
Video encoding formatMJPEG
Model format.xmm (GK7206 NPU-specific format)
Minimum face size48 × 48 pixels (in the detection image)
Cascade thresholdsP-Net: 0.50 / R-Net: 0.70 / O-Net: 0.30

1.3 Directory Structure

app_sample/face_recognize/
├── Makefile                  # Build script
├── src/
│   └── main.c               # Main program (~1800 lines, NPU inference + web service)
├── models/
│   ├── pnet.xmm             # P-Net model (coarse screening, input 12×12)
│   ├── rnet.xmm             # R-Net model (fine screening, input 24×24)
│   └── onet.xmm             # O-Net model (landmarks, input 48×48)
└── web/
    ├── index.html            # Web frontend page
    ├── app.js                # Frontend logic (API polling, drawing detection boxes)
    └── style.css             # Page style

1.4 Common Module

The application depends on the common modules in the app_sample/common/ directory:

app_sample/common/
├── web_server.h                 # Web server header file
└── web_server.c                 # Lightweight HTTP server implementation

The web_server module provides:

FunctionAPIDescription
Static file serviceweb_send_file()Serves HTML/CSS/JS and other static files from the /www/ directory
JSON responsesweb_send_json_ok()Sends HTTP responses in JSON format
MJPEG streamingweb_mjpeg_send_stream()Pushes the MJPEG video stream via callback functions
Custom routesweb_route_handlerHandles project-specific API requests via callback functions
Snapshot serviceweb_snapshot_handlerReturns JPEG snapshots via callback functions

Web Server Configuration

Key configuration parameters (defined in web_server.h):

  • WEB_LISTEN_PORT = 80: HTTP service port
  • WEB_WWW_ROOT = "/www": static file root directory
  • WEB_BACKLOG = 8: maximum number of concurrent connections

2 Build and Deployment

2.1 Prerequisites

Before building this application, make sure the following preparations are complete:

  1. SDK environment ready: set up the cross-compilation toolchain and SDK configuration by following SDK Build
  2. SDK fully built once: the application depends on the SDK's common libraries and headers, so a full build must be executed first to generate the out/ directory
  3. Board driver loaded: make sure the NPU kernel module xm_npu.ko is loaded

2.2 Build the Application

Enter the application directory and build directly with the SDK build system:

# Enter the application directory
cd <SDK_PATH>/app_sample/face_recognize

# Build (the SDK Makefile automatically uses the cross-compilation toolchain)
make clean && make

How the Build Works

The Makefile includes the SDK build rules via include $(SDK_DIR)/build/base.mk and include $(SAMPLE_DIR)/sample_base.mk, which automatically configure the cross compiler, header file paths, and linked libraries. There is no need to set up the toolchain manually.

After a successful build, the executable face_recognize is generated in the current directory.

2.3 Deploy to the Board

Info

The application binary is fairly large, so mount an SD card before running it.

mkdir -p /sd_card
mount /dev/mmcblk1p1 /sd_card

Transfer the following files to the development board via SCP:

Directory Structure Requirements

  1. The web frontend files must be placed in the board's /www/ directory (the default static file root WEB_WWW_ROOT of the web_server module)
  2. The models files must be placed in the application's working directory
# Execute on the development host
# 1. Create the board application directory
ssh root@<board IP> "mkdir -p /sd_card/models /www"

# 2. Transfer the executable
scp face_recognize root@<board IP>:/sd_card/

# 3. Transfer the model files
scp -Or models root@<board IP>:/sd_card/

# 4. Transfer the web frontend files
scp web/* root@<board IP>:/www/

2.4 Run the Application

# Enter the application directory
cd /sd_card

# Add execute permission (first time)
chmod +x face_recognize

# Run
./face_recognize

After startup, the terminal prints the following log:

=== MTCNN Face Detection for GK7206 ===
[init] sensor: 2560x1440 @ 30fps
[npu] Initializing...
[npu] Found 1 device(s)
[npu] Loading P-Net...
[npu] ./models/pnet.xmm loaded
[npu] Loading R-Net...
[npu] ./models/rnet.xmm loaded
[npu] Loading O-Net...
[npu] ./models/onet.xmm loaded
[npu] All models loaded!
[init] Pipeline: VI→VPSS→VENC(MJPEG) + VPSS→NPU(320x180 MTCNN)
[main] System ready. Open http://<board-ip>/ in browser

2.5 Browser Access

Open http://<board IP>/ in a desktop browser to see the live face detection picture:

  • The center of the page shows the MJPEG video stream
  • Detected faces are marked with green rectangles
  • The confidence percentage is shown above each face box
  • Facial landmarks (5 points) are marked with blue dots connected by dashed lines
  • The status bar at the bottom shows the detection FPS and the current face count

2.6 Stop the Application

Press Ctrl+C in the terminal to send the SIGINT signal and exit gracefully. The application stops the NPU inference thread, unloads the models, and releases the video pipeline resources in order.

Note

If Ctrl+C fails to exit gracefully or takes too long, find the PID of the ./face_recognize process and force-terminate it:

ps | grep face_recognize
kill -9 <PID>

3 System Architecture

3.1 Overall Data Flow

The overall architecture follows the classic embedded AI vision pipeline of "video capture → pre-processing → NPU inference → web display":

3.2 Thread Model

The application adopts a multi-threaded architecture with the following thread responsibilities:

ThreadResponsibilityKey operations
Main threadRuns the web server (blocking loop)Receives HTTP requests, dispatches routes, pushes the MJPEG stream
NPU inference threadAcquires frames and runs MTCNN in a loopVPSS frame acquisition → YUV→RGB → MTCNN → update detection results
ISP thread (inside the SDK)Image signal processing3A (AE/AWB/AF), noise reduction, color correction

Threads share the detection results (g_faces array and g_face_count) protected by the pthread_mutex_t g_face_mutex mutex, ensuring data consistency between the NPU thread's writes and the HTTP thread's reads.

4 Internal Execution Logic

4.1 Startup Flow

The main() function starts in 5 stages:

// Stage 1: initialize the video pipeline (VI → VPSS → VENC)
init_system();

// Stage 2: initialize the NPU (load the 3 MTCNN models)
init_npu();

// Stage 3: start the NPU inference thread
g_npu_running = XMEDIA_TRUE;
pthread_create(&g_npu_thread, NULL, npu_inference_thread, NULL);

// Stage 4: register the MJPEG handler and start the web server (blocking)
web_server_set_mjpeg_handler(mjpeg_send_stream, mjpeg_request_stop);
web_server_run(project_route_get);

// Stage 5: clean up resources at shutdown
g_npu_running = XMEDIA_FALSE;
pthread_join(g_npu_thread, NULL);
deinit_npu();
deinit_system();

4.2 Video Pipeline Initialization

init_system() initializes the MPP video pipeline in the following order:

  1. System initialization: configure the VB (Video Buffer) memory pools

    • Pool 0: VI capture buffers (full resolution)
    • Pool 1: VPSS full-resolution / VENC encoding buffers
    • Pool 2: VPSS NPU input buffers (320 × 180)
  2. Module initialization: initialize the VI, VPSS, and VENC modules in order

  3. ISP initialization: configure image signal processing parameters (frame rate, pixel format, resolution, etc.)

  4. VI start: start video input, acquiring images from the sensor

  5. VPSS configuration: configure two output channels

    • ochn0: full-resolution output → sent to VENC for MJPEG encoding
    • ochn1: scaled to 320 × 180 → sent to NPU inference
  6. Binding: VI → VPSS → VENC

VI(pipe=0, chn=0) ──bind──→ VPSS(pipe=0, ochn=0) ──bind──→ VENC(chn=0, MJPEG)
                          └→ VPSS(pipe=0, ochn=1) ──manual acquire──→ NPU inference thread

NPU Frame Acquisition Method

VENC automatically acquires VPSS output frames via bind mode, while the NPU inference thread manually acquires frames from VPSS ochn1 with xmedia_vpss_acquire_ochn_frame(). This is because NPU inference is slower than the video frame rate and unsuitable for bind mode.

4.3 NPU Model Loading

The init_npu() function loads the 3 MTCNN models. Each model is loaded as follows:

// 1. Query the memory size required by the model
xmedia_cl_graph_querysize_from_file(path, &worksize, &weightsize);

// 2. Allocate MMZ memory (physically contiguous; required for NPU DMA access)
mmz_alloc_map("npu_work",  &work_phy,  &work_buf,  worksize);
mmz_alloc_map("npu_weight", &weight_phy, &weight_buf, weightsize);

// 3. Load the model onto the NPU
xmedia_cl_graph_loadmodel_from_file_withmem(&ctx, path,
    work_buf, worksize, weight_buf, weightsize, &graph);

// 4. Get input/output tensor info (two queries: the first gets counts, the second gets details)
xmedia_cl_graph_get_input(graph, 0, &input);           // 1st: get the count
input.tensor = malloc(sizeof(...) * input.num);
xmedia_cl_graph_get_input(graph, input.num, &input);   // 2nd: get the details

// 5. Allocate input/output buffers
mmz_alloc_map("npu_in",  &input_phy,  &input_buf,  inputsize);
mmz_alloc_map("npu_out", &output_phy, &output_buf, outputsize);

// 6. Set tensor addresses and bind
xmedia_cl_graph_set_inout(graph, &input, &output);

Memory Allocation Note

Buffers accessed by the NPU (workspace, weight, input, output) must be allocated with xmedia_mmz_alloc() as physically contiguous memory (MMZ), not ordinary malloc(). This is because the NPU accesses physical memory directly via DMA.

The specifications of the three MTCNN models:

ModelInput sizeOutput TensorDescription
P-Net[1, 3, 12, 12]score(2ch) + bbox(4ch)Coarse screening network; extracts candidates via image pyramid + sliding window
R-Net[1, 3, 24, 24]score(2ch) + bbox(4ch)Fine screening network; crops and infers per candidate
O-Net[1, 3, 48, 48]score(2ch) + bbox(4ch) + landmark(10ch)Output network; refinement + 5-point landmarks

4.4 NPU Inference Thread

npu_inference_thread() is the core inference loop, executing the following steps per frame:

Model inference flow

4.5 MTCNN Detection Algorithm Details

4.5.1 P-Net: Image Pyramid + Sliding-Window Detection

P-Net is a fully convolutional network that takes 12 × 12 image patches and decides whether they contain a face. Since faces vary in size, an image pyramid is built for multi-scale detection:

// Build the pyramid: start from 12/min_face_size, shrink layer by layer
scale = (float)PNET_PATCH_SIZE / PNET_MIN_FACE_SIZE;  // initial: 12/48 = 0.25

while (scaled_image >= 12×12) {
    resize_rgb(src, DET_WIDTH, DET_HEIGHT, scaled, sw, sh);

    // Sliding-window scan
    for (y = 0; y <= sh - 12; y += PNET_STRIDE) {
        for (x = 0; x <= sw - 12; x += PNET_STRIDE) {
            // Extract a 12×12 patch → NPU inference → confidence filtering
            // Map coordinates back to the original image + bbox regression
        }
    }

    scale *= PNET_SCALE_FACTOR;  // 0.707 (= 1/√2, standard MTCNN)
}

Key parameter descriptions:

ParameterValueDescription
PNET_PATCH_SIZE12P-Net input size
PNET_STRIDE6Sliding window stride; 6=fast, 4=balanced, 2=best recall
PNET_MIN_FACE_SIZE48Minimum detectable face size (pixels in the detection image)
PNET_SCALE_FACTOR0.707Pyramid scale factor (1/√2)
FACE_CONF_THRESHOLD0.50P-Net confidence threshold

4.5.2 R-Net: Candidate Fine Screening

R-Net takes the candidate boxes output by P-Net, crops the corresponding regions from the original image, scales them to 24 × 24, and performs a second screening:

for (i = 0; i < pnet_count; i++) {
    // 1. Crop the candidate region and scale it to 24×24
    crop_resize(rgb_img, DET_WIDTH, DET_HEIGHT, &cands[i], crop, 24);

    // 2. Copy to the model input (handling the NCHW layout)
    copy_rgb_to_model_input(&g_rnet, crop, 24, 24);

    // 3. NPU inference
    xmedia_cl_graph_process(g_rnet.graph);

    // 4. Dequantization + confidence filtering + bbox regression
    dequantize_output(&g_rnet, score_idx, s_vals, 2);
    dequantize_output(&g_rnet, box_idx, b_vals, 4);

    if (prob >= 0.70f) {  // R-Net threshold is higher, stricter filtering
        bbox_reg(&cands[out_count], b_vals);
        out_count++;
    }
}
// NMS deduplication
out_count = nms(cands, out_count, 0.50f);

4.5.3 O-Net: Final Refinement + Landmarks

O-Net further refines the box positions on top of R-Net and outputs 5 facial landmarks:

for (i = 0; i < rnet_count; i++) {
    // Crop → 48×48 → NPU inference
    crop_resize(rgb_img, DET_WIDTH, DET_HEIGHT, &cands[i], crop, 48);
    copy_rgb_to_model_input(&g_onet, crop, 48, 48);
    xmedia_cl_graph_process(g_onet.graph);

    // Dequantization: score(2) + bbox(4) + landmark(10)
    dequantize_output(&g_onet, score_idx, s_vals, 2);
    dequantize_output(&g_onet, box_idx, b_vals, 4);
    dequantize_output(&g_onet, lm_idx, l_vals, 10);

    // Landmark regression (the model outputs planar format: [x0..x4, y0..y4])
    for (k = 0; k < 5; k++) {
        landmarks[k*2]     = x1 + l_vals[k]     * bw;  // X coordinate
        landmarks[k*2 + 1] = y1 + l_vals[5 + k] * bh;  // Y coordinate
    }
}

Landmark order: left eye (LE) → right eye (RE) → nose tip (N) → left mouth corner (ML) → right mouth corner (MR).

4.6 Post-processing and Inter-frame Tracking

Shape Filtering

After NMS, filter_implausible() filters implausible detection boxes with the following heuristic rules:

// Reject boxes that are not face-shaped
- Aspect ratio < 0.20 or > 3.0 (note: quantized MTCNN bbox regression systematically narrows boxes)
- Width or height < 6 pixels (too small)
- Area > 90% of the detection image (too large)

Inter-frame Tracking

update_tracks() implements simple IoU-based inter-frame tracking, used to:

  1. Stabilize the display: smooth detection box positions with EMA (exponential moving average), eliminating inter-frame jitter
  2. Delayed disappearance: when a face is temporarily occluded or detection is lost, keep showing it for several frames (MAX_MISS = 3)
  3. Unique ID: each track is assigned a unique ID
// Per-frame processing flow:
// 1. Match new detections against existing tracks using IoU
// 2. Matched → EMA-smooth position update
//    Not matched → create a new track, or miss_count++
// 3. miss_count > MAX_MISS → remove the track

4.7 Output Quantization and Dequantization

The NPU outputs INT8 quantized data, which must be dequantized to floating point before use:

// Dequantization formula
float_value = (int8_value - zero_point) × scale

// Implementation
static void dequantize_output(const npu_model_t *m, int tensor_idx,
                              float *out, int count)
{
    unsigned char *data = get_output_data(m, tensor_idx);
    float scale = m->output.tensor[tensor_idx].quant.scale;
    int zp = m->output.tensor[tensor_idx].quant.zp;
    for (i = 0; i < count; i++)
        out[i] = ((float)data[i] - zp) * scale;
}

4.8 Web Service and API

The application embeds a lightweight HTTP server (web_server.c) that provides the following routes:

RouteMethodFunction
/GETServes the index.html page
/mjpegGETMJPEG video stream (multipart/x-mixed-replace)
/api/facesGETReturns the detection results in JSON format
/api/statusGETReturns brief status information
/style.css, /app.jsGETStatic resource files

JSON format returned by /api/faces:

{
  "count": 2,
  "fps": 18,
  "faces": [
    {
      "x": 0.3125,
      "y": 0.2056,
      "w": 0.1875,
      "h": 0.3333,
      "score": 0.998,
      "landmarks": [
        { "x": 0.35, "y": 0.29 },
        { "x": 0.44, "y": 0.285 },
        { "x": 0.395, "y": 0.34 },
        { "x": 0.36, "y": 0.41 },
        { "x": 0.43, "y": 0.405 }
      ]
    }
  ]
}

All coordinate values are normalized to the 0.0 ~ 1.0 range (relative to the detection image size).

5 How to Write a Similar Application

This section uses face_recognize as a reference template to explain how to develop a custom AI vision application based on the GK7206 NPU.

5.1 Development Steps Overview

Step 1: Prepare the model → train/download a model → convert to .xmm format with the XMTVM tool
Step 2: Create the project → copy the Makefile template → write main.c
Step 3: Initialize the pipeline → VI + VPSS + VENC configuration
Step 4: Load the model → load .xmm via the xmedia_cl_* APIs
Step 5: Inference loop → acquire frame → pre-process → NPU inference → post-process
Step 6: Output results → web display / RTSP streaming / other means

5.2 Step 1: Prepare the Model

Use the XMTVM model conversion tool to convert a trained model (ONNX/PyTorch) into the GK7206 NPU-specific .xmm format:

# Example: convert a model with the XMTVM tool
python3 convert.py --model yolov5s.onnx \
    --input_shape 1,3,320,180 \
    --input_format RGB \
    --quantize int8 \
    --output model.xmm

Model Conversion Notes

  • The model's input_format setting determines the input data format (RGB/YUV, etc.)
  • The quantization method (INT8/UINT8) affects inference accuracy and the post-processing dequantization parameters
  • Specifying (pixel - 127.5) / 128 normalization in the YAML configuration lets the NPU do it internally, reducing CPU load

5.3 Step 2: Create the Project

Create a new project following the face_recognize directory structure:

mkdir -p my_app/src my_app/models my_app/web

Write the Makefile (copy face_recognize/Makefile and modify it):

ifeq ($(CFG_SDK_EXPORT_FLAG),)
        SDK_DIR := $(shell cd $(CURDIR)/../.. && /bin/pwd)
endif

include $(SDK_DIR)/build/base.mk
include $(SAMPLE_DIR)/sample_base.mk

TARGET := my_app                    # Change to your application name

LIBS := -lxmedia_svp -lxmedia_npu $(SAMPLE_LIBS) $(SAMPLE_COMMON_LIB) -lpthread

INCLUDES := $(SAMPLE_INCLUDES)
INCLUDES += -I$(SDK_DIR)/app_sample/common    # If using web_server.c

CFLAGS := $(SAMPLE_CFLAGS) $(LIBS) $(INCLUDES)

SRCS := $(wildcard src/*.c) $(SDK_DIR)/app_sample/common/web_server.c
OBJS := $(patsubst %.c, %.o, $(SRCS))

.PHONY: all clean

all: $(OBJS)
	$(AT)$(CC) -o $(TARGET) $^ $(CFLAGS)

%.o : %.c
	$(AT)$(CC) -c -o $@ $< $(CFLAGS)

clean:
	$(AT)rm -rf $(OBJS) $(TARGET)

5.4 Step 3: Initialize the Video Pipeline

The pipeline initialization code can reuse the init_system() function framework of face_recognize; adjust the following parameters for your needs:

// 1. Adjust the NPU detection resolution to your model's input size
#define DET_WIDTH   320     // Match the model input width
#define DET_HEIGHT  180     // Match the model input height

// 2. VB memory pool configuration (3 pools)
//    - Pool 0: VI capture buffers
//    - Pool 1: VPSS full resolution / VENC
//    - Pool 2: VPSS NPU input (DET_WIDTH × DET_HEIGHT)

// 3. VPSS output channel configuration
//    - ochn0: full resolution → VENC (for web display)
//    - ochn1: scaled to DET_WIDTH×DET_HEIGHT → NPU inference

5.5 Step 4: Load the NPU Model

The model loading code can reuse the load_one_model() function, which encapsulates the complete loading flow:

// Define the model struct (encapsulates graph + tensor + buffers)
typedef struct {
    xmedia_cl_graph graph;
    xmedia_cl_tensor_info_inout input;
    xmedia_cl_tensor_info_inout output;
    void *work_buf, *weight_buf, *input_buf, *output_buf;
    xmedia_u64 work_phy, weight_phy, input_phy, output_phy;
} npu_model_t;

// Load the model
npu_model_t my_model;
load_one_model(&my_model, g_cl_ctx, "./models/my_model.xmm");

5.6 Step 5: Write the Inference Loop

The core pattern of the inference loop is "acquire frame → pre-process → NPU inference → post-process":

void *inference_thread(void *arg)
{
    while (running) {
        // 1. Acquire a frame from VPSS
        xmedia_vpss_acquire_ochn_frame(pipe, ochn, &frame, timeout);

        // 2. Pre-process (depending on the model's needs)
        //    - mmap the frame data to user space
        //    - YUV → RGB conversion (if the model needs RGB)
        //    - resize / crop to the model input size
        //    - handle the data layout (HWC → NCHW, etc.)
        void *frame_y  = xmedia_mmz_map(frame.addr.y_phy_addr, ...);
        void *frame_uv = xmedia_mmz_map(frame.addr.c_phy_addr, ...);
        yuv420sp_to_rgb888(frame_y, frame_uv, rgb_buf, ...);
        copy_rgb_to_model_input(&model, rgb_buf, width, height);

        // 3. Run NPU inference
        xmedia_cl_graph_process(model.graph);

        // 4. Post-process
        //    - read the output buffer and dequantize
        dequantize_output(&model, tensor_idx, output, count);
        //    - parse detection results (NMS, threshold filtering, etc.)
        //    - update shared state (with lock protection)

        // 5. Release the frame
        xmedia_vpss_release_ochn_frame(pipe, ochn, &frame);
        xmedia_mmz_unmap(frame_y);
        xmedia_mmz_unmap(frame_uv);
    }
}

5.7 Key Programming Points

MMZ Memory Management

NPU-related buffers must be allocated from MMZ (Media Memory Zone) as physically contiguous memory:

// Allocate + map
xmedia_u64 phy = xmedia_mmz_alloc(mmz_name, buf_name, size);
void *virt = xmedia_mmz_map(phy, size, cached);

// After use
xmedia_mmz_unmap(virt);
xmedia_mmz_free(phy);

Data Layout Conversion

NPU models usually use the NCHW layout [1, C, H, W], while camera RGB output is HWC. copy_rgb_to_model_input() detects the model layout and converts automatically:

if (model_input_is_nchw(m)) {
    // HWC → NCHW: split into R/G/B channel planes
    for (c = 0; c < 3; c++)
        for (y = 0; y < h; y++)
            for (x = 0; x < w; x++)
                dst[c*plane + y*w + x] = src[(y*w + x)*3 + c];
} else {
    memcpy(dst, src, w * h * 3);  // copy directly
}

Quantized Output Handling

The NPU outputs quantized integer data, which must be dequantized using the output tensor's quantization parameters (scale and zero_point):

float scale = m->output.tensor[idx].quant.scale;
int zp = m->output.tensor[idx].quant.zp;
float value = (float)(int8_data[i] - zp) * scale;

Output Tensor Indexing

When a model has multiple output tensors, identify each output's meaning by its channel count:

// Find the corresponding tensor index by output channel count
int score_idx = find_output_by_ch(&model, 2);    // 2 channels = classification score
int box_idx   = find_output_by_ch(&model, 4);    // 4 channels = bbox regression
int lm_idx    = find_output_by_ch(&model, 10);   // 10 channels = 5-point landmarks

5.8 Troubleshooting

ProblemPossible CauseSolution
NPU fails to load the modelWrong .xmm file path or format mismatchCheck the path, confirm the model input size matches the code
No faces detectedConfidence threshold too high / YUV→RGB conversion errorLower the threshold to debug; dump the first RGB frame to verify color conversion
Detection box positions offsetbbox regression coordinates not mapped correctlyCheck the coordinate mapping from the scaled image to the original
Frame rate too lowP-Net sliding window stride too small / too many pyramid layersIncrease PNET_STRIDE or PNET_MIN_FACE_SIZE
Web page not accessibleFrontend files not in the correct directoryMake sure HTML/JS/CSS are in the web server's static file directory
MMZ allocation failsInsufficient system memoryReduce the number or size of VB pools

6 References

  • NPU Driver and Runtime Library Architecture
  • .xmm Model Loading Details
  • MPP Media Processing Platform
  • SDK Build Guide
  • File Transfer
Edit this page on GitHub
Prev
13 Regional Motion Detection Application