HOME
Shop
  • English
  • 简体中文
HOME
Shop
  • English
  • 简体中文
  • Product Series

    • FPGA+ARM

      • GM-3568JHF

        • Introduction

          • GM-3568JHF Introduction
        • Quick Start

          • Preface
          • Environment Setup
          • Compilation Notes
          • Flashing Guide
          • Debugging Tools
          • Software Update
          • Viewing System Information
          • Test Commands
          • Application Compilation
          • Source Code Access
        • Peripherals & Interfaces

          • USB
          • Display and Touch
          • Ethernet
          • WIFI
          • Bluetooth
          • TF-Card
          • Audio
          • Serial Port
          • CAN
          • RTC
        • Application Development

          • UART Read/Write Demo
          • Key Detection Demo
          • LED Blink Demo
          • MIPI Screen Detection Demo
          • Read USB Device Information Demo
          • FAN Detection Demo
          • FPGA FSPI Communication Demo
          • FPGA DMA Read/Write Demo
          • GPS Debugging Demo
          • Ethernet Test Demo
          • RS485 Read/Write Demo
          • FPGA I2C Read/Write Demo
          • PN532 NFC Card-Reading Demo
          • TF Card Read/Write Demo
        • QT Development

          • ARM64 Cross-Compiler Environment Setup
          • Adding a QT Program to Boot Auto-Start
        • RKNN_NPU Development

          • RK3568 NPU Overview
          • Development Environment Setup
          • Run the Official YOLOv5 Example
        • FPGA Development

          • ARM and FPGA Communication
          • FPGA Development Manual
        • Others

          • Modifying the Root Filesystem
          • System Auto-Start Services
        • Downloads

          • Downloads
      • MB-E30P

        • Introduction

          • MB-E30P Introduction
        • Quick Start

          • Preface
          • Environment Setup
          • Compilation Instructions
          • Flashing Guide
          • Debugging Tools
          • Software Update
          • Viewing Information
          • Test Commands
          • Application Compilation
          • Source Code Acquisition
        • Peripherals & Interfaces

          • USB
          • Display and Touch
          • Ethernet
          • WIFI
          • Bluetooth
          • TF-Card
          • Audio
          • RTC
        • Application Development

          • Key Detection Demo
          • LED Blink Demo
          • MIPI Screen Detection Demo
          • Read USB Device Information Demo
          • FAN Detection Demo
          • FPGA FSPI Communication Demo
          • FPGA DMA Read/Write Demo
          • Ethernet Test Demo
          • FPGA IIC Read/Write Demo
          • PN532 NFC Card Reading Demo
          • TF Card Read/Write Demo
        • QT Development

          • ARM64 Cross-Compiler Environment Setup
          • Adding a QT Program to the Boot Auto-Start Service
        • RKNN_NPU Development

          • RK3568 NPU Overview
          • Development Environment Setup
          • Run the Official YOLOv5 Example
          • Model Conversion In Detail
          • Run Custom Models on the Board
        • FPGA Development

          • ARM and FPGA Communication
          • FPGA Development Manual
        • Others

          • Modifying the Root Filesystem
          • System Auto-Start Service
        • Downloads

          • Downloads
    • ShimetaPi

      • M4-R1

        • Introduction

          • M4-R1 Introduction
        • Quick Start

          • OpenHarmony Overview
          • Image Burning
          • Application Development Quick Start
          • Device Development Quick Start
        • Application Development

          • ArkUI

            • ArkTS Language Overview
            • UI Components - Row Container Introduction
            • UI Components - Column Container Introduction
            • UI Components - Text Component
            • UI Components - Toggle Component
            • UI Components - Slider Component
            • UI Components - Animation Component & Transition Component
          • Documentation

            • OpenHarmony Official Materials
          • Development Notes

            • Full-SDK Replacement Tutorial
            • Introducing and Using Third-Party Libraries
            • HDC Debugging
            • Restore Factory Mode via Command Line
            • Upgrade App to System Permission
          • First App

            • Build Your First ArkTS Application - HelloWorld
          • Demos

            • Serial-Debug-Assistant Application Demo
            • Writing-Board Application Demo
            • Digital Clock Application Demo
            • Wi-Fi Information Acquisition Application Demo
        • Device Development

          • Ubuntu Development

            • Environment Setup
            • Download Source Code
            • Compile Source Code
          • DevEco Device Tool

            • Tool Introduction
            • Development Environment Construction
            • Import the SDK
            • HUAWEI DevEco Tool Function Introduction
        • Kernel Peripherals & Interfaces

          • Guide
          • Device Tree Introduction
          • NAPI Introduction
          • ArkTS Introduction
          • NAPI Development Hands-on Demo
          • GPIO Introduction
          • I2C Communication
          • SPI Communication
          • PWM Control
          • UART Communication
          • TF Card (MicroSD)
          • Screen (Display)
          • Touch
          • Ethernet
          • M.2 SSD
          • Audio
          • WIFI & BT
          • Camera
        • Downloads

          • Downloads
      • M5-R1

        • Introduction

          • M5-R1 Development Docs
        • Quick Start

          • Image Burning
          • Environment Setup
          • Download Source Code
        • Peripherals & Interfaces

          • Raspberry Pi Interfaces
          • GPIO Interface
          • I2C Interface
          • SPI Communication
          • PWM Control
          • Serial Port Communication
          • TF Card
          • Display
          • Touch
          • Audio
          • RTC
          • Ethernet
          • M.2
          • MINI-PCIE
          • Camera
          • WIFI & BT
        • Downloads

          • Downloads
      • Pico-G1

        • Product Overview

          • Product Introduction
          • SDK Version Information
        • Quick Start

          • Development Environment Setup
          • Image Build
          • Image Flashing
          • System Login
          • Network Configuration
          • File Transfer
          • SDK Directory Structure
          • Deploying Your First Application
          • Deploying Your First Driver
          • Mounting an SD Card
        • Peripherals & Interfaces

          • GPIO Control
          • UART Serial Communication
          • I2C Communication
          • SPI Communication
        • MPP Media Development

          • MPP Media Processing Software
          • Image Processing Chain
          • Video Input
          • Image Encoding
        • NPU & AI

          • NPU Driver and Runtime Library Architecture
          • .xmm Model Loading
          • SVP Video Processing
          • AI Noise Reduction (AI_NR)
        • Application Samples

          • Encryption/Decryption Application
          • ADC Acquisition Application
          • Low-Power Application
          • Audio Processing Application
          • Video Encoding Application
          • Video Input Application
          • Video Graphics Subsystem (VGS) Application
          • 08 Region Overlay Application
          • 09 Intelligent Video Engine Application
          • 10 UVC Webcam Application
          • 11 All-in-One Quickstart Application
          • 12 FPN Correction Application
          • 13 Regional Motion Detection Application
          • 14 MTCNN Face Detection Application
        • Expansion Board Peripheral Examples

          • 00 - Pico Expansion Board Peripheral Examples Overview
          • 01 - OLED Display Application
          • 02 - TFT Display Application
          • 03 - MPU6050 Gyroscope Application
          • 04 - ADC Acquisition Application
          • 05 - Passive Buzzer Application
          • 06 - MQ Gas Sensor Application
          • 07 - GPS Positioning Application
          • 08 - SHT20 Temperature & Humidity Application
          • 09 - Ultrasonic Ranging Application
          • 10 - SpO2 Sensor Application
          • 11 - DC Motor Control Application
          • 12 - Servo Control Application
    • OpenHarmony

      • SC-3568HA

        • Introduction

          • SC-3568HA Overview
        • Quick Start Guide

          • OpenHarmony Overview
          • Image Flashing
          • Setting Up the Development Environment
          • Hello World Application and Deployment
        • Application Development

          • ArkUI

            • Introduction to ArkTS Language
            • Introduction to UI Components and Practical Applications (Part 1)
            • Introduction to UI Components and Practical Applications (Part 2)
            • Introduction to UI Components and Practical Applications (Part 3)
          • Expand

            • Getting Started Guide
            • Referencing and Using Third-Party Libraries
            • Application Compilation and Deployment
            • Command-Line Factory Reset
            • System Debugging -- HDC Debugging
            • APP Stability Testing
            • Chapter 7 Application Testing
        • Device Development

          • Environment Setup
          • Download Source Code
          • Compiling Source Code
        • Peripheral And Interface

          • Raspberry Pi interface
          • GPIO Interface
          • I2C Interface
          • SPI communication
          • PWM (Pulse Width Modulation) control
          • Serial port communication
          • TF Card
          • Display Screen
          • Touch
          • Audio
          • RTC
          • Ethernet
          • M.2
          • MINI-PCIE
          • Camera
          • WIFI&BT
          • Raspberry Pi expansion board
        • Downloads

          • Downloads
      • M-K1HSE

        • Introduction

          • M-K1HSE Introduction
        • Quick Start

          • Development environment construction
          • Source code acquisition
          • Compilation Notes
          • Burning Guide
        • Application Development

          • Application Development Environment Setup
          • First Application - Hello World
        • Peripherals and interfaces

          • 01 Audio
          • 02 RS485
          • 03 Display
        • System customization development

          • System transplant
          • System customization
          • Driver Development
          • System Debugging
          • OTA Update
        • Downloads

          • Downloads
    • HVS Camera

      • Quick Start

        • SDK Overview
        • Downloads
        • Your First C++ Program
        • Python Data Analysis
        • MultiVision Studio
      • Development

        • Programming Guides

          • Open Camera
          • Read Events
          • Recording & Replay
          • Event Processing (Denoising)
          • Display & Visualization
          • Tuning
          • Capture APS Image
        • Toolkit SDK

          • Hybrid Vision Toolkit
          • Quick Start
          • C++ API
          • Python API
        • Algorithm

          • Hybrid Vision Algo
          • Hybrid Vision Algo API
          • Windows Algo SDK
        • Samples Overview
        • Applications
      • Fundamentals

        • Event Camera Fundamentals
        • HVS Hybrid Vision
        • Event Visualization
        • Data Formats Reference
        • Glossary
        • Bias & Tuning
        • Video Tutorials
      • USB Cameras

        • HVS Camera Quick Start
        • Networking Capabilities

          • HVS Camera System Architecture
          • EVS Network Server
          • EVS Time Sync
          • Web Window
        • HVS Camera Compatibility Matrix
        • FAQ & Troubleshooting Guide
        • Products

          • CF-NRS1 (Lingguang No.1 Hybrid Vision Camera)
      • MIPI Modules

        • MIPI Module Quick Start
        • Carrier Boards

          • RDK X5 Carrier Board Adaptation
          • Raspberry Pi Carrier Board Adaptation
          • Digua Pi Carrier Board Adaptation
          • ShimeTai Board Carrier Board Adaptation
        • MIPI Module Compatibility Matrix
        • Products

          • EVS_003 Sensor Module
    • AI-model

      • 1684XB-32T

        • Introduction

          • AIBOX-1684XB-32 Introduction
        • Quick Start

          • First Use
          • Network Configuration
          • Disk Usage
          • Memory Allocation
          • Fan Control Strategy
          • Firmware Upgrade
          • Cross Compilation
          • Model Quantization
        • Application Development

          • Development Overview

            • Sophgo SDK Development
            • Sophgo Demo Introduction
          • Large Language Models

            • Deploying Llama3 Example
            • Sophon LLM_api_server Development
            • Deploying MiniCPM-V-2_6
            • Qwen-2-5-VL Image and Video Recognition Demo
            • Qwen3-chat Demo
            • Qwen3-Qwen Agent-MCP Development
            • Qwen3-langchain-AI Agent
          • Deep Learning

            • ResNet (Image Classification)
            • LPRNet (License Plate Recognition)
            • SAM (General Image Segmentation Foundation Model)
            • YOLOv5 (Object Detection)
            • OpenPose (Human Keypoint Detection)
            • PP-OCR (Optical Character Recognition)
        • Downloads

          • Downloads
      • 1684X-416T

        • Introduction

          • AIBOX-1684X-416 Introduction
        • Demo Quick Guide

          • ShimeTai Intelligent Monitoring Demo Quick Usage Guide
      • RDK-X5

        • Introduction

          • RDK-X5 Hardware Introduction
        • Quick Start

          • RDK-X5 Quick Start
        • Application Development

          • AI Online Model Development

            • Experiment 01 - Access Volcengine Doubao AI
            • Experiment 02 - Image Analysis
            • Experiment 03 - Multimodal Visual Analysis & Localization
            • Experiment 04 - Multimodal Image-Text Comparison
            • Experiment 05 - Multimodal Document/Table Analysis
            • Experiment 06 - Camera-based AI Visual Analysis
          • Large Language Models

            • Experiment 01 - Speech Recognition
            • Experiment 02 - Voice Conversation
            • Experiment 03 - Multimodal Image Analysis - Voice
            • Experiment 04 - Multimodal Image Comparison - Voice
            • Experiment 05 - Multimodal Document Analysis - Voice
            • Experiment 06 - Multimodal Vision Application - Voice
          • ROS2 Basics

            • Experiment 01 - Environment Setup
            • Experiment 02 - Create & Build a Workspace Package
            • Experiment 03 - Run ROS2 Topic Communication Node
            • Experiment 04 - ROS2 Camera Application
          • 40-pin IO Development

            • Experiment 01 - GPIO Output (LED Blink)
            • Experiment 02 - GPIO Input
            • Experiment 03 - Button-controlled LED
            • Experiment 04 - PWM Output
            • Experiment 05 - Serial Output
            • Experiment 06 - I2C Experiment
            • Experiment 07 - SPI Experiment
          • USB Module Usage

            • Experiment 01 - USB Voice Module Usage
            • Experiment 02 - Sound Source Localization Module
          • Machine Vision Practice

            • Experiment 01 - Open USB Camera
            • Experiment 02 - Color Recognition
            • Experiment 03 - Gesture Recognition
            • Experiment 04 - YOLOv5 Object Detection
      • RDK-S100

        • Introduction

          • RDK-S100 Hardware Introduction
        • Quick Start

          • RDK-S100 Quick Start
        • Application Development

          • AI Online Model Development

            • Experiment 01 - Access Volcengine Doubao AI
            • Experiment 02 - Image Analysis
            • Experiment 03 - Multimodal Visual Analysis & Localization
            • Experiment 04 - Multimodal Image-Text Comparison
            • Experiment 05 - Multimodal Document/Table Analysis
            • Experiment 06 - Camera-based AI Visual Analysis
          • Large Language Models

            • Experiment 01 - Speech Recognition
            • Experiment 02 - Voice Conversation
            • Experiment 03 - Multimodal Image Analysis - Voice
            • Experiment 04 - Multimodal Image Comparison - Voice
            • Experiment 05 - Multimodal Document Analysis - Voice
            • Experiment 06 - Multimodal Vision Application - Voice
          • ROS2 Basics

            • Experiment 01 - Environment Setup
            • Experiment 02 - Create & Build a Workspace Package
            • Experiment 03 - Run ROS2 Topic Communication Node
            • Experiment 04 - ROS2 Camera Application
          • 40-pin IO Development

            • Experiment 01 - GPIO Output (LED Blink)
            • Experiment 02 - GPIO Input
            • Experiment 03 - Button-controlled LED
            • Experiment 04 - PWM Output
            • Experiment 05 - Serial Output
            • Experiment 06 - I2C Experiment
            • Experiment 07 - SPI Experiment
          • USB Module Usage

            • Experiment 01 - USB Voice Module Usage
            • Experiment 02 - Sound Source Localization Module
          • Machine Vision Practice

            • Experiment 01 - Open USB Camera
            • Experiment 02 - Image Processing Basics
            • Experiment 03 - Object Detection
            • Experiment 04 - Image Segmentation
      • RK1828

        • Introduction

          • M5-182X-A1 AI Edge Box - Product Introduction
          • M5-182X-A1 Hardware Specifications
          • M5-182X-A1 Usage & Safety
        • Quick Start

          • M5-182X-A1 Image Flashing
          • RK182X Hardware Installation & Verification
          • RK182X Development Environment Quick Setup
          • RK182X SDK Overview
          • RK182X Environment Setup in Detail
          • RK182X Quick Start
          • Vendor SDK Data Extraction Record
        • Development Guide

          • ClawChips Architecture and Principles
          • SKILL User Manual
          • RK182X Series LLM Inference (RK1828 Model)
          • RK182X Series CNN Inference (RK1828 Model)
          • Model Conversion
          • RK182X AI Agent Application Development Guide
          • RK182X Industrial Anomaly Detection Application
        • SDK Reference

          • RKNN3-SDK Overview

            • RKNN3 SDK Overview
          • RKNN3-Toolkit

            • RKNN3 Toolkit Installation and Usage
          • RKLLM

            • RKLLM On-Device LLM Inference
          • RK182X Series NPU Overview and Architecture (RK1828 Model)
          • RK182X INT8 Quantized Inference Deployment
          • RK182X MPP Multimedia Framework
          • MPP Details

            • RK182X Video Decoding
            • RK182X Video Encoding
          • NPU Details

            • RKNN Model Conversion
            • RK182X NPU INT8 Quantized Inference
            • RK182X Multi-Model Parallel Inference
          • RGA Details

            • RK182X RGA 2D Graphics Acceleration
          • VPU Details

            • RK182X VPU Codec
        • Hardware Reference

          • RK182X Series Hardware Architecture Overview (RK1828 Model)
          • RK182X Pin Definitions and Multiplexing Configuration
          • RK182X Pin Definitions
          • RK182X Power Management
          • RK182X Clock and PLL Configuration
          • RK182X Clock and Frequency Configuration
        • Tutorials

          • Hello World
          • Hello RK1828 - The First Program
          • RTSP Streaming
          • RTSP Streaming + AI Analysis
          • ShiMetaPi AI Lobster One-Click Deployment
          • PaddleOCR-VL Text Recognition
          • Qwen3-1.7B LLM Text Chat
          • AI Multi-View Inspection (Qwen3-VL Wrapper)
          • YOLOv5 Object Detection
        • Downloads

          • Downloads
        • FAQ

          • FAQ
    • Core-Board

      • C-3568BQ

        • Introduction

          • C-3568BQ Overview
      • C-3588LQ

        • Introduction

          • C-3588LQ Overview
      • GC-3568JBAF

        • Introduction

          • GC-3568JBAF Overview
      • C-K1BA

        • Introduction

          • C-K1BA Overview
    • Software Platform

      • ShiMetaPi Workbench

        • Introduction

          • Product Overview
          • Core Architecture
          • Feature Entries
          • Supported Hardware
          • Release Notes
        • Quick Start

          • Install & Login
          • Connect the Device
          • Set Up the Environment
          • Connect to AIHub
          • First Inference
        • User Guide

          • Workspace Overview
          • Device Manager
          • Model Market
          • One-Click Deploy
          • Vision — SVP
          • Vision - Custom Models
          • shimeta-py IDE
          • Terminal
          • Agent Debug Assistant
          • Settings and Resources
        • FAQ

          • Installation & Login
          • Device Connection
          • Models & Deployment
          • Vision & Runtime
          • Settings & Other
      • ShimetaPi Repository

        • Introduction

          • ShimetaPi Software Repository
        • Pico G1 (GK7206)

          • Quick Start

            • Installation & First Inference
            • shimeta_infer — Image Inference
            • shimeta_camera — Real-time Camera Inference
            • SVP Scene Detection
            • File Transfer & Built-in Model Reference
            • FAQ
          • HTTP API & Python SDK

            • HTTP API Reference
      • Model Fine-tuning Platform

        • Introduction

          • Model Training Platform
        • Quick Start

          • Register & Login
          • Create Your First Model (30-Minute Quick Experience)
        • Training Guide

          • Data Preparation & Annotation
          • Training Parameter Configuration
          • Start & Monitor Training
          • Model Evaluation & Testing
        • Model Deployment

          • Export Model
          • Deploy to Edge Device

Voice LLM Applications

Experiment 01 - Speech Recognition

Experiment preparation:

  1. Install ALSA tools (for recording and playback)

sudo apt-get install alsa-utils

  1. Install the required dependencies

  2. pip install -r requirements.txt

  3. python -m pip install websocket-client

  4. sudo apt-get update && sudo apt-get install -y ffmpeg

  5. Register a iFLYTEK (Xunfei) account

  6. Log in to the iFLYTEK open platform https://www.xfyun.com.cn

  7. Click to enter the console

  8. Register and log in to an account

  9. Create an application

  10. Activate the Speech Recognition - Dictation service

  11. Obtain the four pieces of information: APIID, APISecret, APIKey, and the dictation interface URL

  12. Save the four pieces of information (to be filled into config.py later in the code)

Experiment steps:

  1. Check the voice module usage (ensure the voice module is correctly connected to the RDK board and the speaker)

Run in terminal: arecord -l # Identify the card number and device number of the microphone (note card X and device Y)

Run in terminal: aplay -l # Check the speaker/output device

Run in terminal: sudo arecord -f S16_LE -r 16000 -c 1 -d 5 /tmp/test_mic.wav # Record 5 seconds with the default device, 16k/mono/16bit:

Run in terminal: aplay /tmp/test_mic.wav # Play the audio

  1. Fill the four pieces of information (APIID, APISecret, APIKey, dictation interface URL) into config.py
TOOL
  1. cd AI_online_voice # Enter the package
  2. python examples/01_voice_chat.py # Run the example program; enter r to start testing

Terminal output:

TOOL

Experiment result: Start recording (5 seconds by default; to change the duration, enter r + duration). After recording finishes, the audio is played back, then uploaded to the iFLYTEK dictation large model, and the recognition result is returned to the Linux terminal.

"""
01_voice_chat.py

功能:
- 录制语音
- 使用讯飞 WebSocket API 将语音转为文本
- 在终端打印识别结果(专注于语音转文字)

依赖:
- arecord: 用于录制音频(Linux)
- aplay: 用于播放音频(Linux)
- websocket-client: 用于与讯飞 WebSocket API 通信
- 请在 AI_online_voice/config.py 中填写 XUNFEI_APPID / XUNFEI_API_KEY / XUNFEI_API_SECRET / XUNFEI_WS_URL
"""

import os
import sys
import time
from typing import Optional

# 加入父目录,便于示例脚本直接运行
sys.path.append(os.path.dirname(os.path.dirname(os.path.abspath(__file__))))

# 仅保留音频处理,暂不使用豆包对话
from utils.audio_processor import AudioProcessor

# 讯飞 WebSocket 所需依赖与配置
try:
    import websocket  # websocket-client
except ImportError:
    websocket = None
# 新增:导入超时异常类型用于精细日志
try:
    from websocket import WebSocketTimeoutException
except Exception:
    class WebSocketTimeoutException(Exception):
        pass

import json
import base64
import hmac
import hashlib
import ssl
import wave
from email.utils import formatdate
from urllib.parse import urlparse, quote

from config import (
    XUNFEI_APPID,
    XUNFEI_API_KEY,
    XUNFEI_API_SECRET,
    XUNFEI_WS_URL,
    REQUEST_TIMEOUT,
)


class XunfeiRealtimeSpeechClient:
    """讯飞语音识别(IAT流式WebSocket版)客户端(更新的消息格式与解析)"""

    def __init__(self, app_id: str = None, api_key: str = None, api_secret: str = None, ws_url: str = None):
        self.app_id = app_id or XUNFEI_APPID
        self.api_key = api_key or XUNFEI_API_KEY
        self.api_secret = api_secret or XUNFEI_API_SECRET
        self.ws_url = ws_url or XUNFEI_WS_URL
        self.timeout = REQUEST_TIMEOUT
        self._validate_config()

    def _validate_config(self):
        if not self.app_id or self.app_id == "你的讯飞APPID":
            raise ValueError("请配置正确的讯飞APPID")
        if not self.api_key or self.api_key == "你的讯飞API_KEY":
            raise ValueError("请配置正确的讯飞API_KEY")
        if not self.api_secret or self.api_secret == "你的讯飞API_SECRET":
            raise ValueError("请配置正确的讯飞API_SECRET")
        if websocket is None:
            raise RuntimeError("未安装 websocket-client,请先安装:python -m pip install websocket-client")

    def _rfc1123_date(self) -> str:
        # 生成GMT时间,RFC1123格式
        return formatdate(usegmt=True)

    def _assemble_auth_url(self) -> str:
        """根据APIKey与APISecret生成带鉴权参数的WS URL"""
        parsed = urlparse(self.ws_url)
        host = parsed.netloc
        path = parsed.path
        date = self._rfc1123_date()

        # signature 原始串:
        signature_origin = f"host: {host}\n" + f"date: {date}\n" + f"GET {path} HTTP/1.1"
        # 使用 apiSecret 做 HMAC-SHA256
        signature_sha = hmac.new(self.api_secret.encode("utf-8"), signature_origin.encode("utf-8"), hashlib.sha256).digest()
        signature = base64.b64encode(signature_sha).decode("utf-8")

        # authorization 原始串
        authorization_origin = (
            f"api_key=\"{self.api_key}\", "
            f"algorithm=\"hmac-sha256\", "
            f"headers=\"host date request-line\", "
            f"signature=\"{signature}\""
        )
        authorization = base64.b64encode(authorization_origin.encode("utf-8")).decode("utf-8")

        # 拼接最终URL
        auth_url = (
            f"{self.ws_url}?authorization={quote(authorization)}&date={quote(date)}&host={quote(host)}"
        )
        return auth_url

    def _parse_result_segments(self, result_obj: dict) -> str:
        """解析服务端 data.result.ws 结构为纯文本"""
        try:
            parts = []
            ws_arr = result_obj.get("ws")
            if isinstance(ws_arr, list):
                for ws in ws_arr:
                    cw_arr = ws.get("cw") if isinstance(ws, dict) else None
                    if isinstance(cw_arr, list):
                        for cw in cw_arr:
                            w = cw.get("w") if isinstance(cw, dict) else None
                            if w:
                                parts.append(w)
            return "".join(parts)
        except Exception:
            return ""

    def _safe_json_loads(self, text: str):
        try:
            return json.loads(text)
        except Exception:
            try:
                cleaned = text.strip()
                start = cleaned.find("{")
                end = cleaned.rfind("}")
                if start != -1 and end != -1 and end > start:
                    return json.loads(cleaned[start:end+1])
            except Exception:
                return None

    def transcribe_audio_ws(self, audio_file: str) -> Optional[str]:
        """将音频文件以流式方式发送到讯飞IAT WS接口并获取识别文本"""
        if not os.path.exists(audio_file):
            print(f"音频文件不存在: {audio_file}")
            return None

        # 解析wav
        try:
            wf = wave.open(audio_file, "rb")
        except Exception as e:
            print(f"打开音频文件失败: {e}")
            return None

        framerate = wf.getframerate()
        channels = wf.getnchannels()
        sampwidth = wf.getsampwidth()  # bytes per sample

        # 建议参数:16k, 单声道, 16bit
        if framerate not in (8000, 16000):
            print(f"采样率异常({framerate}),建议使用16k或8k")
        if channels != 1:
            print(f"通道数为{channels},建议使用单声道")
        if sampwidth != 2:
            print(f"位深为{sampwidth*8}bit,建议16bit")

        auth_url = self._assemble_auth_url()
        ws = None
        try:
            ws = websocket.create_connection(
                auth_url,
                timeout=self.timeout,
                sslopt={"cert_reqs": ssl.CERT_NONE},
            )
            ws.settimeout(self.timeout)

            # 计算每帧40ms对应的帧数
            frames_per_chunk = max(1, int(framerate * 0.04))

            # 构建格式字符串,例如 audio/L16;rate=16000;channel=1
            fmt = f"audio/L{sampwidth*8};rate={framerate};channel={channels}"

            # 初始化增量聚合与最终状态标记
            final_text_parts = []
            saw_final_status = False

            # 发送首帧(status=0)
            first_chunk = wf.readframes(frames_per_chunk)
            first_payload = base64.b64encode(first_chunk).decode("utf-8") if first_chunk else ""
            first_frame = {
                "common": {"app_id": self.app_id},
                "business": {
                    "domain": "iat",
                    "language": "zh_cn",
                    "accent": "mandarin",
                    "vinfo": 1,
                    "vad_eos": 2000,
                    "ptt": 0,
                },
                "data": {
                    "status": 0,
                    "format": fmt,
                    "encoding": "raw",
                    "audio": first_payload,
                },
            }
            try:
                ws.send(json.dumps(first_frame, separators=(",", ":")))
            except Exception as e:
                print(f"发送首帧失败: {e}")
                print("可能原因:鉴权失败或 WS URL 错误导致服务端立即关闭连接")
                return None

            # 增强:首帧后循环尝试接收,打印并积累增量结果
            try:
                ws.settimeout(1.0)
                for _ in range(3):
                    try:
                        pre_resp_text = ws.recv()
                    except WebSocketTimeoutException:
                        break
                    if not pre_resp_text:
                        break
                    pre_resp = self._safe_json_loads(pre_resp_text)
                    if not pre_resp:
                        print(f"[首帧返回-非JSON] {pre_resp_text}")
                        break
                    code = pre_resp.get("code")
                    message = pre_resp.get("message")
                    if code is None:
                        header = pre_resp.get("header", {})
                        code = header.get("code", 0)
                        message = header.get("message")
                    data = pre_resp.get("data", {})
                    status = data.get("status")
                    print(f"[首帧返回] code={code}, status={status}, message={message}")
                    if code != 0:
                        desc = message or "识别错误"
                        print(f"识别错误(连接初期): code={code}, message={desc}")
                        return None
                    result = data.get("result")
                    if result:
                        segment = self._parse_result_segments(result)
                        if segment:
                            final_text_parts.append(segment)
                            print(f"[增量结果-首帧] {segment}")
                    if status == 2:
                        saw_final_status = True
                        break
            except Exception as e:
                print(f"[首帧接收日志] {e}")
            finally:
                ws.settimeout(self.timeout)

            # 发送中间帧(status=1)
            while True:
                chunk = wf.readframes(frames_per_chunk)
                if not chunk or saw_final_status:
                    break
                frame = {
                    "common": {"app_id": self.app_id},
                    "data": {
                        "status": 1,
                        "format": fmt,
                        "encoding": "raw",
                        "audio": base64.b64encode(chunk).decode("utf-8"),
                    },
                }
                try:
                    ws.send(json.dumps(frame, separators=(",", ":")))
                except Exception as e:
                    print(f"发送中间帧失败: {e}")
                    print("可能原因:连接已被服务端关闭(鉴权/配置错误、URL错误、参数不匹配)")
                    return None
                # 每次发送后短暂接收,积累增量结果
                try:
                    ws.settimeout(0.5)
                    resp_text_mid = ws.recv()
                    if resp_text_mid:
                        resp_mid = self._safe_json_loads(resp_text_mid)
                        if not resp_mid:
                            print(f"[中间帧返回-非JSON] {resp_text_mid}")
                        else:
                            code_mid = resp_mid.get("code")
                            msg_mid = resp_mid.get("message")
                            if code_mid is None:
                                header_mid = resp_mid.get("header", {})
                                code_mid = header_mid.get("code", 0)
                                msg_mid = header_mid.get("message")
                            data_mid = resp_mid.get("data", {})
                            status_mid = data_mid.get("status")
                            print(f"[中间帧返回] code={code_mid}, status={status_mid}, message={msg_mid}")
                            if code_mid != 0:
                                print(f"识别错误(发送中间帧后): code={code_mid}, message={msg_mid}")
                                return None
                            result_mid = data_mid.get("result")
                            if result_mid:
                                seg_mid = self._parse_result_segments(result_mid)
                                if seg_mid:
                                    final_text_parts.append(seg_mid)
                                    print(f"[增量结果-中间] {seg_mid}")
                            if status_mid == 2:
                                saw_final_status = True
                                break
                except WebSocketTimeoutException:
                    pass
                except Exception as e:
                    print(f"接收中间帧返回失败: {e}")
                    return None
                finally:
                    ws.settimeout(self.timeout)
                time.sleep(0.04)

            # 若尚未收到最终状态,发送结束帧
            if not saw_final_status:
                last_frame = {
                    "common": {"app_id": self.app_id},
                    "data": {
                        "status": 2,
                        "format": fmt,
                        "encoding": "raw",
                        "audio": "",
                    },
                }
                try:
                    ws.send(json.dumps(last_frame, separators=(",", ":")))
                except Exception as e:
                    print(f"发送结束帧失败: {e}")
                    # 即使结束帧发送失败,只要已有增量文本也返回
                    return "".join(final_text_parts) if final_text_parts else None

            # 接收最终结果(容错:超时但已有增量文本则直接返回)
            if not saw_final_status:
                while True:
                    try:
                        resp_text = ws.recv()
                    except Exception as e:
                        print(f"接收结果失败: {e}")
                        return "".join(final_text_parts) if final_text_parts else None
                    if not resp_text:
                        continue
                    resp = self._safe_json_loads(resp_text)
                    if not resp:
                        continue
                    code = resp.get("code")
                    message = resp.get("message")
                    if code is None:
                        header = resp.get("header", {})
                        code = header.get("code", 0)
                        message = header.get("message")
                    if code != 0:
                        desc = message or "识别错误"
                        print(f"识别错误: code={code}, message={desc}")
                        break
                    data = resp.get("data", {})
                    status = data.get("status")
                    result = resp.get("result") or data.get("result")
                    if result:
                        segment = self._parse_result_segments(result)
                        if segment:
                            final_text_parts.append(segment)
                    if status == 2:
                        break
            return "".join(final_text_parts) if final_text_parts else None
        finally:
            try:
                wf.close()
            except Exception:
                pass
            if ws is not None:
                try:
                    ws.close()
                except Exception:
                    pass


class VoiceChatApp:
    """语音对话应用(仅语音转文字与打印)"""

    def __init__(self):
        """初始化应用"""
        self.processor = None
        self.xunfei_ws_client = None
        self.running = False

    def initialize(self) -> bool:
        """初始化客户端和处理器"""
        try:
            self.processor = AudioProcessor()
            self.xunfei_ws_client = XunfeiRealtimeSpeechClient()
            return True
        except Exception as e:
            print(f"初始化失败: {e}")
            return False

    def print_welcome(self):
        """打印欢迎信息"""
        print("\n" + "=" * 50)
        print("语音转文字 - 讯飞 WebSocket API")
        print("=" * 50)
        print("使用说明:")
        print("1. 输入 'r' 或 'record' 开始录音并进行识别(默认5秒)")
        print("2. 输入 'p' 或 'play' <文件> 播放音频文件")
        print("3. 输入 'q' 或 'quit' 退出应用")
        print("4. 输入 'h' 或 'help' 显示帮助信息")
        print("=" * 50 + "\n")

    def print_help(self):
        """打印帮助信息"""
        print("\n" + "=" * 50)
        print("命令列表:")
        print("  r, record [秒数]    - 录制语音 (默认5秒) 并用WebSocket识别,终端打印文本")
        print("  p, play <文件>      - 播放音频文件")
        print("  q, quit             - 退出应用")
        print("  h, help             - 显示帮助信息")
        print("=" * 50 + "\n")

    def handle_command(self, command: str) -> bool:
        """处理命令"""
        parts = command.strip().split()
        if not parts:
            return True

        cmd = parts[0].lower()

        if cmd in ('q', 'quit', 'exit'):
            return False

        elif cmd in ('h', 'help'):
            self.print_help()

        elif cmd in ('r', 'record'):
            # 解析录音时长
            duration = 5
            if len(parts) > 1:
                try:
                    duration = int(parts[1])
                except ValueError:
                    print("无效的时长,使用默认值5秒")

            # 录制音频
            audio_file = self.processor.record(duration)
            if not audio_file:
                print("录音失败")
                return True

            # 新增:录音后强制转换为 16k/1ch/16bit PCM WAV
            converted_file = self.processor.convert_to_wav(audio_file)
            use_file = converted_file or audio_file
            if converted_file:
                print(f"已转换为16k/1ch/16bit: {converted_file}")
            else:
                print("转换失败,使用原始录音进行识别")

            # 新增:打印文件名与完整路径,并先播放音频
            try:
                import os
                file_name = os.path.basename(use_file)
                print(f"原始录音文件: {audio_file}")
                print(f"用于播放与识别的文件: {use_file}")
                print(f"开始播放: {file_name} | {use_file}")
                play_ok = self.processor.play(use_file)
                if not play_ok:
                    print("播放失败,但继续进行识别")
            except Exception as e:
                print(f"播放流程异常: {e},继续进行识别")

            # 使用讯飞WS实时识别
            print("正在进行实时语音识别(WebSocket)...")
            text = self.xunfei_ws_client.transcribe_audio_ws(use_file)

            if text:
                print(f"识别结果: {text}")
            else:
                print("语音识别失败")

        elif cmd in ('p', 'play'):
            if len(parts) < 2:
                print("请提供要播放的音频文件路径")
                return True
            audio_path = parts[1]
            if not os.path.exists(audio_path):
                print(f"音频文件不存在: {audio_path}")
                return True
            self.processor.play(audio_path)

        else:
            print("未知命令,请输入 'h' 或 'help' 查看帮助")

        return True

    def run(self):
        if not self.initialize():
            return

        self.running = True
        self.print_welcome()

        while self.running:
            try:
                command = input("请输入命令 (r/p/h/q): ").strip()
            except (KeyboardInterrupt, EOFError):
                print("\n收到退出信号,正在退出...")
                break

            if not command:
                continue

            if not self.handle_command(command):
                break

        print("应用已退出")


def main():
    app = VoiceChatApp()
    app.run()

if __name__ == "__main__":
    main()
Edit this page on GitHub
Next
Experiment 02 - Voice Conversation