Files
aha/docs/cli.md
T
XiaoYang c582c4cd0a docs: update and optimize README and project documentation system
- Update README.md with improved formatting, logo, badges,
  and comprehensive documentation
- Add README.en.md with English translation of the documentation
- Include detailed quick start guide, CLI reference,
  and supported models table
- Add changelog information highlighting recent features
- Add script directory and adjust script file locations
2026-02-06 16:13:50 +08:00

9.8 KiB

CLI Reference

Complete command-line interface reference for aha.

AHA is a high-performance model inference library based on the Candle framework, supporting various multimodal models including vision, language, and audio models.

aha [COMMAND] [OPTIONS]

Global Options

Option Description Default
-a, --address <ADDRESS> Service listen address 127.0.0.1
-p, --port <PORT> Service listen port 10100
-m, --model <MODEL> Model type (required) -
--weight-path <WEIGHT_PATH> Local model weight path -
--save-dir <SAVE_DIR> Model download save directory ~/.aha/
--download-retries <DOWNLOAD_RETRIES> Download retry count 3
-h, --help Display help information -
-V, --version Display version number -

Commands

cli - Download model and start service (default)

Download the specified model and start an HTTP service. This command is used by default when no subcommand is specified.

Syntax:

aha cli [OPTIONS] --model <MODEL>

Options:

Option Description Default
-a, --address <ADDRESS> Service listen address 127.0.0.1
-p, --port <PORT> Service listen port 10100
-m, --model <MODEL> Model type (required) -
--weight-path <WEIGHT_PATH> Local model weight path (skip download if specified) -
--save-dir <SAVE_DIR> Model download save directory ~/.aha/
--download-retries <DOWNLOAD_RETRIES> Download retry count 3

Examples:

# Download model and start service (default port 10100)
aha cli -m qwen3vl-2b

# Specify port and save directory
aha cli -m qwen3vl-2b -p 8080 --save-dir /data/models

# Use local model (skip download)
aha cli -m qwen3vl-2b --weight-path /path/to/model

# Backward compatible way (equivalent to cli subcommand)
aha -m qwen3vl-2b

run - Direct model inference

Run model inference directly without starting an HTTP service. Suitable for one-time inference tasks or batch processing.

Syntax:

aha run [OPTIONS] --model <MODEL> --input <INPUT> [--input <INPUT2>] --weight-path <WEIGHT_PATH>

Options:

Option Description Default
-m, --model <MODEL> Model type (required) -
-i, --input <INPUT> Input text or file path (model-specific interpretation, supports 1-2 parameters: input1: prompt text, input2: file path) -
-o, --output <OUTPUT> Output file path (optional, auto-generated if not specified) -
--weight-path <WEIGHT_PATH> Local model weight path (required) -

Examples:

# VoxCPM1.5 text-to-speech (single input)
aha run -m voxcpm1.5 -i "太阳当空照" -o output.wav --weight-path /path/to/model

# VoxCPM1.5 read input from file (single input)
aha run -m voxcpm1.5 -i "file://./input.txt" --weight-path /path/to/model

# MiniCPM4 text generation (single input)
aha run -m minicpm4-0.5b -i "你好" --weight-path /path/to/model

# DeepSeek OCR image recognition (single input)
aha run -m deepseek-ocr -i "image.jpg" --weight-path /path/to/model

# RMBG2.0 background removal (single input)
aha run -m RMBG2.0 -i "photo.png" -o "no_bg.png" --weight-path /path/to/model

# GLM-ASR speech recognition (two inputs: prompt text + audio file)
aha run -m glm-asr-nano-2512 -i "请转写这段音频" -i "audio.wav" --weight-path /path/to/model

# Fun-ASR speech recognition (two inputs: prompt text + audio file)
aha run -m fun-asr-nano-2512 -i "语音转写:" -i "audio.wav" --weight-path /path/to/model

# qwen3 text generation (single input)
aha run -m qwen3-0.6b -i "你好" --weight-path /path/to/model

# qwen2.5vl image understanding (two inputs: prompt text + image file)
aha run -m qwen2.5vl-3b -i "请分析图片并提取所有可见文本内容,按从左到右、从上到下的布局,返回纯文本" -i "image.jpg" --weight-path /path/to/model

# Qwen3-ASR speech recognition (single input: audio file)
aha run -m qwen3asr-0.6b -i "audio.wav" --weight-path /path/to/model

serv - Start service

Start HTTP service only, without downloading models. Must specify local model path via --weight-path.

Syntax:

aha serv [OPTIONS] --model <MODEL> --weight-path <WEIGHT_PATH>

Options:

Option Description Default
-a, --address <ADDRESS> Service listen address 127.0.0.1
-p, --port <PORT> Service listen port 10100
-m, --model <MODEL> Model type (required) -
--weight-path <WEIGHT_PATH> Local model weight path (required) -

Examples:

# Start service with local model
aha serv -m qwen3vl-2b --weight-path /path/to/model

# Start with specified port
aha serv -m qwen3vl-2b --weight-path /path/to/model -p 8080

# Specify listen address
aha serv -m qwen3vl-2b --weight-path /path/to/model -a 0.0.0.0

download - Download model

Download the specified model only, without starting the service.

Syntax:

aha download [OPTIONS] --model <MODEL>

Options:

Option Description Default
-m, --model <MODEL> Model type (required) -
-s, --save-dir <SAVE_DIR> Model download save directory ~/.aha/
--download-retries <DOWNLOAD_RETRIES> Download retry count 3

Examples:

# Download model to default directory
aha download -m qwen3vl-2b

# Specify save directory
aha download -m qwen3vl-2b -s /data/models

# Specify download retry count
aha download -m qwen3vl-2b --download-retries 5

# Download MiniCPM4-0.5B model
aha download -m minicpm4-0.5b -s models

Supported Models

Model ID Model Name Description
minicpm4-0.5b OpenBMB/MiniCPM4-0.5B OpenBMB MiniCPM4 0.5B model
qwen2.5vl-3b Qwen/Qwen2.5-VL-3B-Instruct Qwen 2.5 VL 3B model
qwen2.5vl-7b Qwen/Qwen2.5-VL-7B-Instruct Qwen 2.5 VL 7B model
qwen3-0.6b Qwen/Qwen3-0.6B Qwen 3 0.6B model
qwen3vl-2b Qwen/Qwen3-VL-2B-Instruct Qwen 3 VL 2B model
qwen3vl-4b Qwen/Qwen3-VL-4B-Instruct Qwen 3 VL 4B model
qwen3vl-8b Qwen/Qwen3-VL-8B-Instruct Qwen 3 VL 8B model
qwen3vl-32b Qwen/Qwen3-VL-32B-Instruct Qwen 3 VL 32B model
deepseek-ocr deepseek-ai/DeepSeek-OCR DeepSeek OCR model
hunyuan-ocr Tencent-Hunyuan/HunyuanOCR Tencent Hunyuan OCR model
paddleocr-vl PaddlePaddle/PaddleOCR-VL Baidu PaddleOCR VL model
RMBG2.0 AI-ModelScope/RMBG-2.0 RMBG 2.0 background removal model
voxcpm OpenBMB/VoxCPM-0.5B OpenBMB VoxCPM 0.5B speech synthesis model
voxcpm1.5 OpenBMB/VoxCPM1.5 OpenBMB VoxCPM 1.5 speech synthesis model
glm-asr-nano-2512 ZhipuAI/GLM-ASR-Nano-2512 Zhipu AI ASR Nano 2512 speech recognition model
fun-asr-nano-2512 FunAudioLLM/Fun-ASR-Nano-2512 FunAudioLLM ASR Nano 2512 speech recognition model

Common Use Cases

Scenario 1: Quick start inference service

# One command to download and start service
aha -m qwen3vl-2b

Scenario 2: Start service with existing model

# Assuming model is downloaded to /data/models/Qwen/Qwen3-VL-2B-Instruct
aha serv -m qwen3vl-2b --weight-path /data/models/Qwen/Qwen3-VL-2B-Instruct

Scenario 3: Pre-download model

# Download model to specified directory for later use
aha download -m qwen3vl-2b -s /data/models

# Later start with local model
aha serv -m qwen3vl-2b --weight-path /data/models/Qwen/Qwen3-VL-2B-Instruct

Scenario 4: Custom service port and address

# Start service on 0.0.0.0:8080, allow external access
aha -m qwen3vl-2b -a 0.0.0.0 -p 8080

API Endpoints

After the service starts, the following API endpoints are available:

Chat Completion Endpoint

  • Endpoint: POST /chat/completions
  • Function: Multimodal chat and text generation
  • Supported Models: Qwen2.5VL, Qwen3, Qwen3VL, DeepSeekOCR, GLM-ASR-Nano-2512, Fun-ASR-Nano-2512, etc.
  • Format: OpenAI Chat Completion format
  • Streaming Support: Yes

Image Processing Endpoint

  • Endpoint: POST /images/remove_background
  • Function: Image background removal
  • Supported Models: RMBG-2.0
  • Format: OpenAI Chat Completion format
  • Streaming Support: No

Audio Generation Endpoint

  • Endpoint: POST /audio/speech
  • Function: Speech synthesis and generation
  • Supported Models: VoxCPM, VoxCPM1.5
  • Format: OpenAI Chat Completion format
  • Streaming Support: No

Backward Compatibility

To maintain compatibility with older versions, the following two usage methods are equivalent:

# New way (recommended)
aha cli -m qwen3vl-2b

# Old way (backward compatible)
aha -m qwen3vl-2b

Notes

  1. serv subcommand requires --weight-path: Since the serv subcommand does not download models, you must specify the path to an already downloaded model via --weight-path.

  2. Download retry mechanism: By default, retries 3 times, waiting 2 seconds after each failure before retrying. You can adjust the retry count with --download-retries.

  3. Default save directory: Models are saved to ~/.aha/ directory by default, which can be customized via --save-dir or -s parameter.

  4. Port occupation: Ensure the specified port is not occupied before starting the service. The default port is 10100.

  5. Permission issues: If saving to a system directory (such as /data/models), ensure you have the corresponding write permissions.

Getting Help

# View main help
aha --help

# View subcommand help
aha cli --help
aha serv --help
aha download --help

# View version information
aha --version

See Also