chore(docs): restructure README files and add Chinese documentation
This commit is contained in:
-228
@@ -1,228 +0,0 @@
|
|||||||
<p align="center">
|
|
||||||
<img src="assets/img/logo.png" alt="aha logo" width="100"/>
|
|
||||||
</p>
|
|
||||||
|
|
||||||
<p align="center">
|
|
||||||
<!-- <a href="https://github.com/jhqxxx/aha/releases">
|
|
||||||
<img src="https://img.shields.io/github/v/release/jhqxxx/aha" alt="GitHub release (latest by date)">
|
|
||||||
</a>
|
|
||||||
<a href="https://github.com/jhqxxx/aha/actions">
|
|
||||||
<img src="https://img.shields.io/github/actions/workflow/status/jhqxxx/aha/ci.yml" alt="GitHub Actions Workflow Status">
|
|
||||||
</a> -->
|
|
||||||
<a href="https://github.com/jhqxxx/aha/blob/main/LICENSE">
|
|
||||||
<img src="https://img.shields.io/github/license/jhqxxx/aha" alt="GitHub License">
|
|
||||||
</a>
|
|
||||||
<a href="https://github.com/jhqxxx/aha/stargazers">
|
|
||||||
<img src="https://img.shields.io/github/stars/jhqxxx/aha" alt="GitHub Stars">
|
|
||||||
</a>
|
|
||||||
<a href="https://github.com/jhqxxx/aha/issues">
|
|
||||||
<img src="https://img.shields.io/github/issues/jhqxxx/aha" alt="GitHub Issues">
|
|
||||||
</a>
|
|
||||||
</p>
|
|
||||||
|
|
||||||
<p align="center">
|
|
||||||
<a href="README.md">简体中文</a> | <strong>English</strong>
|
|
||||||
</p>
|
|
||||||
|
|
||||||
# aha
|
|
||||||
|
|
||||||
**Lightweight AI Inference Engine — All-in-one Solution for Text, Vision, Speech, and OCR**
|
|
||||||
|
|
||||||
aha is a high-performance, cross-platform AI inference engine built with Rust and the Candle framework. It brings state-of-the-art AI models to your local machine—no API keys, no cloud dependencies, just pure, fast AI running directly on your hardware.
|
|
||||||
|
|
||||||
## Changelog
|
|
||||||
|
|
||||||
### v0.2.0 (2026-02-05)
|
|
||||||
- Added Qwen3-ASR speech recognition model
|
|
||||||
|
|
||||||
### v0.1.9 (2026-01-31)
|
|
||||||
- Added CLI `list` subcommand to show supported models
|
|
||||||
- Added CLI subcommand structure support (`cli`, `serv`, `download`, `run`)
|
|
||||||
- Fixed Qwen3VL thinking startswith bug
|
|
||||||
- Fixed `aha run` multiple inputs bug
|
|
||||||
|
|
||||||
### v0.1.8 (2026-01-17)
|
|
||||||
- Added Qwen3 text model support
|
|
||||||
- Added Fun-ASR-Nano-2512 speech recognition model
|
|
||||||
- Fixed ModelScope Fun-ASR-Nano model load error
|
|
||||||
- Updated audio resampling with rubato
|
|
||||||
|
|
||||||
### v0.1.7 (2026-01-07)
|
|
||||||
- Added GLM-ASR-Nano-2512 speech recognition model
|
|
||||||
- Merged Metal (GPU) support for Apple Silicon
|
|
||||||
- Added dynamic home directory and model download script
|
|
||||||
|
|
||||||
**[View full changelog](docs/changelog.md)** →
|
|
||||||
|
|
||||||
## Quick Start
|
|
||||||
|
|
||||||
### Installation
|
|
||||||
|
|
||||||
```bash
|
|
||||||
git clone https://github.com/jhqxxx/aha.git
|
|
||||||
cd aha
|
|
||||||
cargo build --release
|
|
||||||
```
|
|
||||||
|
|
||||||
**Optional Features:**
|
|
||||||
|
|
||||||
```bash
|
|
||||||
# CUDA (NVIDIA GPU acceleration)
|
|
||||||
cargo build --release --features cuda
|
|
||||||
|
|
||||||
# Metal (Apple GPU acceleration for macOS)
|
|
||||||
cargo build --release --features metal
|
|
||||||
|
|
||||||
# Flash Attention (faster inference)
|
|
||||||
cargo build --release --features flash-attn
|
|
||||||
|
|
||||||
# FFmpeg (multimedia processing)
|
|
||||||
cargo build --release --features ffmpeg
|
|
||||||
|
|
||||||
# Combine multiple features
|
|
||||||
cargo build --release --features "cuda,flash-attn"
|
|
||||||
```
|
|
||||||
|
|
||||||
### CLI Quick Reference
|
|
||||||
|
|
||||||
```bash
|
|
||||||
|
|
||||||
# List all supported models
|
|
||||||
aha list
|
|
||||||
|
|
||||||
# Download model only
|
|
||||||
aha download -m qwen3asr-0.6b
|
|
||||||
|
|
||||||
# Download model and start service
|
|
||||||
aha -m qwen3asr-0.6b
|
|
||||||
|
|
||||||
# Run inference directly (without starting service)
|
|
||||||
aha run -m qwen3asr-0.6b -i "audio.wav"
|
|
||||||
|
|
||||||
# Start service only (model already downloaded)
|
|
||||||
aha serv -m qwen3asr-0.6b -p 10100
|
|
||||||
|
|
||||||
```
|
|
||||||
|
|
||||||
### Chat
|
|
||||||
|
|
||||||
```bash
|
|
||||||
aha serv -m qwen3-0.6b -p 10100
|
|
||||||
```
|
|
||||||
|
|
||||||
Then use the unified (OpenAI-compatible) API:
|
|
||||||
|
|
||||||
```bash
|
|
||||||
curl http://localhost:10100/chat/completions \
|
|
||||||
-H "Content-Type: application/json" \
|
|
||||||
-d '{
|
|
||||||
"model": "qwen3-0.6b",
|
|
||||||
"messages": [{"role": "user", "content": "Hello!"}]
|
|
||||||
}
|
|
||||||
'
|
|
||||||
```
|
|
||||||
|
|
||||||
### Supported Models
|
|
||||||
|
|
||||||
| Category | Models |
|
|
||||||
|----------|--------|
|
|
||||||
| **Text** | Qwen3, MiniCPM4 |
|
|
||||||
| **Vision** | Qwen2.5-VL, Qwen3-VL |
|
|
||||||
| **OCR** | DeepSeek-OCR, Hunyuan-OCR, PaddleOCR-VL |
|
|
||||||
| **ASR** | GLM-ASR-Nano, Fun-ASR-Nano, Qwen3-ASR |
|
|
||||||
| **Audio** | VoxCPM, VoxCPM1.5 |
|
|
||||||
| **Image** | RMBG-2.0 (background removal) |
|
|
||||||
|
|
||||||
## Documentation
|
|
||||||
|
|
||||||
| Document | Description |
|
|
||||||
|----------|-------------|
|
|
||||||
| [Getting Started](docs/getting-started.md) | First steps with aha |
|
|
||||||
| [Installation](docs/installation.md) | Detailed installation guide |
|
|
||||||
| [CLI Reference](docs/cli.md) | Command-line interface |
|
|
||||||
| [API Documentation](docs/api.md) | Library & REST API |
|
|
||||||
| [Supported Models](docs/supported-models.md) | Available AI models |
|
|
||||||
| [Concepts](docs/concepts.md) | Architecture & design |
|
|
||||||
| [Development](docs/development.md) | Contributing guide |
|
|
||||||
| [Changelog](docs/changelog.md) | Version history |
|
|
||||||
|
|
||||||
## Why aha?
|
|
||||||
- **🚀 High-Performance Inference** - Powered by Candle framework for efficient tensor computation and model inference
|
|
||||||
- **🔧 Unified Interface** — One tool for text, vision, speech, and OCR
|
|
||||||
- **📦 Local-First** — All processing runs locally, no data leaves your machine
|
|
||||||
- **🎯 Cross-Platform** — Works on Linux, macOS, and Windows
|
|
||||||
- **⚡ GPU Accelerated** — Optional CUDA support for faster inference
|
|
||||||
- **🛡️ Memory Safe** — Built with Rust for reliability
|
|
||||||
- **🧠 Attention Optimization** - Optional Flash Attention support for optimized long sequence processing
|
|
||||||
|
|
||||||
## Development
|
|
||||||
|
|
||||||
### Using aha as a Library
|
|
||||||
> cargo add aha
|
|
||||||
|
|
||||||
```bash
|
|
||||||
# VoxCPM example
|
|
||||||
use aha::models::voxcpm::generate::VoxCPMGenerate;
|
|
||||||
use aha::utils::audio_utils::save_wav;
|
|
||||||
use anyhow::Result;
|
|
||||||
|
|
||||||
fn main() -> Result<()> {
|
|
||||||
let model_path = "xxx/openbmb/VoxCPM-0.5B/";
|
|
||||||
|
|
||||||
let mut voxcpm_generate = VoxCPMGenerate::init(model_path, None, None)?;
|
|
||||||
|
|
||||||
let generate = voxcpm_generate.generate(
|
|
||||||
"The sun is shining bright, flowers smile at me, birds say early early early".to_string(),
|
|
||||||
None,
|
|
||||||
None,
|
|
||||||
2,
|
|
||||||
100,
|
|
||||||
10,
|
|
||||||
2.0,
|
|
||||||
false,
|
|
||||||
6.0,
|
|
||||||
)?;
|
|
||||||
|
|
||||||
let _ = save_wav(&generate, "voxcpm.wav")?;
|
|
||||||
Ok(())
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
### Extending New Models
|
|
||||||
|
|
||||||
- Create new model file in src/models/
|
|
||||||
- Export in src/models/mod.rs
|
|
||||||
- Add support for CLI model inference in src/exec/
|
|
||||||
- Add tests and examples in tests/
|
|
||||||
|
|
||||||
## Features
|
|
||||||
|
|
||||||
- High-performance inference via Candle framework
|
|
||||||
- Multi-modal model support (vision, language, speech)
|
|
||||||
- Clean, easy-to-use API design
|
|
||||||
- Minimal dependencies, compact binaries
|
|
||||||
- Flash Attention support for long sequences
|
|
||||||
- FFmpeg support for multimedia processing
|
|
||||||
|
|
||||||
## License
|
|
||||||
|
|
||||||
Apache-2.0 — See [LICENSE](LICENSE) for details.
|
|
||||||
|
|
||||||
## Acknowledgments
|
|
||||||
|
|
||||||
- [Candle](https://github.com/huggingface/candle) - Excellent Rust ML framework
|
|
||||||
- All model authors and contributors
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
<p align="center">
|
|
||||||
<sub>Built with ❤️ by the aha team</sub>
|
|
||||||
</p>
|
|
||||||
|
|
||||||
<p align="center">
|
|
||||||
<sub>We're continuously expanding our model support. Contributions are welcome!</sub>
|
|
||||||
</p>
|
|
||||||
|
|
||||||
<p align="center">
|
|
||||||
<sub>If this project helps you, please consider giving us a ⭐ Star!</sub>
|
|
||||||
</p>
|
|
||||||
@@ -1,14 +1,8 @@
|
|||||||
<p align="center">
|
<p align="center">
|
||||||
<img src="assets/img/logo.png" alt="aha logo" width="100"/>
|
<img src="assets/img/logo.png" alt="aha logo" width="120"/>
|
||||||
</p>
|
</p>
|
||||||
|
|
||||||
<p align="center">
|
<p align="center">
|
||||||
<!-- <a href="https://github.com/jhqxxx/aha/releases">
|
|
||||||
<img src="https://img.shields.io/github/v/release/jhqxxx/aha" alt="GitHub release (latest by date)">
|
|
||||||
</a>
|
|
||||||
<a href="https://github.com/jhqxxx/aha/actions">
|
|
||||||
<img src="https://img.shields.io/github/actions/workflow/status/jhqxxx/aha/ci.yml" alt="GitHub Actions Workflow Status">
|
|
||||||
</a> -->
|
|
||||||
<a href="https://github.com/jhqxxx/aha/blob/main/LICENSE">
|
<a href="https://github.com/jhqxxx/aha/blob/main/LICENSE">
|
||||||
<img src="https://img.shields.io/github/license/jhqxxx/aha" alt="GitHub License">
|
<img src="https://img.shields.io/github/license/jhqxxx/aha" alt="GitHub License">
|
||||||
</a>
|
</a>
|
||||||
@@ -21,42 +15,42 @@
|
|||||||
</p>
|
</p>
|
||||||
|
|
||||||
<p align="center">
|
<p align="center">
|
||||||
<a href="README.en.md">English</a> | <strong>简体中文</a>
|
<a href="README.zh-CN.md">简体中文</a> | <strong>English</strong>
|
||||||
</p>
|
</p>
|
||||||
|
|
||||||
# aha
|
# aha
|
||||||
|
|
||||||
**轻量 AI 推理引擎 —— 文本、视觉、语音与 OCR 一站式解决方案**
|
**Lightweight AI Inference Engine — All-in-one Solution for Text, Vision, Speech, and OCR**
|
||||||
|
|
||||||
aha 是一款基于 Rust 和 Candle 框架构建的高性能跨平台 AI 推理引擎。将最先进的 AI 模型带到您的本地机器——无需 API 密钥,无需云依赖,纯粹、快速的 AI 直接在您的硬件上运行。
|
aha is a high-performance, cross-platform AI inference engine built with Rust and the Candle framework. It brings state-of-the-art AI models to your local machine—no API keys, no cloud dependencies, just pure, fast AI running directly on your hardware.
|
||||||
|
|
||||||
## 更新日志
|
## Changelog
|
||||||
|
|
||||||
### v0.2.0 (2026-02-05)
|
### v0.2.0 (2026-02-05)
|
||||||
- 新增 Qwen3-ASR 语音识别模型
|
- Added Qwen3-ASR speech recognition model
|
||||||
|
|
||||||
### v0.1.9 (2026-01-31)
|
### v0.1.9 (2026-01-31)
|
||||||
- 新增 CLI `list` 子命令,显示支持的模型
|
- Added CLI `list` subcommand to show supported models
|
||||||
- 新增 CLI 子命令结构支持(`cli`、`serv`、`download`、`run`)
|
- Added CLI subcommand structure support (`cli`, `serv`, `download`, `run`)
|
||||||
- 修复 Qwen3VL thinking startswith bug
|
- Fixed Qwen3VL thinking startswith bug
|
||||||
- 修复 `aha run` 多输入 bug
|
- Fixed `aha run` multiple inputs bug
|
||||||
|
|
||||||
### v0.1.8 (2026-01-17)
|
### v0.1.8 (2026-01-17)
|
||||||
- 新增 Qwen3 文本模型支持
|
- Added Qwen3 text model support
|
||||||
- 新增 Fun-ASR-Nano-2512 语音识别模型
|
- Added Fun-ASR-Nano-2512 speech recognition model
|
||||||
- 修复 ModelScope Fun-ASR-Nano 模型加载错误
|
- Fixed ModelScope Fun-ASR-Nano model load error
|
||||||
- 使用 rubato 更新音频重采样
|
- Updated audio resampling with rubato
|
||||||
|
|
||||||
### v0.1.7 (2026-01-07)
|
### v0.1.7 (2026-01-07)
|
||||||
- 新增 GLM-ASR-Nano-2512 语音识别模型
|
- Added GLM-ASR-Nano-2512 speech recognition model
|
||||||
- 合并 Metal (GPU) 支持,适用于 Apple Silicon
|
- Merged Metal (GPU) support for Apple Silicon
|
||||||
- 新增动态主目录和模型下载脚本
|
- Added dynamic home directory and model download script
|
||||||
|
|
||||||
**[查看完整更新日志](docs/changelog.zh-CN.md)** →
|
**[View full changelog](docs/changelog.md)** →
|
||||||
|
|
||||||
## 快速开始
|
## Quick Start
|
||||||
|
|
||||||
### 安装
|
### Installation
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
git clone https://github.com/jhqxxx/aha.git
|
git clone https://github.com/jhqxxx/aha.git
|
||||||
@@ -64,104 +58,104 @@ cd aha
|
|||||||
cargo build --release
|
cargo build --release
|
||||||
```
|
```
|
||||||
|
|
||||||
**可选特性:**
|
**Optional Features:**
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
# CUDA (NVIDIA GPU 加速)
|
# CUDA (NVIDIA GPU acceleration)
|
||||||
cargo build --release --features cuda
|
cargo build --release --features cuda
|
||||||
|
|
||||||
# Metal (Apple GPU 加速,适用于 macOS)
|
# Metal (Apple GPU acceleration for macOS)
|
||||||
cargo build --release --features metal
|
cargo build --release --features metal
|
||||||
|
|
||||||
# Flash Attention (更快推理)
|
# Flash Attention (faster inference)
|
||||||
cargo build --release --features flash-attn
|
cargo build --release --features flash-attn
|
||||||
|
|
||||||
# FFmpeg (多媒体处理)
|
# FFmpeg (multimedia processing)
|
||||||
cargo build --release --features ffmpeg
|
cargo build --release --features ffmpeg
|
||||||
|
|
||||||
# 组合多个特性
|
# Combine multiple features
|
||||||
cargo build --release --features "cuda,flash-attn"
|
cargo build --release --features "cuda,flash-attn"
|
||||||
```
|
```
|
||||||
|
|
||||||
### CLI 快速参考
|
### CLI Quick Reference
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
|
|
||||||
# 列出所有支持的模型
|
# List all supported models
|
||||||
aha list
|
aha list
|
||||||
|
|
||||||
# 仅下载模型
|
# Download model only
|
||||||
aha download -m qwen3asr-0.6b
|
aha download -m qwen3asr-0.6b
|
||||||
|
|
||||||
# 下载模型并启动服务
|
# Download model and start service
|
||||||
aha -m qwen3asr-0.6b
|
aha -m qwen3asr-0.6b
|
||||||
|
|
||||||
# 直接运行推理(无需启动服务)
|
# Run inference directly (without starting service)
|
||||||
aha run -m qwen3asr-0.6b -i "audio.wav"
|
aha run -m qwen3asr-0.6b -i "audio.wav"
|
||||||
|
|
||||||
# 仅启动服务(模型已下载)
|
# Start service only (model already downloaded)
|
||||||
aha serv -m qwen3asr-0.6b -p 10100
|
aha serv -m qwen3asr-0.6b -p 10100
|
||||||
|
|
||||||
```
|
```
|
||||||
|
|
||||||
### 对话
|
### Chat
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
aha serv -m qwen3-0.6b -p 10100
|
aha serv -m qwen3-0.6b -p 10100
|
||||||
```
|
```
|
||||||
|
|
||||||
然后使用统一(兼容 OpenAI)的 API:
|
Then use the unified (OpenAI-compatible) API:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
curl http://localhost:10100/chat/completions \
|
curl http://localhost:10100/chat/completions \
|
||||||
-H "Content-Type: application/json" \
|
-H "Content-Type: application/json" \
|
||||||
-d '{
|
-d '{
|
||||||
"model": "qwen3-0.6b",
|
"model": "qwen3-0.6b",
|
||||||
"messages": [{"role": "user", "content": "你好!"}]
|
"messages": [{"role": "user", "content": "Hello!"}]
|
||||||
}'
|
}
|
||||||
|
'
|
||||||
```
|
```
|
||||||
|
|
||||||
|
### Supported Models
|
||||||
|
|
||||||
### 支持的模型
|
| Category | Models |
|
||||||
|
|----------|--------|
|
||||||
| 类别 | 模型 |
|
| **Text** | Qwen3, MiniCPM4 |
|
||||||
|------|------|
|
| **Vision** | Qwen2.5-VL, Qwen3-VL |
|
||||||
| **文本** | Qwen3, MiniCPM4 |
|
|
||||||
| **视觉** | Qwen2.5-VL, Qwen3-VL |
|
|
||||||
| **OCR** | DeepSeek-OCR, Hunyuan-OCR, PaddleOCR-VL |
|
| **OCR** | DeepSeek-OCR, Hunyuan-OCR, PaddleOCR-VL |
|
||||||
| **ASR** | GLM-ASR-Nano, Fun-ASR-Nano, Qwen3-ASR |
|
| **ASR** | GLM-ASR-Nano, Fun-ASR-Nano, Qwen3-ASR |
|
||||||
| **音频** | VoxCPM, VoxCPM1.5 |
|
| **Audio** | VoxCPM, VoxCPM1.5 |
|
||||||
| **图像** | RMBG-2.0 (背景移除) |
|
| **Image** | RMBG-2.0 (background removal) |
|
||||||
|
|
||||||
## 文档
|
## Documentation
|
||||||
|
|
||||||
| 文档 | 描述 |
|
| Document | Description |
|
||||||
|------|------|
|
|----------|-------------|
|
||||||
| [快速入门](docs/getting-started.zh-CN.md) | aha 入门指南 |
|
| [Getting Started](docs/getting-started.md) | First steps with aha |
|
||||||
| [安装指南](docs/installation.zh-CN.md) | 详细安装说明 |
|
| [Installation](docs/installation.md) | Detailed installation guide |
|
||||||
| [CLI 参考](docs/cli.zh-CN.md) | 命令行界面 |
|
| [CLI Reference](docs/cli.md) | Command-line interface |
|
||||||
| [API 文档](docs/api.zh-CN.md) | 库与 REST API |
|
| [API Documentation](docs/api.md) | Library & REST API |
|
||||||
| [支持的模型](docs/supported-models.zh-CN.md) | 可用的 AI 模型 |
|
| [Supported Models](docs/supported-models.md) | Available AI models |
|
||||||
| [核心概念](docs/concepts.zh-CN.md) | 架构与设计 |
|
| [Concepts](docs/concepts.md) | Architecture & design |
|
||||||
| [开发指南](docs/development.zh-CN.md) | 贡献指南 |
|
| [Development](docs/development.md) | Contributing guide |
|
||||||
| [更新日志](docs/changelog.zh-CN.md) | 版本历史 |
|
| [Changelog](docs/changelog.md) | Version history |
|
||||||
|
|
||||||
## 为什么选择 aha?
|
## Why aha?
|
||||||
- **🚀 高性能推理** - 基于 Candle 框架,提供高效的张量计算和模型推理
|
- **🚀 High-Performance Inference** - Powered by Candle framework for efficient tensor computation and model inference
|
||||||
- **🔧 统一接口** — 一个工具搞定文本、视觉、语音和 OCR
|
- **🔧 Unified Interface** — One tool for text, vision, speech, and OCR
|
||||||
- **📦 本地优先** — 所有处理在本地运行,数据不离境
|
- **📦 Local-First** — All processing runs locally, no data leaves your machine
|
||||||
- **🎯 跨平台** — 支持 Linux、macOS 和 Windows
|
- **🎯 Cross-Platform** — Works on Linux, macOS, and Windows
|
||||||
- **⚡ GPU 加速** — 可选 CUDA 支持以获得更快推理
|
- **⚡ GPU Accelerated** — Optional CUDA support for faster inference
|
||||||
- **🛡️ 内存安全** — Rust 构建,稳定可靠
|
- **🛡️ Memory Safe** — Built with Rust for reliability
|
||||||
- **🧠 注意力优化** - 可选 Flash Attention 支持,优化长序列处理
|
- **🧠 Attention Optimization** - Optional Flash Attention support for optimized long sequence processing
|
||||||
|
|
||||||
## 开发
|
## Development
|
||||||
|
|
||||||
### aha 作为库使用
|
### Using aha as a Library
|
||||||
> cargo add aha
|
> cargo add aha
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
# VoxCPM示例
|
# VoxCPM example
|
||||||
use aha::models::voxcpm::generate::VoxCPMGenerate;
|
use aha::models::voxcpm::generate::VoxCPMGenerate;
|
||||||
use aha::utils::audio_utils::save_wav;
|
use aha::utils::audio_utils::save_wav;
|
||||||
use anyhow::Result;
|
use anyhow::Result;
|
||||||
@@ -172,7 +166,7 @@ fn main() -> Result<()> {
|
|||||||
let mut voxcpm_generate = VoxCPMGenerate::init(model_path, None, None)?;
|
let mut voxcpm_generate = VoxCPMGenerate::init(model_path, None, None)?;
|
||||||
|
|
||||||
let generate = voxcpm_generate.generate(
|
let generate = voxcpm_generate.generate(
|
||||||
"太阳当空照,花儿对我笑,小鸟说早早早".to_string(),
|
"The sun is shining bright, flowers smile at me, birds say early early early".to_string(),
|
||||||
None,
|
None,
|
||||||
None,
|
None,
|
||||||
2,
|
2,
|
||||||
@@ -188,43 +182,40 @@ fn main() -> Result<()> {
|
|||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
|
### Extending New Models
|
||||||
|
|
||||||
### 扩展新的模型
|
- Create new model file in src/models/
|
||||||
|
- Export in src/models/mod.rs
|
||||||
|
- Add support for CLI model inference in src/exec/
|
||||||
|
- Add tests and examples in tests/
|
||||||
|
|
||||||
- 在src/models/创建新模型文件
|
## Features
|
||||||
- 在src/models/mod.rs中导出
|
|
||||||
- 在src/exec/中添加支持cli运行模型推理
|
|
||||||
- 在tests/中添加测试和示例
|
|
||||||
|
|
||||||
|
- High-performance inference via Candle framework
|
||||||
|
- Multi-modal model support (vision, language, speech)
|
||||||
|
- Clean, easy-to-use API design
|
||||||
|
- Minimal dependencies, compact binaries
|
||||||
|
- Flash Attention support for long sequences
|
||||||
|
- FFmpeg support for multimedia processing
|
||||||
|
|
||||||
## 特性
|
## License
|
||||||
|
|
||||||
- 基于 Candle 框架的高性能推理
|
Apache-2.0 — See [LICENSE](LICENSE) for details.
|
||||||
- 多模态模型支持(视觉、语言、语音)
|
|
||||||
- 简洁易用的 API 设计
|
|
||||||
- 最小化依赖,紧凑的二进制文件
|
|
||||||
- Flash Attention 支持长序列处理
|
|
||||||
- FFmpeg 支持多媒体处理
|
|
||||||
|
|
||||||
## 许可证
|
## Acknowledgments
|
||||||
|
|
||||||
Apache-2.0 — 详见 [LICENSE](LICENSE)
|
- [Candle](https://github.com/huggingface/candle) - Excellent Rust ML framework
|
||||||
|
- All model authors and contributors
|
||||||
## 致谢
|
|
||||||
|
|
||||||
- [Candle](https://github.com/huggingface/candle) - 优秀的 Rust 机器学习框架
|
|
||||||
- 所有模型作者和贡献者
|
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
<p align="center">
|
<p align="center">
|
||||||
<sub>由 aha 团队用 ❤️ 构建</sub>
|
<sub>Built with ❤️ by the aha team</sub>
|
||||||
</p>
|
</p>
|
||||||
|
|
||||||
<p align="center">
|
<p align="center">
|
||||||
<sub>我们持续扩展支持的模型列表,欢迎贡献!</sub>
|
<sub>We're continuously expanding our model support. Contributions are welcome!</sub>
|
||||||
</p>
|
</p>
|
||||||
|
|
||||||
<p align="center">
|
<p align="center">
|
||||||
<sub>如果这个项目对你有帮助,请给我们一个 ⭐ Star!</sub>
|
<sub>If this project helps you, please consider giving us a ⭐ Star!</sub>
|
||||||
</p>
|
</p>
|
||||||
|
|||||||
+224
@@ -0,0 +1,224 @@
|
|||||||
|
<p align="center">
|
||||||
|
<img src="assets/img/logo.png" alt="aha logo" width="120"/>
|
||||||
|
</p>
|
||||||
|
|
||||||
|
<p align="center">
|
||||||
|
<a href="https://github.com/jhqxxx/aha/blob/main/LICENSE">
|
||||||
|
<img src="https://img.shields.io/github/license/jhqxxx/aha" alt="GitHub License">
|
||||||
|
</a>
|
||||||
|
<a href="https://github.com/jhqxxx/aha/stargazers">
|
||||||
|
<img src="https://img.shields.io/github/stars/jhqxxx/aha" alt="GitHub Stars">
|
||||||
|
</a>
|
||||||
|
<a href="https://github.com/jhqxxx/aha/issues">
|
||||||
|
<img src="https://img.shields.io/github/issues/jhqxxx/aha" alt="GitHub Issues">
|
||||||
|
</a>
|
||||||
|
</p>
|
||||||
|
|
||||||
|
<p align="center">
|
||||||
|
<a href="README.md">English</a> | <strong>简体中文</strong>
|
||||||
|
</p>
|
||||||
|
|
||||||
|
# aha
|
||||||
|
|
||||||
|
**轻量 AI 推理引擎 —— 文本、视觉、语音与 OCR 一站式解决方案**
|
||||||
|
|
||||||
|
aha 是一款基于 Rust 和 Candle 框架构建的高性能跨平台 AI 推理引擎。将最先进的 AI 模型带到您的本地机器——无需 API 密钥,无需云依赖,纯粹、快速的 AI 直接在您的硬件上运行。
|
||||||
|
|
||||||
|
## 更新日志
|
||||||
|
|
||||||
|
### v0.2.0 (2026-02-05)
|
||||||
|
- 新增 Qwen3-ASR 语音识别模型
|
||||||
|
|
||||||
|
### v0.1.9 (2026-01-31)
|
||||||
|
- 新增 CLI `list` 子命令,显示支持的模型
|
||||||
|
- 新增 CLI 子命令结构支持(`cli`、`serv`、`download`、`run`)
|
||||||
|
- 修复 Qwen3VL thinking startswith bug
|
||||||
|
- 修复 `aha run` 多输入 bug
|
||||||
|
|
||||||
|
### v0.1.8 (2026-01-17)
|
||||||
|
- 新增 Qwen3 文本模型支持
|
||||||
|
- 新增 Fun-ASR-Nano-2512 语音识别模型
|
||||||
|
- 修复 ModelScope Fun-ASR-Nano 模型加载错误
|
||||||
|
- 使用 rubato 更新音频重采样
|
||||||
|
|
||||||
|
### v0.1.7 (2026-01-07)
|
||||||
|
- 新增 GLM-ASR-Nano-2512 语音识别模型
|
||||||
|
- 合并 Metal (GPU) 支持,适用于 Apple Silicon
|
||||||
|
- 新增动态主目录和模型下载脚本
|
||||||
|
|
||||||
|
**[查看完整更新日志](docs/changelog.zh-CN.md)** →
|
||||||
|
|
||||||
|
## 快速开始
|
||||||
|
|
||||||
|
### 安装
|
||||||
|
|
||||||
|
```bash
|
||||||
|
git clone https://github.com/jhqxxx/aha.git
|
||||||
|
cd aha
|
||||||
|
cargo build --release
|
||||||
|
```
|
||||||
|
|
||||||
|
**可选特性:**
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# CUDA (NVIDIA GPU 加速)
|
||||||
|
cargo build --release --features cuda
|
||||||
|
|
||||||
|
# Metal (Apple GPU 加速,适用于 macOS)
|
||||||
|
cargo build --release --features metal
|
||||||
|
|
||||||
|
# Flash Attention (更快推理)
|
||||||
|
cargo build --release --features flash-attn
|
||||||
|
|
||||||
|
# FFmpeg (多媒体处理)
|
||||||
|
cargo build --release --features ffmpeg
|
||||||
|
|
||||||
|
# 组合多个特性
|
||||||
|
cargo build --release --features "cuda,flash-attn"
|
||||||
|
```
|
||||||
|
|
||||||
|
### CLI 快速参考
|
||||||
|
|
||||||
|
```bash
|
||||||
|
|
||||||
|
# 列出所有支持的模型
|
||||||
|
aha list
|
||||||
|
|
||||||
|
# 仅下载模型
|
||||||
|
aha download -m qwen3asr-0.6b
|
||||||
|
|
||||||
|
# 下载模型并启动服务
|
||||||
|
aha -m qwen3asr-0.6b
|
||||||
|
|
||||||
|
# 直接运行推理(无需启动服务)
|
||||||
|
aha run -m qwen3asr-0.6b -i "audio.wav"
|
||||||
|
|
||||||
|
# 仅启动服务(模型已下载)
|
||||||
|
aha serv -m qwen3asr-0.6b -p 10100
|
||||||
|
|
||||||
|
```
|
||||||
|
|
||||||
|
### 对话
|
||||||
|
|
||||||
|
```bash
|
||||||
|
aha serv -m qwen3-0.6b -p 10100
|
||||||
|
```
|
||||||
|
|
||||||
|
然后使用统一(兼容 OpenAI)的 API:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
curl http://localhost:10100/chat/completions \
|
||||||
|
-H "Content-Type: application/json" \
|
||||||
|
-d '{
|
||||||
|
"model": "qwen3-0.6b",
|
||||||
|
"messages": [{"role": "user", "content": "你好!"}]
|
||||||
|
}'
|
||||||
|
```
|
||||||
|
|
||||||
|
|
||||||
|
### 支持的模型
|
||||||
|
|
||||||
|
| 类别 | 模型 |
|
||||||
|
|------|------|
|
||||||
|
| **文本** | Qwen3, MiniCPM4 |
|
||||||
|
| **视觉** | Qwen2.5-VL, Qwen3-VL |
|
||||||
|
| **OCR** | DeepSeek-OCR, Hunyuan-OCR, PaddleOCR-VL |
|
||||||
|
| **ASR** | GLM-ASR-Nano, Fun-ASR-Nano,Qwen3-ASR |
|
||||||
|
| **音频** | VoxCPM, VoxCPM1.5 |
|
||||||
|
| **图像** | RMBG-2.0 (背景移除) |
|
||||||
|
|
||||||
|
## 文档
|
||||||
|
|
||||||
|
| 文档 | 描述 |
|
||||||
|
|------|------|
|
||||||
|
| [快速入门](docs/getting-started.zh-CN.md) | aha 入门指南 |
|
||||||
|
| [安装指南](docs/installation.zh-CN.md) | 详细安装说明 |
|
||||||
|
| [CLI 参考](docs/cli.zh-CN.md) | 命令行界面 |
|
||||||
|
| [API 文档](docs/api.zh-CN.md) | 库与 REST API |
|
||||||
|
| [支持的模型](docs/supported-models.zh-CN.md) | 可用的 AI 模型 |
|
||||||
|
| [核心概念](docs/concepts.zh-CN.md) | 架构与设计 |
|
||||||
|
| [开发指南](docs/development.zh-CN.md) | 贡献指南 |
|
||||||
|
| [更新日志](docs/changelog.zh-CN.md) | 版本历史 |
|
||||||
|
|
||||||
|
## 为什么选择 aha?
|
||||||
|
- **🚀 高性能推理** - 基于 Candle 框架,提供高效的张量计算和模型推理
|
||||||
|
- **🔧 统一接口** — 一个工具搞定文本、视觉、语音和 OCR
|
||||||
|
- **📦 本地优先** — 所有处理在本地运行,数据不离境
|
||||||
|
- **🎯 跨平台** — 支持 Linux、macOS 和 Windows
|
||||||
|
- **⚡ GPU 加速** — 可选 CUDA 支持以获得更快推理
|
||||||
|
- **🛡️ 内存安全** — Rust 构建,稳定可靠
|
||||||
|
- **🧠 注意力优化** - 可选 Flash Attention 支持,优化长序列处理
|
||||||
|
|
||||||
|
## 开发
|
||||||
|
|
||||||
|
### aha 作为库使用
|
||||||
|
> cargo add aha
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# VoxCPM示例
|
||||||
|
use aha::models::voxcpm::generate::VoxCPMGenerate;
|
||||||
|
use aha::utils::audio_utils::save_wav;
|
||||||
|
use anyhow::Result;
|
||||||
|
|
||||||
|
fn main() -> Result<()> {
|
||||||
|
let model_path = "xxx/openbmb/VoxCPM-0.5B/";
|
||||||
|
|
||||||
|
let mut voxcpm_generate = VoxCPMGenerate::init(model_path, None, None)?;
|
||||||
|
|
||||||
|
let generate = voxcpm_generate.generate(
|
||||||
|
"太阳当空照,花儿对我笑,小鸟说早早早".to_string(),
|
||||||
|
None,
|
||||||
|
None,
|
||||||
|
2,
|
||||||
|
100,
|
||||||
|
10,
|
||||||
|
2.0,
|
||||||
|
false,
|
||||||
|
6.0,
|
||||||
|
)?;
|
||||||
|
|
||||||
|
let _ = save_wav(&generate, "voxcpm.wav")?;
|
||||||
|
Ok(())
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
|
||||||
|
### 扩展新的模型
|
||||||
|
|
||||||
|
- 在src/models/创建新模型文件
|
||||||
|
- 在src/models/mod.rs中导出
|
||||||
|
- 在src/exec/中添加支持cli运行模型推理
|
||||||
|
- 在tests/中添加测试和示例
|
||||||
|
|
||||||
|
|
||||||
|
## 特性
|
||||||
|
|
||||||
|
- 基于 Candle 框架的高性能推理
|
||||||
|
- 多模态模型支持(视觉、语言、语音)
|
||||||
|
- 简洁易用的 API 设计
|
||||||
|
- 最小化依赖,紧凑的二进制文件
|
||||||
|
- Flash Attention 支持长序列处理
|
||||||
|
- FFmpeg 支持多媒体处理
|
||||||
|
|
||||||
|
## 许可证
|
||||||
|
|
||||||
|
Apache-2.0 — 详见 [LICENSE](LICENSE)
|
||||||
|
|
||||||
|
## 致谢
|
||||||
|
|
||||||
|
- [Candle](https://github.com/huggingface/candle) - 优秀的 Rust 机器学习框架
|
||||||
|
- 所有模型作者和贡献者
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
<p align="center">
|
||||||
|
<sub>由 aha 团队用 ❤️ 构建</sub>
|
||||||
|
</p>
|
||||||
|
|
||||||
|
<p align="center">
|
||||||
|
<sub>我们持续扩展支持的模型列表,欢迎贡献!</sub>
|
||||||
|
</p>
|
||||||
|
|
||||||
|
<p align="center">
|
||||||
|
<sub>如果这个项目对你有帮助,请给我们一个 ⭐ Star!</sub>
|
||||||
|
</p>
|
||||||
Reference in New Issue
Block a user