Automatic Speech Recognition
Transformers
Safetensors
Chinese
English
audio8_asr_infinite
text-generation
streaming
realtime
speech-recognition
audio
custom_code
Instructions to use Edge0/Audio8-ASR-Infinite with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Edge0/Audio8-ASR-Infinite with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("automatic-speech-recognition", model="Edge0/Audio8-ASR-Infinite", trust_remote_code=True)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("Edge0/Audio8-ASR-Infinite", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
Update README.md
Browse files
README.md
CHANGED
|
@@ -31,14 +31,21 @@ With our adapted vLLM build it transcribes unlimited-length audio **24/7** witho
|
|
| 31 |
## Highlights
|
| 32 |
|
| 33 |
- **Super responsive** β the native streaming architecture decodes 12.5 times per second.
|
| 34 |
-
- **Unlimited-length transcription** β a rolling KV
|
| 35 |
-
latency
|
| 36 |
- **Selectable streaming clock** β one text token per clock step
|
| 37 |
(12.5 / 8.3 / 6.25 decisions per second), balancing perception granularity and resource cost.
|
| 38 |
- **Configurable transcription delay** β set how much delay to trade for accuracy.
|
| 39 |
-
- **Semantic VAD** β distinguishes thinking pauses, stuttering and real end of turn, where traditional acoustic VAD fails.
|
| 40 |
- **Bilingual** β Chinese and English.
|
| 41 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 42 |
## Optimized operation points
|
| 43 |
|
| 44 |
The following combinations of frame length and delay are post-trained. Other combinations can be used but performance may not be optimum.
|
|
|
|
| 31 |
## Highlights
|
| 32 |
|
| 33 |
- **Super responsive** β the native streaming architecture decodes 12.5 times per second.
|
| 34 |
+
- **Unlimited-length transcription** β a rolling KV Cache keeps both **memory and
|
| 35 |
+
latency constant**, even in **24/7 operation**.
|
| 36 |
- **Selectable streaming clock** β one text token per clock step
|
| 37 |
(12.5 / 8.3 / 6.25 decisions per second), balancing perception granularity and resource cost.
|
| 38 |
- **Configurable transcription delay** β set how much delay to trade for accuracy.
|
| 39 |
+
- **Semantic VAD** β distinguishes thinking pauses, stuttering and real end of turn, where traditional acoustic VAD usually fails.
|
| 40 |
- **Bilingual** β Chinese and English.
|
| 41 |
|
| 42 |
+
## See Audio8-ASR-Infinite in action
|
| 43 |
+
|
| 44 |
+
The checkpoint has a native context of 30 seconds. But with Rolling KV Cache, it can transcribe 24/7 nonstop.
|
| 45 |
+
|
| 46 |
+
<video controls playsinline width="100%" preload="metadata"
|
| 47 |
+
src="https://huggingface.co/Edge0/Audio8-ASR-Infinite/resolve/main/Audio8-Asr-Infinite-Demo.mp4"></video>
|
| 48 |
+
|
| 49 |
## Optimized operation points
|
| 50 |
|
| 51 |
The following combinations of frame length and delay are post-trained. Other combinations can be used but performance may not be optimum.
|