Skip to content
 
 

Repository files navigation

🎙️ Lazy Speech To Text Converter

Forked from webrtc-speech-to-text, with a Vue 3 frontend, multi-vendor transcription (Whisper, Google, Azure, Baidu, Xunfei), audio recording, and file management.

Demo Screenshot

Transform speech to text effortlessly with WebRTC and AI


✨ Features

Feature Description
🎤 Real-time Streaming WebRTC-based audio capture with low latency
🌍 99+ Languages Powered by Whisper AI - works offline
🔒 Privacy First Local processing - your audio never leaves your machine
📱 Cross-platform Works on Chrome, Firefox, and Safari
🎛️ Flexible Options Record only, transcribe only, or both
📊 Visual Feedback Real-time audio waveform visualization
🔐 User Authentication Simple login system for access control

🚀 Quick Start

# 1. Install Whisper (one-time)
pip install whisper-ctranslate2

# 2. Build the project
make

# 3. Run the server
./webrtc-transcriber

# 4. Open browser
open http://localhost:9070

That's it! No cloud accounts, no API keys, no configuration needed.

See Runbook for prerequisites, troubleshooting, and alternative vendors.


🏗️ Architecture

┌─────────────┐     WebRTC      ┌─────────────────┐     Audio      ┌──────────────┐
│   Browser   │ ◄─────────────► │  Go Server      │ ─────────────► │  Whisper AI  │
│  (Vue 3)    │                 │  (Pion WebRTC)  │                │  (or Cloud)  │
└─────────────┘                 └─────────────────┘                └──────────────┘
       │                               │                                  │
       │    DataChannel               │                                  │
       ◄──────────────────────────────┼──────────────────────────────────┘
              (Transcription Results)
Layer Technology
Backend Go + Pion WebRTC + Opus
Frontend Vue 3 + TypeScript + Vite + Tailwind CSS
Transcription Whisper (default), Google, Azure, Baidu, Xunfei

See Architecture for component diagrams, call chains, and module dependencies.


⚙️ Usage

./webrtc-transcriber [options]

  --vendor string     whisper | google | azure | baidu | xunfei | recorder (default "whisper")
  --model string      tiny | base | small | medium | large (default "small")
  --language string   en | zh | ja | auto | ... (default "auto")
  --output string     Output directory (default "recordings")
  --http.port string  HTTP port (default "9070")

Copy env.example to .env for cloud vendor credentials and account configuration.

See Data & API for full HTTP API contracts and Conventions for configuration details.


📚 Documentation

The doc/ directory contains a comprehensive Project Knowledge Base built with Sphinx + MyST Markdown.

cd doc/
pip install -r requirements.txt   # One-time setup
make html                          # Build HTML documentation
make serve                         # Build with live reload
Document Description
Project Overview Purpose, boundaries, tech stack, deployment model
Repository Map Directory structure, entry points, naming conventions
Architecture Component diagram, call chains, module dependencies
Workflows Real-time transcription, file management, vendor selection
Data & API HTTP API contracts, DataChannel messages, Go interfaces
Conventions Code style, error handling, config management
Runbook Build, run, debug, and troubleshoot
Testing Test strategy, coverage targets, critical path checklist
AI Guide Orientation guide for AI assistants

Vendor Setup Guides


⚠️ Disclaimer

This project is a proof of concept and should not be deployed in production without implementing proper security measures.

📄 License

MIT - see LICENSE for details.


Made with ❤️ by Walter Fan

About

ASR using WebRTC

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages