Speech2Text - Audio Transcription
Advanced speech-to-text transcriber using HuggingFace transformers and Facebook's Hubert model for accurate audio transcription.
🎯 What This Project Does
- 🎙️ Multi-Format Support - Process MP3, MP4, WAV audio and video files
- 🧠 AI-Powered Transcription - Uses Facebook's Hubert model for accurate speech recognition
- 📝 Downloadable Transcripts - Get clean, formatted text output
- ⏱️ Timestamped Segments - View transcription with time markers
🌐 Live Demo
🚀 How to Use
Using the App
- Upload Audio/Video - Support for .mp3, .mp4, .wav formats
- Watch Processing - Real-time transcription progress
- View Results - Timestamped segments and full transcript
- Download Text - Export transcript in clean text format
Local Development
BASH
# Clone and setup
git clone https://github.com/aishwaryaj7/speech2text.git
cd speech2text
# Install dependencies
pip install -r requirements.txt
# Run the app
streamlit run src/app.py
🔧 Key Features
- Multi-Format Support: MP3, MP4, WAV audio and video files
- AI-Powered Transcription: Facebook's Hubert model for accurate speech recognition
- Timestamped Segments: Precise time markers for each phrase
- Real-time Processing: Live transcription with progress indicators
- Download Options: Export transcripts in multiple formats
- User-Friendly Interface: Intuitive Streamlit interface
🛠️ Tech Stack
AI/ML: HuggingFace Transformers, Hubert Model, Wav2Vec2 Audio: Torchaudio, Librosa, Pydub Frontend: Streamlit Deployment: Streamlit Cloud
🤝 Skills Demonstrated
- Speech Recognition & Audio Processing
- HuggingFace Transformers & Model Deployment
- Streamlit Development & Cloud Deployment
- Signal Processing & Feature Extraction
- User Experience Design