Back to Projects
🎙️ Speech Recognition🤗 HuggingFace⚡ Hubert Model🎵 Torchaudio🎈 Streamlit

Speech2Text - Audio Transcription

Advanced speech-to-text transcriber using HuggingFace transformers and Facebook's Hubert model for accurate audio transcription.


🎯 What This Project Does

  • 🎙️ Multi-Format Support - Process MP3, MP4, WAV audio and video files
  • 🧠 AI-Powered Transcription - Uses Facebook's Hubert model for accurate speech recognition
  • 📝 Downloadable Transcripts - Get clean, formatted text output
  • ⏱️ Timestamped Segments - View transcription with time markers

🌐 Live Demo

📺 Demo Video


🚀 How to Use

Using the App

  1. Upload Audio/Video - Support for .mp3, .mp4, .wav formats
  2. Watch Processing - Real-time transcription progress
  3. View Results - Timestamped segments and full transcript
  4. Download Text - Export transcript in clean text format

Local Development

BASH
# Clone and setup
git clone https://github.com/aishwaryaj7/speech2text.git
cd speech2text

# Install dependencies
pip install -r requirements.txt

# Run the app
streamlit run src/app.py

🔧 Key Features

  • Multi-Format Support: MP3, MP4, WAV audio and video files
  • AI-Powered Transcription: Facebook's Hubert model for accurate speech recognition
  • Timestamped Segments: Precise time markers for each phrase
  • Real-time Processing: Live transcription with progress indicators
  • Download Options: Export transcripts in multiple formats
  • User-Friendly Interface: Intuitive Streamlit interface

🛠️ Tech Stack

AI/ML: HuggingFace Transformers, Hubert Model, Wav2Vec2 Audio: Torchaudio, Librosa, Pydub Frontend: Streamlit Deployment: Streamlit Cloud


🤝 Skills Demonstrated

  • Speech Recognition & Audio Processing
  • HuggingFace Transformers & Model Deployment
  • Streamlit Development & Cloud Deployment
  • Signal Processing & Feature Extraction
  • User Experience Design