Doc Parser - PDF Processing & Table Querying
Smart PDF extraction tool that converts documents to structured JSON/Markdown with AI-powered table querying.
🎯 What This Project Does
- 📄 Complete PDF Extraction - Text, images, and metadata from any PDF
- 🔄 Format Conversion - Transform content to structured JSON or Markdown
- 🤖 AI-Powered Querying - Ask questions about tables using OpenAI's API
- 💾 Download Everything - Export extracted data, images, and formatted content
🌐 Live Demo
🚀 How to Use
Using the App
- Upload PDF - Drag and drop any PDF document
- Choose Format - Select JSON or Markdown conversion
- Extract Content - Get text, images, and metadata automatically
- Query Tables - Ask AI questions about tabular data
- Download Results - Export everything in your preferred format
Local Development
BASH
# Clone and setup
git clone https://github.com/aishwaryaj7/doc_parser.git
cd doc_parser
# Install dependencies
pip install -r requirements.txt
# Run the app
streamlit run src/app.py
🔧 Key Features
- Complete PDF Extraction: Text, images, tables, and metadata
- Multiple Output Formats: JSON and Markdown conversion
- AI-Powered Querying: OpenAI integration for table analysis
- Drag-and-Drop Interface: Intuitive file upload
- Real-time Processing: Immediate feedback and progress indicators
- Download Options: Export data in multiple formats
🛠️ Tech Stack
Processing: PyMuPDF, Python AI: OpenAI API (gpt-3.5-turbo-instruct) Frontend: Streamlit Deployment: Streamlit Cloud
🤝 Skills Demonstrated
- Document Processing & PDF Parsing
- AI Integration & Natural Language Querying
- Streamlit Development & Cloud Deployment
- Data Transformation & Format Conversion
- User Experience Design