Back to Projects
📄 PyMuPDF🤖 OpenAI API🔄 JSON/Markdown🎈 Streamlit

Doc Parser - PDF Processing & Table Querying

Smart PDF extraction tool that converts documents to structured JSON/Markdown with AI-powered table querying.


🎯 What This Project Does

  • 📄 Complete PDF Extraction - Text, images, and metadata from any PDF
  • 🔄 Format Conversion - Transform content to structured JSON or Markdown
  • 🤖 AI-Powered Querying - Ask questions about tables using OpenAI's API
  • 💾 Download Everything - Export extracted data, images, and formatted content

🌐 Live Demo

📺 Demo Video


🚀 How to Use

Using the App

  1. Upload PDF - Drag and drop any PDF document
  2. Choose Format - Select JSON or Markdown conversion
  3. Extract Content - Get text, images, and metadata automatically
  4. Query Tables - Ask AI questions about tabular data
  5. Download Results - Export everything in your preferred format

Local Development

BASH
# Clone and setup
git clone https://github.com/aishwaryaj7/doc_parser.git
cd doc_parser

# Install dependencies
pip install -r requirements.txt

# Run the app
streamlit run src/app.py

🔧 Key Features

  • Complete PDF Extraction: Text, images, tables, and metadata
  • Multiple Output Formats: JSON and Markdown conversion
  • AI-Powered Querying: OpenAI integration for table analysis
  • Drag-and-Drop Interface: Intuitive file upload
  • Real-time Processing: Immediate feedback and progress indicators
  • Download Options: Export data in multiple formats

🛠️ Tech Stack

Processing: PyMuPDF, Python AI: OpenAI API (gpt-3.5-turbo-instruct) Frontend: Streamlit Deployment: Streamlit Cloud


🤝 Skills Demonstrated

  • Document Processing & PDF Parsing
  • AI Integration & Natural Language Querying
  • Streamlit Development & Cloud Deployment
  • Data Transformation & Format Conversion
  • User Experience Design