Sitelet https://github.com/asincole/napi-pdf-parser
Skip to content

Repository files navigation

PDF Parser

A high-performance PDF parser built with Rust and Node.js, featuring intelligent text extraction with automatic OCR fallback powered by PaddleOCR v5.

Features

  • Multiple parsing strategies: Native extraction, OCR-only, or hybrid mode
  • Intelligent fallback: Automatically switches to OCR when native extraction yields insufficient text
  • Multi-format output: Extracts text, images, and structured metadata
  • High performance: Built with Rust and pdfium for fast PDF rendering
  • OCR support: PaddleOCR v5 models for accurate text recognition
  • Cross-platform: Supports macOS and Linux (x64/ARM64)

Prerequisites

  • Node.js >= 12.22.0 (see package.json for exact version support)
  • Rust (latest stable version)
  • OCR Models (see Models section)

Installation

# Install dependencies
pnpm install

# Build the native addon
pnpm build

Models

This parser requires PaddleOCR v5 models for OCR functionality. Download the following files from HuggingFace PP-OCRv5 Collection:

Required Files

Place these files in the models/ directory:

  1. pp-ocrv5_mobile_det.onnx - Text detection model
  2. pp-ocrv5_mobile_rec.onnx - Text recognition model
  3. ppocrv5_dict.txt - Character dictionary

Model Variants

Mobile Models (Recommended for most use cases):

  • Faster inference, smaller size (~4.6MB det + ~15.8MB rec)
  • Suitable for real-time processing
  • Good accuracy for standard documents

Server Models (For higher accuracy):

  • Larger size (~84MB det + ~80MB rec)
  • Better for complex layouts, handwritten text
  • Higher computational requirements

Download links:

Quick Start

import { parse_pdf, parse_pdf_native, parse_pdf_ocr } from './dist/index'

// Hybrid mode (recommended) - tries native extraction first, falls back to OCR
const result = await parse_pdf('/path/to/document.pdf', './output')

console.log(result.metadata)
// {
//   pdf_name: "document",
//   text_length: 5234,
//   has_real_words: true,
//   image_count: 5,
//   pdf_size_bytes: 1048576,
//   used_ocr: false
// }

API Reference

parse_pdf(pdf_path: string, output_dir: string): Promise<PdfParseResult>

Hybrid mode - Intelligently combines native extraction with OCR fallback.

  • First attempts native text extraction
  • Falls back to OCR if:
    • Extracted text < 500 characters, OR
    • No real words detected (alphabetic tokens ≥ 2 characters)
  • Best for unknown PDF types

Example:

const result = await parse_pdf('./invoice.pdf', './parsed')

console.log(result.text_path) // ./parsed/invoice/text.txt
console.log(result.images_dir) // ./parsed/invoice/images/
console.log(result.metadata.used_ocr) // true or false

parse_pdf_native(pdf_path: string, output_dir: string): Promise<PdfParseResult>

Native extraction only - Extracts text and images using pdfium.

  • Fast text extraction from text-based PDFs
  • Renders pages to JPEG images
  • No OCR processing
  • Best for digitally-created PDFs

Example:

const result = await parse_pdf_native('./report.pdf', './output')
// Uses only native extraction, never falls back to OCR

parse_pdf_ocr(pdf_path: string, output_dir: string): Promise<void>

OCR-only mode - Renders all pages and performs OCR.

  • Renders each page to JPEG
  • Runs PaddleOCR inference
  • Outputs text and structured CSV
  • Best for scanned documents or image-based PDFs

Example:

await parse_pdf_ocr('./scanned-doc.pdf', './ocr-output')
// All pages processed through OCR pipeline

Output Types

PdfParseResult

interface PdfParseResult {
  pdf_path: string // Original PDF path
  output_path: string // Output directory path
  text_path: string // Path to extracted text file
  images_dir: string // Path to images directory
  metadata: PdfMetadata // Document metadata
}

PdfMetadata

interface PdfMetadata {
  pdf_name: string // PDF filename without extension
  text_length: number // Total characters extracted
  has_real_words: boolean // Whether real words were detected
  image_count: number // Number of page images rendered
  pdf_size_bytes: number // Original PDF file size
  used_ocr: boolean // Whether OCR was used
}

Output Structure

After parsing, the output directory contains:

output_dir/
└── document_name/
    ├── text.txt          # Extracted text content
    ├── metadata.json     # Document metadata
    ├── data.csv          # OCR results (when OCR used)
    └── images/
        ├── 0.jpg         # Page 0 rendered as JPEG
        ├── 1.jpg         # Page 1
        └── ...

Files

  • text.txt: Plain text extraction (newline-separated pages)
  • metadata.json: JSON file with PdfMetadata structure
  • data.csv: OCR results with confidence scores (CSV format)
  • images/: Rendered PDF pages as JPEGs (2000px max dimension)

Parsing Strategies

When to Use Each Function

Function Use Case Speed Accuracy
parse_pdf_native Digital PDFs with embedded text ⚡⚡⚡ Fastest ✓ Good for text PDFs
parse_pdf_ocr Scanned documents, images ⚡ Slow (OCR) ✓✓ Best for scans
parse_pdf Unknown/mixed PDF types ⚡⚡ Adaptive ✓✓ Optimal

Fallback Logic

The parse_pdf function uses this heuristic:

  1. Extract text natively using pdfium
  2. Check if text contains real words (alphabetic tokens ≥ 2 chars)
  3. If text < 500 chars OR no real words detected:
    • Render all pages to images
    • Run OCR on each page
    • Replace text with OCR results
  4. Mark used_ocr flag in metadata

Performance & Limitations

Constraints

  • Maximum PDF size: 25 MB (26,214,400 bytes)
  • Maximum image size: 10 MB per rendered page
  • OCR confidence threshold: 0.7 (text below this is filtered)
  • Image resolution: 2000px max width/height

Performance Characteristics

  • Native extraction: ~50-200ms per page (text-only)
  • Image rendering: ~100-300ms per page
  • OCR processing: ~500-2000ms per page (depends on complexity)
  • Memory usage: Proportional to PDF size and page count

Optimization Tips

  1. Use parse_pdf_native for known text-based PDFs
  2. Use mobile OCR models for faster inference
  3. Process large PDFs in batches
  4. Consider server models only for complex documents

Platform Support

Platform Architecture Support
macOS x86_64 (Intel) ✅ Supported
macOS aarch64 (Apple Silicon) ✅ Supported
Linux x86_64 ✅ Supported
Linux aarch64 (ARM64) ✅ Supported
Windows x86_64 ❌ Not in build targets

The native addon is prebuilt for these platforms via GitHub Actions CI.

Development

Building

# Development build
pnpm build:debug

# Production build (optimized)
pnpm build

# Format code
pnpm format

# Lint
pnpm lint

Testing

# Run tests
pnpm test

Benchmarking

# Run benchmarks
pnpm bench

Project Structure

.
├── src/
│   ├── lib.rs                  # NAPI bindings and main functions
│   ├── parse_pdf_native.rs     # Native extraction logic
│   └── pdf_utils.rs            # OCR and utility functions
├── models/                     # OCR model files (not in repo)
├── dist/                       # Built native addon
└── __test__/                   # Test files

Technical Details

Dependencies

  • napi-rs: Node.js N-API bindings for Rust
  • pdfium-render: PDF rendering using Google's pdfium
  • oar-ocr: ONNX runtime for OCR inference
  • image: Image processing (JPEG encoding/decoding)
  • turbojpeg: Fast JPEG compression

OCR Configuration

Default OCR settings (can be modified in src/pdf_utils.rs):

TextDetectionConfig {
    limit_side_len: 736,
    score_threshold: 0.3,
    box_threshold: 0.6,
    unclip_ratio: 2.0,
}

Troubleshooting

"OCR model not found" error

Ensure all three model files exist in models/:

  • pp-ocrv5_mobile_det.onnx
  • pp-ocrv5_mobile_rec.onnx
  • ppocrv5_dict.txt

"PDF is too large" error

The PDF exceeds the 25 MB limit. Consider:

  • Splitting the PDF into smaller files
  • Increasing MAX_PDF_BYTE_LENGTH in src/parse_pdf_native.rs (requires rebuild)

Poor OCR quality

Try:

  1. Switching to server models (higher accuracy)
  2. Adjusting OCR_CONFIDENCE_THRESHOLD in src/pdf_utils.rs
  3. Tuning detection config parameters

Build fails on macOS

Ensure you have:

  • Xcode Command Line Tools: xcode-select --install
  • Latest Rust: rustup update

License

MIT

Credits

About

High-performance PDF parsing for Node.js, powered by Rust: pdfium extraction via napi-rs bindings, with ONNX OCR fallback for scanned documents.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages