A high-performance PDF parser built with Rust and Node.js, featuring intelligent text extraction with automatic OCR fallback powered by PaddleOCR v5.
- Multiple parsing strategies: Native extraction, OCR-only, or hybrid mode
- Intelligent fallback: Automatically switches to OCR when native extraction yields insufficient text
- Multi-format output: Extracts text, images, and structured metadata
- High performance: Built with Rust and pdfium for fast PDF rendering
- OCR support: PaddleOCR v5 models for accurate text recognition
- Cross-platform: Supports macOS and Linux (x64/ARM64)
- Node.js >= 12.22.0 (see
package.jsonfor exact version support) - Rust (latest stable version)
- OCR Models (see Models section)
# Install dependencies
pnpm install
# Build the native addon
pnpm buildThis parser requires PaddleOCR v5 models for OCR functionality. Download the following files from HuggingFace PP-OCRv5 Collection:
Place these files in the models/ directory:
- pp-ocrv5_mobile_det.onnx - Text detection model
- pp-ocrv5_mobile_rec.onnx - Text recognition model
- ppocrv5_dict.txt - Character dictionary
Mobile Models (Recommended for most use cases):
- Faster inference, smaller size (~4.6MB det + ~15.8MB rec)
- Suitable for real-time processing
- Good accuracy for standard documents
Server Models (For higher accuracy):
- Larger size (~84MB det + ~80MB rec)
- Better for complex layouts, handwritten text
- Higher computational requirements
Download links:
import { parse_pdf, parse_pdf_native, parse_pdf_ocr } from './dist/index'
// Hybrid mode (recommended) - tries native extraction first, falls back to OCR
const result = await parse_pdf('/path/to/document.pdf', './output')
console.log(result.metadata)
// {
// pdf_name: "document",
// text_length: 5234,
// has_real_words: true,
// image_count: 5,
// pdf_size_bytes: 1048576,
// used_ocr: false
// }Hybrid mode - Intelligently combines native extraction with OCR fallback.
- First attempts native text extraction
- Falls back to OCR if:
- Extracted text < 500 characters, OR
- No real words detected (alphabetic tokens ≥ 2 characters)
- Best for unknown PDF types
Example:
const result = await parse_pdf('./invoice.pdf', './parsed')
console.log(result.text_path) // ./parsed/invoice/text.txt
console.log(result.images_dir) // ./parsed/invoice/images/
console.log(result.metadata.used_ocr) // true or falseNative extraction only - Extracts text and images using pdfium.
- Fast text extraction from text-based PDFs
- Renders pages to JPEG images
- No OCR processing
- Best for digitally-created PDFs
Example:
const result = await parse_pdf_native('./report.pdf', './output')
// Uses only native extraction, never falls back to OCROCR-only mode - Renders all pages and performs OCR.
- Renders each page to JPEG
- Runs PaddleOCR inference
- Outputs text and structured CSV
- Best for scanned documents or image-based PDFs
Example:
await parse_pdf_ocr('./scanned-doc.pdf', './ocr-output')
// All pages processed through OCR pipelineinterface PdfParseResult {
pdf_path: string // Original PDF path
output_path: string // Output directory path
text_path: string // Path to extracted text file
images_dir: string // Path to images directory
metadata: PdfMetadata // Document metadata
}interface PdfMetadata {
pdf_name: string // PDF filename without extension
text_length: number // Total characters extracted
has_real_words: boolean // Whether real words were detected
image_count: number // Number of page images rendered
pdf_size_bytes: number // Original PDF file size
used_ocr: boolean // Whether OCR was used
}After parsing, the output directory contains:
output_dir/
└── document_name/
├── text.txt # Extracted text content
├── metadata.json # Document metadata
├── data.csv # OCR results (when OCR used)
└── images/
├── 0.jpg # Page 0 rendered as JPEG
├── 1.jpg # Page 1
└── ...
- text.txt: Plain text extraction (newline-separated pages)
- metadata.json: JSON file with
PdfMetadatastructure - data.csv: OCR results with confidence scores (CSV format)
- images/: Rendered PDF pages as JPEGs (2000px max dimension)
| Function | Use Case | Speed | Accuracy |
|---|---|---|---|
parse_pdf_native |
Digital PDFs with embedded text | ⚡⚡⚡ Fastest | ✓ Good for text PDFs |
parse_pdf_ocr |
Scanned documents, images | ⚡ Slow (OCR) | ✓✓ Best for scans |
parse_pdf |
Unknown/mixed PDF types | ⚡⚡ Adaptive | ✓✓ Optimal |
The parse_pdf function uses this heuristic:
- Extract text natively using pdfium
- Check if text contains real words (alphabetic tokens ≥ 2 chars)
- If text < 500 chars OR no real words detected:
- Render all pages to images
- Run OCR on each page
- Replace text with OCR results
- Mark
used_ocrflag in metadata
- Maximum PDF size: 25 MB (26,214,400 bytes)
- Maximum image size: 10 MB per rendered page
- OCR confidence threshold: 0.7 (text below this is filtered)
- Image resolution: 2000px max width/height
- Native extraction: ~50-200ms per page (text-only)
- Image rendering: ~100-300ms per page
- OCR processing: ~500-2000ms per page (depends on complexity)
- Memory usage: Proportional to PDF size and page count
- Use
parse_pdf_nativefor known text-based PDFs - Use mobile OCR models for faster inference
- Process large PDFs in batches
- Consider server models only for complex documents
| Platform | Architecture | Support |
|---|---|---|
| macOS | x86_64 (Intel) | ✅ Supported |
| macOS | aarch64 (Apple Silicon) | ✅ Supported |
| Linux | x86_64 | ✅ Supported |
| Linux | aarch64 (ARM64) | ✅ Supported |
| Windows | x86_64 | ❌ Not in build targets |
The native addon is prebuilt for these platforms via GitHub Actions CI.
# Development build
pnpm build:debug
# Production build (optimized)
pnpm build
# Format code
pnpm format
# Lint
pnpm lint# Run tests
pnpm test# Run benchmarks
pnpm bench.
├── src/
│ ├── lib.rs # NAPI bindings and main functions
│ ├── parse_pdf_native.rs # Native extraction logic
│ └── pdf_utils.rs # OCR and utility functions
├── models/ # OCR model files (not in repo)
├── dist/ # Built native addon
└── __test__/ # Test files
- napi-rs: Node.js N-API bindings for Rust
- pdfium-render: PDF rendering using Google's pdfium
- oar-ocr: ONNX runtime for OCR inference
- image: Image processing (JPEG encoding/decoding)
- turbojpeg: Fast JPEG compression
Default OCR settings (can be modified in src/pdf_utils.rs):
TextDetectionConfig {
limit_side_len: 736,
score_threshold: 0.3,
box_threshold: 0.6,
unclip_ratio: 2.0,
}Ensure all three model files exist in models/:
pp-ocrv5_mobile_det.onnxpp-ocrv5_mobile_rec.onnxppocrv5_dict.txt
The PDF exceeds the 25 MB limit. Consider:
- Splitting the PDF into smaller files
- Increasing
MAX_PDF_BYTE_LENGTHinsrc/parse_pdf_native.rs(requires rebuild)
Try:
- Switching to server models (higher accuracy)
- Adjusting
OCR_CONFIDENCE_THRESHOLDinsrc/pdf_utils.rs - Tuning detection config parameters
Ensure you have:
- Xcode Command Line Tools:
xcode-select --install - Latest Rust:
rustup update
MIT
- OCR powered by PaddleOCR v5
- PDF rendering via pdfium
- Built with napi-rs