Falaah Falaah AI

Document Processing & AI Extraction

Understand how Muin's AI processes documents through a four-stage pipeline — detecting types, extracting fields, and making all content instantly searchable.

When you upload a document to Muin, powerful AI processes it automatically. This guide explains what happens behind the scenes and how to get the best extraction results.

How Muin Processes Documents

Every document goes through a four-stage pipeline:

Upload → Detection → Extraction → Indexing
   │          │           │           │
Received  Type ID'd   Data pulled  Searchable

Typical processing time: 5-30 seconds per document, depending on complexity.


Stage 1: Document Reception

When a document arrives (via upload, cloud sync, or email):

  1. File validation - Format and size checked
  2. Virus scan - Security verification
  3. Queue assignment - Prioritized for processing
  4. Tenant isolation - Associated with your organization

Stage 2: Document Type Detection

Muin examines the document to determine what kind of business document it is.

How Detection Works

Our AI analyzes multiple signals:

  • Visual layout - Tables, headers, formatting patterns
  • Key phrases - “Invoice”, “Due Date”, “Total Due”
  • Document structure - Where amounts and dates appear
  • Logo recognition - Known vendor templates

Supported Document Types

Type What Muin Looks For
Invoice Vendor info, line items, totals, due date
Contract Parties, signature blocks, terms, dates
Certificate Issuing body, certification type, expiry
Purchase Order PO number, items, quantities, delivery
Receipt Merchant, items, amounts, payment method
Expense Report Categories, amounts, approvals
Policy Document Policy number, effective dates, coverage
W-9/Tax Form Tax ID, name, address
Bank Statement Account info, transactions, balances
Form/Application Fields, checkboxes, signatures

Automatic vs. Manual Assignment

Automatic (>90% confidence): Muin assigns the type without asking. Most common documents are detected automatically.

Review required (below 90% confidence): A badge indicates “Type: Unconfirmed” - click to select the correct type.

Overriding detection:

  1. Open the document
  2. Click the document type badge
  3. Select the correct type
  4. Extraction re-runs with new type

Stage 3: AI Field Extraction

Based on document type, Muin extracts relevant business data.

Extraction by Document Type

Invoices:

  • Vendor name and address
  • Invoice number and date
  • Due date and payment terms
  • Line items (description, quantity, price)
  • Subtotal, tax, and total
  • PO reference number
  • Bank/payment details

Contracts:

  • Party names and roles
  • Effective date and term
  • Expiration/renewal date
  • Key obligations
  • Liability and indemnification
  • Termination clauses
  • Signature status

Certificates:

  • Issuing authority
  • Certificate holder
  • Certification type
  • Issue date
  • Expiration date
  • Coverage amounts (for insurance)
  • Certificate number

Purchase Orders:

  • PO number
  • Vendor information
  • Line items and quantities
  • Unit prices and totals
  • Delivery date and address
  • Payment terms
  • Approver information

Confidence Scores

Each extracted field has a confidence score:

Score Indicator Meaning
95-100% Green checkmark High confidence, likely correct
80-94% Yellow indicator Good confidence, verify if important
Below 80% Red flag Low confidence, review recommended

Fields below 80% confidence are highlighted for your review.

Reviewing Extractions

To review and correct extracted data:

  1. Open the document
  2. The extraction panel shows on the right
  3. Fields are displayed with confidence indicators
  4. Click any field to edit
  5. Corrections are saved immediately

Your corrections improve future extractions. Muin learns from your feedback to better handle similar documents.


Stage 4: Content Indexing

After extraction, Muin indexes the document for search.

Text Extraction

  • Native PDFs: Text extracted directly
  • Scanned documents: OCR converts images to text
  • Handwriting: Basic OCR, lower accuracy
  • Tables: Structure preserved for querying

Semantic Indexing

Beyond keyword matching, Muin understands meaning:

  • “payment terms” matches “Net 30” and “due upon receipt”
  • “vendor info” finds company names even without “vendor”
  • Questions like “when is this due?” find due dates

Search Availability

After indexing completes:

  • Document appears in search results
  • Muin Chat can answer questions about it
  • Filters include the document
  • Workflows can reference extracted data

Processing Status Indicators

Monitor document processing from the Document Hub:

Status Meaning
Uploading File transfer in progress
Queued Waiting to process
Processing AI extraction running
Ready Complete, fully searchable
Review Needs type confirmation or field review
Failed Processing error, see troubleshooting

Click any status to see details.


What to Do If Processing Fails

Common Failure Reasons

Issue Solution
Password-protected PDF Remove password and re-upload
Corrupted file Try exporting to new PDF
Image too small Use minimum 300 DPI scan
Unsupported language Contact support for language requests
Complex formatting Try simplified version

Reprocessing Documents

To retry processing:

  1. Open the failed document
  2. Click Reprocess in the document toolbar
  3. Optionally, set the document type manually first
  4. Processing restarts

Manual Data Entry

For documents that won’t process:

  1. Set the document type manually
  2. Click Edit Fields
  3. Enter data manually
  4. Click Save

The document is then searchable with your entered data.


Custom Document Types

Create document types specific to your business.

Creating a Custom Type

  1. Navigate to Settings → Document Types
  2. Click Create Custom Type
  3. Enter type name and description
  4. Define extraction fields:
    • Field name
    • Field type (text, number, date, currency)
    • Required or optional
    • Where to look (header, body, footer)
  5. Save the custom type

Custom Type Examples

Insurance Certificate:

  • Insured name
  • Insurance company
  • Policy number
  • Effective date
  • Expiration date
  • Coverage type (GL, Auto, Workers Comp)
  • Coverage amount

Project Proposal:

  • Client name
  • Project name
  • Proposed amount
  • Start date
  • Duration
  • Key deliverables

Improving Extraction Accuracy

Document Quality

Best practices:

  • Use 300+ DPI for scans
  • Ensure good contrast (dark text, light background)
  • Avoid shadows and glare
  • Keep documents flat when scanning

Prefer digital PDFs: Native PDFs (not scanned) extract more accurately than images.

Consistent Formatting

Muin learns from patterns:

  • Consistent vendor invoice templates improve accuracy
  • Standard document layouts extract better
  • Training improves over time

Feedback Loop

Every correction you make helps:

  1. Muin tracks which extractions you correct
  2. Patterns are identified across corrections
  3. Future similar documents extract more accurately
  4. Accuracy improves continuously

Processing Limits

Standard Limits

Metric Limit
Max file size 50 MB
Max pages 100 pages
Max batch 20 files, 10 MB each
Concurrent processing 10 documents

Enterprise Limits

Contact sales for higher limits including:

  • Larger file sizes
  • More concurrent processing
  • Priority queue access
  • Custom OCR models

Next Steps

Now that you understand document processing: