Muin is in private beta.Watch the public release announcement —talk to us.
Falaah Falaah AI

Document Processing & AI Extraction

Understand how Muin's AI processes documents through a four-stage pipeline — detecting types, extracting fields, and making all content instantly searchable.

When you upload a document to Muin, powerful AI processes it automatically. This guide explains what happens behind the scenes and how to get the best extraction results.

How Muin Processes Documents

Every document goes through a four-stage pipeline:

Upload → Detection → Extraction → Indexing
   │          │           │           │
Received  Type ID'd   Data pulled  Searchable

Typical processing time: 5-30 seconds per document, depending on complexity.


Stage 1: Document Reception

When a document arrives (via upload, cloud sync, or email):

  1. File validation - Format and size checked
  2. Virus scan - Security verification
  3. Queue assignment - Prioritized for processing
  4. Tenant isolation - Associated with your organization

Stage 2: Document Type Detection

Muin examines the document to determine what kind of business document it is.

How Detection Works

Our AI analyzes multiple signals:

  • Visual layout - Tables, headers, formatting patterns
  • Key phrases - “Invoice”, “Due Date”, “Total Due”
  • Document structure - Where amounts and dates appear
  • Logo recognition - Known vendor templates

Supported Document Types

TypeWhat Muin Looks For
InvoiceVendor info, line items, totals, due date
ContractParties, signature blocks, terms, dates
CertificateIssuing body, certification type, expiry
Purchase OrderPO number, items, quantities, delivery
ReceiptMerchant, items, amounts, payment method
Expense ReportCategories, amounts, approvals
Policy DocumentPolicy number, effective dates, coverage
W-9/Tax FormTax ID, name, address
Bank StatementAccount info, transactions, balances
Form/ApplicationFields, checkboxes, signatures

Automatic vs. Manual Assignment

Automatic (>90% confidence): Muin assigns the type without asking. Most common documents are detected automatically.

Review required (below 90% confidence): A badge indicates “Type: Unconfirmed” - click to select the correct type.

Overriding detection:

  1. Open the document
  2. Click the document type badge
  3. Select the correct type
  4. Extraction re-runs with new type

Stage 3: AI Field Extraction

Based on document type, Muin extracts relevant business data.

Extraction by Document Type

Invoices:

  • Vendor name and address
  • Invoice number and date
  • Due date and payment terms
  • Line items (description, quantity, price)
  • Subtotal, tax, and total
  • PO reference number
  • Bank/payment details

Contracts:

  • Party names and roles
  • Effective date and term
  • Expiration/renewal date
  • Key obligations
  • Liability and indemnification
  • Termination clauses
  • Signature status

Certificates:

  • Issuing authority
  • Certificate holder
  • Certification type
  • Issue date
  • Expiration date
  • Coverage amounts (for insurance)
  • Certificate number

Purchase Orders:

  • PO number
  • Vendor information
  • Line items and quantities
  • Unit prices and totals
  • Delivery date and address
  • Payment terms
  • Approver information

Confidence Scores

Each extracted field has a confidence score:

ScoreIndicatorMeaning
95-100%Green checkmarkHigh confidence, likely correct
80-94%Yellow indicatorGood confidence, verify if important
Below 80%Red flagLow confidence, review recommended

Fields below 80% confidence are highlighted for your review.

Reviewing Extractions

To review and correct extracted data:

  1. Open the document
  2. The extraction panel shows on the right
  3. Fields are displayed with confidence indicators
  4. Click any field to edit
  5. Corrections are saved immediately

Your corrections improve future extractions. Muin learns from your feedback to better handle similar documents.


Stage 4: Content Indexing

After extraction, Muin indexes the document for search.

Text Extraction

  • Native PDFs: Text extracted directly
  • Scanned documents: OCR converts images to text
  • Handwriting: Basic OCR, lower accuracy
  • Tables: Structure preserved for querying

Semantic Indexing

Beyond keyword matching, Muin understands meaning:

  • “payment terms” matches “Net 30” and “due upon receipt”
  • “vendor info” finds company names even without “vendor”
  • Questions like “when is this due?” find due dates

Search Availability

After indexing completes:

  • Document appears in search results
  • Muin Chat can answer questions about it
  • Filters include the document
  • Workflows can reference extracted data

Processing Status Indicators

Monitor document processing from the Document Hub:

StatusMeaning
UploadingFile transfer in progress
QueuedWaiting to process
ProcessingAI extraction running
ReadyComplete, fully searchable
ReviewNeeds type confirmation or field review
FailedProcessing error, see troubleshooting

Click any status to see details.


What to Do If Processing Fails

Common Failure Reasons

IssueSolution
Password-protected PDFRemove password and re-upload
Corrupted fileTry exporting to new PDF
Image too smallUse minimum 300 DPI scan
Unsupported languageContact support for language requests
Complex formattingTry simplified version

Reprocessing Documents

To retry processing:

  1. Open the failed document
  2. Click Reprocess in the document toolbar
  3. Optionally, set the document type manually first
  4. Processing restarts

Manual Data Entry

For documents that won’t process:

  1. Set the document type manually
  2. Click Edit Fields
  3. Enter data manually
  4. Click Save

The document is then searchable with your entered data.


Custom Document Types

Create document types specific to your business.

Creating a Custom Type

  1. Navigate to SettingsDocument Types
  2. Click Create Custom Type
  3. Enter type name and description
  4. Define extraction fields:
    • Field name
    • Field type (text, number, date, currency)
    • Required or optional
    • Where to look (header, body, footer)
  5. Save the custom type

Custom Type Examples

Insurance Certificate:

  • Insured name
  • Insurance company
  • Policy number
  • Effective date
  • Expiration date
  • Coverage type (GL, Auto, Workers Comp)
  • Coverage amount

Project Proposal:

  • Client name
  • Project name
  • Proposed amount
  • Start date
  • Duration
  • Key deliverables

Improving Extraction Accuracy

Document Quality

Best practices:

  • Use 300+ DPI for scans
  • Ensure good contrast (dark text, light background)
  • Avoid shadows and glare
  • Keep documents flat when scanning

Prefer digital PDFs: Native PDFs (not scanned) extract more accurately than images.

Consistent Formatting

Muin learns from patterns:

  • Consistent vendor invoice templates improve accuracy
  • Standard document layouts extract better
  • Training improves over time

Feedback Loop

Every correction you make helps:

  1. Muin tracks which extractions you correct
  2. Patterns are identified across corrections
  3. Future similar documents extract more accurately
  4. Accuracy improves continuously

Processing Limits

Standard Limits

MetricLimit
Max file size50 MB
Max pages100 pages
Max batch20 files, 10 MB each
Concurrent processing10 documents

Enterprise Limits

Contact sales for higher limits including:

  • Larger file sizes
  • More concurrent processing
  • Priority queue access
  • Custom OCR models

Next Steps

Now that you understand document processing: