Document Processing & AI Extraction
Understand how Muin's AI processes documents through a four-stage pipeline — detecting types, extracting fields, and making all content instantly searchable.
When you upload a document to Muin, powerful AI processes it automatically. This guide explains what happens behind the scenes and how to get the best extraction results.
How Muin Processes Documents
Every document goes through a four-stage pipeline:
Upload → Detection → Extraction → Indexing
│ │ │ │
Received Type ID'd Data pulled Searchable
Typical processing time: 5-30 seconds per document, depending on complexity.
Stage 1: Document Reception
When a document arrives (via upload, cloud sync, or email):
- File validation - Format and size checked
- Virus scan - Security verification
- Queue assignment - Prioritized for processing
- Tenant isolation - Associated with your organization
Stage 2: Document Type Detection
Muin examines the document to determine what kind of business document it is.
How Detection Works
Our AI analyzes multiple signals:
- Visual layout - Tables, headers, formatting patterns
- Key phrases - “Invoice”, “Due Date”, “Total Due”
- Document structure - Where amounts and dates appear
- Logo recognition - Known vendor templates
Supported Document Types
| Type | What Muin Looks For |
|---|---|
| Invoice | Vendor info, line items, totals, due date |
| Contract | Parties, signature blocks, terms, dates |
| Certificate | Issuing body, certification type, expiry |
| Purchase Order | PO number, items, quantities, delivery |
| Receipt | Merchant, items, amounts, payment method |
| Expense Report | Categories, amounts, approvals |
| Policy Document | Policy number, effective dates, coverage |
| W-9/Tax Form | Tax ID, name, address |
| Bank Statement | Account info, transactions, balances |
| Form/Application | Fields, checkboxes, signatures |
Automatic vs. Manual Assignment
Automatic (>90% confidence): Muin assigns the type without asking. Most common documents are detected automatically.
Review required (below 90% confidence): A badge indicates “Type: Unconfirmed” - click to select the correct type.
Overriding detection:
- Open the document
- Click the document type badge
- Select the correct type
- Extraction re-runs with new type
Stage 3: AI Field Extraction
Based on document type, Muin extracts relevant business data.
Extraction by Document Type
Invoices:
- Vendor name and address
- Invoice number and date
- Due date and payment terms
- Line items (description, quantity, price)
- Subtotal, tax, and total
- PO reference number
- Bank/payment details
Contracts:
- Party names and roles
- Effective date and term
- Expiration/renewal date
- Key obligations
- Liability and indemnification
- Termination clauses
- Signature status
Certificates:
- Issuing authority
- Certificate holder
- Certification type
- Issue date
- Expiration date
- Coverage amounts (for insurance)
- Certificate number
Purchase Orders:
- PO number
- Vendor information
- Line items and quantities
- Unit prices and totals
- Delivery date and address
- Payment terms
- Approver information
Confidence Scores
Each extracted field has a confidence score:
| Score | Indicator | Meaning |
|---|---|---|
| 95-100% | Green checkmark | High confidence, likely correct |
| 80-94% | Yellow indicator | Good confidence, verify if important |
| Below 80% | Red flag | Low confidence, review recommended |
Fields below 80% confidence are highlighted for your review.
Reviewing Extractions
To review and correct extracted data:
- Open the document
- The extraction panel shows on the right
- Fields are displayed with confidence indicators
- Click any field to edit
- Corrections are saved immediately
Your corrections improve future extractions. Muin learns from your feedback to better handle similar documents.
Stage 4: Content Indexing
After extraction, Muin indexes the document for search.
Text Extraction
- Native PDFs: Text extracted directly
- Scanned documents: OCR converts images to text
- Handwriting: Basic OCR, lower accuracy
- Tables: Structure preserved for querying
Semantic Indexing
Beyond keyword matching, Muin understands meaning:
- “payment terms” matches “Net 30” and “due upon receipt”
- “vendor info” finds company names even without “vendor”
- Questions like “when is this due?” find due dates
Search Availability
After indexing completes:
- Document appears in search results
- Muin Chat can answer questions about it
- Filters include the document
- Workflows can reference extracted data
Processing Status Indicators
Monitor document processing from the Document Hub:
| Status | Meaning |
|---|---|
| Uploading | File transfer in progress |
| Queued | Waiting to process |
| Processing | AI extraction running |
| Ready | Complete, fully searchable |
| Review | Needs type confirmation or field review |
| Failed | Processing error, see troubleshooting |
Click any status to see details.
What to Do If Processing Fails
Common Failure Reasons
| Issue | Solution |
|---|---|
| Password-protected PDF | Remove password and re-upload |
| Corrupted file | Try exporting to new PDF |
| Image too small | Use minimum 300 DPI scan |
| Unsupported language | Contact support for language requests |
| Complex formatting | Try simplified version |
Reprocessing Documents
To retry processing:
- Open the failed document
- Click Reprocess in the document toolbar
- Optionally, set the document type manually first
- Processing restarts
Manual Data Entry
For documents that won’t process:
- Set the document type manually
- Click Edit Fields
- Enter data manually
- Click Save
The document is then searchable with your entered data.
Custom Document Types
Create document types specific to your business.
Creating a Custom Type
- Navigate to Settings → Document Types
- Click Create Custom Type
- Enter type name and description
- Define extraction fields:
- Field name
- Field type (text, number, date, currency)
- Required or optional
- Where to look (header, body, footer)
- Save the custom type
Custom Type Examples
Insurance Certificate:
- Insured name
- Insurance company
- Policy number
- Effective date
- Expiration date
- Coverage type (GL, Auto, Workers Comp)
- Coverage amount
Project Proposal:
- Client name
- Project name
- Proposed amount
- Start date
- Duration
- Key deliverables
Improving Extraction Accuracy
Document Quality
Best practices:
- Use 300+ DPI for scans
- Ensure good contrast (dark text, light background)
- Avoid shadows and glare
- Keep documents flat when scanning
Prefer digital PDFs: Native PDFs (not scanned) extract more accurately than images.
Consistent Formatting
Muin learns from patterns:
- Consistent vendor invoice templates improve accuracy
- Standard document layouts extract better
- Training improves over time
Feedback Loop
Every correction you make helps:
- Muin tracks which extractions you correct
- Patterns are identified across corrections
- Future similar documents extract more accurately
- Accuracy improves continuously
Processing Limits
Standard Limits
| Metric | Limit |
|---|---|
| Max file size | 50 MB |
| Max pages | 100 pages |
| Max batch | 20 files, 10 MB each |
| Concurrent processing | 10 documents |
Enterprise Limits
Contact sales for higher limits including:
- Larger file sizes
- More concurrent processing
- Priority queue access
- Custom OCR models
Next Steps
Now that you understand document processing:
- Searching Documents - Find processed documents
- Organizing Documents - Folders, tags, favorites
- Building Workflows - Automate based on extractions