THE PROBLEM
MOST DOCUMENT DATA NEVER REACHES YOUR LMS IN A USABLE FORM.
Lending workflows depend on data extracted from third-party documents — valuation reports, income tax returns, bank statements, salary certificates, audit reports. Most institutions extract this data manually: a credit analyst reads the document and types values into fields. This process is slow, error-prone, and impossible to audit at scale.
Template-based OCR is an incomplete fix. Valuation reports from different agencies use different layouts. A rule written for one format breaks on the next. The result is fragile automation that still requires manual intervention — the worst of both worlds.
⏱
TURNAROUND DELAYS
Manual extraction adds hours to loan processing queues, especially in secured lending where document volumes are high.
⚠️
DATA ENTRY ERRORS
Transposition errors in valuation figures or property details can propagate into underwriting decisions.
🔄
FORMAT DEPENDENCY
Rule-based extraction breaks when a vendor changes their report layout — requiring engineering effort to fix.
🔍
NO AUDIT TRAIL
Manual entry leaves no record of what was extracted from which version of which document.
HOW IT WORKS
CONFIGURE ONCE. EXTRACT FROM ANY DOCUMENT.
The extraction config is reusable — set it up once for a document type and run it on every new document of that type.
1
CREATE AN EXTRACTION CONFIG
Name the config, select a model, and some other parameters. This becomes your reusable extraction template.
LLM CONFIG
2
DEFINE YOUR PROMPTS
Add field keys and natural-language extraction prompts — e.g. fair_market_value → "What is the fair market value of the property?"
NO-CODE · CONFIGURABLE
3
UPLOAD THE DOCUMENT
Select or upload the document. The system parses the text and passes it to the LLM with your configured prompts.
DOCUMENT UPLOAD
4
RECEIVE STRUCTURED OUTPUT
Get a structured key-value result for each field. Fields not found in the document return "Not Available" — no fabrication.
JSON OUTPUT
KEY CAPABILITIES
WHAT THE SYSTEM DOES — AND HOW.
📄
FORMAT-INDEPENDENT EXTRACTION
Works on documents from any agency or vendor, regardless of layout. No template matching — the LLM reads and understands context.
⚙️
CONFIGURABLE FIELD PROMPTS
Define each field as a natural-language question. Add or modify fields without engineering changes — via the UI.
🛡️
HONEST "NOT AVAILABLE" HANDLING
When a field is absent, the system returns "Not Available" rather than inferring or fabricating a value. Extraction results are reliable.
📋
STRUCTURED JSON OUTPUT
Results are returned as structured key-value data, ready for downstream use in LOS/LMS workflow without further parsing.
🗂️
FULL AUDIT TRAILS
Every extraction is logged — document, config, run, output. Available directly from the extraction config view for compliance review.
🔁
REUSABLE EXTRACTION CONFIGS
A config defined for a document type runs on every new document of that type without reconfiguration.
USE CASES
CONFIGURED FOR ANY DOCUMENT IN THE LENDING WORKFLOW.
Doc Extractor works across all major document categories in lending. A single extraction config covers all agencies and formats within a category.
🪪
KYC & IDENTITY
15+ document types
Aadhaar
PAN
Driving Licence
Passport
Voter ID
Form 60
CKYC Document
🏦
FINANCIAL & LOAN
35+ document types
Bank Statement
Loan Agreement
Sanction Letter
CIBIL / Credit Report
ITR (All Types)
Salary Slip
Form 16
🪪
COMPANY & BUSINESS
20+ document types
CA Certified Financials
GST Certificate
Partnership Deed
GST Returns (3B)
MOA / AOA
Audit Report
MSME / Udyam
🏠
PROPERTY & VALUATION
30+ document types
Valuation Report
Registered Sale Deed
Encumbrance Certificate
Mutation Certificate
Utility Bills
Jamabandi / Khasra
Property Tax Receipt
⚖️
LEGAL & OPERATIONS
20+ document types
Legal Opinion Report
Power of Attorney
FI Report
PDC / NACH
NOC
Affidavit
PD Report
100+ DOCUMENT TYPES
New types added from the dashboard — no development needed.
Works on PDFs, scanned PDFs, and images.
PLATFORM INTEGRATION
PART OF THE LENDING STACK, NOT SEPARATE FROM IT.
Doc Extractor is embedded within the OneFin platform. Extracted data flows directly into loan application records — no copy-paste, no API wiring between systems.
🔗
CONNECTED TO YOUR LOS AND LMS DATA LAYER
Extraction outputs link to the loan application in context. Extracted fields can populate collateral records, update underwriting inputs, or trigger downstream workflow steps — depending on your configuration.
LOS
LMS
DOCUMENT VAULT
RULE ENGINE
RECON
WHO IT'S FOR
RELEVANT ACROSS THE LOAN PROCESSING CHAIN.
OPERATIONS / CREDIT
CREDIT AND OPERATIONS TEAMS
Eliminates manual data entry from document review. Extraction results are available immediately after document upload — no waiting on back-office processing.
CTO / TECHNOLOGY
TECHNOLOGY LEADERS
No external AI vendor dependency. No custom ML pipeline to build or maintain. Extraction configs are managed in the UI. New document types are configurable without engineering.
CRO / RISK
RISK AND COMPLIANCE TEAMS
Extracted data comes with a full audit trail. Fields absent from the document return "Not Available" — no silent fabrication. Every extraction is logged for regulatory review.

