The Apryse Summer 2026 Release: OUT NOW

On-premises ICR that turns handwritten forms into structured, searchable data

OCR reads print. Handwriting ICR reads handwriting. Apryse ICR uses neural networks that analyze individual writing styles rather than matching against fixed templates, turning documents with written text into searchable PDFs and structured JSON with word-level position data. It runs as an add-on to the Apryse Server SDK, inside your own infrastructure, with no external APIs.

Deployment & integration​​​​‌‍​‍​‍‌‍‌​‍‌‍‍‌‌‍‌‌‍‍‌‌‍‍​‍​‍​‍‍​‍​‍‌​‌‍​‌‌‍‍‌‍‍‌‌‌​‌‍‌​‍‍‌‍‍‌‌‍​‍​‍​‍​​‍​‍‌‍‍​‌​‍‌‍‌‌‌‍‌‍​‍​‍​‍‍​‍​‍​‍‌​‌‌​‌‌‌‌‍‌​‌‍‍‌‌‍​‍‌‍‍‌‌‍‍‌‌​‌‍‌‌‌‍‍‌‌​​‍‌‍‌‌‌‍‌​‌‍‍‌‌‌​​‍‌‍‌‌‍‌‍‌​‌‍‌‌​‌‌​​‌​‍‌‍‌‌‌​‌‍‌‌‌‍‍‌‌​‌‍​‌‌‌​‌‍‍‌‌‍‌‍‍​‍‌‍‍‌‌‍‌​​‌‌‍​‍‌‍​‌‍‌‍​​‍‌‍‌‍​​‍​​‌‌‍​​‍‌‌‍‌‌‌‍​​​‌‍‌‌​‍‌​‌​​‌​‌‍‌​​‌​​‍‌‌‍​‍‌‍​‌‌‍‌‌‌‍​‌​‍‌​‌​​‌‌‍​‌​‍​‌‍‌​​​​‌‍‌​​‍​​‌‌​‍​​​​​‍‌​‍‌‌​‌‍‌‌​​‌‍‌‌​‌‌​​‌‍​‌‌‍‌‌‍‌‌​‍‌​​‌‍​‌‌‌​‌‍‍​​‌‌​​‌‍​‌‌‍‌‌‍‌‌‌​​‍‌‌‌‌‍‍‌‌‍​‌‍‌​‌‍‌‌‌​‍​‍‌‌​‌‌‌​​‍‌‌‌‍‍‌‍‌‌‌‍‌​‍‌‌​​‌​‌​​‍‌‌​​‌​‌​​‍‌‌​​‍​​‍​​​​​​‍‌‍​‍‌‍​‍​​‌‍‌​​​‍‌‍​​​‍‌‍‌​​​​​‍‌‌​​‍​​‍​‍‌‌​‌‌‌​‌​​‍‍‌‍​‌‍‌‍‍‌‌​‌‍‌‌‌‍‍‌‌​​‍‌‌​‌‌‌​​‍‌‌‌‍‍‌‍‌‌‌‍‌​‍‌‌​​‌​‌​​‍‌‌​​‌​‌​​‍‌‌​​‍​​‍‌‍‌​‌‍​​‍​‌‍​‍​‍​‌‍​‌​‌‍​‌​​​‌‍‌‌‌‍‌‍​‍​​‍‌‌​​‍​​‍​‍‌‌​‌‌‌​‌​​‍‍‌‍​‌‍‍​‌‍‍‌‌‍​‌‍‌​‌​‍‌‍‌‌‌‍‍​‍‌‌​‌‌‌​​‍‌‌‌‍‍‌‍‌‌‌‍‌​‍‌‌​​‌​‌​​‍‌‌​​‌​‌​​‍‌‌​​‍​​‍‌‍‌‌‌‍‌‍​‍​​‌‌‍‌​​‌‍​​​​‌​‌‍‌‍‌‍​​‍‌‌‍‌‍​‍‌‌​​‍​​‍​‍‌‌​‌‌‌​‌​​‍‍‌‌​‌‍‌‌‌‍​‌‌​​‌‍​‍‌‍​‌‌​‌‍‌‌‌‌‌‌‌​‍‌‍​​‌​‍‌‌​​‍‌​‌‍‌​‌‌​‌‌‌‌‍‌​‌‍‍‌‌‍​‍‌‍‌‍‍‌‌‍‌​​‌‌‍​‍‌‍​‌‍‌‍​​‍‌‍‌‍​​‍​​‌‌‍​​‍‌‌‍‌‌‌‍​​​‌‍‌‌​‍‌​‌​​‌​‌‍‌​​‌​​‍‌‌‍​‍‌‍​‌‌‍‌‌‌‍​‌​‍‌​‌​​‌‌‍​‌​‍​‌‍‌​​​​‌‍‌​​‍​​‌‌​‍​​​​​‍‌​‍‌‍‌‌​‌‍‌‌​​‌‍‌‌​‌‌​​‌‍​‌‌‍‌‌‍‌‌​‍‌‍‌​​‌‍​‌‌‌​‌‍‍​​‌‌​​‌‍​‌‌‍‌‌‍‌‌‌​​‍‌‌‌‌‍‍‌‌‍​‌‍‌​‌‍‌‌‌​‍​‍‌‌​‌‌‌​​‍‌‌‌‍‍‌‍‌‌‌‍‌​‍‌‌​​‌​‌​​‍‌‌​​‌​‌​​‍‌‌​​‍​​‍​​​​​​‍‌‍​‍‌‍​‍​​‌‍‌​​​‍‌‍​​​‍‌‍‌​​​​​‍‌‌​​‍​​‍​‍‌‌​‌‌‌​‌​​‍‍‌‍​‌‍‌‍‍‌‌​‌‍‌‌‌‍‍‌‌​​‍‌‌​‌‌‌​​‍‌‌‌‍‍‌‍‌‌‌‍‌​‍‌‌​​‌​‌​​‍‌‌​​‌​‌​​‍‌‌​​‍​​‍‌‍‌​‌‍​​‍​‌‍​‍​‍​‌‍​‌​‌‍​‌​​​‌‍‌‌‌‍‌‍​‍​​‍‌‌​​‍​​‍​‍‌‌​‌‌‌​‌​​‍‍‌‍​‌‍‍​‌‍‍‌‌‍​‌‍‌​‌​‍‌‍‌‌‌‍‍​‍‌‌​‌‌‌​​‍‌‌‌‍‍‌‍‌‌‌‍‌​‍‌‌​​‌​‌​​‍‌‌​​‌​‌​​‍‌‌​​‍​​‍‌‍‌‌‌‍‌‍​‍​​‌‌‍‌​​‌‍​​​​‌​‌‍‌‍‌‍​​‍‌‌‍‌‍​‍‌‌​​‍​​‍​‍‌‌​‌‌‌​‌​​‍‍‌‌​‌‍‌‌‌‍​‌‌​​‍​‍‌‌

Installs into the SDK you already run

Handwriting ICR runs natively in the languages and environments your team already uses, on Windows, Linux, and macOS. The module is driven through HandwritingICRModule and HandwritingICROptions, available in C++, C# (.NET and .NET Framework), Java, and Node.js, with full samples in Python, C#, C++, Go, Java, Node.js, PHP, Ruby, VB, and Objective-C.

Installation is a file-level step rather than a package install: expand the ICR archive directly into your existing SDK directory (for example into PDFNetC64/ for the 64-bit C/C++ package, overwriting files where required). Every file in the archive's Lib folder must be present. If the SDK cannot locate the module, register its folder with PDFNet.AddResourceSearchPath(). PDFNet.Initialize() must be called before any ICR operation.

Making a handwritten PDF searchable is a single call:

Trusted for document workflows where accuracy, control, and deployment flexibility matter.

autodesk logo
boeing logo
notability logo
docusign logo
egress logo
microsoft logo
thomson_reuters logo
encode logo
ibm logo
autodesk logo
boeing logo
notability logo
docusign logo
egress logo
microsoft logo
thomson_reuters logo
encode logo
ibm logo

From handwriting to structured data

Copied to clipboard

Intelligent Character Recognition extracts handwritten text from images and image-based PDFs, the content standard OCR cannot parse reliably. Rather than matching characters against fixed templates, the module uses neural networks and machine learning to analyze individual writing styles and extract meaning from highly unstructured input.

Capability
What it does
Handwriting recognition
Machine-learning-based interpretation of handwritten characters, adapting to individual writing styles rather than a fixed template set.
Unstructured input handling
Sources with no consistent layout: medical forms, insurance claims, historical and archival documents, logistics and shipping records.
Searchable PDF output
Adds an invisible, selectable text layer to an image-based PDF, making handwritten content searchable and copyable.
Structured JSON output
Returns text and metadata nested as pages, paragraphs, lines, and words, with coordinates, length, font size, and orientation for each word.
External ICR results
Apply JSON produced by a different OCR or ICR engine through the same API.
Zonal recognition
Restrict recognition to inclusion zones, or exclude regions such as a signature block, per page.
Post-processing hooks
Extract results, run spell checking or allowlist and blocklist comparison, then re-apply the corrected values to the document.
Page targeting
Process a subset of pages rather than the whole document.
Self-hosted
Runs entirely inside your infrastructure through the Server SDK. No external APIs, no document data exposure.

What comes out

Copied to clipboard

ICR produces two things: a PDF with an invisible, selectable text layer, or structured JSON you can use without touching the PDF at all.

Output
Format
Produced by
Searchable PDF
PDF with invisible selectable text layer
HandwritingICRModule.ProcessPDF
Structured text and coordinates
JSON
HandwritingICRModule.GetICRJsonFromPDF

Structured output

The JSON is nested: pages contain paragraphs, which contain lines, which contain words. Each page carries its number, DPI, and coordinate origin (top-left by default, or PDF-style bottom-left). Each word carries its bounding box x and y, length, font size, text, and orientation. Each line optionally carries a bounding box.

That structure is what makes handwriting recognition reviewable rather than a black box. Every recognized value can be traced to a location on the page, so a misread name or figure can be surfaced, highlighted, and corrected before it reaches a system of record.

Because the output is a PDF, the rest of the Server SDK picks it up from there. Converting a searchable PDF to PDF/A is part of the base package. Converting it to DOCX, XLSX, PPTX, or HTML uses the Structured Output Module, a separate add-on on the same SDK.

Extract, correct, then re-apply

Applying raw recognition output straight to the document is one call. With handwriting it rarely should be. Handwritten glyphs vary widely between writers, values escape the boxes a form designer intended, and pen colour and pressure change from page to page. That makes the correction step matter more than it does for print.

The API is built for that. Extract the results as JSON, run whatever validation your domain calls for: a spell checker, a lookup against expected patient or customer names, an allowlist or blocklist comparison and write the corrected values back.

The same entry point accepts JSON generated by a different OCR or ICR engine, so an existing recognition pipeline can be kept while Apryse handles validation, placement, and searchable-PDF generation.

When to consider ICR

When the content you need is handwritten and arriving as an image: intake forms, claims, delivery confirmations, archived records, and standard OCR is returning unusable results on those fields. ICR closes the gap that keeps otherwise-digital workflows dependent on manual data entry. ICR does not read printed or machine-generated text; use the OCR Module for that. Most real document sets contain both, and the two modules are commonly enabled on the same SDK. To pull structured fields, tables, and key-value pairs from the recognized content, with preprocessing such as deskewing, despeckling, and rotation cleanup applied first and then add Smart Data Extraction.

Inside a recognition pass

ICR is an add-on module to the Apryse Server SDK. You download the module, expand it into your SDK directory, and call it from the same API you already use for the rest of your document pipeline.

Hand it the page

Pass an image or an image-based PDF containing handwritten content. Recognition quality tracks input quality: resolution, contrast, scan noise, and handwriting clarity all affect the result.

Narrow the scope

Optionally restrict processing to a subset of pages, to inclusion zones, or away from regions such as signature blocks.

Recognize

Neural network models interpret the handwritten characters and their positions on the page. Processing runs inside your environment. No page content is sent to an external API.

Correct, then take the output

Extract the JSON, validate it against dictionaries or expected values, and re-apply. Write the results back as an invisible text layer for a searchable PDF, or carry the JSON forward for routing and analytics.

Typical workflow

Handwritten form or scanned record → page and zone selection → recognition → JSON extraction → validation and human review → corrected results re-applied → searchable PDF and structured data → claims system, EHR, records platform, or extraction pipeline.

Four ways to get handwriting into a system

Teams adding ICR to an application usually weigh four options: manual data entry, a cloud ICR API, a general-purpose AI model, or a standard OCR. Apryse ICR is the embedded SDK approach.

Over manual data entry

Copied to clipboard

Handwritten fields are still routinely keyed in by hand or sent to an outsourced team, which sets a floor on cost, turnaround, and error rate that no amount of workflow tooling can move. ICR converts transcription into an automated step with a validation pass, so people review exceptions instead of typing every form.

Over a cloud handwriting API

Copied to clipboard

A cloud service sends every page: intake forms, claims, clinical notes, to infrastructure you do not control. Apryse ICR runs inside your own application on your own servers. Handwritten content is disproportionately personal and regulated, so for most teams evaluating this, that is the deciding factor rather than a preference.

Over a general-purpose AI model

Copied to clipboard

A general-purpose model can transcribe handwriting, but it infers text and layout across the whole page in one pass, which makes output variable between runs and hard to audit on a document you may later have to defend. ICR is purpose-built for character recognition and returns positional data alongside the text, so every recognized value traces back to a location on the page.

Over standard OCR

Copied to clipboard

OCR is built for printed, machine-generated characters and degrades on handwriting, because its underlying assumption of consistent glyph shapes does not hold. ICR is trained for that variability instead of against it. The two are complementary: OCR for the printed body of a form, ICR for the handwritten fields.

Running ICR first is also the cheaper way to use a model. Handing an LLM a page image spends tokens on pixels; handing it the recognized text spends them on words, and a page of text costs a fraction of what the same page costs as an image. Recognition is a fixed local cost paid once on your own hardware, while image tokens are billed on every page of every run.

ICR for real-world documents

Watch a handwritten form move through recognition to a searchable PDF and structured JSON, without leaving the deployment environment.

What it runs on, and how it gets there

Copied to clipboard

Handwriting ICR is embedded in your application and runs on infrastructure controlled by your organization.

Attribute
Details
Integration model
Add-on module for the Apryse Server SDK, embedded in your application.
SDK availability
Server SDK only. Web and mobile applications call it server-side.
Supported operating systems
Windows, Linux, and macOS (x64). ARM and 32-bit are not supported.
Languages and bindings
.NET, Java, Python, Node.js, C++, Go, PHP, Ruby, Objective-C
Installation
Expand the module archive directly into the existing SDK directory. All files in the archive's Lib folder must be present. Register the folder with PDFNet.AddResourceSearchPath() if needed.
Input formats
Images and image-based PDFs containing handwritten content.
Output
Searchable PDF with a selectable text layer, and JSON with word-level text and position data.
Hardware
CPU only. No GPU required.
Sample application
HandwritingICRTest, included in the main SDK download.
Licensing
Add-on package license to the Apryse Server SDK. Trial keys have unlimited access to all add-on modules.

Documents are processed inside your environment and are not sent to Apryse or a third party for processing.

ICR Use Cases

ICR applies across nearly every industry that still runs partly on paper.

Sanity Image

Banking, financial services, and insurance (BFSI)

Loan applications, cheques, account forms, claims in document-driven processes where handwritten fields sit inside otherwise structured forms that standard OCR cannot parse.

Sanity Image

Legal and records management

Archival correspondence, annotated files, and historical documents that need to become searchable inside controlled infrastructure.

Sanity Image

Healthcare and life sciences

Handwritten patient intake forms, prescriptions, and clinical notes that create bottlenecks, slow data entry, and error rates in patient data workflows.

Sanity Image

Government and public sector

Census records, permit applications, and physical archives held in storage and not yet digitized.

Sanity Image

Logistics and trade

Handwritten shipment records and delivery confirmations slowing supply chains that are otherwise fully digital.

Frequently asked questions

Handwriting ICR is a self-hosted add-on module for the Apryse Server SDK that extracts handwritten text from images and image-based PDFs. It produces searchable PDFs with selectable text layers, and JSON containing text and word-level position data. It runs inside an application, on infrastructure controlled by the customer.

OCR recognizes printed, machine-generated characters. ICR recognizes handwriting, using neural networks that analyze individual writing styles rather than matching against fixed character templates. Documents containing both are typically processed with both modules on the same SDK.

No. ICR runs entirely within your infrastructure through the Server SDK. There are no external API calls and no document data exposure.

No. Printed and machine-generated text should be processed with the OCR Module. Enable both modules if your documents contain both.

Pages, paragraphs, lines, and words. Each page carries its number, DPI, and coordinate origin. Each word carries bounding box coordinates, length, font size, text content, and orientation. Line-level bounding boxes are also available.

Yes. Pages can be restricted to a subset with SetPages, and zones set per page: inclusion zones limit recognition to the regions you list, ignore zones exclude them. Zone coordinates are in PDF user space with the origin at the bottom left, and rotate with the page.

Accuracy depends on the source: handwriting clarity, image resolution, contrast, and layout all affect the result, and strikethroughs or corrections affect results with even the best models. For production workflows, extract the JSON and run a validation pass: spell checking, dictionary lookup, or allowlist comparison, before results are trusted downstream, with human review on high-stakes fields. The fastest way to assess it is to run your own documents through a trial.

Yes. JSON produced by a different OCR or ICR engine can be applied to a PDF through the same API, using the documented page and word structure.

ICR is a Server SDK module. Web and mobile front ends call it server-side, and the resulting searchable PDF is then viewed or annotated in the Apryse Web SDK or Mobile SDK.

Windows, Linux, and macOS on x64. 32-bit platforms are not supported for add-on modules. Node.js and Python packages are available for Windows and Linux; other languages install by expanding the module archive into the SDK directory. ARM and Apple M architecture are coming soon.

Handwriting ICR is licensed as an add-on package to the Apryse Server SDK and requires a separate module download. Trial keys have unlimited access to all add-on modules during evaluation. Contact Apryse for pricing based on your deployment and the capabilities you need.

Run it against your own forms

Load your own handwritten forms, claims, and archived records and check the output inside your own environment - on the documents you actually receive, not a clean sample set.