The Apryse Summer 2026 Release: OUT NOW

Home

All Blogs

Exporting OCR Data to JSON with Apryse

Published January 30, 2025

Updated August 12, 2026

Read time

4 min

email
linkedIn
twitter
link

Exporting OCR Data to JSON with Apryse

Sanity Image

Garry Klooesterman

Senior Technical Content Creator

Sanity Image

Summary: OCR technology has revolutionized the handling of printed documents by converting images of text into machine-readable data. However, this data is often unstructured and hard to analyze. Converting OCR output into a structured format like JSON makes the text searchable, accessible, and easy to integrate with various applications. This enables efficient data analysis, automation, and archiving, unlocking the value of previously inaccessible information.

Introduction

Copied to clipboard

Optical Character Recognition (OCR) technology has transformed how we handle printed documents. Converting images of text into a machine-readable form often results in unstructured and difficult to analyze data. A powerful solution is to structure this OCR output into a format like JSON (JavaScript Object Notation). OCR to JSON extraction makes text searchable, accessible, and easy to integrate with various applications, enabling efficient data analysis, automated entry, archiving, and more. Ultimately, OCR to JSON unlocks the value of previously inaccessible information.

In this blog, we’ll explore JSON, exporting data to JSON using an OCR SDK, how the Apryse OCR SDK will be your top choice, and the process to export OCR to JSON.

JSON

Copied to clipboard

JSON, a lightweight data-interchange format, is easily usable by humans and machines. Built on a subset of JavaScript, and drawing familiar conventions from languages like C, C++, C#, Java, JavaScript, Python, and others, JSON offers an ideal and flexible way to exchange data. Its straightforward structure requires no specialized knowledge or tools for analysis and interpretation.

Apryse OCR

Copied to clipboard

The Apryse OCR Module adds enterprise-grade text recognition to the Apryse Server SDK, ready to transform your document workflows.

Key Features

  • Handles various document types, including tables, multiple languages, and different text orientations.
  • Includes preprocessing steps such as deskew and despeckle that you apply before recognition.
  • Efficiently handles high-volume workflows.
  • Easily integrated into existing systems and server environments.
  • Can be fine-tuned for optimal performance by adjusting processing parameters.
  • Runs entirely on your own infrastructure, so documents never leave your environment.

Developers can fine-tune the OCR engine for optimal performance by adjusting processing parameters to meet the needs and requirements of any project.

From reduced costs and increased efficiency to improved accessibility and customized data output, the Apryse OCR offers a comprehensive suite of benefits for your business. If you’re looking for a practical example, see how to create a .NET Core cross-platform OCR application.

OCR Engines

Copied to clipboard

The Apryse Server SDK offers more than one OCR module, so you can match the recognition engine to your document set rather than accept a single fixed path. The default module is built on deep learning neural networks and covers a broad language set including CJK, Arabic, Cyrillic, and Latin-alphabet languages, the right starting point for most projects, and the one this blog uses. Alternative modules are available for workloads with existing accuracy baselines. If only one module is installed, the SDK uses it automatically; with more than one present, you select through the OCR options object. See the add-on modules documentation for the current list and installation steps.

Export OCR to JSON

Copied to clipboard

JSON’s structured format makes it ideal for organized data storage and seamless exchange between applications and systems. When you export OCR to JSON, you not only preserve the extracted text but also include valuable metadata – information about the text’s appearance with the document, such as font-size, orientation, and more. This rich data set simplifies storage, access, and future use. Furthermore, the extraction process itself can be implemented with minimal code.

Step 2: Now let’s use OCR to extract the data to a JSON file.

Output Attributes

Copied to clipboard

Now that we have our JSON file with the extracted OCR data, we can see the various attributes of the data. Output consists of nested arrays for pages, paragraphs, lines, and words. Pages have the following metadata:

Blog image

Figure 1: Page metadata in JSON

Each word has the following metadata:

Blog image

Figure 2: Word metadata in JSON

Sample JSON output

Copied to clipboard

Here we can see what the JSON output from OCR extraction could look like.

Blog image

Figure 3: Sample JSON output

So, what do I do with the data?

Copied to clipboard

Great question! Now that our data has been extracted, organized, and stored in a JSON file, we can explore its potential uses, such as:

  • Storing the data for later use.
  • Using the data in another application, database, or system.
  • Applying corrected or processed data back onto the original input document.

Conclusion

Copied to clipboard

In conclusion, by utilizing the robust Apryse Server SDK's OCR Module to extract OCR data to JSON, businesses can unlock valuable information trapped within static formats. Streamlined by the OCR capabilities, this process makes data searchable and accessible and facilitates streamlined integration with various applications and workflows. When using Apryse’s OCR to JSON functionality, you can maximize efficiency, improve accessibility, and realize the full potential of your information.

Check out a demo of Apryse OCR module in action. Contact our sales team for any questions.

Need help setting up? Join our Discord community for support and discussions.