Our Fall Release Arrives October 7th: Get a Sneak Peek Now

Server SDK - Smart Data Extraction

Reduce LLM Token Costs

The fastest way to extract data with an LLM is also one of the most expensive, because you pay for every token it reads and writes. Watch how extracting specific data first changes the math, then see the difference at your volume. In this demo, a 30-page prospectus goes to an LLM in its entirety, then again after extracting only the key data, with token counts and inference cost compared side by side. Want to estimate your own costs?

Data Extraction Approaches Comparison & Associated Costs

When an analyst wants to pull a few key figures out of an investment prospectus (the expense ratio, the minimum investment, the top holdings), the fastest method is to hand the whole document to an LLM and ask. It may not be great at recognizing document structure or layout, but it returns the data on the first try, and it feels like magic.

Reaching for the model by default is one of the most expensive habits you can build. The bill hides in plain sight.

You pay twice. First, for everything the model must process to find a handful of relevant figures. Then again for the response it generates. Two meters running, both billing for content you never needed.

For one analyst pulling one prospectus, that waste is a rounding error. But the moment a developer embeds that same extraction process into a pipeline, set to run unattended across endless documents, the rounding error compounds. Across a million documents a year, it's a line item.

How LLM Token Pricing Affects AI Processing Costs

Copied to clipboard

Cost = (input tokens × input price) + (output tokens × output price) + any special charges (reasoning, caching, retrieval).

When working with an LLM, the unit of compute is the token. The price per unit is your per-token rate.

You don't get to negotiate the per-token price. You do get to decide how many tokens you send, and that's where nearly all the recoverable waste lives.

Understanding Input and Output Token Costs

Copied to clipboard

On a typical "dump it in" call, the tokens fall into two piles, and neither is doing work for you.

The input pile is everything you paid the model to read. A web page is mostly navigation, scripts, styling, and markup. A PDF transformed to text carries elements that repeat on every page, such as headers and footers, plus whole irrelevant sections. You pay the per-token rate to push all of that through the model so it can find the small part that mattered.

The output pile is what you paid the model to write. Ask in plain language, get plain language back: "Sure! Based on the document, the values you're looking for are…". You wanted three fields of structured data. You paid for a paragraph.

Reduce LLM Token Usage with Structured Data Extraction

Copied to clipboard

Right-sizing your model helps. A lower per-token rate helps. But the largest, most reliable savings come from there being fewer tokens to pay for, because you decided what the model needed to see before you ever called it.

Most documents have structure (tables, sections, fields, layout) that code can resolve without ever calling the model. A deterministic engine reads a page's geometry, finds the table, and isolates the relevant section, all for zero tokens. Save the LLM for the reasoning: comparing values, summarizing them, answering something that spans them. Hand it the structured result, not the raw document, and bind it to a tight schema so it returns the fields and nothing else.

Apryse Smart Data Extraction works this way. It reads a page's geometry with a deterministic layout engine, not an AI one, so it reconstructs a document's structure the same way every time. Purpose-built models then interpret what that structure means, trained specifically on high-stakes documents like contracts and forms rather than general use. The expensive, token-burning work of finding what's relevant happens first, for the cost of ordinary CPU time.

That matters for PII. Send the LLM only the data it needs, and the cheapest path and the most private path become the same one. The tokens you never send are tokens that never leave your environment. When documents arrive as scans, the same parsing step runs on self-hosted OCR, so recognition happens on your infrastructure before extraction begins.

How much can you save?

LLM Token Cost Calculator

Take a 30-page financial prospectus. An LLM doesn't need the whole thing, it needs a handful of fields and clauses buried inside it. Adjust the volume, token counts, and pricing below to compare sending the raw document with extracting first.

  • Extraction happens on-prem.
  • Only the extracted data leaves your environment, not the full document.
  • This is inference cost only. It excludes SDK license and server cost

Apryse helps organizations optimize documents before they reach the AI layer. By extracting structure, relevant content and contextual metadata, Apryse enables AI systems to work with less data, reduce token consumption and deliver more accurate responses at lower cost.

Calculator inputs

Pricing Breakdown
See Calculation Assumptions

Cost Per Full Document

Input

$0.0300

Output

$0.0080

Total Daily Cost

$380

Cost Per Extracted Data

Input

$0.0034

Output

$0.0054

Total Daily Cost

$88

Estimated Savings

(Full Document - Extracted Data = Savings)

Daily

$380 - $88 = $292

Monthly

$11,400 - $2,640 = $8,760

Yearly

$136,800 - $31,680 = $105,120

Bottom Line

Sending 10,000 documents daily to Open Ai's GPT-5.6 Sol, with the token pricing being $0.0040 input / $0.0040 output per 1,000 tokens, is estimated to cost $380 per day without extraction and $88 per day with extraction.

Estimates are illustrative only and depend on user inputs and stated assumptions. Actual costs and savings may vary. Apryse does not guarantee savings or results.

How Apryse Smart Data Extraction Reduces AI Costs

Sending the raw document means paying a markup on noise with every document. That gap isn't a rounding error. It's the difference between an AI feature that's cheap to run at scale and one that eventually eats its own margin.

LLMs are precision instruments, not bulk processors. Feed them the smallest, cleanest input that does the job, and let deterministic code do the heavy, repetitive structural work it was always better at.

Smart Data Extraction is an add-on module for the Apryse Server SDK. Get your free trial key at dev.apryse.com.