The extended deadline [1] for individual returns is October 15. For many firms the months between that deadline and January are the only stretch of the year when staff have time to change how source documents get into the return.
This piece is for tax partners and operations leads at firms of roughly 10 to 150 people who still have preparers keying numbers from K-1 packages, brokerage statements, and bank statements. It describes how we would set up extraction with a review queue and a source trail, and how to decide whether you need anything custom at all. We have not published CPA engagements; this is our approach, built from public vendor documentation and IRS and FTC guidance.
Where the keying time goes
Standard one-page forms are rarely the problem. The time goes into documents with variable layouts and long attachments.
A partnership Schedule K-1 carries coded items whose detail often sits outside the form. The IRS partner instructions [2] list codes A through X and ZZ for box 11 alone, and box 22 is checked when the partnership attaches a statement for multiple activities. Those statements differ by preparer and fund administrator. State schedules and footnotes add more.
Brokerage packages have their own version of this, with summary pages followed by lot-level detail. Bank statements for clients with business activity add another layer, because the preparer often needs transactions or totals that no tax form summarizes.
Before choosing a tool, count where the hours actually went last season. Pull a sample of returns and note, by document type, how many pages were keyed and how long review took. That count decides most of what follows.
Check what your tax software already offers
Several vendors already cover common documents, and some of that coverage arrived this year.
Wolters Kluwer announced general availability [3] of AI document extraction in CCH Axcess Scan on June 29, 2026, for firms on or moving to CCH Axcess Tax. The company says it handles W-2s and 1099s through K-1 supplemental data and supporting statements. Its feature page [4] describes a K-1 Manager where reviewers validate, adjust, and reconcile extracted data before it reaches the return. The efficiency figures in the announcement are the vendor’s own, from early users, and we have not seen them independently measured.
Thomson Reuters lists SurePrep [5] as working with CCH Axcess Tax, GoSystem Tax RS, Lacerte, and UltraTax CS. K1x [6] lists integrations with CCH Axcess, ProSystem fx, ONESOURCE, GoSystem RS, UltraTax, and GruntWorx, along with APIs for sending K-1s in and pulling structured data out.
If one of these covers your documents and your tax software, run a pilot on it first. Custom work earns its cost in narrower cases:
- document types the packaged tools skip, such as bank statements that need transaction detail;
- a mapping from extracted fields to your own workpapers or input conventions that the vendor cannot represent;
- a tax software and document management combination no vendor connects end to end;
- volume high enough that per-document pricing exceeds the cost of running your own pipeline.
Give every value a source trail
A reviewer should be able to click any extracted number and land on the page and box it came from. Without that, review means reopening the PDF and searching, which is most of the work you were trying to remove.
Store each extracted field as a record, not just as a value in an export file.
| Field | Purpose |
|---|---|
| Client and tax year | Ties the value to an engagement |
| Document ID and hash | Proves which file was read |
| Page and region | Lets the reviewer jump to the box or table row |
| Target field | The input in the tax software the value maps to |
| Extracted value | What the system read |
| Confidence | How sure the extractor was, per field |
| Check results | Which validation rules passed or failed |
| Reviewer action | Accepted, corrected, or rejected, with the corrected value |
| Reviewer and time | Who signed off and when |
This record does double duty. It is the audit trail when a partner asks where a number came from in March, and it is the data you use to measure accuracy and tune thresholds next year.
Route uncertain fields to a review queue
Field-level confidence from a model is a useful signal and an unreliable guarantee. Pair it with deterministic checks that a tax reviewer would recognize.
Checks we would start with:
- totals on a brokerage summary equal the sum of the detail lines;
- the partnership EIN and partner name match the client file;
- each coded K-1 item has the supporting statement the code implies;
- no value moves sharply from last year’s return for the same entity without an explanation;
- a box checked for an attached statement with no statement found.
Any field that fails a check or falls below its confidence threshold goes to the queue. Everything else is presented as accepted but still visible, so a reviewer can spot-check. Thresholds should be set per field type. A misread state code on a K-1 deserves a lower bar for review than a misread page number.
The queue itself should be dull. Show the value, the source image cropped to the region, the failed check, and three buttons. Preparers should not need a second screen.
Measure accuracy on your own documents first
Vendor accuracy figures come from someone else’s documents. Yours include your clients’ funds, your clients’ banks, and the scans your clients actually send.
Build a test set from last season. Take a few hundred documents across the types that consumed the most time, with the values your staff keyed and reviewed as the answer key. Then run the extractor and score it field by field, by document type.
Report two numbers for each document type. The first is the error rate on fields the system accepted without review, which is what reaches the return unchecked. The second is the share of fields sent to the queue, which tells you how much human time remains. A system that sends everything to review is safe and useless. A system that sends nothing is fast and dangerous. The thresholds you choose set where you sit between them, and the test set lets you choose with evidence.
Keep part of the set aside and never tune against it, so the final score reflects documents the system has not seen.
Keep the data inside your security program
Tax data is covered by federal security rules, and an extraction pipeline is a new system that handles it.
The FTC’s Safeguards Rule guidance [7] names tax preparation firms as covered financial institutions. It requires encrypting customer information at rest and in transit, reviewing access controls, logging authorized users’ activity, and disposing of customer information no later than two years after its last use to serve the customer. It also requires contracts that set security expectations for service providers and a way to monitor them. IRS Publication 4557 [8] restates these obligations for tax professionals, and Publication 5708 [9] gives a template for the written information security plan. The IRS reminded preparers in August 2026 [10] that the plan is required.
In practice, we would deploy the pipeline in the firm’s own cloud account and add it to the plan’s inventory of systems that hold client data. Model endpoints need the same scrutiny as any vendor. Microsoft states that prompts and completions sent to models sold by Azure [11] are not used to train foundation models, but that abuse monitoring can store flagged prompts for human review unless the customer is approved for modified abuse monitoring. AWS states that model providers on Bedrock [12] have no access to customer prompts and completions. Read the current terms for whichever service you use, record the settings in the plan, and turn off any stored-conversation features the pipeline does not need.
Getting values into the return
The last step is the one most likely to stall a project, so confirm it before building anything else.
CCH Axcess publishes an Open Integration platform [13] with a Tax API Kit, a developer portal, and consulting from Wolters Kluwer. That gives a documented path for writing reviewed values into returns. For UltraTax and Lacerte, check with your vendor which import formats your edition accepts before you design the export, and test a full round trip on a closed return from last year.
Whatever the path, only reviewed values should cross into the tax software, and each should carry its record ID so the return can be traced back to the source page.
When not to build this
Skip custom extraction if a packaged tool already covers your document mix and your tax software. A pilot on the vendor product during the off-season is cheaper and faster.
Low volume is another reason. A firm that sees a few dozen complex K-1 packages a year will spend more maintaining a pipeline than keying them.
Hold off when no one at the firm will own the review queue and the thresholds after launch. Extraction quality drifts as funds change administrators and banks change statement layouts, and someone has to notice.
If the firm’s written security plan is out of date, fix that first, because the pipeline has to fit inside it.
If your firm has a document mix that the packaged tools do not handle well, we can help you scope it on your own documents before January. Read more about our document and data extraction work.
Sources
This article describes how we would approach the work. It is drawn from IRS and FTC guidance and public vendor documentation, and it does not report results from a client engagement or independently measured vendor accuracy.
- IRS: Get an extension to file your tax return
- IRS: Partner’s Instructions for Schedule K-1 (Form 1065)
- Wolters Kluwer: CCH Axcess Expert AI announcement (June 2026)
- Wolters Kluwer: CCH Axcess Scan features
- Thomson Reuters: SurePrep
- K1x: integrations and API
- FTC: Safeguards Rule, what your business needs to know
- IRS: Publication 4557, Safeguarding Taxpayer Data
- IRS: Publication 5708, written information security plan template
- IRS: Security Summit reminder on written information security plans (August 2026)
- Microsoft: data, privacy and security for Azure AI Foundry models
- AWS: data protection in Amazon Bedrock
- Wolters Kluwer: CCH Axcess Open Integration
