Grants are the lifeblood of the nonprofit sector: 87% of organizations rely on grants from foundations and 46% rely on grants from the government.1 But this revenue comes at a cost: ImpactGraph data shows that more than 90% of grants are saddled with arduous restrictions and compliance requirements.2 Each additional requirement is another item that organizations need to track, which creates a silent administrative tax that goes unfunded. Most nonprofits do not have the capacity to handle this work, which means they may not even apply in the first place. Although movements like trust-based philanthropy have made it easier, these limits still prevent dollars from flowing to the people who could have the greatest impact on their communities.

This challenge seems like a logical place for new artificial intelligence tools like Large Language Models (LLMs), which power platforms like ChatGPT and Claude. These tools have massive potential due to their ability to read and interpret entire documents in seconds. Early research also suggests it has incredible promise to streamline accounting work.3 When it comes to nonprofit-specific grant accounting, however, frontier models frequently make mistakes. ImpactGraph research shows that OpenAI’s second largest model GPT 5.6 Sol correctly extracts requirements from grant agreements only 35% of the time.

ImpactGraph’s mission since day 1 has been to rethink the way that nonprofit leaders do financial work, and AI tools play a large part in how we’re able to help organizations do more with less. When it comes to grants management, we help in a few ways:

  1. Review and extract key details from grant agreements like payments, tasks, and compliance requirements
  2. Proactively suggest grant usage and optimize spending
  3. Report on their grants in an intuitive way to development teams, funders, and board members

Our grant extraction pipeline uses a human-in-the-loop approach to prevent hallucinations. While this approach cuts down the time spent processing grants, it has two downsides. First, organizations still need to retain an expert to review the grant. Second, that expert still has to read the entire grant to ensure the model neither hallucinated nor missed important information.

Our goal is to develop an AI system so trustworthy that even non-experts can use it. We want this system to enable people to rapidly process several documents at once, relying on their judgement only on tasks where AI is known to be weak. To develop this capability we need to know the answer to two questions:

  1. What kind of information from grant agreements does AI struggle to understand?
  2. What can be done to make its information extraction more accurate?

Our research yielded three extraction tasks where AI performance falters and several strategies to bolster its performance on the most difficult extraction tasks. With these strategies implemented, we were able to boost AI’s performance by 38% and 49% on the two most difficult extraction tasks (expenditure restrictions and requirements).

When does AI fail?

We answered this question by hand-labeling a sample of grant agreements with key information such as the grant name, the award amount, payment dates, restrictions, and requirements. We then asked the frontier models to do the same and scored their answers against ours.

Bar chart titled 'Where AI Models Fail to Read Grant Agreements' comparing Claude (Opus 5) and ChatGPT (GPT-5.6 Sol) F1-scores across seven grant agreement elements. Both models score below 0.6 on funding restrictions, funding requirements, and grant purpose, and above 0.65 on grant name, funder name, payment schedule, and grant type.

From this data we identified two categories of information:

  1. Explicitly stated (payment schedules, grant names, funder names, grant type)
  2. Implicitly stated (funding restrictions, funding requirements, grant purpose)

It’s the implicit information where AI performance degrades. This makes sense: an AI model can essentially copy and paste a payment date table from a document, but identifying what counts as a restriction on expenditures requires it to reason about the text it sees. Consider this example from a grant agreement:

“Reporting – Grantee shall report on activities on a quarter basis through [the grantor’s] quarterly reporting form with a narrative description and photos.”

Claude interpreted this statement as a single reporting requirement. But the correct interpretation is to iterate four reports per year until the end date of the grant period. It’s reasoning where these models struggle, and this is where we looked to create performance improvements.

Solution 1: An Expert Layer

We taught the models how to reason about implicit statements by implementing an expert layer before it executes the grant extraction. The expert layer instructs the model how to reason about the decisions it needs to make when extracting information from a grant agreement. We identified candidate instructions for the expert layer by classifying the errors GPT made when trying to execute the information extraction task and then A/B tested those candidates against baseline GPT performance.

Our experiments yielded over 100 beneficial reasoning instructions. Here’s a sample of the broad categories the instructions fit into, with an example of each:

  1. Date inference – Generate every instance of a recurring report from a start date, end date, and frequency word (“annual”, “quarterly”).

    Example: “Report quarterly through the grant period” on a two-year grant becomes eight dated reporting requirements, not one.

  2. Consequential reasoning – Define items by the effect they have on the grantor and grantee, not by where they appear in the document.

    Example: An indirect cost rate needs budget context. Sometimes it’s included in the award amount, sometimes it’s added on top. Either way, it isn’t a direct line-item expense.

  3. Exposed reasoning – Justify each choice before returning it.

    Example: The model explains that it created a net asset restriction because the grant is multi-year and future-year funds aren’t spendable yet.

  4. Brightlining – Check whether a candidate item belongs to an explicitly included or excluded class before keeping it.

    Example: Boilerplate like “Grantee must maintain proper financial records” appears in nearly every agreement. It’s legally important, but not a restriction that the recipient needs to track.

  5. Elementwise classification – When two classes share a fuzzy boundary, look for the specific elements that tell them apart.

    Example: A matching grant could be interpreted as both a requirement and a restriction. We’re able to resolve this and handle the unique revenue recognition.

Our expert layer demonstrates impressive results: a 34.5% relative improvement in F1-score over the best model for restrictions and a 48.6% relative improvement for requirements.

Box plot titled 'The Expert Layer Significantly Improves Performance on Implicit Extractions' comparing GPT-5.6 Sol with and without the expert layer. F1-score rises from 0.55 to 0.74 for funding restrictions, 0.35 to 0.53 for funding requirements, and 0.54 to 0.75 for grant purpose.

Solution 2: VLM Parsing

Accurate reasoning depends on accurate data. Therefore, we looked for ways to improve the data fed to AI models. The process of turning an image of a grant agreement into text is called parsing. Our original parser used ML technology called OCR, and we built a new parser based on an AI-driven VLM (Visual Language Model). What’s important about a VLM is that it is capable of understanding spatial structure in a document: whitespace, tables, paragraph breaks, bullet points, etc. Our hypothesis was that this spatial structure contains important information that bolsters the capability of LLMs to understand the grant agreement.

Box plot titled 'VLM Parsing Significantly Reduces Hallucinated Funding Restrictions' comparing ImpactGraph with Textract against ImpactGraph with LlamaParse on funding restrictions. Precision rises from 0.54 to 0.62 while recall stays level at 0.67 versus 0.69. Predictions were generated by GPT mini, so the scores are not comparable to the expert-layer chart.

The evidence points to our hypothesis being correct: we saw a 15.3% increase in the precision of our restrictions extractions without any tradeoff in recall. In other words, the VLM parser helped the model hallucinate significantly less often without missing any additional information.

What this means for nonprofit accounting

Accelerating grant agreement processing is essential to letting nonprofit organizations focus on their missions. AI can help via a human-in-the-loop approach, but the unreliability of frontier models means it still requires the attention of experts. ImpactGraph’s expert layer greatly improves the trustworthiness of AI models, moving them much closer to becoming a true partner for organizations. If you’re an ImpactGraph customer, you should start noticing the effects of these improvements soon. In the meantime we will continue to push forward the horizon of AI capability in nonprofit accounting.

Want to try it for free on your next grant agreement? Get started with ImpactGraph.

Citations

  1. https://candid.org/blogs/diversifying-revenue-sources-where-nonprofits-find-funding/ ↩

  2. Estimated from a sample of 60 nonprofit grant agreements randomly selected from public and proprietary sources. “Arduous” means we did not count boilerplate restrictions or requirements commonly found in all grant agreements. ↩

  3. Choi, J.H. and Xie, C.L. (2026), Human + AI in Accounting: Early Evidence from the Field. Journal of Accounting Research, 64: 1333-1373. https://doi.org/10.1111/1475-679x.70052 ↩