Paola Ganum is an Economist, and Tohid Atashbar a Senior Economist, both at the International Monetary Fund
The debate over artificial intelligence in economics has moved rapidly from novelty to practical use. Economists are already asking whether large language models (LLMs) can help with literature reviews, forecasting support, data extraction, and document analysis. Korinek (2023) argues that generative AI could become a powerful complement to economic research, while previous VoxEU columns show how LLMs are being used to navigate large research literatures and to perform complex text analysis at scale (Garg and Fetzer 2025, Bermejo et al 2024).
At the same time, broader debate about AI’s economic impact has become more concrete, with firms themselves increasingly expecting productivity gains even as near-term effects remain uneven (Yotzov et al 2026).
But one question matters especially for public institutions: can current LLMs do more than summarise text? Can they support expert judgement in policy surveillance tasks that require reading long reports, tracing arguments across sections, and deciding whether the analysis is genuinely strong?
In recent work, we study exactly that question by testing whether current LLMs can assess macrofinancial coverage in IMF Article IV staff reports (Ganum and Atashbar 2026). These reports are long, technical documents that bring together macroeconomic developments, financial sector vulnerabilities, risk assessment, and policy advice. Reviewing how well they integrate macrofinancial issues is a demanding exercise even for trained economists. Our goal is not to ask whether LLMs can replace economists, but whether they can reliably assist them.
Why this is a hard test for AI
This task is harder than ordinary text classification. It is not enough for a model to spot terms like ‘banking sector’, ‘vulnerabilities’, or ‘macroprudential policy’. A good review must judge whether risks are clearly identified, whether they are integrated into the baseline macroeconomic discussion, and whether the policy advice actually follows from the risks described.
That makes this a useful real-world test of frontier AI systems. Much of the existing literature evaluates LLMs on narrower tasks such as sentiment, classification, or extraction. Our setting instead asks models to perform a multi-step professional review: read a full report, answer factual yes/no questions, assign qualitative ratings, and justify those ratings.
Many policy institutions face similar demands: long reports, repeated review exercises, and a mix of routine extraction and difficult judgment. LLMs are becoming good enough to support the former. They are not yet dependable enough to own the latter
What we did
We compared several GPT-family models against human economist benchmarks using published IMF Article IV staff reports from 2016 to 2024, excluding the pandemic years 2020 and 2021. The sample covers 543 reports across advanced economies, emerging markets, and low-income countries. For each report, economists had already answered a detailed questionnaire on macrofinancial coverage. We then asked different GPT models to answer the same questionnaire.
The questionnaire had two parts. First, there were three qualitative rating areas: macrofinancial coverage in the baseline, assessment of risks and vulnerabilities, and mapping from risks to policy advice. Second, there were 40 binary yes/no questions on more concrete issues, such as whether the report discusses banking sector vulnerabilities, financial integrity, analytical tools, or financial sector policies.
We tested GPT-4o, GPT-4.1, o1 (a.k.a. GPT-o1), GPT-5, and GPT-5.5. We also refined the prompt, adding examples of stronger and weaker economist write-ups so the model had a clearer benchmark for what high and low ratings should look like.
The main result: sseful assistance, not expert replacement
The headline result is encouraging, though nuanced. On binary yes/no questions, the more advanced models perform fairly well. In 2024, average exact match rates against human economists were roughly 76% to 81% across the better-performing models. On structured factual questions, then, today’s models can already provide useful support.
On qualitative ratings, performance is lower but still meaningful. In 2024, average accuracy on ratings reached about 71% to 75% for the stronger models. Exact agreement on the precise rating remained low, but near-agreement was much better: for GPT-o1 and GPT-4.1, the model’s overall rating was within half a point of the human rating in roughly 70% to 74% of cases, and within one point in around 97% to 98% of cases. We obtained similar results with GPT-5 and GPT-5.5 models, where differences between model and human ratings did not exceed one point in 91% of cases.
This pattern matters. It suggests that LLMs are often directionally right even when they are not perfectly aligned with the human benchmark. For institutional workflows, that can still be valuable. A model that is usually ‘close enough’ may help with first-pass reviews, consistency checks, or identifying reports that deserve closer human attention.
Newer models are clearly better
A second finding is the pace of improvement across model generations. Our early tests with GPT-4o showed modest performance and a strong tendency to rate reports too favourably. Accuracy on rating questions was much weaker, especially in earlier years. But newer models did substantially better. GPT-4.1, o1, GPT-5, and GPT-5.5 all performed better on the rating task, with o1 and GPT-5.5 showing especially strong gains. In 2024, rating accuracy reached 77% for GPT-5.5, 75% for o1, 72% for GPT-5, and 71% for GPT-4.1, compared with 59% for GPT-4o.
This mirrors a broader point in the recent AI literature: model choice matters, and it may matter even more in tasks that require long-context reading or multi-step reasoning. Some models are better generalists; others are better at deliberate reasoning or tracking detail across long documents. For applied economic work, that means institutions should not think of ‘using AI’ as a single decision. The practical question is which model, for which task, under which prompt design.
Where LLMs still fall short
The clearest weakness is not hallucination in the usual sense. In our setting, the main problem is over-optimism. Across models and years, LLMs tended to assign higher and less dispersed ratings than human economists. Even when newer models improved, they still leaned toward generosity.
In effect, the models were more willing than human reviewers to give credit for mentioning macrofinancial issues, even when the discussion lacked depth or when the links to policy advice were weak.
This matters because surveillance is not only about whether a topic appears in the report. It is about whether the analysis is integrated and substantive. Human reviewers seem better at making that distinction.
The second weakness is performance on open-ended judgment tasks. Our regression analysis shows that model-human matches are much more likely on factual and simple questions than on complex or rating-based ones. The predicted probability of a match was about 80% for factual questions and simple questions, but only about 7% for rating questions.
That is a sharp reminder that these systems are still better at structured extraction than at nuanced professional evaluation.
What this means for policy institutions
The right lesson is neither hype nor dismissal. On one side, our findings suggest that LLMs can already save time on routine elements of document review. They appear especially useful for extracting objective information, checking coverage of standard issues, and perhaps providing a second opinion. In settings where staff face hundreds of long reports, those gains can be important.
On the other side, current models are not ready to replace economists in high-stakes review work. They still struggle when the task requires deeper contextual judgment, interpretation of ambiguous wording, or a more critical reading of whether the argument really hangs together. That is exactly where expert oversight remains essential.
The most promising approach is therefore human-machine complementarity. LLMs can act as preliminary reviewers, helping staff process long documents more efficiently and more consistently. Human economists can then focus on the hardest part: deciding whether the underlying analysis is genuinely convincing and well-integrated.
That broader lesson likely applies beyond our setting. Many policy institutions face similar demands: long reports, repeated review exercises, and a mix of routine extraction and difficult judgment. LLMs are becoming good enough to support the former. They are not yet dependable enough to own the latter.
References
Bermejo, V, A Gago, R Galvez, and N Harari (2024), “Generative AI as a replacement for human coders in large-scale complex text analysis: New evidence from large language models”, VoxEU, 24 November 2024.
Ganum, P and T Atashbar (2026), “How Effectively Can Current LLMs Analyze Macrofinancial Issues?”, IMF Working Paper WP/26/35.
Garg, P and T Fetzer (2025), “Leveraging large language models for large-scale information retrieval in economics”, VoxEU, 19 February 2025.
Korinek, A (2023), “Generative AI for Economic Research: Use Cases and Implications for Economists”, Journal of Economic Literature 61(4): 1281–1317.
Yotzov I, JM Barrero, N Bloom, P Bunn, S Davis, K Foster, A Jalca, B Meyer, P Mizen, M Navarrete, P Smietanka, G Thwaites, and B Z Wang (2026), “Firms predict an AI productivity boom is coming”, VoxEU, 12 March 2026.
Authors’ note: The views expressed in this column represent only our own and should not be attributed to the International Monetary Fund, its Executive Board, or its management. This article was originally published on VoxEU.org.
