Part of the process: Taking AI integration to the next level
This guest opinion piece is the fifth in a series from Björn Fastabend, bringing us a regulator’s perspective alongside a wealth of experience in digital reporting. Björn is head of the XBRL collection and processing unit at BaFin, Germany’s Federal Financial Supervisory Authority, where he supervises all related activities and leads the implementation of strategic initiatives. He is also a member of XBRL International’s Board of Directors, and previously chaired our Best Practices Board.
The views and opinions expressed in this publication are those of the author. They do not purport to reflect the opinions or views of BaFin.
The next logical step
In this series I’ve outlined how regulators can start experimenting with AI now, without waiting for fully explainable AI to become available. The approach is simple: leverage AI to do what it does best, exploration, and let supervisors focus on verifying findings of interest using business intelligence (BI) tools. I discussed a simple workflow, allowing users to “chat” with their data using natural language – making the benefits of AI accessible to even the least tech-savvy supervisors. This can deliver enhanced insights over BI tools alone, but only scratches the surface of what is possible when employing AI in the regulatory supervisory process.
Thus far we have been assuming that the analyst initiates the conversation, making the conscious decision on when to interact with their data. But what if they didn’t have to? Wouldn’t it be so much easier and more logical to automate AI involvement? Each and every filing gets analyzed automatically, with no human trigger involved. The AI system becomes simply part of the process, suggesting areas for investigation ready for when the human supervisors get to work
While this approach would automate the AI component, the goal should never be to cut out human decision making. Whatever ways we incorporate AI into the analysis process, supervisors will still need to verify AI findings using BI tools to maintain regulatory rigor, and to supply their own expertise to interpret and contextualize the results. The only difference from the approach I’ve previously outlined would be that AI gets to work in the background without requiring activation. Let’s have a look at how this might work.
The pipeline alternative
Automation makes sense. Think about it: we are dealing with XBRL data, a format that is inherently designed to be processed automatically. The great strength of AI is its ability to process vast quantities of data. A powerful way to leverage the combination of XBRL and AI is to enable automated analysis upstream, as a routine early step in the data processing pipeline.
After the XBRL data has been disseminated into the database and is on its way to the data warehouse, an AI could automatically process and analyze the data, flag any findings and finish by preparing a review for the user. At the same time, providing chat capabilities for supervisors to interact with their data using natural language processing remains a valuable use case for AI. The beauty is that these approaches are not mutually exclusive. Supervisors may choose to follow automated processing with further interaction as they pursue leads and explore their own ideas.
Picture the following: It’s Monday morning and a banking supervisor comes into their office. A major bank has made its latest report – and, before the supervisor has finished their coffee, a combination of AI and machine learning processes has analyzed the data and created a report. The supervisor does not start from scratch. They already have an overview of potential issues and items to follow up on. The AI process has taken care of the heavy lifting of comparing peer group data and finding outliers. The supervisor will now be able to use this report as a baseline, perhaps asking AI further questions, and verifying the findings using BI tools as usual, spending less time on manual analysis and more on points of interest.
This approach not only frees resources by automating core AI analysis without a manual trigger, but implicitly strengthens the analytic process, since all reports are examined, not only those that supervisors choose to interact with. As an added bonus, it also elegantly circumvents a very human problem.
User adoption solved
I believe that AI-aided analysis, combining AI for exploration and BI for verification, will prove not only viable but a powerful asset for supervisors. But there’s a catch I have not previously discussed. Humans vary widely in their engagement and capabilities. Some supervisors will be keen to maximize their use of AI; others will be moderately interested but lack the time and energy to prioritize new ways of working; and others will prefer to avoid change entirely.
A system that provides a baseline automated analysis of every filing will make it easier for all supervisors to reap the benefits of AI, without needing every individual to take a proactive approach. It will also ensure that every report goes through a minimum level of AI analysis – while allowing further interaction with the data as desired.
At the same time, this solution adds further capabilities to the analysis process, enabling for example detection of outliers, narrative inconsistencies, threshold breaches, and peer group deviations. These are much less viable using only the workflow I outlined previously. Automation is a win-win for all involved: a smoother path to leveraging AI, and more comprehensive insights.
Challenges
Before you pick up the phone to call your IT department and ask them to implement this automated processing, let’s take a look at the potential pitfalls.
First and foremost, such a system needs built-in counter measures against false positive results. The risk of AI hallucination and the potentially large volumes of data being processed make this a significant challenge. One possible solution is to employ a “fact-check AI”: after the first AI system analyzes the data, a second AI is used to reverse engineer its findings. It can check if the values reported exist in the database, assess plausibility, and attach a confidence score to each finding. Only those findings that surpass a given score are then shown to the user. While it would not be impossible for a small number of errors to sneak through, this approach seems likely to prevent most false positives reaching the user, lending the system necessary functionality and credibility.
It is important to note that not all analysis operations require a large language model (LLM). Outlier detection, for example, is very well suited to machine learning (ML), which is far less likely to hallucinate and thus to generate false positive findings. A successful system is likely to combine both types of AI model: ML excels at identifying what is statistically unusual, while LLMs provide the ability to interpret why that might be significant. Using the XBRL taxonomy to supply context, an LLM can translate a statistical outlier into a human-readable finding, explaining which reported concept looks anomalous, how it compares to expected, peer and past disclosures, and why it might warrant a closer look.
That combination of ML precision and LLM interpretation is key to a genuinely useful pipeline rather than an interesting experiment. This value is enabled by the semantic richness of XBRL data and underlying taxonomies, which give the LLM the context it requires to make meaningful conclusions. As my colleague Gerhard Gerlach has said – and I will repeat until it sticks – “Provide an AI with an XBRL taxonomy and change the analysis from guessing to knowing”!
When designing this automated processing, it is important to consider timing. Each individual filing could be analyzed on submission in timeseries comparisons against previous filings by the same company, and perhaps checks against baseline expectations. For peer-group comparisons, however, it is necessary to wait to receive all relevant filings for the reporting period. Since the taxonomy dictates the structure of the filings, XBRL provides not only rich semantic information but also essential comparability in such analyses.
A further consideration, and one that is non-negotiable in a regulatory context, is the audit trail, which inevitably gains complexity with the addition of automation. If an AI system flags an anomaly that ultimately leads to regulatory action, the path from raw data to that decision must be fully traceable and defensible. This means logging not just the AI’s findings, but the inputs it worked from, the confidence scores assigned, and the verification steps taken by the supervisor. The human verification step via BI tools remains essential, not just as a quality check, but as the legally defensible layer of the process.
Don’t wait for perfect
The workflow I have outlined here is not a distant vision. The technologies exist today: XBRL data in structured databases, ML anomaly detection, LLMs that can reason over taxonomy context, BI tools that provide verification and legal defensibility. All that’s needed is the decision to start.
Most regulators can expect increasing volume and granularity of data in the coming years. The question is not whether AI should play a role in processing that data, but whether regulators will shape that role or find themselves scrambling to catch up.
The pipeline I have described here is one approach, but it will not be perfect on day one. Confidence thresholds will need calibration. Audit trail infrastructure will need investment. There will be false positives that slip through and findings that miss the mark. That is not a reason to wait; it is a reason to start now, and make mistakes while the stakes are still manageable.
As I have argued throughout this series: AI explores, BI verifies, humans decide. Automation simply shifts the bulk of AI analysis upstream, so that by the time the supervisor arrives on Monday morning the preliminary work is done and potential insights are waiting to be explored.

