Choosing AI Models for Archival Work: From Static Lists to Continuous Evidence

Archives operate in an era where AI tools are evolving rapidly and selecting the right technology, whether identifying names in multilingual documents, detecting sensitive personal information, or improving searchability, has remained a fragmented and overwhelming process. ArchXAI addresses this challenge by providing a continuously updated, transparent evidence base that helps archivists move beyond guesswork and hype. Rather than asking “which AI model is best?” this project asks the more practical and useful question: “under what conditions does a particular approach work, and where does it fail?”

ArchXAI has launched a public website that compares AI models and methods for use in archival work. The site is updated continuously as new tools, tests, and results become available. The resource is available at ArchXAI Technology Updates.

The aim is to follow developments in artificial intelligence and natural language processing and to identify which approaches are genuinely useful for archives. The work examines tasks such as identifying people, places, organisations, and dates in documents; classifying images; detecting sensitive personal information; analysing tone; and improving search.

The site is not simply a list of available tools. It is an open and evolving evidence base. It explains what has been tested, why each task matters for archival work, how the results were measured, and what limitations remain. The website includes an explanation of the testing methods, pages devoted to individual topics, and a development blog where new benchmark results and observations are published.

Current Research Findings

The strongest evidence currently covers four areas.

Identifying names, places, organisations, and dates

Named entity recognition, often shortened to NER, is a method for automatically identifying important names and references in text.

The results suggest that specialised AI models designed for this task remain the most reliable starting point for routine multilingual indexing. Large language models can also be useful, but they are generally slower and may be better suited to difficult cases or to adding further information.

Detecting sensitive personal information

ArchXAI has also compared methods for finding and managing personally identifiable information, such as names, addresses, identification numbers, and other sensitive details.

The evaluation considers not only how accurately the systems detect this information, but also how well they support archivists in reviewing results and deciding what should be hidden, anonymised, or made available.

Analysing tone and sentiment

AI can be used to identify whether language appears positive, negative, neutral, formal, hostile, or emotional.

However, the current recommendation is cautious. Existing datasets and categories may not reflect the complexity of archival documents, historical language, or the context in which records were created. More testing is therefore needed before these methods can be used confidently in archival work.

Improving search

ArchXAI has tested how AI-based semantic search could help users find records that are related in meaning, even when they do not contain exactly the same words as the search query.

The main finding is that semantic search is useful, but it should not replace traditional search methods. The most effective archival search systems are likely to combine several approaches, including keywords, metadata, names, dates, numbers, and AI-based similarity search.

Why this approach matters

Choosing technology for archival work cannot be treated as a one-time decision. AI models change quickly, available datasets vary in quality, and the needs of archives differ according to language, collection type, legal requirements, and institutional practice.

A continuously updated and transparent comparison framework reflects this reality. It provides a shared reference point for understanding how different methods perform in real archival settings. It also makes the testing process visible and identifies areas where the available evidence is still incomplete.

Most importantly, it shifts the discussion away from the simple question, “Which model is best?”

A more useful question is:

Under what conditions does a particular approach work, and where does it fail?

The ArchXAI Technology Updates site helps answer that question through an open, revisable, and evidence-based process.

Writer: Paul Johannes, National Archives of Estonia