Archipanion Extract – Using AI to help make digitised collections more searchable

Digitisation has made vast cultural heritage collections more accessible, but the real challenge is turning scanned documents into information that can be meaningfully searched and explored. Building on its work to make cultural heritage collections more searchable and accessible, Archipanion is now inviting institutions to explore its new Archipanion Extract platform through an early-access programme.

Archipanion Extract showing information extracted from Franklin D. Roosevelt’s message to Winston Churchill, dated 9 February 1942. Example message courtesy of the Franklin D. Roosevelt Presidential Library & Museum, Hyde Park, New York.

Many cultural heritage institutions have invested significant resources in digitising their collections and making them available online. However, digitisation and online publication do not automatically make collection-wide searching possible.

In text-based archival records, this often presents as the ‘page-image’ problem: even when documents have been scanned, the information researchers need, such as names, places, dates and subjects, remains embedded in scanned page images and cannot be searched across a collection in a structured way.

AI metadata extraction addresses this problem. It converts scanned archival records into structured data by extracting strategically chosen fields, such as names, dates or other high-value descriptive information. The extracted fields can then be integrated into an existing catalogue system or searched through a dedicated interface, allowing staff and researchers to search across the collection rather than examining records one by one. Until now, applying this process has required specialist technical support. Archipanion Extract puts the process into the hands of cultural heritage professionals themselves.

Archipanion Extract showing information extracted from Franklin D. Roosevelt’s message to Winston Churchill, dated 20 February 1942. Each extracted field can be checked against the original scan. Example message courtesy of the Franklin D. Roosevelt Presidential Library & Museum, Hyde Park, New York.

How Archipanion Extract works

Upload a collection, describe what you want extracted, check the results against the original scans and export the structured data for use in your catalogue or search interface. The result is a collection you can search using the information you have chosen to extract.

The six steps in Archipanion Extract, from uploading a collection to exporting the extracted data. Example documents courtesy of the Franklin D. Roosevelt Presidential Library & Museum, Hyde Park, New York.

Choosing what matters

One of the most important ideas behind Archipanion Extract is that cultural heritage professionals decide what information is worth extracting from a collection.

This matters because there is no single set of metadata fields that is equally useful for every collection. The questions a researcher might want to ask of diplomatic correspondence are different from those they might ask of a postcard collection. What deserves to become structured data therefore depends on the records, their context and the ways in which people are likely to use them.

Archipanion Extract showing the fields the user asked AI to extract from a comic postcard. Example postcard from the Tichnor Brothers Postcard Collection, Boston Public Library.

This makes the work as much curatorial as technical. A good AI metadata extraction process therefore starts with a question: what would we like people to be able to find across the whole collection that they cannot easily find today?

Early access as a partnership

Archipanion is now offering institutions early access to Archipanion Extract.

Participating institutions will be able to trial the platform on a limited number of their own collections, free of charge and with hands-on support from the Archipanion team. The early-access programme is collaborative: participants are invited not only to use the platform but also to share feedback and suggestions to help improve it.

What participants will gain

By the end of the trial, participants will have completed full extractions on their selected collections and exported the extracted data as CSV or JSON files ready for use in their own catalogues and systems. They will also have gained the practical knowledge to continue the same work on further collections within their own institutions.