Designing Explainable AI Systems for Digital Archives
User Research
Usability Testing
Helping researchers understand and critically evaluate AI-generated information
UI/UX
Tools
Figma, Figjam, Google Forms, Microsoft Teams
Role
UX Researcher & UX Designer
Project Overview
Background
Digital archives are managing rapidly growing collections that are increasingly difficult to process manually with limited staff and resources. To manage this scale, institutions are turning to AI to support tasks such as generating metadata, summarising records, and improving how information is organised and discovered.
Aim
Although concerns around AI transparency in digital archives are well documented, practical design-led solutions remain limited. This project explored how explainable design could address this gap by helping researchers recognise AI involvement, understand its role through clear explanations, and make more informed decisions when interpreting archival information.
Problem
While automation helps archives make more of their collections accessible, it introduces a new challenge for researchers. AI can influence the information they discover and use, yet its involvement is often invisible or poorly explained, making it difficult to understand how information was produced or judge its reliability.
Contribution
I led the project from initial research through to design and evaluation. Through literature review, user research, prototyping, and usability testing, I developed and validated three reusable interaction patterns that improve transparency, support user understanding, and encourage more critical engagement with AI-generated information.
Define
Affinity Mapping
Empathy Mapping
How Might we' (HMW)
Design Process
Brainstorming
Crazy Eights
Priority Matrix
Develop
Think-aloud usability testing
Evaluation Criteria
Deliver
Discover
Literature Review
Online Survey
User Interviews
Understanding The Problem
Discover
Define
Before designing a solution, I needed to understand both how AI is currently being used in digital archives and how researchers perceive AI-generated information. I conducted user research to explore participants' awareness of AI, their expectations around transparency, and the challenges they face when evaluating archival information. These insights established a clear understanding of the problem space and informed the direction of the design.
Survey and User Interviews
I combined an online survey with semi-structured interviews to explore users' experiences, attitudes, and expectations. The survey provided a broad understanding of existing behaviours and perceptions, while the interviews offered deeper insight into the reasoning behind participants' responses.
The research explored:
Current awareness - how familiar participants were with AI and its role in digital archives.
Expectations of transparency - what users wanted to know when AI contributed to archival information.
Trust and reliability - how AI involvement influenced confidence in archival records and the role of human oversight.
User agency - whether participants wanted opportunities to question, challenge, or report AI-generated information.
Synthesising Findings
I synthesised findings from the survey and interviews using affinity mapping, grouping related observations and responses to identify recurring themes across the research. I then developed an empathy map to bring together what users say, think, feel, and do when encountering AI-generated archival information.
Together, these activities revealed connections between users' awareness of AI, expectations around transparency, perceptions of trust, need for human oversight, and desire for greater control. These themes formed the basis of the key insights that guided the design direction.
Key Insights
Three key themes emerged from the research, revealing how users currently understand AI in digital archives and what they need to engage with AI-generated information more confidently and critically. These findings highlighted opportunities to improve awareness, transparency, and user agency throughout the research experience.
-
Users were generally familiar with AI-based tools in other contexts, but their understanding of how AI is used within digital archives varied considerably.
Awareness of AI use in digital archives was uneven.
Users often lacked visibility into where and how AI had influenced archival information.
This lack of visibility created uncertainty around how records were produced and interpreted.
33.3% of participants were aware that AI was used in digital archives, 33.3% had heard about its use but lacked understanding, and 31% were previously unaware.
“I want to be aware and informed so I can critically think about any biases that might be there.”
-
Users consistently wanted to know when AI had contributed to archival information, particularly when its involvement could influence how they interpreted or relied on a record.
Knowing when AI had been used helped users judge reliability and decide when further verification was needed.
Users actively distinguished between AI-generated and human-created information.
Clear disclosure supported more informed and critical research decisions.
Over 90% of participants wanted to be informed about AI involvement, either always or when it directly influenced how information was presented.
“I don’t mind if it is AI, but it feels important to know the difference between a summary the authors generated and one an AI generated.”
“I would always want to know if AI was used on something I am looking at.”
-
Users recognised that AI systems can make mistakes and wanted ways to respond when they encountered information they believed might be inaccurate.
Awareness of AI's limitations encouraged more cautious and critical use.
Users wanted the ability to question, flag, or correct AI-generated information.
Visible human review increased confidence and provided reassurance that responsibility did not sit solely with the AI system.
Over 90% of participants felt that being able to flag or correct perceived errors in AI-generated information would be useful.
“AI is wrong sometimes! It’s good to know so that you’re not blaming the wrong thing.”
“I would feel more confident using the information if it had been reviewed.”
HMW... support users in forming accurate mental models of AI processes, capabilities, and limitations within digital archives?
HMW... support user agency by allowing users to question, correct, or challenge AI-generated information?
HMW... provide meaningful visibility into human–AI collaboration within archival systems?
Participants:
42 Survey Responses
4 Semi-structured Interviews
How Might We?
To translate these findings into actionable design opportunities, I reframed the key user needs as three How Might We questions. These provided a clear focus for ideation while ensuring potential solutions remained grounded in the research.
Designing the Solution
Develop
With a clearer understanding of users' needs, I began exploring potential solutions to address the key opportunities identified through research.
Concept Exploration
Using the How Might We questions as a starting point, I explored a range of potential solutions through brainstorming and rapid sketching. Ideas ranged from simple indicators and contextual explanations to feedback mechanisms and more detailed AI transparency features.
Sketching allowed me to quickly iterate on layouts, interactions, and user flows before moving into digital wireframes.
Designing Within Existing Archive Experiences
Rather than designing an entirely new digital archive, I wanted the prototype to reflect the systems researchers were already familiar with. I carried out an interface audit of existing digital archives relevant to the project, including The National Archives, the British Library, and Europeana, comparing common patterns across their information architecture, terminology, layouts, and interactions.
Using these findings, I mapped a typical archive structure and created a simplified sitemap focused on the research journey of searching for and viewing records. This provided a familiar foundation for integrating the proposed solutions into realistic points within the archival research experience.
Prioritising Ideas
I used a prioritisation matrix to evaluate the concepts based on their potential user impact and feasibility of implementation.
This helped narrow the initial ideas to three concepts that addressed the strongest needs identified through research while remaining realistic to integrate into existing archival systems:
AI Activity Indicator - making AI involvement visible and explaining where it has contributed.
Explanation Tooltips - providing contextual information about AI-generated content when users need it.
Error & Feedback System - allowing researchers to report potential inaccuracies and making human review visible.
Developing the Design Patterns
With the overall structure established, I developed the three prioritised concepts into explainable design patterns integrated across the search and record-viewing experience.
Given the information-dense nature of digital archive interfaces, the patterns were designed to be lightweight and unobtrusive. Explanations could be accessed on demand, allowing researchers to explore additional context when needed without cluttering the interface or disrupting their existing research workflow.
The three patterns address the core needs identified through research: making AI involvement visible, supporting understanding, and giving researchers greater agency over AI-generated information.
Design Pattern 1 | AI Activity Indicator
Design Pattern 2 | Error and Feedback System
Design Pattern 3 | Explanation Tooltips
What it is
An expandable indicator integrated into the record detail page
Signals when AI has contributed to record information
How it works
Indicates which fields are AI-generated or influenced
Explains why AI was used and communicates relevant limitations
Provides access to further information and a way to report potential errors
Why
Users often lacked awareness of when AI is involved
Participants wanted clear disclosure to avoid confusion and misattribution of errors
UX value
Makes AI visible at the point where information is interpreted
Supports accurate mental models without overwhelming users
What it is
A structured feedback form embedded within the record detail page
Allows users to report potential errors in AI-generated content
How it works
Users select the field in question and describe the issue
Optional space allows users to suggest corrections or provide supporting sources
An ‘Under review’ label indicates that flagged information is awaiting human review
Why
Users wanted greater control and the ability to respond to uncertain or inaccurate information
Being able to flag errors was associated with increased confidence and reassurance
UX value
Gives users a clear way to question and respond to AI-generated information
Makes human oversight and accountability visible within the interface
What it is
Contextual information icons positioned alongside AI-generated or influenced information
Integrated into both search results and record detail pages
How it works
Explains why AI-generated information or search results are being shown
Provides context about how information was generated and what it was based on
Uses a high-level confidence indicator alongside a brief explanation
Why
Users preferred simple, contextual explanations of AI involvement
Additional information needed to be accessible without disrupting the research workflow
UX value
Provides relevant explanations at the point they are needed
Supports informed judgement without adding unnecessary interface clutter
Evaluating the Concept
Deliver
With the three design patterns developed as a proof of concept, I conducted usability testing to evaluate how effectively they supported researchers in recognising, understanding, and critically engaging with AI-generated information.
Understanding
Could users understand AI's role, limitations, and the explanations provided?
Critical Engagement
Users' ability to reflect on, question, or evaluate AI-generated information.
Think-Aloud Usability Testing
I conducted think-aloud usability testing with five participants, using scenario-based tasks that guided them through the search results and record detail pages. Participants were encouraged to verbalise their thoughts as they interacted with the prototype, allowing me to observe how they interpreted the design patterns and the information presented.
The evaluation was structured around four criteria:
Visibility & Clarity
The extent to which the interface made the presence of AI visible and clear to users.
Findings - Visibility & Understanding
Testing showed that participants generally recognised AI involvement, particularly once they reached the record detail page. However, the meaning of individual indicators wasn't always immediately clear. Some users recognised that elements were interactive without initially understanding that they specifically related to AI.
— “I knew I could click it, but I didn’t realise at first that it was about AI.”
Understanding developed through repeated interaction and consistent visual cues. Once participants had encountered the indicators and tooltips, they were generally able to explain what they represented and describe how AI had contributed to the information they were viewing.
Plain-language explanations were sufficient for most participants. Rather than wanting technical detail, users valued enough context to understand what AI had done, why it had been used, and where uncertainty might exist.
— “I don’t need loads of detail… just enough to understand what’s going on.”
Challenge Identified:
Some participants were unsure whether confidence scores represented an assessment made by the AI system or human judgement, highlighting the need to communicate their source and meaning more clearly.
— “I wasn’t sure if that score came from the system or from a person.”
Findings - Critical Engagement & Agency
Awareness of AI had a noticeable impact on how participants interpreted information. Knowing that AI was involved encouraged users to approach the content more cautiously and consider whether further verification might be needed, particularly when the information was important to their research.
— “If I know AI’s been involved, I’d probably look a bit closer at it.”
AI-generated summaries and metadata were generally viewed as useful starting points for orienting research or narrowing down results, rather than information participants would rely on without question. This suggested that making AI involvement visible could support more considered use without discouraging users from engaging with AI-generated information altogether.
The Error & Feedback System also supported participants' sense of control. Although most didn't expect to use it frequently, knowing they could report potential inaccuracies was reassuring, while the ‘Under review’ state helped make human oversight and accountability visible.
— “I might not use it all the time, but it’s reassuring that it’s there.”
Challenge Identified:
Participants were less certain about what would happen after submitting feedback, including whether they would see the outcome of a review or be informed if information changed.
— “Would this just go away, or would something else happen?”
Outcomes
The proof of concept demonstrated that lightweight explainable design patterns could support greater transparency without requiring users to understand the technical workings of AI. Participants developed practical understandings of AI's role, approached AI-generated information with greater consideration, and valued having ways to question potential inaccuracies.
Testing also highlighted opportunities for further development, particularly around communicating uncertainty more clearly and providing greater visibility into what happens after users submit feedback.
User Agency
The extent to which users could maintain an active role when interacting with AI-generated information.
Conclusion
Reflection
This project explored how explainable design could make the use of AI in digital archives more transparent and understandable. The research showed that users don't necessarily need technical explanations of AI, but they do need enough context to recognise its involvement, understand its limitations, and make informed judgements about the information they encounter.
The proof of concept demonstrated how lightweight, contextual design patterns could support these needs while fitting within existing archival research workflows. Testing also reinforced the importance of user agency and visible human oversight, particularly when AI-generated information may be uncertain or inaccurate.
Next Steps
Usability testing highlighted several opportunities to further develop and validate the design patterns. These would focus on:
Refining uncertainty communication - making the source and meaning of confidence indicators clearer.
Closing the feedback loop - showing users what happens after an error is reported and communicating the outcome of human review.
Testing in real archival contexts - evaluating the patterns with a larger and more diverse group of researchers using live archival collections.
Exploring scalability - investigating how the patterns could adapt to different archive platforms, AI processes, and types of AI-generated information.
These next steps reflect the areas of uncertainty identified during testing, particularly around confidence indicators and post-submission feedback.
This project challenged me to design for a problem where more information doesn't necessarily mean greater transparency. One of my biggest learnings was that explainability is as much about deciding what information to show, when to show it, and how users can act on it as it is about explaining how AI works.
Working with an emerging and technically complex subject also strengthened my ability to translate research into clear, practical design decisions. I learned to focus on the information users actually needed to understand AI’s role, helping them make informed decisions and maintain an active role in the experience.
Thanks for reading!
Interested in the research behind the project?
Read the full Master's thesis (PDF)