IIPC Web Archiving Conference 2026 report from the UK Web Archive
Reflections from UK Web Archive colleagues who attended the 2026 International Internet Preservation Consortium (IIPC) Web Archiving Conference in April 2026.
10 June 2026Reflections from UK Web Archive colleagues who attended the 2026 International Internet Preservation Consortium (IIPC) Web Archiving Conference in April 2026.
10 June 2026Blog series UK Web Archive
Author Helena Byrne, Curator of Web Archives

The official conference banner for the International Internet Preservation Consortium Web Archiving Conference and General Assembly 2026.
This year’s IIPC General Assembly and Web Archiving Conference took place at the Royal Library of Belgium (KBR) in Brussels from 20-23 April, 2026.
UK Web Archive colleagues from the Bodleian Libraries, British Library and the National Library of Scotland attended the Web Archiving Conference both as delegates and presenters. There was a packed programme with a variety of presentation forms and workshops which shared best practices and innovative projects from the world of web archiving. In this blog post the attendees report highlights of their conference experience.
Overall, this was a highly rewarding experience. The event was exceptionally well organised, sincere thanks to the organising team and Brussels proved to be a wonderful and culturally rich host city.
My focus during the conference was primarily on sessions covering technological progress, specifically tool development and the growing role of AI in web archiving. A particular highlight was the opportunity to speak directly with Jefferson Bailey, which gave me valuable
perspective I couldn't have gained remotely.
One unexpected but welcome thread throughout the conference was the intersection of science and the arts. Dr. Andrea Kocsis (National Library of Scotland) offered a compelling illustration of how these disciplines connect within the web archiving space, a reminder that this field draws on a broader range of expertise than is sometimes assumed.
The conference also prompted useful reflection on our own work. One concrete takeaway was recognising the need to integrate Jupyter Notebooks into our technology stack, something I hadn't prioritised before but now see as worth pursuing.
The closing keynote panel, ‘Web Archiving for Accountability: New Frontiers in Open Source Intelligence and Digital Evidence’, was an inspiring note to end on. It reinforced the broader significance of this field and left me genuinely motivated about where it is heading.
The IIPC WAC was hands down the most motivating and inspiring event of the year. I research web archives, and I feel that the ideas I gained in Brussels will fuel my productivity for another rotation. First of all, I was shocked by the keynote panel on the fragility of open source software in web archiving. I find it important to ring the bells that, without financial support for open-source infrastructure, our preservation efforts are at risk of collapsing. People often don’t consider that open-source development also requires labour and a team. For example, Browsertrix, the widely used, revolutionary crawling service, is maintained by a handful of people, including Tessa Walsh, whom I had the absolute pleasure to connect with at the conference.
Overall, this year was all about connections and collaboration for me: I exchanged ideas with David Mahoney, who developed the Wasteback Machine, which quantifies the completeness of archived pages, and we designed a philosophy of technology paper with Nora Ramsey, Assistant Web Archivist at BL. It was not the first time that I found myself sharing a panel with Jon Tonnesse of the National Library of Norway on the problems of usability and accessibility of web archives - therefore, we also decided to formalise our agreement as a collaboration.
In the name of accessibility, I presented Digital Ghosts, my research into creative approaches to web archives, and I was delighted by how much it resonated with the audience. After the talk, the Internet Archive, the National Library of France, and the Royal Danish Library all expressed interest in implementing my recommendations and in collaborating. These are just cherry-picked highlights from the event, and I have no room to mention all the incredible people I encountered, but I will cherish the moment of sunbathing with them on the secret rooftop terrace of the Royal Belgian Library while discussing the burning issues of web preservation.
The 2026 IIPC WAC and GA meeting was another invaluable opportunity to share challenges and breakthroughs we have encountered since the 2025 conference. This event was quite a busy one. I was involved in the pre-conference workshop GLAM Labs & Jupyter Notebooks. On day one of the conference I was co-chair of the Short Talks session with Sharon Healy. On the second day of the conference I co-presented with my colleagues Nicola Bingham and Nora Ramsey in the Collections as Data: Workflows & Use Cases and the Archiving Blogs & News sessions. The final day was the General Assembly. During the opening session I shared an update on the Celebrating Digital Innovation in Sporting Heritage Award 2025 that was given to the IIPC in recognition of the achievement of preserving the online heritage of Olympic events since 2010 and Paralympic events since 2012.I also presented an update to the Content Development Group on the progress of the 2026 Winter and Paralympics collection. Not only were the presentations incredibly helpful to navigate through shared problems but the biggest benefit were the corridor conversations with colleagues face to face and some of these conversations will turn into collaborations over the coming months.
The IIPC Web Archiving Conference offered a fascinating look at how rapidly the web archiving landscape is changing — and how archivists are constantly adapting. One of my biggest takeaways was the growing tension between open web preservation and increasingly hostile online environments. Several speakers highlighted how AI-driven bot detection, CAPTCHAs, pay-per-crawl models, and platform lock-downs are making large-scale web archiving far more difficult than even a few years ago.
I was particularly struck by discussions around social media archiving. Platforms like Instagram, TikTok, and X now require sophisticated browser-based approaches just to capture public content, with archivists balancing rate limits, logouts, and ever-changing site structures.
Another standout theme was sustainability. Many of the tools underpinning global web archives rely on surprisingly small teams and fragile funding models. Despite these challenges, the conference felt optimistic, with collaborative open-source development and AI-assisted discovery tools pointing toward an innovative future for preserving the web.
The 2026 IIPC Web Archiving Conference at the Royal Library of Belgium (KBR) may well be my new favourite conference. I came away inspired by the workshops and talks and full of ideas.
The GLAM Labs and Jupyter Notebooks workshop provided practical, hands-on, and relevant computational approaches for working with web archive datasets. It provided accessible and sustainable ways to query and extract web archive data from WARC files. Of particular interest to me were the Browsertrix sessions. As we use Browsertrix to collect social media for our election collections, I was keen to hear the latest updates from Webrecorder and see how others are approaching social media archiving. I got to connect with Anders Klindt Myrvoll (Royal Danish Library) on uniqueBrowsertrix workflows to archive particularly difficult sites. I was impressed by SolrWayback, an open source tool for browsing and exploring harvested ARC/WARC files, and am interested in how we can adjust to use Browsertrix's WACZ files with the application. As my background is in open access repositories, I also had a great chat with Kathryn Cassidy (Digital Repository of Ireland) on incorporating web archive data in repositories.
As co-chair of the IIPC Training Working Group alongside Claire Newing (National Archives) and Camille Lawrence (Webrecorder), we reviewed the consortium's web archiving training materials, which aim to provide a complete beginner's level training course in web archiving. I was happy to see how we can update our training to increase engagement with web archiving groups and ensure the materials continue to serve the community well.
A strong sense of direction was given on the first day, in the Jupyter Notebooks for Digitized and Born-Digital Collections workshop. Ten workbooks allowed us to progressively query .warc files and datasets, including sets prepared by the National Library of Scotland. We used these to extract and reuse text, and try retrieval augmented generation on the digital collection, A Medical History of British India.
Andrea Kocsis presented Lessons from the Digital Ghosts exhibition, about a real highlight for NLS last year, an exhibition of metadata, mainly recorded by curator Trevor Thomson as he collected the web in Scotland for preservation. It explored ideas of time, movement and loss while attempting to come to grips with the sheer scale of the publishing environment. Dr Kocsis pointed out that a lack of awareness is the most important block to accessing web archives, but also that creativity is a great way to open our collections to the public.
I was happy to see the Memento protocol return in David Mahoney’s presentation. It focused on means of advancing sustainable web practices, like the Wasteback machine, which estimates emissions of websites over time. They find that since 2010, average desktop pages have grown by 523.2% and mobile pages by 1700%.
Lastly, I am grateful to our hosts, KBR, who gave us tours of their building. It retains its original furniture – including card catalogues with neat note-making shelves. I am hopeful that we can develop such an intuitive research environment for readers of web material.
Another excellent IIPC Web Archiving Conference. The Royal Library of Belgium (KBR) provided a wonderful setting, and the thoughtfulness and generosity of the hosts and IIPC colleagues created a welcoming and collaborative atmosphere throughout the week.
The conference theme of sustainability felt especially timely and relevant. Across many sessions there was a strong recognition that web archiving now faces significant additional pressures, not only in terms of sustainable technical infrastructure and resourcing, but also around growing barriers to our harvesting of web content as website owners tighten access controls and implement protections against automated bots and crawlers. It was valuable to hear how institutions across the international community are responding to these shared challenges.
The social media archiving sessions were especially relevant to our current work at the Bodleian, as we continue exploring how to develop our collecting in this area. It is inspiring to see what can be done with those large-scale social media datasets acquired over several years before the landscape became considerably more restrictive. In particular, the work presented by Institut National de L’Audiovisuel (INA) on a discovery and access interface for their Twitter corpus (which includes 200M photos and 25M videos as well as 1.2 billion tweets) was extremely impressive, demonstrating the research potential of these collections if we can but secure their preservation.
Another great conference put together by the IIPC, and this time with the Royal Library of Belgium in Brussels!
The first session of the conference focussed on AI and I had expected that many following sessions would include AI, so I was surprised how little came up. AI is still a fast moving, rapidly changing technology, so it is difficult to foresee its applicability and real world opportunities in web archiving. At the British Library, we are still working on understanding the risks and opportunities AI might bring, so maybe this is a common position across the community at the moment.
There were many interesting and informative presentations, but the last day’s discussions about tools support is worth noting here. The IIPC WA community has recognised over recent years that the support for WA software is very reliant on individual institutions, or even individual people. This is a difficult and risky situation as it puts too much strain on them. Therefore, the IIPC community has established a Technical Committee to attempt to address this concern. At present, the remit of this group is in development, but there is an ambition and intention to make progress with ensuring tools support exists, initially for the most important pieces of web archiving software this year.
For more details about IIPC activities visit the IIPC website and the IIPC blog.

This blog post is part of the UK Web Archive series. The UK Web Archive was established in 2004 to collect, make accessible and preserve web resources of scholarly and cultural importance from the UK domain.
The collection is selective, built on nominations from subject specialists and other external experts. The British Library prioritises websites that:

You can access millions of collection items for free. Including books, newspapers, maps, sound recordings, photographs, patents and stamps.