OpenWebSearch.eu and LUMI AI Factory: powering Dataset-as-a-Service with a European Open Web Index | LUMI AI Factory

The Open Web Index (OWI) keeps on giving well beyond the official project timeline. Thanks to a collaboration with LUMI AI Factory, researchers and innovators can now profit from Open Web Index data when experimenting with AI applications. The LUMI AI team recently posted a dedicated blog post with details on the collaboration. LUMI AI Factory’s Dataset-as-a-Service (LUMI AIF DaaS) packages curated, high-value datasets together with HPC resources, tooling and operational support. This allows researchers and innovators to use web-scale data without building and operating crawling/indexing pipelines first. As part of that offering, the OWS datasets are made available inside the LUMI environment — ingested, formatted and connected to the compute stack so users can immediately start running experiments.

Read the full article here: https://lumi-ai-factory.eu/articles/openwebsearch/

Remaining active: We provided feedback for various European Commission consultations

After the successful completion of the OpenWebSearch.EU project in April this year we are continuing the operation, maintanance and further development of the Open Web Index (OWI) and a wider European Web Data Infrastructure under the Open Web Search (OWS) initiative.
The initiative unites most of the former project partners and brings in further parties from the research and tech domain. The goal is not only to continue with various research projects surrounding the OWI, but to further promote the need for a public-private web data infrastructure for Europe. With this mission in mind, we have provided feedback to various EC consultations in the past months.

Spreading the word in Europe

The OWS initiative supported by Partner Project PriDI (Privacy Enhanced Digital Infrastructures) actively contributed to various Public Consultations conducted by the European Commission.

Public Consultations enable citizens and businesses to share their views on new EU policies and existing laws. The Commission analyses the feedback and contributions received and takes them into account for fine tuning its initiatives. Providing feedback in the Commission’s consultations is a chance to actively shape EU policies and strategic European goals. It furthermore strengthens the visibility of the OWS Initiative and its aim for developing an open European Web Data Infrastructure on European level.

So far, in 2026 we provided feedback in the following consultations:

  • In January 2026 we participated in the consultation concerning the so-called Digital Decade Policy Programme.

The Digital Decade policy programme is commonly agreed by the European Commission, Parliament and Council and serves as a digital strategy. The Programme outlines a vision for Europe’s digital transformation, sets out concrete targets and objectives and provides a governance framework and mechanisms for collaboration, both between the Commission and Member States and among Member States.

In this consultation the OWS Initiative calls for building a European Web Data Infrastructure that provides start-ups, SMEs and industry with large-scale access to high-quality web data. We describe, why the European Web Data Infrastructure is an inevitable resource for a neutral and open internet and how it can help to ensure freedom of choice.

You can find the contribution under this link: https://ec.europa.eu/info/law/better-regulation/have-your-say/initiatives/15514-Review-of-the-Digital-Decade-policy-programme/F33363480_en

  • In March 2026 the European Commission conducted a so-called Digital Fitness Check.

With the Digital Fitness Check, the European Commission aims to elaborate how the digital rulebook affects and supports the competitiveness objective of the EU. The goal is to assess how different laws work together, identifying synergies, good practices, and any remaining gaps, overlaps and inconsistencies. Its aim is to ensure that the overall system remains effective, proportionate and fit for the future, and delivers on the EU’s high standard of protection of fundamental rights.

In this consultation we call for clarifying European copyright and data protection law as well as liability aspects. Europe must enable European digital companies to build competitive digital services themselves on the basis of truly sovereign large-scale access to high-quality Web data independent from overseas services.

The contribution is available under this link: https://ec.europa.eu/info/law/better-regulation/have-your-say/initiatives/15554-Digital-fitness-check-testing-the-cumulative-impact-of-the-EUs-digital-rules/F33376983_en

  • In June 2026 the OWS Initiative provided feedback in a consultation for a better copyright environment for European creativity and innovation.

In 2026 the European Commission is due to review the Directive 2019/790 EC on Copyright in the Digital Single Market. For this purpose the Commission inter alia intends to assess the need to further improve the copyright framework to address new challenges raised by market and technology developments.

Our contribution in this consultation reflected our experiences and findings from the OWS project and the PriDI project with regard to copyright. Complex copyright and database laws give rise to liability risks related to building, redistributing and using the European Web Data Infrastructure, especially in a commercial context. Data collected through crawling the web inevitably includes copyrighted material. We therefore call to resolve the legal uncertainty surrounding the crawling, processing, indexing, and redistribution of publicly accessible Web content for the development and operation of independent European Web Data Infrastructures.

You can access our contribution via this link: https://ec.europa.eu/info/law/better-regulation/have-your-say/initiatives/18173-Targeted-initiative-for-a-better-copyright-environment-for-European-creativity-and-innovation/F33504259_en

  • In July 2026 the Commission opened a consultation related to European Partnerships to be implemented as joint undertakings.

The goal is to strengthen European Partnerships by streamlining implementation modes and using a portfolio approach to identify a limited set of coherent partnerships, with objectives closely aligned with the European Competitiveness Fund.

The OWS Initiative is working intensively to bring the European Web Data Infrastructure into the European Competitiveness Fund. We therefore used this consultation for calling to strengthen the EUs digital leadership objectives through establishing a European Web Data Infrastructure JU as well as dedicated base funding in the upcoming Multiannual Financial Framework (MFF) / European Competitiveness Fund (ECF).

The full contribution is available under this link: https://ec.europa.eu/info/law/better-regulation/have-your-say/initiatives/16413-Horizon-Europe-European-Partnerships-to-be-implemented-as-joint-undertakings/F33508399_en

Influencing and shaping European policies and existing laws remains one of our key topics and we continue looking for possibilities to spread the word and highlight the need for a European Web Data Infrastructure. For enhancing the topic’s visibility and weight, we call our community to support us by also contributing to relevant EC consultations referencing the OWS project and the existing European Web Data Infrastructure.

Let’s reclaim the web together!

OWS x PriDI goes EBDVF

European Big Data Value Forum 2026 takes place in Galway on 30 September and 1 October, bringing together industry professionals, business developers, researchers and policy-makers from all over Europe and other regions of the world to advance policy actions and industrial and research activities in the areas of Data and AI. 

The Open Web Index (OWI) was developed by the OpenWebSearch.eu (OWS) team between the years 2022 and 2026 as part of a Horizon Europe funded research project. It spans many languages and domains and facilitates access to web material that researchers need to build, evaluate and improve information retrieval and large-scale AI systems.

PriDI is a project that started in parallel to the creation of the OWI. It is focussed on assessing the legal requirements for privacy-enhancing digital infrastructures. The project team supports the further development of an open European web search infrastructure at the interface between law, society and business informatics. The aim is to anchor fundamental values such as privacy and data protection in the sense of „values by design“ in an open European web index that is currently under development.

The joint OWS/PriDI team looks forward to welcoming interested parties at at EBDVF end of September. 

See you in Galway this autumn?

Register for #ossym 2026 now!

The 8th International Open Search Symposium (#ossym) brings together experts and practitioners from across computer science, law and regulation, ethics, economics, and civil society.  .

#ossym2026 will take place as a hybrid event in Berlin from 7–9 October 2026, hosted at CODE University of Applied Sciences and supported by the German Aerospace Center (DLR). This year’s conference theme is “Boosting Digital Sovereignty.”
Join in for inspiring talks, thoughtful discussions, and the latest research on open and distributed web search, open web data infrastructure, and their role in AI.

Program focus: open web search + AI in the context of sovereignty

A special focus lies on how an open web data infrastructure can potentially counter the current concentration of web services, value loss, and regulatory pressure. Key questions include:

  • How can open web search and web data infrastructure strengthen digital diversity?
  • Which technical and economic concepts improve acceptance and resilience?
  • Which ethical and legal questions must be addressed to make this approach work?

Keynotes

Confirmed* keynote speakers include:

  • Stéphane Beemelmans (DLR)
  • Peter Gehle (DLR)
  • Julian Kunkel (GWDG / University of Göttingen)
  • Aline Blankertz (Structural Integrity)
  • Stefan Mesken (DeepL)
  • Hannes Taubenböck (DLR)

Tracks and sessions

#ossym2026 features a hybrid format with scientific talks, poster sessions, panels, demonstrations, and informal discussion spaces. The conference includes science tracks on:

  • Crawling
  • Retrieval
  • Spatial modeling
  • AI in Science
  • Web data driven applications

And sessions around:

  • Ethics + Society
  • Legal aspects of open web search & AI training
  • User behavior and preferences in search and AI

Workshops and community formats

Workshops and community activities include (among others):

  • Evaluating design patterns for legally compliant Open Web Index and Search Engine systems
  • Disinformation tools
  • Open Web Data Working Group of EuroHPC
  • Student project session / index access tutorial
  • Additional workshop opportunities (details may be updated on the website)

Side Event at Weizenbaum Institut

This year, a special  panel discussion (in German language) & networking dinner, organized by Open Search Foundation will take place in collaboration with Weizenbaum Institut. The theme of the evening being: “Souveränität im Netz – Wie eine europäische Webdateninfrastruktur Freiheit, Wirtschaft und Demokratie stärkt”
Find more details here.

Conference chairs and venue

The symposium is hosted at CODE University of Applied Sciences, Berlin.
The conference chairs are Prof. Dr. Michael Granitzer, Prof. Dr. Christian Guetl, Dr. Megi Sharikadze, Dr. Stefan Voigt, and Dr. Andreas Wagner.

Registration

Registration for onsite and online participation is open and free of charge. Register here.
The conference programm* is available here.

*The conference programm and keynote speaker announcements are published under reservation of the right to make alterations

42 months in the making: OpenWebSearch.EU project end results look promising!

While the project officially ended on 28 February 2026 with very positive results, the journey continues under the European Open Web Search initiative. But first things first…

How it started and where we are at

In September 2022 a consortium of 14 European organizations*, spread across 7 countries, have joined into the Horizon Europe funded project OpenWebSearch.EU, with the goal to conceptualize, develop, test and evaluate a first ever publicly federated European Open Web Index. 

A Web index per se is a necessary fundament for any search engine (including LLM based agentic search applications) to draw information from. Without a well-structured Web index that processes and stores up-to-date crawls, thereby mapping out the World Wide Web, search practically cannot operate. An “open” web index refers to the standards of open source technologies, meaning that the index is accessible publicly and transparent in its nature.

Why that matters?

Europe, with its longstanding and multi-cultural history, has not yet managed to create a sovereign European web data infrastructure that is inclusive of the socio-cultural and linguistic diversity that makes up the European identity. Instead, accessing information online is still largely dependent on overseas Big Tech services that do not fully represent the digital needs and views of European societies. In addition, these services, as by their commercial nature, are not free of biases, e.i. in political, cultural and commercial regards.

But why has Europe not yet created a web index of its own then?

There have been some commercial companies that have made attempts at creating indices in the past and present, and some of them have had a certain level of success. But maintaining an index is a costly endeavor. The enormous need for computing, storage and maintenance requires solid funding.

Another component is the lack of a true pan-European ambition. Solutions on national levels might work on a certain level, but they so far do not suffice to cater to pan-European needs. Europe, by its history, is not one homogenous continent, but a collective of individual countries, languages and cultures that should benefit from a European Web index equally. 

Birthing a truly European search infrastructure, that meets people’s expectations, will ideally be curated collaboratively and across country borders. This is where the OpenWebSearch.EU project is unique in its federated design. Additionally, ethical aspects are included by design. This means to foster inclusion, diversity, transparency, while honoring copyrights and guaranteeing privacy protection.

So, what project objectives and milestones have been reached?

The biggest milestone by far and a key project result has been the successful launch of the Open Web Index (OWI) prototype in May 2025.

The index has been developed by teams from five technical workpackages, including team members from the University of Passau, CERN, German Aerospace Center, A1 Slovenia, Radboud University,  Graz University of Technology and Webis Group. The tech-related infrastructure and governance tasks were managed by teams from LRZ, German Aerospace Center, CERN, IT4I and CSC – IT Center for Science. 

The OWI can be accessed publicly via the OWI dashboard, which allows users to download index shards under a dedicated research license, conduct an ethical self-assessment and get regular index statistics updates: https://openwebindex.eu

Users can download index data via the Lexis Platform (provided by IT4I) or the OWI command line tool OWILIX, both of which are offered via the dashboard. 

Curious, how index data can be used?

Various Third-Party Partner projects give insights into practical applications and future ideas. Once the OWI was ready for early adoption trials, nine technical projects, two infrastructure projects, one economic and three legal projects were selected and financially supported via three distinct OpenWebSearch.EU open calls. 

The results of these projects can be found here: https://openwebsearch.eu/third-party-projects/

Some highlights include the successful implementation of an argument search engine to assess trustworthy health data with fairness-conscious ranking, as well as an innovative LLM based crawling method to outperform traditional PageRank heuristics.
Yet another interesting project is the vertical search engine Nooon, that gives access to disability related data and makes results sharable in trustworthy ways.

How the OWI can boost European economy

To back up the immense Marco-economic potential of an Open Web Index for Europe, the 2024 „Market potential assessment of OpenWebSearch.eu“ study was conducted by third-party partner MRC. 

The project team analyzed the economic and societal impact of an open Web index, using both top-down and bottom-up methods to ensure a comprehensive analysis of different scenarios, including qualitative feedback from potential future users.

Key findings indicate that the OWI could achieve a return on investment within four years of operation, with a projected net benefit of around €4.5 billion over a decade. These benefits are believed to derive from economic gains and societal improvements such as strengthening European digital sovereignty and global techno­logical competitiveness across a wide range of industries and use cases. Find the full study for download here: https://openwebsearch.eu/the-project/research-results/market-potential-assessment-of-an-european-open-web-index/

Some lessons learned

Despite many successful results, some challenges have been present as well.
The current scope of the OWI is not yet able to compete with commercial search indexes, such as Google or Microsoft Bing. And while directly competing with these commercial players has never been the idea, scaling up accordingly is a necessity to make the Open Web Index a tool for digital sovereignty.

The OpenWebSearch.EU team is still working to launch the first search API and operational search frontend for the index and legal questions regarding commercial licenses and EU regulatory work still need to be solved for commercial use.

However, the project illustrates and gives evidence that an open Web index on a federated public-private infrastructure is more than doable. This approach adds to sovereignty in web search, AI and web analytics and can be scaled up, if financial backing is provided and if more infrastructure partners are willing to join. 

What’s next?

As mentioned in the beginning, the end of the project marks the start the next ventures in the exciting journey of the Open Web Search initiative. 

A majority of the project partners are currently working on continuation set-ups. The initiative is largely coordinated by the Open Search Foundation with the University of Passau in the technical driver’s seat. The goal is to keep growing, gain reliable financial and legal support from political stakeholders in Brussels and European Member States and to further grow the community of developers and users.

An important part of the OpenWebSearch.EU project were the community building activities as well as communication and dissemination of research results. The community will be kept up to date via the Mattermost community platform and via OpenWebSearch.eu and OpenSearchFOundation.org Newsletters.

Putting people first was and is the ultimate goal. 

While the idea of „European Open Web Search and trustworthy AI as a public good“ is still the mission, it takes active support and engagement from the public to guarantee success.

Is Europe ready to become sovereign in the Web? Now, that that a proof of concept is here, will the Europe take the next step?

It is up to European citizens, politicians, and decision makers to make the call now! A big thanks to everyone who has supported the OpenWebSearch.EU project in the past 3,5 years! Many thanks  to Horizon Europe and Next Generation Internet! And a special thanks to the entire OpenWebSearch.EU teams for walking this exciting walk together! 

ows.eu consortium meeting at Ostrava (IT4I)

*University of Passau (Germany), Leibniz Supercomputing Centre (Germany), Radboud University (The Netherlands), Leipzig University (Germany), Graz University of Technology (Austria) German Aerospace Center (Germany), IT4I (Czech Republic), CERN (Switzerland), Open Search Foundation (Germany), A1 Slovenija (Slovenia), CSC – IT Center for Science (Finland), nl Net Foundation (The Netherlands), Suma-ev (Germany), Webis Group (Germany)

“Wo bleibt das europäische Google oder Facebook?” | arte.tv

Our project was recently featured in a German arte.tv report about European alternatives to overseas BigTech web services. The video highlights our commitment to strengthening European digital sovereignty in the world wide web.
The report provides insights from Prof. Dr. Ir. Djoerd Hiemstra, Professor of Federated Search and Head of the Information Retrieval research group at Radboud University, one of our consortium partners. Djoerd introduced the Open Web Index in its current state and the role it could play in creating powerful European search solutions.

Skip to minute 4:16 to hear Djoerd‘s insights:

Alternatively, watch the video directly on Arte.tv: https://www.arte.tv/de/videos/121620-127-A/wo-bleibt-das-europaeische-google-oder-facebook/

OWS.EU Partner in Focus: CERN

CERN, the European Organization for Nuclear Research, is one of the world’s largest and most respected centres for scientific research. Its business is fundamental physics, finding out what the Universe is made of and how it works. Within the OpenWebSearch.EU (OWS) project CERN plays a crucial role not only with regards to supercomputing infrastructure, but also via its contributions to ethical and legal assessments, as well as project management and communications support. 

The CERN project team is led on by Andreas Wagner, IT Solutions Architect and complemented by Noor Afshan Fathima with whom we spoke in her role as Data Infrastructure Engineer, about the project progress thus far.

Noor Afshan Fathima, CERN, Section: IT-PW-WA, Research Fellow

Please describe your organisation’s tasks in the project. What is your field of expertise that you bring to the project?

CERN contributes across 6 workpackages – WP1 (Fill in Crawlers), WP4 (Search Applications), WP5 (Federated Data Infrastructures), and WP6 (Ethical, Legal, and Societal Aspects) – bringing expertise in infrastructure engineering, development of science search applications, ELSA considerations, and governance of the federated open search infrastructure. It also contributes to WP7 – Dissemination and Communication, WP8 – Project Management.

In WP1, we developed two purpose-built authenticated web crawlers for CERN’s internal web estate: cern-owler (Java/Playwright for HTML, producing 125 WARC archives totalling 3.1 GB) and owler_auth_pdf (Python/Tika for PDFs, extracting 2,211 documents from 90+ domains). Together they delivered 3.3 GB of content across 287 files to AccGPT (Accelerator GPT), CERN’s experimental AI-powered chatbot for the accelerator complex, covering 25,292 seed URLs across 182 CERN domains. We also participate in project-level coordination and contribute to the governance of the federated open search infrastructure, drawing on CERN’s institutional experience in managing large-scale, multi-partner scientific collaborations.

In WP4, we developed two complementary POC search applications demonstrating MOSAIC’s flexibility. The first is an institutional search engine built from custom-crawled WARC archives of CERN’s public web content, fed through the full OWS preprocessing pipeline (resilipipe, open-web-indexer, lucene-ciff) into MOSAIC, indexing 4,352 documents from 6,738 crawled pages. The second is Nooon, a vertical search engine for disability-related knowledge, built from a 2.59-million-document OWI (Open Web Index) slice extracted via the command line interface OWILIX. Nooon is designed to support HR and Diversity & Inclusion offices in fair hiring and inclusive policy development.

In WP5, we operate and document the production server fleet that underpins the Open Web Index at CERN. This includes the URL Frontier coordination service — where we drove the migration from OpenSearch to ScyllaDB to handle 94.7 million operations per day across 6.68 billion URL records — the web crawling infrastructure processing up to 3 TB of content per day, an iRODS data federation spanning four sites across five European data centers (CERN, LRZ, DLR, CSC Finland), load balancers, metrics collection, and the application hosting servers. Our systematic documentation methodology, developed specifically for this project, covers discovery, deep-dive analysis, checklist completion, and academic chapter creation for each server.

In WP6, we contribute to the ethical, legal, and societal dimensions of the project. This includes work on ELSA (Ethical, Legal, and Societal Aspects) as they apply to open web search — particularly around privacy-preserving information retrieval for vulnerable populations, knowledge sovereignty, and the responsible handling of disability-related data. Our OSSYM 2025 publication and CERN preprint on empirical ethics in disability information retrieval directly address these concerns. We also contributed to the governance of the federated data infrastructure.

In WP7 (Dissemination and Communication), we have contributed to raising the visibility of the OpenWebSearch.EU project through major CERN communication channels. Three feature articles were published on home.cern and in the CERN Courier: “A European project to make web search more open and ethical” and “Ethical, open and non-commercial: Open Web Search project designed to provide Europe with an alternative” on the CERN news site, and “Towards an unbiased digital world” in the CERN Courier. These articles reached CERN’s global audience of researchers, engineers, and policy-makers, highlighting both the technical infrastructure and the ethical dimensions of building a European open web index. Beyond written dissemination, we have presented the project at multiple international venues including OSSYM 2024 and 2025, CS3 2025, EGI 2024, and the Cambridge Forum on AI, contributing to community building around open search infrastructure.

How is the project progressing? Which major milestones did you achieve?

The project is progressing well, with all CERN-side deliverables on track. Our major milestones include:

URL Frontier evolution: We completed the migration of the URL coordination service from OpenSearch to ScyllaDB, resolving critical performance bottlenecks caused by JVM garbage collection pauses and write-heavy workloads (99.88% writes). The production ScyllaDB deployment now handles 24.3 billion total operations with zero failures and continuous uptime, storing 5.07 TB across the crawl state database.

Authenticated crawling and AccGPT delivery: We developed two purpose-built crawlers — cern-owler (Java/Playwright for HTML) and owler_auth_pdf (Python/Tika for PDFs) — capable of navigating CERN’s Keycloak SSO. Together they delivered 287 files totalling 3.3 GB to AccGPT’s S3-based knowledge base, covering 25,292 seed URLs across 182 CERN domains.

Search application deployments: The institutional search engine indexes 4,352 documents from 6,738 crawled pages through the complete OWS pipeline. Nooon serves 2.59 million disability-focused documents through MOSAIC, demonstrating the OWI-to-vertical-search workflow at scale.

iRODS data federation: We established the CERN node in a five-site iRODS federation (CERN ↔ LRZ ↔ DLR ↔ IT4I ↔ CSC Finland), which in future enabling cross-institutional data sharing for the Open Web Index across three European countries.

Infrastructure documentation: We completed comprehensive documentation for 8+ production servers using our systematic five-phase methodology, producing deliverable-ready chapters for D5.3 covering the full infrastructure stack from load balancers to database clusters.

Publications:CERN’s work in the project has produced a substantial publication record. As first author, six papers span disability information retrieval and infrastructure architecture: two at OSSYM 2025 (Knowledge Sovereignty in Disability IR; Architecting the URL Frontier datastore), two at OSSYM 2024 (Federated Data Infrastructure for the Open Web Search; Architecting the OpenSearch service at CERN), one accepted at the Cambridge Forum Journal on AI: Culture and Society (empirical ethics, article in progress), and one submitted to SEASON — the Search Engines and Society Network (Ethical Privacy in Disability Data Retrieval). As co-author, contributions include the Springer book chapter on the Open Web Index (2024), the JASIST journal article on Open Web Index impact (2023), federated infrastructure papers at CS3 2025 and EGI 2024/2025, plus two Zenodo deliverables (Pilot Infrastructure Launch; Training Material for Partners). A companion preprint is deposited at CERN’s document server (CERN-OPEN-2025-004). Three feature articles were published on home.cern and in the CERN Courier as part of WP7 dissemination.

What are the challenges you have been facing (regarding your tasks)?

Authenticated crawling at institutional scale. CERN’s web estate sits behind a Keycloak Single Sign-On layer that conventional crawlers cannot penetrate. Building browser-based authentication into the crawling pipeline — using Playwright to handle OAuth2 flows, session tokens, and cookie management at scale — required significant engineering effort and careful coordination with various CERN’s teams.

Network access complexity. CERN’s network security model requires two-hop SSH access (desktop → lxplus gateway → target server) with different authentication patterns per server type. Communication between servers and workstations require staging through intermediate nodes, which was a very interesting challenge to work on.

Which milestones do you plan to achieve in the remaining months?

In the remaining project period, we are focusing on completing and polishing our deliverable contributions and extending the search applications:

D5.3 completion: Finalise the remaining server documentation chapters and integrate all CERN infrastructure sections into the consolidated deliverable, including the URL Frontier evolution narrative, crawler infrastructure, and federation topology.

D4.4 integration: Complete the CERN search applications section with final evaluation results, upload the Zenodo reproducibility artifact package, and integrate figures and cross-references into the consolidated document.

Full estate crawling: Extend the institutional search from the current 6,738-page public subset to CERN’s full 25,000+ page web estate, integrating authenticated content into the MOSAIC index with appropriate access controls.

Nooon enhancements: Implement topic-level Curlie filtering for semantic corpus construction beyond keyword matching, and explore cross-corpus comparison capabilities (e.g., Disability in Employment vs. Disability in Education) tailored to HR and D&I workflows.

Frontier integration: Connect the authenticated crawlers to CERN’s URL Frontier infrastructure for continuous, scheduled crawling rather than the current manual campaign-based approach.

What makes the OWS project special to you?

The OWS project represents something genuinely rare: the attempt to build a public, European alternative to the commercial search infrastructure that shapes how billions of people access information. Working on this at CERN feels especially fitting — the web was born here, and now we are contributing to ensuring it remains open and searchable by everyone, not just by those who can afford to build their own index.

What makes it personally meaningful is the Nooon component. Building a search engine specifically for disability-related knowledge – one that surfaces voices and resources that mainstream search systematically underrepresents – connects the project’s technical ambitions to real human outcomes. When an HR professional can discover evidence-based accommodation guidelines or a disability advocate can find peer-reviewed employment research through an open, privacy-preserving infrastructure, that is the kind of impact that motivates the work.

The project also demonstrates that European research institutions can collaborate on infrastructure at scale. The iRODS federation across five sites in three countries, the shared URL Frontier coordinating billions of URLs, the OWILIX tooling that lets anyone extract a thematic slice of the web – these are building blocks for digital sovereignty that go beyond any single institution’s capability.

Do you already have plans for the time after the project ends?

Yes, several strands of work are designed to continue beyond the project timeline:

Nooon and fair hiring: Nooon is supported through the Open Search Foundation’s Ethics working group and CERN’s Disability Network within the Diversity and Inclusion programme. There is active interest from HR departments in exploring fair hiring tools built on open search infrastructure. We plan to extend Nooon to multilingual and multimodal corpora incorporating lived-experience contributions from disabled people and caregivers, with client-side preprocessing to protect sensitive employment data.

AccGPT integration: The authenticated crawling infrastructure continues to support AccGPT’s knowledge base requirements independently of the OWS project. The 3.3 GB already delivered serves as the foundation, with plans to extend coverage to CERN’s full web estate and establish continuous crawl schedules.

Infrastructure sustainability: The MOSAIC deployments on open-science-search serve as reference implementations for institutional search at CERN. The documented infrastructure and the systematic methodology we developed for server documentation provide templates that other institutions can adapt for their own open search deployments.

Open science artifacts: All reproducibility artifacts – seed inventories, crawler source code, pipeline outputs, and evaluation data – will be publicly available on Zenodo, ensuring that our contributions remain accessible and reproducible for the broader research community working on open web search.

Thank you for the interview!

Read more about CERN: https://home.cern/

Watch our interview with Noor about the search engine Nooon:

From shop counter to online catalogue: Inside the DTCommerce project

A Slovenian team set out to build open-source tools that help small retailers go digital easily, by importing product descriptions from a spreadsheet into an online shop – with AI-enhanced descriptions and images, in just a few clicks

For small to medium sized Brick-and-Mortar retailers, the move from physical shops to e-commerce is a long and cost intensive process. These businesses typically have an accounting system with a list of products, perhaps a supplier’s website with technical specifications, and neither the time nor the budget to manually write product descriptions, source images, and populate an online shop for hundreds or thousands of items. The result is that many small retailers either delay their digital transition or end up with online catalogues that are sparse, poorly described, and unappealing to customers.

The main challenge is not a lack of products but a lack of digital product content. A physical shop’s inventory usually exists as a list of names, SKU (stock keeping units) codes, and prices in accounting software. An online shop needs other specifications : well-crafted product descriptions, high-quality images, metadata, and engaging presentations. Creating this content  quickly becomes a substantial undertaking.

The DTCommerce project, carried out by the Slovenian company ZenLab under the European OpenWebSearch.eu project funding, set out to solve this problem with an automated solution. The idea is simple: take the product list a shop already has, find the corresponding product information on the web, enhance the descriptions using AI, and deliver the result as a ready-to-use online shop – with minimal manual effort.

The Approach: Automated Extraction and AI Enhancement

The DTCommerce system operates in two stages. The first stage is a web crawling process that, given a list of product URLs from supplier or manufacturer websites, automatically extracts the key product information, such as: title, description, imagery, price, and technical specifications. The crawler is built on Scrapy, a well-established open-source web scraping framework, and includes support for structured data formats (JSON-LD) as well as domain-specific extractors for particular target sites.

The second stage is where AI comes in. The raw product descriptions extracted from supplier websites are often technical, dry, and written for a trade audience rather than end consumers. DTCommerce feeds these descriptions to an AI language model (Perplexity AI’s sonar-pro), which rephrases them into clearer, more engaging copy while preserving every technical detail – dimensions, model numbers, and specifications. The original description is retained alongside the enhanced version, so nothing is lost. The result is a set of enriched product records in a standardised format, ready to be imported into an e-commerce system.

From Pipeline to Plugin: A Few Clicks to a Full Shop

To make the pipeline usable for non-technical shop owners, the ZenLab team built a WordPress/WooCommerce plugin that wraps the entire workflow into a simple administrative interface. The process works as follows: the shop owner exports a product list from their accounting software as an Excel file and uploads it to the plugin. The plugin creates basic product entries in WooCommerce, sends them to the enrichment service, and automatically populates each product page with enhanced descriptions and images – all without requiring the shop owner to edit a single product manually.

An Honest Detour: When the Open Web Index Didn’t Have What Was Needed

DTCommerce was originally designed to use the Open Web Index (OWI) as its primary data source for finding product information across the web. The vision was that a shop owner could provide a product name or SKU code, and the system would search the OWI to find matching products on supplier and manufacturer websites, automatically retrieving descriptions and images.

In practice, the specific e-commerce sites that the project’s use cases required were not present in the OWI at the time of development. This is not surprising: the OWI is still being built, and its coverage of niche commercial sites – particularly smaller B2B suppliers – is not yet comprehensive. The team adapted by switching to direct web scraping of predefined supplier URLs, which allowed the project to deliver its core functionality on schedule.

Why the project matters

DTCommerce addresses a real and widespread problem. Across Europe, millions of small retailers face pressure to establish an online presence but lack the resources to do so effectively. By automating the most labour-intensive part of the process – creating digital product content – the project lowers the barrier to entry in a meaningful way. The fact that the tools are open source and built on widely used platforms (WordPress, WooCommerce, Scrapy) means they are accessible to a broad audience and can be adapted to different markets and product domains.

The project also illustrates a type of application that open web search infrastructure is well suited to support. The ability to search an open web index for product information – matching a local shop’s inventory against the broader web – is precisely the kind of use case that depends on open, non-proprietary access to web data. As the OWI matures, tools like DTCommerce stand to benefit directly. The project overall also demonstrates both the potential of the OWI-based approach and its current practical limits.

Final Outlook

The DTCommerce activities will follow with further development of tools compatible also with other e-commerce platforms. The tool will remain as open source, the company will be developing and automated portal for data exchange and enrichment, available on demand for various e-commerce integrations.

Find the full project report here: https://zenodo.org/records/18300935

The DT Commerce project was funded under the OpenWebSearch.eu initiative (Horizon Europe, Grant Agreement 101070014, Call #2).

 

The case for Neural Crawling: Inside the FUN project

A research team from Pisa and Glasgow proposes that AI language models should decide which web pages to download – and shows why this matters for the future of search

Before a search engine can find anything, it must first build a collection of web pages to search through. This collection is assembled by a crawler – a piece of software that systematically visits web pages, follows links, and downloads content. The decisions the crawler makes about which pages to prioritise determine, in a very direct way, what the search engine will eventually be able to find.

For over two decades, the dominant approach to crawling prioritisation has been PageRank and related link-analysis methods: pages that are linked to by many other important pages are assumed to be important themselves. This was a reasonable assumption in the era of keyword search. But search is changing. Users increasingly ask questions in natural language rather than typing keywords, and automated systems like retrieval-augmented generation (RAG) pipelines issue their own queries to search engines. These new kinds of queries demand pages with rich, coherent, meaningful content – and there is no guarantee that such pages are also the most popular or the most linked-to.

The FUN project – Focused Neural Crawling – funded under the European OpenWebSearch.EU project and carried out by researchers at the University of Pisa and the University of Glasgow, tackles this problem head-on. It proposes a new paradigm: instead of using link popularity to decide what to crawl, use AI language models to estimate the semantic quality of web pages and prioritise accordingly.

Why crawling matters more than you might think

It is easy to focus on the visible parts of a search engine – the ranking algorithms, the interface, the speed of results – and overlook the crawler. But the crawler is the primary content filter in the entire search pipeline. It decides what gets downloaded, stored, and indexed. Everything that happens downstream – indexing, ranking, retrieval – operates only on the content the crawler has already collected. A sophisticated ranking algorithm cannot compensate for a poor crawling strategy: if valuable pages were never downloaded, they simply do not exist as far as the search engine is concerned.

The web is vast, and no crawler can download everything. Choices must be made, and the heuristics that guide those choices shape the quality of the entire search corpus. Traditional heuristics like PageRank assume that a page’s importance can be inferred from its position in the web’s link structure. This works well when search queries are short keyword strings and when the most popular pages tend to be the most useful. But the FUN team argues that this assumption is increasingly outdated.

The Shift: From link popularity to semantic quality

The core idea behind neural crawling is straightforward: instead of asking “How popular is this page?”, the crawler asks “How likely is this page to contain content that would be useful for answering a search query?” To answer this question, the system uses a neural quality estimator – a small language model that has been trained to predict, from the text of a document alone, whether that document is likely to be relevant to any query. The model does not need to know what the query will be; it assesses the intrinsic quality of the text itself: its coherence, informativeness, and semantic richness.

There is an obvious practical problem: the crawler needs to decide whether to prioritise a page before it has downloaded that page. It cannot read the text of a page it has not yet fetched. The FUN team addresses this with two quality propagation strategies. The first is based on the observation that web pages tend to link to other pages of similar quality. If a high-quality page links to an unknown page, there is a reasonable probability that the unknown page is also of decent quality. The crawler can therefore use the quality of already-downloaded pages as a proxy for the likely quality of the pages they link to.
The second strategy works at the domain level: pages within the same domain tend to have similar quality. Once the crawler has downloaded a few pages from a domain, it can estimate the quality of the domain as a whole and use that estimate to prioritise other pages from the same domain.

What the experiments show

The team tested their approach through large-scale simulations on ClueWeb22-B, a web corpus of 87 million pages, using two different sets of test queries. One set consisted of traditional keyword queries; the other consisted of natural language questions.

The results are striking. On natural language queries, the neural crawling strategies consistently outperformed PageRank in both the quality of the crawled corpus and the effectiveness of downstream search results. The domain-level strategy (DomQ) was particularly strong, building corpora that led to substantially better retrieval performance. On traditional keyword queries, the neural strategies performed comparably to PageRank – they did not lose ground on the type of search that PageRank was designed for.

The efficiency results were also notable. Neural crawlers collected relevant pages faster than PageRank in the early stages of the crawl, meaning they built useful search corpora more quickly and with less wasted bandwidth downloading low-value pages. This matters in practice, because crawling the web is expensive in terms of network resources, storage, and computing time.

A key finding underpinning the domain-level approach is that the semantic quality of a web page is strongly correlated with the average quality of other pages on the same domain (Pearson correlation of 0.649). By contrast, the equivalent correlation for PageRank scores is much weaker (0.272). In other words, knowing that a domain tends to host high-quality content is a much better predictor of individual page quality than knowing that a domain is well-linked.

Why it matters

The FUN project is significant for the OpenWebSearch.EU project in a very direct way. Running an open European web index requires crawling decisions – and the quality of those decisions determines the quality of the index. If an open web index is built using traditional crawling heuristics, it inherits the biases of those heuristics: a preference for popular, well-linked content at the expense of semantically rich but less connected pages. Neural crawling offers a way to build corpora that are better suited to the modern demands of natural language search and AI-powered information retrieval.

To make this practical, the FUN team produced not just research findings but usable software. Their quality scoring tools are compatible with the OWS parquet file format and integrated into Resilipipe, the open-source content analysis framework used by OpenWebSearch.eu.

The FUN project demonstrates that the way we crawl the web should evolve alongside the way we search it. As search queries become more conversational and AI systems become major consumers of search infrastructure, the assumption that link popularity is the best guide for crawling priorities is no longer sufficient. Neural quality estimation offers a complementary – and in many cases superior – signal.

What’s next

Future work could explore combining neural and link-based signals in hybrid strategies, using ensembles of quality estimators that assess different dimensions of page quality (spam, machine-generated content, factual accuracy), and evaluating neural crawling with more advanced retrieval models beyond BM25. The approach could also be adapted for other tasks that depend on corpus quality, such as building high-quality training data for large language models.

To read the full technical report, go here: https://zenodo.org/records/17359141

The FUN project was funded under the OpenWebSearch.EU project (Horizon Europe, Grant Agreement 101070014, Call #2).

How Dutch municipalities are sharing Search Intelligence to serve citizens better: Inside the CIFFIL Service project

The CIFFIL Service project shows that open web index standards can help small municipalities improve their search quality by accessing results from larger ones

Search engines work best when they have a lot of data to learn from. The more documents in a collection, the better the system can distinguish between common words and genuinely informative ones – and therefore the better it can identify what is relevant to a query. This is a well-known principle in information retrieval, and it creates an obvious problem for anyone who needs to search a small collection of documents: the search results are simply not as good as they could be.

The CIFFIL Service project, funded under the European OpenWebSearch.EU project, tackled exactly this problem – in a setting with direct consequences for citizens. Spinque, a Dutch search technology company, builds search systems for municipalities that allow council members and residents to search through publicly available government documents. Some of these municipal collections are small, containing fewer than 10,000 documents, and are full of domain-specific jargon. The result is that search quality suffers. The CIFFIL project set out to fix this by allowing municipalities to share their search index data with one another through an open standard.

The problem: Small collections, unreliable statistics

Most search engines use some variant of a ranking algorithm called BM25. At its core, BM25 judges the relevance of a document to a query by looking at how often the query terms appear in the document and how rare those terms are across the collection as a whole. Terms that do not appear in a lot of documents signal relevance.

This is where small collections often fall short. When a collection has only a few thousand documents, the estimates of how common or rare a term is very unreliable. The ranking algorithm, relying on these skewed statistics, makes poor decisions about what is relevant. The result for the user is a search experience that feels hit-or-miss.

The solution: Sharing index data through an open solution

The CIFFIL team’s approach is simple. If a small municipality’s search system suffers from unreliable statistics because its collection is too small, why not supplement those statistics with data from a larger municipality that deals with similar types of documents? After all, Dutch municipal documents share a common vocabulary of administrative, legal, and policy language.

The technical mechanism for this sharing is the Common Index File Format, or CIFF – an open standard developed in the information retrieval research community for exchanging inverted index data between systems. An inverted index is the core data structure behind a search engine: it maps every term in a collection to the documents in which that term appears, along with statistics such as how often it appears and in how many documents.

Spinque integrated CIFF support into its search platform, Spinque Desk. This involved building a CIFF reader (to import index data), a CIFF writer (to export index data), and – critically – a modified BM25 ranking component that can combine the statistics from a local collection with those from an external CIFF index. When a small municipality’s search system uses this combined approach, it effectively “borrows” the larger municipality’s understanding of which terms are common and which are rare, while still searching its own documents.

Proof of concept

The team implemented tests for the functionalities implemented for this project. Specifically, they did manual testing by doing experiments using CIFF exports to see if they could replicate effectiveness results on open datasets. Additionally, they implemented unit tests to ensure the parser and writer were producing indexes according to the CIFF specifications.

The results were clear. The small collection performed substantially worse than the baseline, confirming that skewed statistics degrade search quality. But when the small collection borrowed statistics from the larger one, performance not only recovered but actually slightly exceeded the baseline – because the small collection, now ranked with accurate statistics, contained a higher concentration of relevant documents.

In practice

The project created CIFF indices for four major Dutch municipalities: Amsterdam, Utrecht, Nijmegen, and Almere. A live deployment was initiated for the municipality of Nieuwegein, a smaller city near Utrecht, using the Utrecht index as the background collection. Evaluation of the real-world impact on user experience is ongoing.

All of the CIFF tools developed during the project have been released as open-source software, and the export service ensures that published indices are automatically updated when the underlying data changes.

Why it matters

The CIFFIL project illustrates a principle that is central to the OpenWebSearch.eu core idea: that open, interoperable standards can enable forms of cooperation that proprietary systems cannot. By sharing index statistics through CIFF, municipalities can improve their search quality without sharing their actual documents, without depending on a single commercial provider, and without each needing to build a large collection of their own. It is a form of search infrastructure as a public good.

The approach is also notable for its simplicity. It does not require neural models, large language models, or expensive computational resources. It works by making better use of data that already exists, through a well-understood ranking algorithm and an open file format.

What’s next

The immediate priorities are completing the open publication of all four municipal indices, conducting user-experience evaluations in the live deployments, and publishing the experimental findings as a research paper. Longer-term, the approach could be extended to additional municipalities and to other domains where small document collections need better search – such as cultural heritage institutions, local archives, or specialised libraries. The underlying principle – that sharing standardised index data can improve search quality without centralising control – has broad applicability wherever open, cooperative search infrastructure is valued.

Find the full project report here: https://zenodo.org/records/17750643

The CIFFIL Service project was funded under the OpenWebSearch.EU project (Horizon Europe, Grant Agreement 101070014, Call #2).