SOTAVerified

Dr Web: a modern, query-based web data retrieval engine

2025-02-18Code Available0· sign in to hype

Ylli Prifti, Alessandro Provetti, Pasquale De Meo

Code Available — Be the first to reproduce this paper.

Reproduce

Code

Abstract

This article introduces the Data Retrieval Web Engine (also referred to as doctor web), a flexible and modular tool for extracting structured data from web pages using a simple query language. We discuss the engineering challenges addressed during its development, such as dynamic content handling and messy data extraction. Furthermore, we cover the steps for making the DR Web Engine public, highlighting its open source potential.

Tasks

Reproductions