A copy of this work was available on the public web and has been preserved in the Wayback Machine. The capture dates from 2018; you can also visit <a rel="external noopener" href="https://link.springer.com/content/pdf/10.1007%2F978-3-642-31753-8_38.pdf">the original URL</a>. The file type is <code>application/pdf</code>.
Clustering Visually Similar Web Page Elements for Structured Web Data Extraction
[chapter]
<span title="">2012</span>
<i title="Springer Berlin Heidelberg">
<a target="_blank" rel="noopener" href="https://fatcat.wiki/container/2w3awgokqne6te4nvlofavy5a4" style="color: black;">Lecture Notes in Computer Science</a>
</i>
We propose a novel approach for extraction of structured web data called ClustVX. It clusters visually similar web page elements by exploiting their visual formatting and structural features. Clusters are then used to derive extraction rules. The experimental evaluation results of ClustVX system on three publicly available benchmark data sets outperform state-of-the-art structured data extraction systems.
<span class="external-identifiers">
<a target="_blank" rel="external noopener noreferrer" href="https://doi.org/10.1007/978-3-642-31753-8_38">doi:10.1007/978-3-642-31753-8_38</a>
<a target="_blank" rel="external noopener" href="https://fatcat.wiki/release/v5inxjfvyfe6xghcwfkcouxaxa">fatcat:v5inxjfvyfe6xghcwfkcouxaxa</a>
</span>
<a target="_blank" rel="noopener" href="https://web.archive.org/web/20180726144532/https://link.springer.com/content/pdf/10.1007%2F978-3-642-31753-8_38.pdf" title="fulltext PDF download" data-goatcounter-click="serp-fulltext" data-goatcounter-title="serp-fulltext">
<button class="ui simple right pointing dropdown compact black labeled icon button serp-button">
<i class="icon ia-icon"></i>
Web Archive
[PDF]
<div class="menu fulltext-thumbnail">
<img src="https://blobs.fatcat.wiki/thumbnail/pdf/80/32/80328aad0112379fc95d60d6281041866eae0018.180px.jpg" alt="fulltext thumbnail" loading="lazy">
</div>
</button>
</a>
<a target="_blank" rel="external noopener noreferrer" href="https://doi.org/10.1007/978-3-642-31753-8_38">
<button class="ui left aligned compact blue labeled icon button serp-button">
<i class="external alternate icon"></i>
springer.com
</button>
</a>