Web Data Extraction and Integration in Domain

2018 ◽

Vol 9 (3) ◽

pp. 752 ◽

Cited By ~ 2

Author(s):

Ily Amalina Ahmad Sabri ◽

Mustafa Man

Keyword(s):

Data Extraction ◽

Extraction Process ◽

Structured Data ◽

Web Pages ◽

Web Page ◽

Web Data ◽

Web Documents ◽

Web Extraction ◽

Comparison Time ◽

The Web

<p>Web data extraction is the process of extracting user required information from web page. The information consists of semi-structured data not in structured format. The extraction data involves the web documents in html format. Nowadays, most people uses web data extractors because the extraction involve large information which makes the process of manual information extraction takes time and complicated. We present in this paper WEIDJ approach to extract images from the web, whose goal is to harvest images as object from template-based html pages. The WEIDJ (Web Extraction Image using DOM (Document Object Model) and JSON (JavaScript Object Notation)) applies DOM theory in order to build the structure and JSON as environment of programming. The extraction process leverages both the input of web address and the structure of extraction. Then, WEIDJ splits DOM tree into small subtrees and applies searching algorithm by visual blocks for each web page to find images. Our approach focus on three level of extraction; single web page, multiple web page and the whole web page. Extensive experiments on several biodiversity web pages has been done to show the comparison time performance between image extraction using DOM, JSON and WEIDJ for single web page. The experimental results advocate via our model, WEIDJ image extraction can be done fast and effectively.</p>

Download Full-text

State-of-the-art web data extraction systems for online business intelligence

Informacijos mokslai ◽

10.15388/im.2013.0.1595 ◽

2013 ◽

Vol 64 ◽

pp. 145-155

Author(s):

Tomas Grigalis ◽

Antanas Čenys

Keyword(s):

Business Intelligence ◽

State Of The Art ◽

Data Extraction ◽

Web Pages ◽

Web Data ◽

Web Data Extraction ◽

Technological Advances ◽

Online Business ◽

A Company ◽

Support Decision Making

The success of a company hinges on identifying and responding to competitive pressures. The main objective of online business intelligence is to collect valuable information from many Web sources to support decision making and thus gain competitive advantage. However, the online business intelligence presents non-trivial challenges to Web data extraction systems that must deal with technologically sophisticated modern Web pages where traditional manual programming approaches often fail. In this paper, we review commercially available state-of-the-art Web data extraction systems and their technological advances in the context of online business intelligence.Keywords: online business intelligence, Web data extraction, Web scrapingŠiuolaikinės iš tinklalapių duomenis renkančios ir verslo analitikai tinkamos sistemos (anglų k.)Tomas Grigalis, Antanas Čenys Santrauka Šiuolaikinės verslo organizacijos sėkmė priklauso nuo sugebėjimo atitinkamai reaguoti į nuolat besikeičiančią konkurencinę aplinką. Internete veikiančios verslo analitinės sistemos pagrindinis tikslas yra rinkti vertingą informaciją iš daugybės skirtingų internetinių šaltinių ir tokiu būdu padėti verslo organizacijai priimti tinkamus sprendimus ir įgyti konkurencinį pranašumą. Tačiau informacijos rinkimas iš internetinių šaltinių yra sudėtinga problema, kai informaciją renkančios sistemos turi gerai veikti su itin technologiškai sudėtingais tinklalapiais. Šiame straipsnyje verslo analitikos kontekste apžvelgiamos pažangiausios internetinių duomenų rinkimo sistemos. Taip pat pristatomi konkretūs scenarijai, kai duomenų rinkimo sistemos gali padėti verslo analitikai. Straipsnio pabaigoje autoriai aptaria pastarųjų metų technologinius pasiekimus, kurie turi potencialą tapti visiškai automatinėmis internetinių duomenų rinkimo sistemomis ir dar labiau patobulinti verslo analitiką bei gerokai sumažinti jos išlaidas.

Download Full-text

Web Harvesting

Advances in Data Mining and Database Management - Web Usage Mining Techniques and Applications Across Industries ◽

10.4018/978-1-5225-0613-3.ch014 ◽

2017 ◽

pp. 351-378 ◽

Cited By ~ 1

Author(s):

B. Umamageswari ◽

R. Kalpana

Keyword(s):

Web Mining ◽

Data Extraction ◽

Web Pages ◽

Extraction Techniques ◽

Web Data ◽

Web Data Extraction ◽

Region Extraction ◽

Universal Applicability ◽

Web Harvesting

Web mining is done on huge amounts of data extracted from WWW. Many researchers have developed several state-of-the-art approaches for web data extraction. So far in the literature, the focus is mainly on the techniques used for data region extraction. Applications which are fed with the extracted data, require fetching data spread across multiple web pages which should be crawled automatically. For this to happen, we need to extract not only data regions, but also the navigation links. Data extraction techniques are designed for specific HTML tags; which questions their universal applicability for carrying out information extraction from differently formatted web pages. This chapter focuses on various web data extraction techniques available for different kinds of data rich pages, classification of web data extraction techniques and comparison of those techniques across many useful dimensions.

Download Full-text

Web Data Extraction with Hierarchical Clustering and Rich Features

Applied Mechanics and Materials ◽

10.4028/www.scientific.net/amm.55-57.1003 ◽

2011 ◽

Vol 55-57 ◽

pp. 1003-1008

Author(s):

Yong Quan Dong ◽

Xiang Jun Zhao ◽

Gong Jie Zhang

Keyword(s):

Hierarchical Clustering ◽

Data Extraction ◽

Web Pages ◽

Web Data ◽

Clustering Techniques ◽

Feature Vectors ◽

Web Data Extraction ◽

Novel Approach ◽

Data Elements ◽

Content Feature

A novel approach is proposed to automatically extract data records from detail pages using hierarchical clustering techniques. The approach uses the information of the listing pages to identify the content blocks in detail pages, which narrows the scopes of Web data extraction. Meanwhile, it also makes full use of the structure and content features to cluster content feature vectors. Finally, it aligns data elements of multiple details pages to extract the data records. Experiment results on test beds of real web pages show that the approach can achieve high extraction accuracy and outperforms the existing techniques substantially.

Download Full-text

Efficient Methodology for Deep Web Data Extraction

Turkish Journal of Computer and Mathematics Education (TURCOMAT) ◽

10.17762/turcomat.v12i1s.1769 ◽

2021 ◽

Vol 12 (1S) ◽

pp. 286-293

Author(s):

Shilpa Deshmukh, Et. al.

Keyword(s):

Information Extraction ◽

Data Extraction ◽

Deep Web ◽

Web Pages ◽

Web Data ◽

Web Information Extraction ◽

Web Data Extraction ◽

Web Information ◽

Assessment Measure ◽

Enormous Number

Deep Web substance are gotten to by inquiries submitted to Web information bases and the returned information records are enwrapped in progressively created Web pages (they will be called profound Web pages in this paper). Removing organized information from profound Web pages is a difficult issue because of the fundamental mind boggling structures of such pages. As of not long ago, an enormous number of strategies have been proposed to address this issue, however every one of them have characteristic impediments since they are Web-page-programming-language subordinate. As the mainstream two-dimensional media, the substance on Web pages are constantly shown routinely for clients to peruse. This inspires us to look for an alternate path for profound Web information extraction to beat the constraints of past works by using some fascinating normal visual highlights on the profound Web pages. In this paper, a novel vision-based methodology that is Visual Based Deep Web Data Extraction (VBDWDE) Algorithm is proposed. This methodology basically uses the visual highlights on the profound Web pages to execute profound Web information extraction, including information record extraction and information thing extraction. We additionally propose another assessment measure amendment to catch the measure of human exertion expected to create wonderful extraction. Our investigations on a huge arrangement of Web information bases show that the proposed vision-based methodology is exceptionally viable for profound Web information extraction.

Download Full-text

Semantic educational data extraction using structur-al domain relationships

International Journal of Engineering & Technology ◽

10.14419/ijet.v7i1.2.9060 ◽

2017 ◽

Vol 7 (1.2) ◽

pp. 175

Author(s):

Manchikatla Srikanth

Keyword(s):

Adaptive Learning ◽

Data Extraction ◽

Dynamic Environment ◽

Mining Industry ◽

Specific Area ◽

Learning System ◽

Web Pages ◽

Web Data ◽

Educational Content ◽

Web Data Extraction

In the mining industry', some of the domains are most popular and it plays a vital role in the specific area. Educational Mining and Web-Data Extraction are the two important factors play a leading role in mining industry. The main objective of the proposed system is to extract the related contents from web using semantic (relating to meaning in language or logic) principles as well as to allow the providers to dynamically generate the web pages for educational content and allow the users to search and extract the data from server based on content. The main model of this system is to illustrate the adaptive learning system. For demonstration we consider the semantic principles for Educational content over dynamic environment. This site allows the providers to create web pages related to educational content dynamically and this will be getting approved by the Administrator to live in process. Once the site is live the users can search for the exact content present into the site based on semantic principles. The proposed model is designed for dynamic web data extraction and content analysis from the extracted data due to educational principles. In the proposed system Semantic Web Extraction (SWE) procedures are highly analyzed and utilized for content manipulations. Energetic data extraction scheme for users based on educational content rather than header, title, meta tags and descriptions.

Download Full-text

Web Harvesting

The Dark Web ◽

10.4018/978-1-5225-3163-0.ch010 ◽

2018 ◽

pp. 199-226 ◽

Cited By ~ 1

Author(s):

B. Umamageswari ◽

R. Kalpana

Keyword(s):

Web Mining ◽

Data Extraction ◽

Web Pages ◽

Extraction Techniques ◽

Web Data ◽

Web Data Extraction ◽

Region Extraction ◽

Universal Applicability ◽

Web Harvesting

Web mining is done on huge amounts of data extracted from WWW. Many researchers have developed several state-of-the-art approaches for web data extraction. So far in the literature, the focus is mainly on the techniques used for data region extraction. Applications which are fed with the extracted data, require fetching data spread across multiple web pages which should be crawled automatically. For this to happen, we need to extract not only data regions, but also the navigation links. Data extraction techniques are designed for specific HTML tags; which questions their universal applicability for carrying out information extraction from differently formatted web pages. This chapter focuses on various web data extraction techniques available for different kinds of data rich pages, classification of web data extraction techniques and comparison of those techniques across many useful dimensions.

Download Full-text

Web Data Extraction Based on Tag Path Clustering

Advanced Materials Research ◽

10.4028/www.scientific.net/amr.756-759.1590 ◽

2013 ◽

Vol 756-759 ◽

pp. 1590-1594

Author(s):

Gui Li ◽

Cheng Chen ◽

Zheng Yu Li ◽

Zi Yang Han ◽

Ping Sun

Keyword(s):

Data Extraction ◽

Structured Data ◽

Web Pages ◽

Web Data Extraction ◽

Web Document ◽

Automatic Methods ◽

Simple Extraction ◽

Web Document Clustering ◽

Fully Automatic ◽

The Web

Fully automatic methods that extract structured data from the Web have been studied extensively. The existing methods suffice for simple extraction, but they often fail to handle more complicated Web pages. This paper introduces a method based on tag path clustering to extract structured data. The method gets complete tag path collection by parsing the DOM tree of the Web document. Clustering of tag paths is performed based on introduced similarity measure and the data area can be targeted, then taking advantage of features of tag position, we can separate and filter record, finally complete data extraction. Experiments show this method achieves higher accuracy than previous methods.

Download Full-text

A deep web data extraction model for web mining: a review

Indonesian Journal of Electrical Engineering and Computer Science ◽

10.11591/ijeecs.v23.i1.pp519-528 ◽

2021 ◽

Vol 23 (1) ◽

pp. 519

Author(s):

Ily Amalina Ahmad Sabri ◽

Mustafa Man

Keyword(s):

Web Mining ◽

Data Extraction ◽

Deep Web ◽

Quality Data ◽

Web Pages ◽

Web Data ◽

Comprehensive Overview ◽

Web Data Extraction ◽

Proposed Model ◽

Extraction Model

The World Wide Web has become a large pool of information. Extracting structured data from a published web pages has drawn attention in the last decade. The process of web data extraction (WDE) has many challenges, dueto variety of web data and the unstructured data from hypertext mark up language (HTML) files. The aim of this paper is to provide a comprehensive overview of current web data extraction techniques, in termsof extracted quality data. This paper focuses on study for data extraction using wrapper approaches and compares each other to identify the best approach to extract data from online sites. To observe the efficiency of the proposed model, we compare the performance of data extraction by single web page extraction with different models such as document object model (DOM), wrapper using hybrid dom and json (WHDJ), wrapper extraction of image using DOM and JSON (WEIDJ) and WEIDJ (no-rules). Finally, the experimentations proved that WEIDJ can extract data fastest and low time consuming compared to other proposed method.<br /><div> </div>

Download Full-text