NDLI: The Mining and Extraction of Primary Informative Blocks and Data Objects from Systematic Web Pages

Content Provider	IEEE Xplore Digital Library
Author	Yi-Feng Tseng Hung-Yu Kao
Copyright Year	2006
Description	Author affiliation: Dept. of Comput. Sci. & Inf. Eng., Nat. Cheng Kung Univ., Tainan (Yi-Feng Tseng; Hung-Yu Kao)
Abstract	With the fast development of Internet, the Web has already been an enormous database so far, which contains extremely abundant information. Most of Web pages are represented their content by using a list of objects, such as search engine results, product information of shopping Web sites and so on, and these objects form the primary information of each page. In this paper, we focus on the issues of mining primary information and the constituted object groups. The system is divided into three major phases: (1) By transforming each Web page into corresponding tree structures, our system can visit all regions of the Web page in an efficient way, and detects the informative parts. (2) We design and quantize several novel features according to the characters of regions of a Web page. (3) A weighting model is proposed that calculates the important degree of each region, we then extract the primary information of the Web pages. The experimental result proves our system can be applied to a large number of Web pages with different themes and styles to find the correct primary information and the list of corresponding objects
Sponsorship	IEEE Comput. Soc. WIC ACM
Starting Page	370
Ending Page	373
File Size	460541
Page Count	4
File Format	PDF
ISBN	0769527477
DOI	10.1109/WI.2006.167
Language	English
Publisher	Institute of Electrical and Electronics Engineers, Inc. (IEEE)
Publisher Date	2006-12-18
Publisher Place	China
Access Restriction	Subscribed
Rights Holder	Institute of Electrical and Electronics Engineers, Inc. (IEEE)
Subject Keyword	DOM Humans Data engineering HTML Data mining Information analysis Computer science Web pages Search engines Information Extraction Block Importance Internet Books
Content Type	Text
Resource Type	Article

Sl.	Authority	Responsibilities	Communication Details
1	Ministry of Education (GoI), Department of Higher Education	Sanctioning Authority	https://www.education.gov.in/ict-initiatives
2	Indian Institute of Technology Kharagpur	Host Institute of the Project: The host institute of the project is responsible for providing infrastructure support and hosting the project	https://www.iitkgp.ac.in
3	National Digital Library of India Office, Indian Institute of Technology Kharagpur	The administrative and infrastructural headquarters of the project	Dr. B. Sutradhar bsutra@ndl.gov.in
4	Project PI / Joint PI	Principal Investigator and Joint Principal Investigators of the project	Dr. B. Sutradhar bsutra@ndl.gov.in Prof. Saswat Chakrabarti will be added soon
5	Website/Portal (Helpdesk)	Queries regarding NDLI and its services	support@ndl.gov.in
6	Contents and Copyright Issues	Queries related to content curation and copyright issues	content@ndl.gov.in
7	National Digital Library of India Club (NDLI Club)	Queries related to NDLI Club formation, support, user awareness program, seminar/symposium, collaboration, social media, promotion, and outreach	clubsupport@ndl.gov.in
8	Digital Preservation Centre (DPC)	Assistance with digitizing and archiving copyright-free printed books	dpc@ndl.gov.in
9	IDR Setup or Support	Queries related to establishment and support of Institutional Digital Repository (IDR) and IDR workshops	idr@ndl.gov.in

Data Extraction Based on Index Path in Web

Web information extraction

WISE: a visual tool for automatic extraction of objects from World Wide Web

Removing non-informative blocks from the web pages

Web informative content block detecting based on entropy and parent-child relationship in DOM

A Novel Method to Extract Informative Blocks from Web Pages

Extraction of Informative Blocks from Web Pages

ViWER- data extraction for search engine results pages using visual cue and DOM Tree

Automatic identification of informative sections of Web pages

The Mining and Extraction of Primary Informative Blocks and Data Objects from Systematic Web Pages

Similar Documents

Data Extraction Based on Index Path in Web

Web information extraction

WISE: a visual tool for automatic extraction of objects from World Wide Web

Removing non-informative blocks from the web pages

Web informative content block detecting based on entropy and parent-child relationship in DOM

A Novel Method to Extract Informative Blocks from Web Pages

Extraction of Informative Blocks from Web Pages

ViWER- data extraction for search engine results pages using visual cue and DOM Tree

Automatic identification of informative sections of Web pages

The Mining and Extraction of Primary Informative Blocks and Data Objects from Systematic Web Pages