A Framework For Extracting Information From Web Using VTD-XML‘s XPath

Journal Title: International Journal on Computer Science and Engineering - Year 2012, Vol 4, Issue 3

Abstract

The exponential growth of WWW (World Wide Web) is the cause for vast pool of information as well as several challenges posed by it, such as extracting potentially useful and unknown information from WWW. Many websites are built with HTML, because of its unstructured layout, it is difficult to obtain effective and precise data from web using HTML. The advent of XML (Extensible Markup Language) proposes a better solution to extract useful knowledge from WWW. Web Data Extraction based on XML Technology solves this problem because XML is a general purpose specification for exchanging data over the Web. In this paper, a framework is suggested to extract the data from the web. Here the semi-structured data in the web page is transformed into well-structured data using standard XML technologies and the new parsing technique called extended VTD-XML (Virtual Token Descriptor for XML) along with Xpath implementation has been used to extract data from the well-structured XML document.

Authors and Affiliations

C. Subhashini , Dr. Arti Arya

Keywords

Related Articles

Generating Membership Values And Fuzzy Association Rules From Numerical Data

The most important task in the design of fuzzy classification systems is to find a set of fuzzy rules from training data to deal with a specific classification problem. In this paper, a method to generate fuzzy rules fro...

Modelisation of the maintenanceproduction couple by a graph of scheduling

The problem of scheduling, both maintenance and production, was a work of many articles. The algorithms genetics, ant, and multi-agent, gives solutions approaches with this problem. Modelled the problem in a simple and c...

Secured, Authenticated Communication Model for Dynamic Multicast Groups

Secure Multicast networks forms the backbone for many web and multimedia applications such as Interactive TV, Teleconference etc. The main challenge for secure multicast is scalability, efficiency and authenticity. A co...

An Algorithm for Frequent Pattern Mining Based On Apriori

Frequent pattern mining is a heavily researched area in the field of data mining with wide range of applications. Mining frequent patterns from large scale databases has emerged as an important problem in data mining and...

Simulation and Analysis of Digital Video Watermarking Using MPEG-2

Quantization Index Modulation (QIM) is an important method for embedding digital watermark signal with information. This technique achieves very efficient tradeoffs among watermark embedding rate, the amount of embedding...

Download PDF file
  • EP ID EP129958
  • DOI -
  • Views 115
  • Downloads 0

How To Cite

C. Subhashini, Dr. Arti Arya (2012). A Framework For Extracting Information From Web Using VTD-XML‘s XPath. International Journal on Computer Science and Engineering, 4(3), 463-468. https://europub.co.uk/articles/-A-129958