IDEA GROUP PUBLISHING 14 Chan,701 Chen, & Ho E. Chocolate Avenue, Suite 200, Hershey PA 17033-1240, USA 16*'$! Tel: 717/533-8845; Fax 717/533-8661; URL-http://www.idea-group.com Chapter II A Duplicate Chinese Document Image Retrieval System Based on Line Segment Feature in Character Image Block Yung-Kuan Chan, National Huwei Institute of Technology, Taiwan Tung-Shou Chen, National Taichung Institute of Technology, Taiwan Yu-An Ho, National Taichung Institute of Technology, Taiwan ABSTRACT With the rapid progress of digital image technology, the management of duplicate document images is also emphasized widely. As a result, this paper suggests a duplicate Chinese document image retrieval (DCDIR) system, which uses the ratio of the number of black pixels to that of white pixels on the scanned line segments in a character image block as the feature of the character image block. Experimental results indicate that the system can indeed effectively and quickly retrieve the desired duplicate Chinese document image from a database. INTRODUCTION With the widespread popularity of electronic documents as well as electronic books, more and more people have started reading these new sorts of publications. Most ThisCopyright chapter appears the book, Systems and Content-Based Retrieval edited by © 2004,in Idea GroupMultimedia Inc. Copying or distributing in print or Image electronic forms, without written Sagarmay Deb. ofCopyright © 2004, Group Inc. Copying or distributing in print or electronic forms permission Idea Group Inc. isIdea prohibited. without written permission of Idea Group Inc. is prohibited. Duplicate Chinese Document Image Retrieval System 15 of the electronic books and electronic documents are stored in an image format. As for traditional newspapers or magazines, they can also be transformed into digital image formatted data by using the tool like a scanner, and then saved in a computer for the purpose of data backup. Meanwhile, by combining with the characteristic of rapid data processing of a computer, the data of the duplicate documents can be quickly acquired and retrieved. Because the technology of Optical Character Recognition (OCR) has been improved to a degree, a document image can be transformed to become the data with Big5 code format, and only the Big5 codes are stored so as to reduce the required memory space. However, using the technique of OCR may cause some errors in document contents because of recognition failure. Moreover, in order to preserve the original contents and states of a document, transforming a document into an image and storing the image in a database is sometimes necessary. The primary function of a duplicate document image retrieval (DDIR) system is to transform the data in document format into those represented in digital image format. This transformation would be completed by using the tool like a scanner. Then, the DDIR system will store these images and their corresponding features in a database for data backup purpose. Here the document image is called a duplicate document image. When intending to retrieve a duplicate document image, users can also use the tool like a scanner to input the first several text lines of the original document into the system, and to create a query document image and figure out the feature of the image. The DDIR system finally transmits the users the duplicate document image whose image feature is similar to that of the query document image. In the aspect of extracting the feature from a document image, many techniques have been proposed (Caprari, 2000; Doermann, Li, & Kia, 1997; Peng et al., 2001). Caprari (2001) used the method of document template to measure the overlapping ratio of black pixels and white pixels between two document images when their two templates overlap. Generally, the higher the overlapping ratio is, the more similar these two documents are. Doermann, Li and Kia (1997) encoded the heights of English letter shapes and used these codes as the feature of the English document image. Meanwhile, by comparing and matching these feature codes, the duplicate document images that are stored in the database can be searched and retrieved. Besides, Peng et al. (2001) applied the block sizes and locations to be the feature of a document image, and suggested a component block list matching technique based on this feature for the duplicate document image retrieval. The features mentioned above are only suitable for stating the characteristic of an English document image. The characteristics of Chinese characters are different from those of English ones, and the strokes and shapes of Chinese characters are much more complicated than those of English characters. Therefore, this paper suggests a line segment feature, which is appropriate for stating the characteristics of the Chinese character image blocks in a Chinese document image, and applies this feature to construct a duplicate Chinese document image retrieval (DCDIR) system. The next section will give a brief review of some related works. The following section will describe the line segment feature in a character image block proposed in this chapter. Then, the chapter will delineate how to make use of this proposed feature to create a DCDIR system. This is followed with a statement on the experimental and analyzed results. This is followed by a discussion on the future trends of the duplicate document image retrieval system. At last, the conclusions will be given. Copyright © 2004, Idea Group Inc. Copying or distributing in print or electronic forms without written permission of Idea Group Inc. is prohibited. 8 more pages are available in the full version of this document, which may be purchased using the "Add to Cart" button on the publisher's webpage: www.igi-global.com/chapter/duplicate-chinese-documentimage-retrieval/27053 Related Content 3D DMB Player and Its Reliable 3D Services in T-DMB Systems Cheolkon Jung and Licheng Jiao (2012). Depth Map and 3D Imaging Applications: Algorithms and Technologies (pp. 434-450). www.irma-international.org/chapter/dmb-player-its-reliable-services/60279/ Spontaneous Facial Expression Analysis and Synthesis for Interactive Facial Animation Yongmian Zhang, Jixu Chen, Yan Tong and Qiang Ji (2011). Computer Vision for Multimedia Applications: Methods and Solutions (pp. 20-37). www.irma-international.org/chapter/spontaneous-facial-expression-analysissynthesis/48311/ Enhancing Robustness in Speech Recognition using Visual Information Omar Farooq and Sekharjit Datta (2012). Speech, Image, and Language Processing for Human Computer Interaction: Multi-Modal Advancements (pp. 149-171). www.irma-international.org/chapter/enhancing-robustness-speechrecognition-using/65058/ Relevant Feature Subset Selection from Ensemble of Multiple Feature Extraction Methods for Texture Classification Bharti Rana, Akanksha Juneja and Ramesh Kumar Agrawal (2015). International Journal of Computer Vision and Image Processing (pp. 48-65). www.irma-international.org/article/relevant-feature-subset-selection-fromensemble-of-multiple-feature-extraction-methods-for-textureclassification/151507/ Science of Emoticons: Research Framework and State of the Art in Analysis of kaomoji-type Emoticons Michal Ptaszynski, Jacek Maciejewski, Pawel Dybala, Rafal Rzepka, Kenji Araki and Yoshio Momouchi (2012). Speech, Image, and Language Processing for Human Computer Interaction: Multi-Modal Advancements (pp. 234-260). www.irma-international.org/chapter/science-emoticons-research-frameworkstate/65062/
© Copyright 2026 Paperzz